Normalizing flows form a flexible class of invertible transformations for learning probability distributions and can be interpreted as flows on spaces of measures. In this talk, we introduce a theoretical framework in which normalizing flows are viewed as approximations of optimal transport maps, constructed via neural differential equations with a control-linear structure. Within this framework, we establish approximation results showing that suitably constrained neural ODEs can approximate optimal transport maps between absolutely continuous measures. To obtain a tractable finite-dimensional optimization problem, the transport is approximated using discrete empirical measures; consistency as the number of atoms grows is ensured by a $\Gamma$-convergence result. However, the optimal transport plans associated with these discrete measures encode information only in an $L^2$-type topology. This leads to a mismatch with the underlying approximation results for diffeomorphisms, which are formulated in a stronger topology, namely uniform convergence on compact sets, thereby hindering a direct application of these results. We discuss ongoing work aimed at bridging this gap by incorporating risk measures into the optimization problem, providing a principled way to interpolate between $L^2$ and $C^0$ topologies. Finally, we outline future directions toward deriving quantitative estimates, with the goal of expressing the approximation error of the optimal transport map in terms of the Wasserstein distance between discrete empirical measures and their continuous counterparts. (Joint with Sara Farinelli)