SCI Publications

2025


A. Salinas, I. Sohail, V. Pascucci, P. Stefanakis, S. Amjad, A. Panta, R. Schigas, T. Chun-Yiu Chui, N. Duboc, M. Farrokhabadi, R. Stull. “Climate Data for Power Systems Applications: Lessons in Reusing Wildfire Smoke Data for Solar PV Studies,” Subtitled “arXiv:2509.09888v2,” 2025.

ABSTRACT

Data reuse is using data for a purpose distinct from its original intent. As data sharing becomes more prevalent in science, enabling effective data reuse is increasingly important. In this paper, we present a power systems case study of data repurposing for enabling data reuse. We define data repurposing as the process of transforming data to fit a new research purpose. In our case study, we repurpose a geospatial wildfire smoke forecast dataset into a historical dataset. We analyze its efficacy toward analyzing wildfire smoke impact on solar photovoltaic energy production. We also provide documentation and interactive demos for using the repurposed dataset. We identify key enablers of data reuse including metadata standardization, contextual documentation, and communication between data creators and reusers. We also identify obstacles to data reuse such as risk of misinterpretation and barriers to efficient data access. Through an iterative approach to data repurposing, we demonstrate how leveraging and expanding knowledge transfer infrastructures like online documentation, interactive visualizations, and data streaming directly address these obstacles. The findings facilitate big data use from other domains for power systems applications and grid resiliency. 



A. Samanta, Y. Jiang, R. Stutsman, R.B. Roy. “Not Just Fast, But Also Sustainable: Rethinking Network Routing,” In MAIoT '25: Proceedings of the Middleware for Autonomous AIoT Systems in the Computing Continuum , ACM, 2025.

ABSTRACT

The carbon footprint of networking poses a significant environmental concern. In this paper, we show how transmission latency-focused traditional routing can be sub-optimal in terms of carbon footprint. Based on the dynamic factors that affect a network’s carbon footprint and transmission latency, we design a carbon-aware routing solution. We evaluate our solution on real backbone Internet topologies to show that there is an opportunity to minimize the network’s carbon emissions with only modest latency penalties.



A. Samanta, R. Stutsman, R.B. Roy. “GridGreen: Integrating Serverless Computing in HPC Systems for Performance and Sustainability,” In SoCC '25: Proceedings of the 2025 ACM Symposium on Cloud Computing, ACM, 2025.

ABSTRACT

We present GridGreen, a scheduling framework that improves the sustainability and performance of scientific workflow execution by integrating serverless computing with traditional on-premise high performance computing (HPC) clusters. GridGreen allocates workflow components across HPC and serverless environments leveraging spatio-temporal variation of carbon intensity and component execution characteristics. It incorporates component-level optimization, speculative pre-warming, I/O-aware data management, and fallback adaptation to jointly minimize carbon footprint and service time under user-defined cost constraints. Our evaluations on large-scale bioinformatics workflows across leadership-class HPC facilities and cloud-based serverless regions demonstrate that GridGreen achieves robust, cost-effective execution while improving carbon efficiency.



A. Samanta, Y. Jiang, R. Stutsman, R.B. Roy. “Water Footprint of Datacenter Applications: Methodological Implications of Manufacturing, Operational, and Decommissioning Phases,” In SoCC '25: Proceedings of the 2025 ACM Symposium on Cloud Computing , ACM, 2025.

ABSTRACT

Rising computational demands have made cloud datacenters’ water footprint a critical concern. We demonstrate how different water footprint accounting methodologies – incorporating operational, manufacturing, and decommissioning water consumption, impact measurements and highlight the need for methodology standardization for water-aware operations. Our analysis reveals opportunities for water-aware scheduling in datacenters by considering regional water variations and lifecycle impacts.



C. Scully-Allison, K. Williams, S. Brink, O. Pearce, K. Isaacs. “A Tale of Two Models: Understanding Data Workers' Internal and External Representations of Complex Data,” Subtitled “arXiv:2501.09862v2,” 2025.

ABSTRACT

Data workers may have a different mental model of their data than the one reified in code. Understanding the organization of their data is necessary for analyzing data, be it through scripting, visualization, or abstract thought. More complicated organizations, such as tables with attached hierarchies, may tax people’s ability to think about and interact with data. To better understand and ultimately design for these situations, we conduct a study across a team of ten people working with the same reified data model. Through interviews and sketching, we probed their conception of the data model and developed themes through reflexive data analysis. Participants had diverse data models that differed from the reified data model, even among team members who had designed the model, resulting in parallel hazards limiting their ability to reason about the data. From these observations, we suggest potential design interventions for data analysis processes and tools.



C. Scully-Allison, K. Menear, K. Potter, A. McNutt, K.E. Isaacs, D. Duplyakin. “Same Data, Different Audiences: Using Personas to Scope a Supercomputing Job Queue Visualization,” 2025.

ABSTRACT

Domain-specific visualizations sometimes focus on narrow, albeit important, tasks for one group of users. This focus limits the utility of a visualization to other groups working with the same data. While tasks elicited from other groups can present a design pitfall if not disambiguated, they also present a design opportunity -- development of visualizations that support multiple groups. This development choice presents a trade off of broadening the scope but limiting support for the more narrow tasks of any one group, which in some cases can enhance the overall utility of the visualization. We investigate this scenario through a design study where we develop \textitGuidepost, a notebook-embedded visualization of supercomputer queue data that helps scientists assess supercomputer queue wait times, machine learning researchers understand prediction accuracy, and system maintainers analyze usage trends. We adapt the use of personas for visualization design from existing literature in the HCI and software engineering domains and apply them in categorizing tasks based on their uniqueness across the stakeholder personas. Under this model, tasks shared between all groups should be supported by interactive visualizations and tasks unique to each group can be deferred to scripting with notebook-embedded visualization design. We evaluate our visualization with nine expert analysts organized into two groups: a "research analyst" group that uses supercomputer queue data in their research (representing the Machine Learning researchers and Jobs Data Analyst personas) and a "supercomputer user" group that uses this data conditionally (representing the HPC User persona). We find that our visualization serves our three stakeholder groups by enabling users to successfully execute shared tasks with point-and-click interaction while facilitating case-specific programmatic analysis workflows. 



S. Sebbio, L. Carnevale, D. Balouek, M. Parashar, M. Villari. “Data-Driven Operational Artificial Intelligence for Computing Continuum: A Natural Disaster Management Use Case,” In 2025 IEEE 25th International Symposium on Cluster, Cloud and Internet Computing Workshops (CCGridW), pp. 92-99. 2025.

ABSTRACT

The increasing frequency of documented natural disasters can be attributed to advances in communication technologies, such as satellites, the Internet, and smart devices that facilitate better disaster reporting. This is coupled with an actual rise in the occurrence of such events and improved documentation of their impacts. These trends underscore the pressing need for scalable and intelligent technological solutions to efficiently process large datasets, allowing informed decision-making and effective disaster response. This study presents a Computing Continuum framework that integrates intelligence across cloud, edge and deep edge tiers for efficient disaster data processing. A significant characteristic is the incorporation of Artificial Intelligence for IT Operations (AIOps), which leverages machine learning and analytics to facilitate dynamic resource management and adaptive system modeling, thereby addressing the intricate challenges posed by disaster scenarios. The architecture encompasses an AI-driven framework for monitoring and managing service, network, and infrastructure layers, tailoring policies to specific disaster needs. The proposed framework is applied to wildfire management, leveraging an AI Operation Manager to coordinate sensor-equipped drones for real-time data acquisition and processing. Operating at the deep edge tier, these drones transmit environmental data to edge and cloud infrastructures for analysis. This multi-tiered approach improves situational awareness, disaster response, and resource utilization.



M. Shao, S. Joshi. “Domain-Shift Immunity in Deep Deformable Registration via Local Feature Representations,” Subtitled “arXiv:2512.23142v1,” 2025.

ABSTRACT

Deep learning has advanced deformable image registration, surpassing traditional optimization-based methods in both accuracy and efficiency. However, learning-based models are widely believed to be sensitive to domain shift, with robustness typically pursued through large and diverse training datasets, without explaining the underlying mechanisms. In this work, we show that domain-shift immunity is an inherent property of deep deformable registration models, arising from their reliance on local feature representations rather than global appearance for deformation estimation. To isolate and validate this mechanism, we introduce UniReg, a universal registration framework that decouples feature extraction from deformation estimation using fixed, pre-trained feature extractors and a UNet-based deformation network. Despite training on a single dataset, UniReg exhibits robust cross-domain and multi-modal performance comparable to optimization-based methods. Our analysis further reveals that failures of conventional CNN-based models under modality shift originate from dataset-induced biases in early convolutional layers. These findings identify local feature consistency as the key driver of robustness in learning-based deformable registration and motivate backbone designs that preserve domain-invariant local features. 



H. Shrestha, J. Wilburn, B. Bollen, A.M. McNutt, A. Lex, L. Harrison. “ReVISitPy: Python Bindings for the reVISit Study Framework,” In Eurovis 2025, 2025.

ABSTRACT

User experiments are an important part of visualization research, yet they remain costly, time-consuming to create, and difficult to prototype and pilot. The process of prototyping a study-from initial design to data collection and analysis-often requires the use of multiple systems (e.g. webservers and databases), adding complexity. We present reVISitPy, a Python library that enables visualization researchers to design, pilot deployments, and analyze pilot data entirely within a Jupyter notebook. Re- VISitPy provides a higher-level Python interface for the reVISit Domain-Specific Language (DSL) and study framework, which traditionally relies on manually authoring complex JSON configuration files. As study configurations grow larger, editing raw JSON becomes increasingly tedious and error-prone. By streamlining the configuration, testing, and preliminary analysis workflows, reVISitPy reduces the overhead of study prototyping and helps researchers quickly iterate on study designs before full deployment through the reVISit framework.



X. Tang, X. Li, T. Tasdizen. “Dynamic Scale for Transformer,” In Medical Imaging with Deep Learning, 2025.

ABSTRACT

To enhance the hierarchical transformer with fixed embedding sizes, we propose a dynamic CLS token that aggregates information from CLS tokens across all layers, each embedded with varying receptive fields, by leveraging a squeeze-and-excitation module. This architecture offers a more flexible approach to utilizing multi-scale features in transformers. Code is available at: https://github.com/xiaoyatang/DynamicCLS.git.



X. Tang, B. Zhang, M.M. Ho, B.S. Knudsen, T. Tasdizen. “DuoFormer: Leveraging Hierarchical Representations by Local and Global Attention Vision Transformer,” Subtitled “arXiv:2506.12982,” 2025.

ABSTRACT

Despite the widespread adoption of transformers in medical applications, the exploration of multi-scale learning through transformers remains limited, while hierarchical representations are considered advantageous for computer-aided medical diagnosis. We propose a novel hierarchical transformer model that adeptly integrates the feature extraction capabilities of Convolutional Neural Networks (CNNs) with the advanced representational potential of Vision Transformers (ViTs). Addressing the lack of inductive biases and dependence on extensive training datasets in ViTs, our model employs a CNN backbone to generate hierarchical visual representations. These representations are adapted for transformer input through an innovative patch tokenization process, preserving the inherited multi-scale inductive biases. We also introduce a scale-wise attention mechanism that directly captures intra-scale and inter-scale associations. This mechanism complements patch-wise attention by enhancing spatial understanding and preserving global perception, which we refer to as local and global attention, respectively. Our model significantly outperforms baseline models in terms of classification accuracy, demonstrating its efficiency in bridging the gap between Convolutional Neural Networks (CNNs) and Vision Transformers (ViTs). The components are designed as plug-and-play for different CNN architectures and can be adapted for multiple applications. The code is available at https://github.com/xiaoyatang/DuoFormer.git.



L.C.R. Tanner, A. Busatto, J.A. Bergquist, W.W. Good, B. Zenger, G. Plank, A. Narayan, K. Gillette, R.S. MacLeod. “Uncertainty quantification via polynomial chaos expansion of myocardial fibre orientation and cardiac activation patterns,” In The Journal of Physiology, Wiley, 2025.

ABSTRACT

Predictive models and computational simulations of cardiac electrophysiology depend on precise anatomical representations, including the local myocardial fibre structure. However, obtaining patient-specific fibre information is challenging. In addition, the influence of physiological variability in fibre orientation on cardiac activation simulations is poorly understood. We implemented rule-based algorithms to generate fibres and robust uncertainty quantification methods to determine model output variability with respect to ventricular activation sequences. We used polynomial chaos, which reduces computational demands by using an emulator to approximate the underlying forward model. Our study examined activation sequences in response to nine stimuli and five metrics quantifying essential features of the activation sequence. The results indicated that the primary fibre orientation impacts the overall spread of activation, which could impact more complex patterns of activation; however, there is minimal impact on the location of discrete activation features, such as breakthrough sites. For free wall stimuli, the standard deviation (STD) was highest near the stimulus site, diminishing with distance. Apical stimuli showed complementary STD patterns, with epicardial pacing maximizing STD in the right basal area and endocardial pacing in the left. Ventricular junction stimuli exhibited symmetrical STD patterns, low near the stimulus but increasing sharply towards the apex, peaking on the left in the apical region. Furthermore, variability in the imbrication or helix angle did not impact the activation sequences. We conclude that in many relevant modelling contexts, the variability in myocardial fibre orientation can play an important role in the resulting activation sequences and should be accounted for.



L.C.R. Tanner, A. Busatto, T. Grandits, J.A. Bergquist, B. Zenger, S. Pezzuto, G. Plank, R.S. MacLeod, K. Gillette. “Reconstructing ventricular activation sequences from epicardial data: Insights from Geodesic Back-Propagation optimization in porcine models,” In Computers in Biology and Medicine, Vol. 198, 2025.

ABSTRACT

Cardiac digital twins (CDTs) are emerging as powerful tools in personalized medicine, providing subject-specific models to simulate and understand cardiac function. A central challenge in constructing CDTs is accurately personalizing the structure and function of the His-Purkinje system (HPS), which determines ventricular activation. In this study, we leveraged a novel modified Geodesic-BP method to infer early activation sites (EASs) from epicardial activation times. The EASs can then serve as a surrogate for Purkinje-myocardial junctions, facilitating anterograde ventricular activation. We used both experimental porcine (N 5) and synthetic (N  5) datasets of epicardial activation times measured or assigned to locations on an electrode sock. For both datasets, we optimized for initial estimates of 5, 50, 100, and 200 EASs and assessed output variability by repeating the inference process 10 times. Assessments were based on matching the predicted activation times at both the epicardial locations and from intracardiac measurements made with multielectrode needles, and throughout the ventricular myocardium for the synthetic dataset. The algorithm could consistently recover global ventricular activation patterns from epicardial data alone. For the experimental dataset, the minimum and maximum mean absolute differences were 0.19 ms and 3.86 ms on the epicardial sock and 2.65 ms and 10.69 ms for the needles. For the synthetic dataset, the corresponding values were 0.13 ms and 2.81 ms on the sock and 2.42 ms and 14.07 ms ms throughout the ventricular myocardium. However, discrepancies between the epicardial surface and intramural myocardium, overfitting of EASs, and variability across repeated runs revealed key limitations. These findings highlight both the overall potential and current limitations of inferring EASs using the proposed optimization approach. They demonstrate the feasibility of deriving informative activation patterns from limited data, while underscoring the need to incorporate stronger physiological priors and anatomical constraints. Ultimately, our results motivate future efforts to refine simulation-based personalization frameworks, improve robustness, and enhance the physiological realism of CDTs for more accurate and reliable applications.



W. Tao, S. Joshi, R. Whitaker. “Integrated Model Selection and Scalability in Functional Data Analysis through Bayesian Learning,” Subtitled “Preprints.org,” 2025.
DOI: 10.20944/preprints202503.0658.v1

ABSTRACT

Functional data, including one-dimensional curves and higher-dimensional surfaces, have become increasingly prominent across scientific disciplines. They offer a continuous perspective that captures subtle dynamics and richer structures compared to discrete representations, thereby preserving essential information and facilitating more natural modeling of real-world phenomena, especially in sparse or irregularly sampled settings. A key challenge lies in identifying low-dimensional representations and estimating covariance structures that capture population statistics effectively. We propose a novel Bayesian framework with a nonparametric kernel expansion and a sparse prior, enabling direct modeling of measured data and avoiding the artificial biases from regridding. Our method, Bayesian scalable functional data analysis (BSFDA), automatically selects both subspace dimensionalities and basis functions, reducing computational overhead through an efficient variational optimization strategy. We further propose a faster approximate variant that maintains comparable accuracy but accelerates computations significantly on large-scale datasets. Extensive simulation studies demonstrate that our framework outperforms conventional techniques in covariance estimation and dimensionality selection, showing resilience to high dimensionality and irregular sampling. The proposed methodology proves effective for multidimensional functional data and showcases practical applicability in biomedical and meteorological datasets. Overall, BSFDA offers an adaptive, continuous, and scalable solution for modern functional data analysis across diverse scientific domains.



W. Tao, T. J. Somorin, J. Kueper, A. Dixon, N. Kass, N. Khan, K. Iyer, J. Wagoner, A. Rogers, R. Whitaker, S. Elhabian, J. A. Goldstein. “Quantifying Sagittal Craniosynostosis Severity: A Machine Learning Approach With CranioRate,” In The Cleft Palate Craniofacial Journal, Sage, 2025.
DOI: 10.1177/105566562513473

ABSTRACT

Objective

To develop and validate machine learning (ML) models for objective and comprehensive quantification of sagittal craniosynostosis (SCS) severity, enhancing clinical assessment, management, and research.
 

Design

A cross-sectional study that combined the analysis of computed tomography (CT) scans and expert ratings.
 

Setting

The study was conducted at a children's hospital and a major computer imaging institution. Our survey collected expert ratings from participating surgeons.
 

Participants

The study included 195 patients with nonsyndromic SCS, 221 patients with nonsyndromic metopic craniosynostosis (CS), and 178 age-matched controls. Fifty-four craniofacial surgeons participated in rating 20 patients head CT scans.
 

Interventions

Computed tomography scans for cranial morphology assessment and a radiographic diagnosis of nonsyndromic SCS.
 

Main Outcomes

Accuracy of the proposed Sagittal Severity Score (SSS) in predicting expert ratings compared to cephalic index (CI). Secondary outcomes compared Likert ratings with SCS status, the predictive power of skull-based versus skin-based landmarks, and assessments of an unsupervised ML model, the Cranial Morphology Deviation (CMD), as an alternative without ratings.
 

Results

The SSS achieved significantly higher accuracy in predicting expert responses than CI (P < .05). Likert ratings outperformed SCS status in supervising ML models to quantify within-group variations. Skin-based landmarks demonstrated equivalent predictive power as skull landmarks (P < .05, threshold 0.02). The CMD demonstrated a strong correlation with the SSS (Pearson coefficient: 0.92, Spearman coefficient: 0.90, P < .01).
 

Conclusions

The SSS and CMD can provide accurate, consistent, and comprehensive quantification of SCS severity. Implementing these data-driven ML models can significantly advance CS care through standardized assessments, enhanced precision, and informed surgical planning.



Y. Tatari, H.T. Nguyen, A. Arzani, P. Newell . “Investigation of particle transport in geothermal systems using integrated CFD–DEM and data-driven approaches,” In Geothermics, Vol. 136, Elsevier, 2025.

ABSTRACT

Geothermal systems provide continuous, low-carbon energy by harnessing the Earth’s renewable heat, but their efficiency can be hindered by issues such as limited heat production and thermal breakthrough. A promising approach to overcome these issues is to inject polymer-based microcapsules into fractures to modify permeability, which requires a clear understanding of particle transport within the fracture network. To do so, this study used Computational Fluid Dynamics–Discrete Element Method (CFD–DEM) simulations combined with machine learning (ML) to capture particle transport behavior under varying temperature conditions. The dataset comprises 45 CFD–DEM test cases, which are generated by coupling OpenFOAM (for fluid dynamics) and LIGGGHTS (for particle tracking), enabling detailed modeling of thermo-hydro processes and particle interactions. To assess their influence on transport behavior, key parameters include particle diameter (), formation temperature (), and particle volume fraction (). Supervised learning models, including random forest and interpretable decision tree classifiers, were trained to classify flow blockage. Feature importance analysis identified and  as the most critical factors impacting the particle transports. To avoid sealing progression, the accumulation of low-velocity particles over time was fit with a sigmoid function. Results show that higher particle concentrations and larger diameters reduce transport efficiency, while elevated inlet velocities enhance particle mobility and prolong transport through the fracture. This interpretable data-driven approach, grounded in CFDEM simulations, offers a predictive tool for particle transport in fractures subject to complex geothermal environments.



M. Taufer, R. Mihalcea, M. Turk, D. Lopresti, A. Wierman, K. Butler, S. Koenig, D. Danks, W. Gropp, M. Parashar, Y. Gil, B. Regli, R. Rajaraman, D. Jensen, N. Bliss, M. Maher. “Now More Than Ever, Foundational AI Research and Infrastructure Depends on the Federal Government,” Subtitled “arXiv:2506.14679,” 2025.

ABSTRACT

Leadership in the field of AI is vital for our nation's economy and security. Maintaining this leadership requires investments by the federal government. The federal investment in foundation AI research is essential for U.S. leadership in the field. Providing accessible AI infrastructure will benefit everyone. Now is the time to increase the federal support, which will be complementary to, and help drive, the nation's high-tech industry investments.



T. Transue, B. Wang. “Learning Decentralized Swarms Using Rotation Equivariant Graph Neural Networks,” Subtitled “arXiv:2502.17612,” 2025.

ABSTRACT

The orchestration of agents to optimize a collective objective without centralized control is challenging yet crucial for applications such as controlling autonomous fleets, and surveillance and reconnaissance using sensor networks. Decentralized controller design has been inspired by self-organization found in nature, with a prominent source of inspiration being flocking; however, decentralized controllers struggle to maintain flock cohesion. The graph neural network (GNN) architecture has emerged as an indispensable machine learning tool for developing decentralized controllers capable of maintaining flock cohesion, but they fail to exploit the symmetries present in flocking dynamics, hindering their generalizability. We enforce rotation equivariance and translation invariance symmetries in decentralized flocking GNN controllers and achieve comparable flocking control with 70% less training data and 75% fewer trainable weights than existing GNN controllers without these symmetries enforced. We also show that our symmetry-aware controller generalizes better than existing GNN controllers. Code and animations are available at github.com/Utah-Math-Data-Science/Equivariant-Decentralized-Controllers.



J.P. Tuffour, R. Ewing, G. Tian. “Augmenting Low-Carbon Sustainable Cities Using Transit-Linked Mixed-Use Designs: A Quasi-Experimental Regression of Travel and Emission Outcomes,” In Sustainable Cities and Society, Vol. 135, Elsevier, 2025.

ABSTRACT

Cities today continue to struggle with the converging twin crises of rising vehicular air pollution and emission-intensive land use, prompting a renewed urgency for climate-smart urban designs such as mixed-use (MXDs) and transit-oriented developments (TODs). Yet, not all developments are designed to have a transit-focused orientation, and even if they are, individual lifestyle choices confound their true effects. This study investigates whether those MXDs situated near high-frequency transit (a design typology we call Transit-Oriented Mixed-Use Developments (TOD-MXDs)) represent a more sustainable antidote to the environmental and mobility crises afflicting cities. Leveraging the largest geo-spatially harmonized database of individual-level and built environment characteristics across 36 diverse U.S. regions, we employ a mix of quasi-experimental design and advanced multivariate regression modeling to estimate travel characteristics and emissions outcomes. Results showed that although TOD-MXDs remain relatively scarce, they consistently outperform their non-transit-oriented counterparts, generating about 20 % less VMT, higher walking (87.6 %) and biking (83 %) trip shares, and significantly lower CO2e emissions, even after accounting for residential self-selection and the full suite of 7-D built environment variables. Furthermore, a unit increase in the presence of transit around mixed-use designs was associated with a ∼13.8 % reduction in VMT, suggesting that coupling land-use diversity with transit integration produces a synergistic effect that advances low-carbon and sustainable city outcomes. As conventional planning and policymakers strive for climate-resilient and healthy communities, this study makes a bold case: TOD-MXDs are not just better, they may be essential to the future of climate adaptation in cities through built-environment design.



J. Turnage, M. Lowery, J. Jakeman, Z. Morrow, A. Narayan. “An Optimal Weighted Least-Squares Method for Operator Learning,” Subtitled “arXiv:2512.11168v1,” 2025.

ABSTRACT

We consider the problem of learning an unknown, possibly nonlinear operator between separable Hilbert spaces from supervised data. Inputs are drawn from a prescribed probability measure on the input space, and outputs are (possibly noisy) evaluations of the target operator. We regard admissible operators as square-integrable maps with respect to a fixed approximation measure, and we measure reconstruction error in the corresponding Bochner norm. For a finite-dimensional approximation space   of dimension  , we study weighted least squares estimators in   and establish probabilistic stability and accuracy bounds in the Bochner norm. We show that there exist sampling measures and weights - defined via an operator-level Christoffel function - that yield uniformly well-conditioned Gram matrices and near-optimal sample complexity, with a number of training samples   on the order of  . We complement the analysis by constructing explicit operator approximation spaces in cases of interest: rank-one linear operators that are dense in the class of bounded linear operators, and rank-one polynomial operators that are dense in the Bochner space under mild assumptions on the approximation measure. For both families we describe implementable procedures for sampling from the associated optimal measures. Finally, we demonstrate the effectiveness of this framework on several benchmark problems, including learning solution operators for the Poisson equation, viscous Burgers' equation, and the incompressible Navier-Stokes equations.