SCI Publications
2025
M.M. Mohammad, N. Baret, R. Mayerhofer, A. McNutt, P. Rosen.
Towards Scalable Visual Data Wrangling via Direct Manipulation, Subtitled arXiv:2512.18405, 2025.
Data wrangling - the process of cleaning, transforming, and preparing data for analysis - is a well-known bottleneck in data science workflows. Existing tools either rely on manual scripting, which is error-prone and hard to debug, or automate cleaning through opaque black-box pipelines that offer limited control. We present Buckaroo, a scalable visual data wrangling system that restructures data preparation as a direct manipulation task over visualizations. Buckaroo enables users to explore and repair data anomalies - such as missing values, outliers, and type mismatches - by interacting directly with coordinated data visualizations. The system extensibly supports user-defined error detectors and wranglers, tracks provenance for undo/redo, and generates reproducible scripts for downstream tasks. Buckaroo maintains efficient indexing data structures and differential storage to localize anomaly detection and minimize recomputation. To demonstrate the applicability of our model, Buckaroo is integrated with the \textitHopara pan-and-zoom engine, which enables multi-layered navigation over large datasets without sacrificing interactivity. Through empirical evaluation and an expert review, we show that Buckaroo makes visual data wrangling scalable - bridging the gap between visual inspection and programmable repairs.
Z. Morrow, M. Penwarden, B. Chen, A. Javeed, A. Narayan, J. Jakeman.
SUPN: Shallow Universal Polynomial Networks, Subtitled arXiv:2511.21414v1, 2025.
Deep neural networks (DNNs) and Kolmogorov-Arnold networks (KANs) are popular methods for function approximation due to their flexibility and expressivity. However, they typically require a large number of trainable parameters to produce a suitable approximation. Beyond making the resulting network less transparent, overparameterization creates a large optimization space, likely producing local minima in training that have quite different generalization errors. In this case, network initialization can have an outsize impact on the model's out-of-sample accuracy. For these reasons, we propose shallow universal polynomial networks (SUPNs). These networks replace all but the last hidden layer with a single layer of polynomials with learnable coefficients, leveraging the strengths of DNNs and polynomials to achieve sufficient expressivity with far fewer parameters. We prove that SUPNs converge at the same rate as the best polynomial approximation of the same degree, and we derive explicit formulas for quasi-optimal SUPN parameters. We complement theory with an extensive suite of numerical experiments involving SUPNs, DNNs, KANs, and polynomial projection in one, two, and ten dimensions, consisting of over 13,000 trained models. On the target functions we numerically studied, for a given number of trainable parameters, the approximation error and variability are often lower for SUPNs than for DNNs and KANs by an order of magnitude. In our examples, SUPNs even outperform polynomial projection on non-smooth functions.
Q.C. Nguyen, R. Doumbia, T.T. Nguyen, X. Yue, H. Mane, J. Merchant, T. Tasdizen, M. Alirezaei, P. Dipankar, D. Li, P. Sai Priya Mullaputi A. Alibilli, Y. Hswen, X. He.
Changes in the Neighborhood Built Environment and Chronic Health Conditions in Washington, DC, in 2014-2019: Longitudinal Analysis, In JMIR Form Res, Vol. 9, pp. e74195. 2025.
Background: Google Street View (GSV) images offer a unique and scalable alternative to in-person audits for examining neighborhood built environment characteristics. Additionally, most prior neighborhood studies have relied on cross-sectional designs.
Objective: This study aimed to use GSV images and computer vision to examine longitudinal changes in the built environment, demographic shifts, and health outcomes in Washington, DC, from 2014 to 2019.
Methods: In total, 434,115 GSV images were systematically sampled at 100 m intervals along primary and secondary road segments. Convolutional neural networks, a type of deep learning algorithm, were used to extract built environment features from images. Census tract summaries of the neighborhood built environment were created. Multilevel mixed-effects linear models with random intercepts for years and census tracts were used to assess associations between built environment changes and health outcomes, adjusting for covariates, including median age, percentage male, percentage Hispanic, percentage African American, percentage college educated, percentage owner-occupied housing, and median household income.
Results: Washington, DC, experienced a shift toward higher-density housing, with non-single-family homes rising from 66% to 72% of the housing stock. Single-lane roads increased from 37% to 42%, suggesting a shift toward more sustainable and compact urban forms. Gentrification trends were reflected in a rise in college-educated residents (16%-41%), a US $17,490 increase in the median household income, and a US $159,600 increase in property values. Longitudinal analyses revealed that increased construction activity was associated with lower rates of obesity, diabetes, high cholesterol, and cancer, while growth in non-single-family housing was correlated with reductions in the prevalence of obesity and diabetes. However, neighborhoods with higher proportions of African American residents experienced reduced construction activity.
Conclusions: Washington, DC, has experienced significant urban transformation, marked by substantial changes in neighborhood built environments and demographic shifts. Urban development is associated with reduced prevalence of chronic conditions. These findings highlight the complex interplay between urban development, demographic changes, and health, underscoring the need for future research to explore the broader impacts of neighborhood built environment changes on community composition and health outcomes. GSV imagery, along with advances in computer vision, can aid in the acceleration of neighborhood studies.
D. De Novi, L. Carnevale, D. Balouek, M. Parashar, M. Villari.
Predictive Resource Management in the Computing Continuum: Transfer Learning from Virtual Machines to Containers using Transformers, In UCC '25: Proceedings of the 18th IEEE/ACM International Conference on Utility and Cloud Computing , IEEE, 2025.
Efficient workload forecasting is a key enabler of modern AIOps (Artificial Intelligence for IT Operations), supporting proactive and autonomous resource management across the computing continuum, from edge environments to large-scale cloud infrastructures. In this paper, we propose a Temporal Transformer architecture for CPU utilization prediction, designed to capture both short-term fluctuations and long-range temporal dependencies in workload dynamics. The model is first pretrained on a large-scale Microsoft Azure VM dataset and subsequently fine-tuned on the Alibaba container dataset, enabling effective transfer learning across heterogeneous virtualization environments. Experimental results demonstrate that the proposed approach achieves high predictive accuracy while maintaining a compact model size and inference times compatible with real-time operation. Qualitative analyses further highlight the model’s ability to reproduce workload patterns with high fidelity. These findings indicate that the proposed Temporal Transformer constitutes a lightweight and accurate forecasting component for next-generation AIOps pipelines, suitable for deployment across both cloud and edge intelligence scenarios.
I. Osinomumu, U. Reddy, G. Doretto, D. Adjeroh.
A Survey of Machine Learning Techniques in Flavor Prediction and Analysis, In Trends in Food Science & Technology, Vol. 166, Elsevier, 2025.
Background
Scope and approach
Key findings and conclusions
Significance and novelty
T. A. J. Ouermi, E. Li, K. Moreland, D. Pugmire, C. R. Johnson,, T. M. Athawale.
Efficient Probabilistic Visualization of Local Divergence of 2D Vector Fields with Independent Gaussian Uncertainty, In 2025 IEEE Workshop on Uncertainty Visualization: Unraveling Relationships of Uncertainty, AI, and Decision-Making, 2025.
This work focuses on visualizing uncertainty of local divergence of two-dimensional vector fields. Divergence is one of the fundamental attributes of fluid flows, as it can help domain scientists analyze potential positions of sources (positive divergence) and sinks (negative divergence) in the flow. However, uncertainty inherent in vector field data can lead to erroneous divergence computations, adversely impacting downstream analysis. While Monte Carlo (MC) sampling is a classical approach for estimating divergence uncertainty, it suffers from slow convergence and poor scalability with increasing data size and sample counts. Thus, we present a two-fold contribution that tackles the challenges of slow convergence and limited scalability of the MC approach. (1) We derive a closed-form approach for highly efficient and accurate uncertainty visualization of local divergence, assuming independently Gaussian-distributed vector uncertainties. (2) We further integrate our approach into Viskores, a platform-portable parallel library, to accelerate uncertainty visualization. In our results, we demonstrate significantly enhanced efficiency and accuracy of our serial analytical (speed-up up to $1946 \times$) and parallel Viskores (speed-up up to 19698X) algorithms over the classical serial MC approach. We also demonstrate qualitative improvements of our probabilistic divergence visualizations over traditional mean-field visualization, which disregards uncertainty. We validate the accuracy and efficiency of our methods on wind forecast and ocean simulation datasets.
E.N. Paccione, E. Kwan, J.A. Bergquist, A.S. Shah, B. A. Orkild, B. Hunt, M. L. Salter, J. Mendes, E. DiBella, A. E. Arai, T. J. Bunch, R. S. MacLeod, J. N. Kunz, Y. Hitchcock, K. E. Kokeny, S. Lloyd, R. Ranjan.
Radiation Induces Diffuse Extracellular Remodeling of Healthy Myocardium in a Dose and Time Dependent Manner Without a Dense Ablative Effect, In Heart Rhythm, Elsevier, pp. 1547-5271. 2025.
DOI: https://doi.org/10.1016/j.hrthm.2025.06.041
Background
Objective
Methods
Results
Conclusions
A. Panta, A. Gooch, G. Scorzelli, M. Taufer, V. Pascucci.
Scalable Climate Data Analysis: Balancing Petascale Fidelity and Computational Cost, In 2025 IEEE 25th International Symposium on Cluster, Cloud and Internet Computing Workshops (CCGridW), pp. 245-248. 2025.
The growing resolution and volume of climate data from remote sensing and simulations pose significant storage, processing, and computational challenges. Traditional compression or subsampling methods often compromise data fidelity, limiting scientific insights. We introduce a scalable ecosystem that integrates hierarchical multiresolution data management, intelligent transmission, and ML-assisted reconstruction to balance accuracy and efficiency. Our approach reduces storage and computational costs by 99%, lowering expenses from 100,00 to 24 while maintaining a Root Mean Square (RMS) error of 1.46 degrees Celsius. Our experimental results confirm that even with significant data reduction, essential features required for accurate climate analysis are preserved. Validated on petascale NASA climate datasets, this solution enables cost-effective, high-fidelity climate analysis for research and decision-making.
A. Panta, A. Sahistan, X. Huang, A.A. Gooch, G. Scorzelli, H. Torres, P. Klein, G. A Ovando-Montejo, P. Lindstrom, V. Pascucci.
Expanding Access to Science Participation: A FAIR Framework for Petascale Data Visualization and Analytics, In IEEE Trans Vis Comput Graph, IEEE, 2025.
The massive data generated by scientists daily serve as both a major catalyst for new discoveries and innovations, as well as a significant roadblock that restricts access to the data. Our paper introduces a new approach to removing big data barriers and democratizing access to petascale data for the broader scientific community. Our novel data fabric abstraction layer allows user-friendly querying of scientific information while hiding the complexities of dealing with file systems or cloud services. We enable FAIR (Findable, Accessible, Interoperable, and Reusable) access to datasets such as NASA's petascale climate datasets. Our paper presents an approach to managing, visualizing, and analyzing petabytes of data within a browser on equipment ranging from the top NASA supercomputer to commodity hardware like a laptop. Our novel data fabric abstraction utilizes state-of-the art progressive compression algorithms and machinelearning insights to power scalable visualization dashboards for petascale data. The result provides users with the ability to identify extreme events or trends dynamically, expanding access to scientific data and further enabling discoveries. We validate our approach by improving the ability of climate scientists to visually explore their data via three fully interactive dashboards. We further validate our approach by deploying the dashboards and simplified training materials in the classroom at a minorityserving institution. These dashboards, released in simplified form to the general public, contribute significantly to a broader push to democratize the access and use of climate data.
M. Parashar.
Autonomic Computing Rebooted: Taming the Computing Continuum, In Transactions on Autonomous and Adaptive Systems, ACM, 2025.
DOI: hps://doi.org/10.1145/3768320
Technological advances and rapid service deployments have resulted in a pervasive and interconnected computing continuum, which is enabling new classes of application formulations and workflows, delivering novel services to consumers, and is becoming a core engine for discovery, innovation, and economic growth. The computing continuum is also unleashing new system and application management challenges, which must be addressed before its potential and promises are truly realized. Autonomic computing can provide the abstractions and mechanisms essential to effectively harnessing the computing continuum, but must evolve to address these new challenges. This paper is a call to action for rebooting autonomics to enable us to harness the computing continuum.
T. Patel, T.A.J. Ouermi, T. Athawale,, C.R. Johnson.
Fast HARDI Uncertainty Quantification and Visualization with Spherical Sampling, In Computer Graphics Forum, Vol. 44, No. 3, pp. 1--12. 2025.
In this paper, we study uncertainty quantification and visualization of orientation distribution functions (ODF), which corresponds to the diffusion profile of high angular resolution diffusion imaging (HARDI) data. The shape inclusion probability (SIP) function is the state-of-the-art method for capturing the uncertainty of ODF ensembles. The current method of computing the SIP function with a volumetric basis exhibits high computational and memory costs, which can be a bottleneck to integrating uncertainty into HARDI visualization techniques and tools. We propose a novel spherical sampling framework for faster computation of the SIP function with lower memory usage and increased accuracy. In particular, we propose direct extraction of SIP isosurfaces, which represent confidence intervals indicating spatial uncertainty of HARDI glyphs, by performing spherical sampling of ODFs. Our spherical sampling approach requires much less sampling than the state-of-the-art volume sampling method, thus providing significantly enhanced performance, scalability, and the ability to perform implicit ray tracing. Our experiments demonstrate that the SIP isosurfaces extracted with our spherical sampling approach can achieve up to 8164× speedup, 37282× memory reduction, and 50.2% less SIP isosurface error compared to the classical volume sampling approach. We demonstrate the efficacy of our methods through experiments on synthetic and human-brain HARDI datasets.
A. C. Peterson, M. R. Requist, J. C. Benna, J. R. Nelson, S. Elhabian, C. de Cesar Netto, T. C. Beals, A. L. Lenz.
Talar Morphology of Charcot-Marie-Tooth Patients With Cavovarus Feet, In Foot & Ankle International, Sage Publications, 2025.
DOI: 10.1177/10711007241309915
PubMed ID: 39937093
Background:
Methods:
Results:
Conclusion:
P. Ramonetti, M. Floca, K. O'Laughlin, A. Gupta, M. Parashar, I. Altintas.
National Data Platform's Education Hub, Subtitled arXiv:2510.12820v1, 2025.
As demand for AI literacy and data science education grows, there is a critical need for infrastructure that bridges the gap between research data, computational resources, and educational experiences. To address this gap, we developed a first-of-its-kind Education Hub within the National Data Platform. This hub enables seamless connections between collaborative research workspaces, classroom environments, and data challenge settings. Early use cases demonstrate the effectiveness of the platform in supporting complex and resource-intensive educational activities. Ongoing efforts aim to enhance the user experience and expand adoption by educators and learners alike.
A. Roberts, J. Marquez, K.H. NG, K. Mickelson, A. Panta, G. Scorzelli, A. Gooch, P. Cushman, M. Fritts, H. Neog, V. Pascucci, M. Taufer.
The Making of a Community Dark Matter Dataset with the National Science Data Fabric, Subtitled arXiv:2507.13297v1, 2025.
Dark matter is believed to constitute approximately 85% of the universe’s matter, yet its fundamental nature remains elusive. Direct detection experiments, though globally deployed, generate data that is often locked within custom formats and non-reproducible software stacks, limiting interdisciplinary analysis and innovation. This paper presents a collaboration between the National Science Data Fabric (NSDF) and dark matter researchers to improve accessibility, usability, and scientific value of a calibration dataset collected with Cryogenic Dark Matter Search (CDMS) detectors at the University of Minnesota. We describe how NSDF services were used to convert data from a proprietary format into an open, multi-resolution IDX structure; develop a web-based dashboard for easily viewing signals; and release a Python-compatible CLI to support scalable workflows and machine learning applications. These contributions enable broader use of high-value dark matter datasets, lower the barrier to entry for new collaborators, and support reproducible, cross-disciplinary research.
R. Basu Roy, T. Patel, B. Li, S. Samsi, V. Gadepally, D. Tiwari.
GreenMix: Energy-Efficient Serverless Computing via Randomized Sketching on Asymmetric Multi-Cores, In Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis, Association for Computing Machinery, 2025.
ISBN: 9798400714665
DOI: 10.1145/3712285.3759861
GreenMix is motivated by the renewed interest in asymmetric multi-core processors and the emergence of the serverless computing model. Asymmetric multi-cores offer better energy and performance trade-offs by placing different core types on the same die. However, existing serverless scheduling techniques do not leverage these benefits. GreenMix is the first serverless work to reduce energy and serverless keep-alive costs while meeting QoS targets by leveraging asymmetric multi-cores. GreenMix employs randomized sketching, tailored for serverless execution and keep-alive, to perform within 10% of the optimal solution in terms of energy efficiency and keep-alive cost reduction. GreenMix’s effectiveness is demonstrated through evaluations on clusters of ARM big.LITTLE and Intel Alder Lake asymmetric processors. It outperforms competing state-of-the-art schedulers, offering a novel approach for energy-efficient serverless computing.
S. Saha, S. Joshi, R. Whitaker.
ARD-VAE: A Statistical Formulation to Find the Relevant Latent Dimensions of Variational Autoencoders, Subtitled arXiv:2501.10901, 2025.
The variational autoencoder (VAE) is a popular, deep, latent-variable model (DLVM) due to its simple yet effective formulation for modeling the data distribution. Moreover, optimizing the VAE objective function is more manageable than other DLVMs. The bottleneck dimension of the VAE is a crucial design choice, and it has strong ramifications for the model’s performance, such as finding the hidden explanatory factors of a dataset using the representations learned by the VAE. However, the size of the latent dimension of the VAE is often treated as a hyperparameter estimated empirically through trial and error. To this end, we propose a statistical formulation to discover the relevant latent factors required for modeling a dataset. In this work, we use a hierarchical prior in the latent space that estimates the variance of the latent axes using the encoded data, which identifies the relevant latent dimensions. For this, we replace the fixed prior in the VAE objective function with a hierarchical prior, keeping the remainder of the formulation unchanged. We call the proposed method the automatic relevancy detection in the variational autoencoder (ARD-VAE). We demonstrate the efficacy of the ARD-VAE on multiple benchmark datasets in finding the relevant latent dimensions and their effect on different evaluation metrics, such as FID score and disentanglement analysis.
S. Saha, S. Joshi, R. Whitaker.
Disentanglement Analysis in Deep Latent Variable Models Matching Aggregate Posterior Distributions, Subtitled arXiv:2501.15705, 2025.
Deep latent variable models (DLVMs) are designed to learn meaningful representations in an unsupervised manner, such that the hidden explanatory factors are interpretable by independent latent variables (aka disentanglement). The variational autoencoder (VAE) is a popular DLVM widely studied in disentanglement analysis due to the modeling of the posterior distribution using a factorized Gaussian distribution that encourages the alignment of the latent factors with the latent axes. Several metrics have been proposed recently, assuming that the latent variables explaining the variation in data are aligned with the latent axes (cardinal directions). However, there are other DLVMs, such as the AAE and WAE-MMD (matching the aggregate posterior to the prior), where the latent variables might not be aligned with the latent axes. In this work, we propose a statistical method to evaluate disentanglement for any DLVMs in general. The proposed technique discovers the latent vectors representing the generative factors of a dataset that can be different from the cardinal latent axes. We empirically demonstrate the advantage of the method on two datasets.
S. Saha, R. Whitaker.
AdaSemSeg: An Adaptive Few-shot Semantic Segmentation of Seismic Facies, Subtitled arXiv:2501.16760, 2025.
Automated interpretation of seismic images using deep learning methods is challenging because of the limited availability of training data. Few-shot learning is a suitable learning paradigm in such scenarios due to its ability to adapt to a new task with limited supervision (small training budget). Existing few-shot semantic segmentation (FSSS) methods fix the number of target classes. Therefore, they do not support joint training on multiple datasets varying in the number of classes. In the context of the interpretation of seismic facies, fixing the number of target classes inhibits the generalization capability of a model trained on one facies dataset to another, which is likely to have a different number of facies. To address this shortcoming, we propose a few-shot semantic segmentation method for interpreting seismic facies that can adapt to the varying number of facies across the dataset, dubbed the AdaSemSeg. In general, the backbone network of FSSS methods is initialized with the statistics learned from the ImageNet dataset for better performance. The lack of such a huge annotated dataset for seismic images motivates using a self-supervised algorithm on seismic datasets to initialize the backbone network. We have trained the AdaSemSeg on three public seismic facies datasets with different numbers of facies and evaluated the proposed method on multiple metrics. The performance of the AdaSemSeg on unseen datasets (not used in training) is better than the prototype-based few-shot method and baselines.
A. Sahistan, S. Zellmann, N. Morrical, V. Pascucci, I. Wald.
Multi-Density Woodcock Tracking: Efficient & High-Quality Rendering for Multi-Channel Volumes, In Eurographics Symposium on Parallel Graphics and Visualization, Eurographics, 2025.
Volume rendering techniques for scientific visualization have increasingly transitioned toward Monte Carlo (MC) methods in recent years due to their flexibility and robustness. However, their application in multi-channel visualization remains underexplored. Traditional compositing-based approaches often employ arbitrary color blending functions, which lack a physical basis and can obscure data interpretation. We introduce multi-density Woodcock tracking, a simple and flexible extension of Woodcock tracking for multi-channel volume rendering that leverages the strengths of Monte Carlo methods to generate high-fidelity visuals. Our method offers a physically grounded solution for inter-channel color blending and eliminates the need for arbitrary blending functions. We also propose a unified blending modality by generalizing Woodcock’s distance tracking method, facilitating seamless integration of alternative blending functions from prior works. Through evaluation across diverse datasets, we demonstrate that our approach maintains real-time interactivity while achieving high-quality visuals by accumulating frames over time.
S.A. Sakin, K.E. Isaacs.
Managing Data for Scalable and Interactive Event Sequence Visualization, Subtitled arXiv:2508.03974, 2025.
Parallel event sequences, such as those collected in program execution traces and automated manufacturing pipelines, are typically visualized as interactive parallel timelines. As the dataset size grows, these charts frequently experience lag during common interactions such as zooming, panning, and filtering. Summarization approaches can improve interaction performance, but at the cost of accuracy in representation. To address this challenge, we introduce ESeMan (Event Sequence Manager), an event sequence management system designed to support interactive rendering of timeline visualizations with tunable accuracy. ESeMan employs hierarchical data structures and intelligent caching to provide visualizations with only the data necessary to generate accurate summarizations with significantly reduced data fetch time. We evaluate ESeMan's query times against summed area tables, M4 aggregation, and statistical sub-sampling on a variety of program execution traces. Our results demonstrate ESeMan provides better performance, achieving sub-100ms fetch times while maintaining visualization accuracy at the pixel level. We further present our benchmarking harness, enabling future performance evaluations for event sequence visualization.
Page 9 of 154
