SCI Publications
2026
S.F. Ahmed, G. Rasineni, F. Koehler, A.Z.B. Aziz, M. Wang, A. Gyulassy, B. Summa, J. Q. Brown, V. Pascucci, S. Y. Elhabian.
PC-MIL: Decoupling Feature Resolution from Supervision Scale in Whole-Slide Learning, Subtitled arXiv:2604.12100v1, 2026.
Whole-slide image (WSI) classification in computational pathology is commonly formulated as slide-level Multiple Instance Learning (MIL) with a single global bag representation. However, slide-level MIL is fundamentally underconstrained: optimizing only global labels encourages models to aggregate features without learning anatomically meaningful localization. This creates a mismatch between the scale of supervision and the scale of clinical reasoning. Clinicians assess tumor burden, focal lesions, and architectural patterns within millimeter-scale regions, whereas standard MIL is trained only to predict whether "somewhere in the slide there is cancer." As a result, the model's inductive bias effectively erases anatomical structure. We propose Progressive-Context MIL (PC-MIL), a framework that treats the spatial extent of supervision as a first-class design dimension. Rather than altering magnification, patch size, or introducing pixel-level segmentation, we decouple feature resolution from supervision scale. Using fixed 20x features, we vary MIL bag extent in millimeter units and anchor supervision at a clinically motivated 2mm scale to preserve comparable tumor burden and avoid confounding scale with lesion density. PC-MIL progressively mixes slide- and region-level supervision in controlled proportions, enabling explicit train-context x test-context analysis. On 1,476 prostate WSIs from five public datasets for binary cancer detection, we show that anatomical context is an independent axis of generalization in MIL, orthogonal to feature resolution: modest regional supervision improves cross-context performance, and balanced multi-context training stabilizes accuracy across slide and regional evaluation without sacrificing global performance. These results demonstrate that supervision extent shapes MIL inductive bias and support anatomically grounded WSI generalization.
M. Alishiri, A. Arzani.
Self-explainable Operator Learning for Discovering Spatial Patterns in Functional Data, Subtitled arXiv:2607.02203, 2026.
Operator learning has emerged as a powerful tool for modeling complex physical systems in functional spaces. However, their neural network-based architectures make them opaque models, obscuring the reasoning behind their predictions. In this work, we introduce a self-explainable operator learning framework that overcomes this challenge by reformulating operator learning as a linear combination of generalized functional linear models expressed through integral equations. Exploiting the additive decomposability of these integral equations, we divide the input domain into subdomains and compute localized integrals to evaluate the contribution of each region to the final prediction. This decomposition enables direct interpretability where the model explains both inputs and outputs by linking specific input regions to corresponding output patterns, thereby revealing which spatial features drive predictions. We demonstrate the framework on function-to-scalar and function-to-function mappings in fluid flow problems involving blood flow and unsteady aerodynamics. The results show that the operator most often prioritizes regions with strong feature gradients, providing physically meaningful insight into the model's decision-making process. Comparisons with established post-hoc explainability methods demonstrate qualitative agreement while highlighting the key advantage of the proposed approach: explainability is embedded directly within the operator structure itself and does not require an external tool. Therefore, our framework provides a mathematically transparent and physically interpretable approach to uncover relationships within data, fostering trust in machine learning for scientific applications by enabling more informed data-driven analysis of physical systems.
O. Alter, E. Newman, S. P. Ponnapalli, J. W. Tsai.
Quantum mechanics-based multitensor AI/ML uniquely able to discover, validate, and interpret predictors from small-cohort noisy high-dimensional multiomic data, In APL Quantum, Vol. 3, AIP, 2026.
Prediction in medicine remains limited. Previously, by using our “comparative spectral decompositions” of two matrices and, separately, two third-order tensors, we demonstrated accurate, precise, actionable, and interpretable tumor whole-genome and, separately, whole-transcriptome predictors—of patients’ survival, treatment responses, and drug targets—in different cancers. Here, we introduce a unified framework that generalizes these exact and structure-preserving algorithms to multiple tensors of any order to model real-world data that measure multiple aspects of interrelated phenomena. We prove properties (e.g., existence and uniqueness) and define metrics (e.g., the “multitensor joint Shannon entropy” and the “multitensor comparative angular distance”) necessary to derive, test, and explain a model. We highlight the novel connection to the quantum mechanical concept of “entanglement” in addition to that of “superposition.” We illustrate the framework in the discovery and validation of two novel predictors in neuroblastoma—each with three entangled representations—in the tumor and blood genomes and tumor transcriptome, where the result of the measurement of any one representation approximately determines the results of the measurements of the other two. Finally, we show that in every representation, the two predictors combined are consistently more accurate than the best standard-of-care biomarker (i.e., the one-gene test for tumor MYCN amplification), and we interpret them in terms of known and new disease mechanisms and drug targets.
T. M. Athawale, K. Moreland, D. Pugmire, C. R. Johnson, P. Rosen, M. Norman, A. Georgiadou,, A. Entezari.
MAGIC: Marching Cubes Isosurface Uncertainty Visualization for Gaussian Uncertain Data with Spatial Correlation, In TVCG, IEEE, 2026.
In this paper, we study the propagation of data uncertainty through the marching cubes algorithm for isosurface visualization for correlated uncertain data. Consideration of correlation has been shown paramount for avoiding errors in uncertainty quantification and visualization in multiple prior studies. Although the problem of isosurface uncertainty with spatial data correlation has been previously addressed, there are two major limitations to prior treatments. First, there are no analytical formulations for uncertainty quantification of isosurfaces when the data uncertainty is characterized by a Gaussian distribution with spatial correlation. Second, as a consequence of the lack of analytical formulations, existing techniques resort to a Monte Carlo sampling approach, which is expensive and difficult to integrate into visualization tools. To address these limitations, we present a closed-form framework to efficiently derive uncertainty in marching cubes level-sets for Gaussian uncertain data with spatial correlation (MAGIC). To derive closed-form solutions, we leverage the Hinkley’s derivation on the ratio of Gaussian distributions. With our analytical framework, we achieve a significant speed-up and enhanced accuracy of uncertainty quantification over classical Monte Carlo methods. We further accelerate our analytical solutions using many-core processors to achieve speed-ups up to 585× and integrability with production visualization tools for broader impact. We demonstrate the effectiveness of our correlation-aware uncertainty framework through experiments on meteorology, urban flow, and astrophysics simulation datasets.
P. Atkins, M. Parashar.
Exploring the Social Life of Data: Finding Data You Can Trust, Subtitled arXiv.2608.11395, 2026.
Artificial intelligence is changing the scale and tempo of scientific inquiry. Models can now search, integrate, and reason over data far beyond data repositories familiar to any individual researcher. Yet this expansion creates a prior problem: before a model can produce a trustworthy scientific result, it must locate data that are appropriate for the question, sufficiently reliable for the intended analysis, and accompanied by enough context to support responsible interpretation. As data becomes increasingly abundant, the challenge of finding data has been overcome by the challenge of finding data that you can trust. This paper explores how the social and empirical evidence that accumulates when data are used in research can be used, analogous to social trust networks, to determine fit for purpose and trust. Specifically, the paper explores data-usage graphs as a new layer of scientific data infrastructure. A data-usage graph connects datasets to the publications, people, institutions, topics, software, models, workflows, and other datasets through which they are produced and used. These connections reveal the \it social life of data: who has relied on a source, for which questions, in what combinations, with which methods, and with what observable impact. They can turn scattered traces of practice into data-usage descriptors that complement conventional metadata and support judgments of trust and fitness for purpose. The central claim is not that popularity establishes trust, but that this can be grown with appropriate contextual history. Usage evidence must therefore be combined with production quality, provenance, governance, semantic clarity, and community validation. The feasibility and value of data usage graphs is demonstrated by implementing the prototype data insights discovery service within the National Data Platform (NDP).
A.Z.B. Aziz, S.F. Ahmed, G. Rasineni, M. Wang, O. Hatipoglu, M. Ricci, M. Shaw, G. Li, J. Q. Brown, V. Pascucci, S. Elhabian.
SIMPLER: H&E-Informed Representation Learning for Structured Illumination Microscopy, Subtitled arXiv:2604.10334v1, 2026.
Structured Illumination Microscopy (SIM) enables rapid, high-contrast optical sectioning of fresh tissue without staining or physical sectioning, making it promising for intraoperative and point-of-care diagnostics. Recent foundation and large-scale self-supervised models in digital pathology have demonstrated strong performance on section-based modalities such as Hematoxylin and Eosin (H&E) and immunohistochemistry (IHC). However, these approaches are predominantly trained on thin tissue sections and do not explicitly address thick-tissue fluorescence modalities such as SIM. When transferred directly to SIM, performance is constrained by substantial modality shift, and naive fine-tuning often overfits to modality-specific appearance rather than underlying histological structure. We introduce SIMPLER (Structured Illumination Microscopy-Powered Learning for Embedding Representations), a cross-modality self-supervised pretraining framework that leverages H&E as a semantic anchor to learn reusable SIM representations. H&E encodes rich cellular and glandular structure aligned with established clinical annotations, while SIM provides rapid, nondestructive imaging of fresh tissue. During pretraining, SIM and H&E are progressively aligned through adversarial, contrastive, and reconstruction-based objectives, encouraging SIM embeddings to internalize histological structure from H&E without collapsing modality-specific characteristics. A single pretrained SIMPLER encoder transfers across multiple downstream tasks, including multiple instance learning and morphological clustering, consistently outperforming SIM models trained from scratch or H&E-only pretraining. Importantly, joint alignment enhances SIM performance without degrading H&E representations, demonstrating asymmetric enrichment rather
D. Balouek, L. Carnevale, M. Villari, M. Parashar.
Managing Compute-Communication Tradeoffs in Urgent Edge-Cloud Systems: Case studies on Disaster Recovery and Earthquake Early Warning, In Frontiers in Complex Systems , Frontiers, 2026.
Urgent Computing applications demand time-critical processing under strict latency, reliability, and scalability constraints. Earthquake Early Warning (EEW) systems exemplify this challenge, where milliseconds can determine the effectiveness of life-saving alerts. This paper investigates how Edge Computing Middleware can support such urgent workloads across the computing continuum, spanning edge devices, fog layers, and cloud infrastructures. We first propose a reference architecture, derived from a survey of Edge-based Middleware systems, to enable adaptive orchestration, low-latency stream processing, and dynamic workload placement across heterogeneous resources. To validate this architecture, we present two complementary case studies that systematically explore the compute–communication tradeoffs central to urgent computing: (1) Disaster Recovery, which focuses on communication-centric challenges such as network latency, data transfer costs, and placement policies; and (2) Earthquake Early Warning, which examines compute-centric optimizations, including hyperparameter tuning and sensor selection for real-time seismic analysis. Together, these case studies demonstrate how architectural decisions and workload-specific optimizations interact to meet the stringent requirements of urgent applications. Finally, we identify key research challenges in prediction and automated resource management for managing trade-offs across the computing continuum. By integrating architectural design, empirical tradeoff evaluation, and dataset-driven analysis, this work contributes a systematic framework for engineering urgent computing systems, with direct implications for next-generation Earthquake Early Warning infrastructures and other time-critical applications.
W. Bangerth, C. R. Johnson, D. K. Njeru, B. van Bloemen Waanders.
Estimating and using information in inverse problems, In Inverse Problems and Imaging, Vol. 24, pp. 1--33. 2026.
ISSN: 1930-8337
DOI: 10.3934/ipi.2026003
In inverse problems, one attempts to infer spatially variable functions from indirect measurements of a system. To practitioners of inverse problems, the concept of "information" is familiar when discussing key questions such as which parts of the function can be inferred accurately and which cannot. For example, it is generally understood that we can identify system parameters accurately only close to detectors, or along ray paths between sources and detectors, because we have "the most information" for these places.
Although referenced in many publications, the "information" that is invoked in such contexts is not a well understood and clearly defined quantity. Herein, we present a definition of information density that is based on the variance of coefficients as derived from a Bayesian reformulation of the inverse problem. We then discuss three areas in which this information density can be useful in practical algorithms for the solution of inverse problems, and illustrate the usefulness in one of these areas – how to choose the discretization mesh for the function to be reconstructed – using numerical experiments.
S. Bhandari, N. Khan, A. Novotna, T. Jeong, L. Bowman, M. Hernandez, T. Somorin, V. Govani, J. Goldstein, S. Elhabian.
TRACE: Artifact-Robust Statistical Shape Modeling from Imperfect Surface Scans-A Case Study in Craniosynostosis 3D Photography, Subtitled arXiv:2608.22131, 2026.
Craniosynostosis severity analysis increasingly relies on statistical shape models (SSMs) to quantify cranial morphology, but most existing workflows depend on computed tomography or heavily curated three-dimensional (3D) photographs. Raw clinical 3D photographs provide a radiation-free and repeatable alternative, yet often contain shoulders, hands, hair, clothing, scanner noise, and incomplete boundaries that corrupt correspondences. We introduce the Template-constrained Robust Artifact-aware Correspondence Estimation (TRACE) framework, an unsupervised method for constructing SSMs directly from artifact-contaminated clinical 3D head photographs. TRACE predicts sparse anatomically corresponding head-surface control points from the raw point cloud, refines them through a coarse-to-fine Surface-Aware Deformation cascade, and uses thin-plate spline warping to deform a clean template mesh into a subject-specific head reconstruction. This template-constrained formulation keeps dense correspondences on clinically relevant head anatomy while suppressing non-head artifacts. The correspondence module is decoupled from the point-cloud encoder, enabling the same deformation pipeline to be paired with different backbones, including PointNet, DGCNN, and Point Transformer V3. Across all backbones, TRACE substantially improves surface sampling, topology preservation, and shape-model quality over prior SSM methods, providing a scalable foundation for photograph-based craniosynostosis shape analysis and a framework that may extend to other artifact-contaminated surface scans when an appropriate clean template is available.
T. Bidone.
Rethinking Contractility in Active Cytoskeletal Matter, In Biophysical Journal, 2026.
R.T. Black, S.A. Maas, W. Wu, J. Maheshwari, T. Kolev, J.A. Weiss, M.A. Jolley.
An open-source computational framework for immersed fluid-structure interaction modeling using FEBio and MFEM, Subtitled arXiv:2601.08266v1, 2026.
Fluid-structure interaction (FSI) simulation of biological systems presents significant computational challenges, particularly for applications involving large structural deformations and contact mechanics, such as heart valve dynamics. Traditional ALE methods encounter fundamental difficulties with such problems due to mesh distortion, motivating immersed techniques. This work presents a novel open-source immersed FSI framework that strategically couples two mature finite element libraries: MFEM, a GPU-ready and scalable library with state-of-the-art parallel performance developed at Lawrence Livermore National Laboratory, and FEBio, a nonlinear finite element solver with sophisticated solid mechanics capabilities designed for biomechanics applications developed at the University of Utah. This coupling creates a unique synergy wherein the fluid solver leverages MFEM's distributed-memory parallelization and pathway to GPU acceleration, while the immersed solid exploits FEBio's comprehensive suite of hyperelastic and viscoelastic constitutive models and advanced solid mechanics modeling targeted for biomechanics applications. FSI coupling is achieved using a fictitious domain methodology with variational multiscale stabilization for enhanced accuracy on under-resolved grids expected with unfitted meshes used in immersed FSI. A fully implicit, monolithic scheme provides robust coupling for strongly coupled FSI characteristic of cardiovascular applications. The framework's modular architecture facilitates straightforward extension to additional physics and element technologies. Several test problems are considered to demonstrate the capabilities of the proposed framework, including a 3D semilunar heart valve simulation. This platform addresses a critical need for open-source immersed FSI software combining advanced biomechanics modeling with high-performance computing infrastructure.
J. Bond, J. Pake, C. David, A. McNutt, T.S. Muwonge, D. Orchard, R. Perera.
Literate Execution, Subtitled arXiv:2604.26967v2, 2026.
\emphLiterate programming, introduced by Knurth, interleaves code and prose so that a program can be read as both executable and explanatory text. We propose \emphliterate execution, which inverts this relationship: rather than embedding code within a static narrative, we treat documentation -- and other expository elements such as visualisations -- as first-class artefacts that can be computed alongside a running program and then integrated into a view of its execution. We explore this idea through Fluid, a programming language with a provenance-tracking runtime that records fine-grained dependencies between inputs and outputs. These provenance relationships can be surfaced as interactions that allow readers to explore how intermediate values contribute to a result. By integrating visualisation, provenance, and exposition, literate execution aims to make programs more explorable and self-explanatory, and explorable explanations easier to program.
K. Borkiewicz, J. Li, J.A. Levine, K.E. Isaacs.
May (A) I Beautify Your Visualization? Expert Judgments of Acceptable Aesthetic Alterations, Subtitled arXiv:2607.00239, 2026.
In 3D visualizations of natural phenomena, improving aesthetics can provide measurable benefits, but often involves transformations that affect how the data is perceived. As a growing range of tools - including AI-based methods - make visual design and modification more accessible, it is increasingly important to understand trade offs and concerns when making these changes. We conducted an expert survey (N=95) with visualization researchers, practitioners, and domain scientists, investigating reactions to fifteen alterations spanning presentation-level adjustments (e.g., lighting, camera position) and data-level modifications (e.g., removing errors, filling gaps), applied by both humans and AI systems. Results show differences in perceived acceptability are driven by the transformation's meaning, regardless of whether it operates at the presentation or data level. Additionally, certain modifications were consistently judged as more permissible than others regardless of human or AI authorship. While this relative ordering remains largely stable, AI-generated transformations are consistently rated as less acceptable than identical human-produced changes. These results reveal a distinction between more permissible and more sensitive alterations, and suggest the need for both designers and AI-assisted visualization tools to incorporate constraints and guardrails that reflect these differences.
D. Brown, Y. Huang, S.H. Wang, B. Wang.
Second Order Drifting Models, Subtitled arXiv:2608.07924, 2026.
Drifting models are a recent class of one-step generative models that evolve the model distribution during training using a predefined sample-based drift field. Although they avoid iterative inference, their kernel-based drift fields induce frequency-dependent training dynamics: In the linearized regime, each Fourier mode of the density residual decays at a rate determined by the kernel spectrum, leading to slow recovery of fine-scale structure. We propose Second-Order Drifting Models, which lift drifting dynamics into phase space by augmenting generated samples with artificial velocity variables. We show that the resulting density perturbations obey accelerated second-order dynamics in Fourier space, connecting drifting models to the celebrated Nesterov acceleration from optimization theory. This provides a principled mechanism for mitigating the spectral stiffness of first-order drifting while preserving one-step inference. We derive a practical semi-implicit training algorithm and evaluate it on synthetic distribution matching, sequential data generation, and robotic control. Across these settings, the second-order drifting model improves convergence behavior and achieves competitive or superior performance over first-order drifting baselines.
A. Busatto, L.C.R. Tanner, J.A. Bergquist, G. Plank, K. Gillette, A. Narayan, R.S. MacLeod.
Uncertainty quantification of conduction velocity in models of cardiac spread of activation, In Med Biol Eng Comput, Springer Nature, 2026.
This study quantified the effect of conduction velocity (CV) variability on cardiac electrical activation patterns, a key factor for cardiac digital twins. We examined how myocardial and endocardial longitudinal, transverse, and sheet CVs influence ventricular activation across multiple pacing sites. Three porcine biventricular heart models, each including a fast-conducting endocardial layer, were used to simulate electrical activation with an eikonal approach. Uncertainty quantification with polynomial chaos expansion systematically varied six CV parameters within physiological ranges. In total, 1,868 simulations from eight ventricular pacing sites were analyzed for activation time, variability, and global sensitivities. Myocardial longitudinal CV showed the greatest influence on activation timing (global sensitivity up to 0.98). Endocardial-layer longitudinal CV was similarly important for endocardial stimuli, while transverse and sheet CVs had minimal effects. Activation-time variability reached 15 ms, increasing with distance from the pacing origin. Longitudinal CVs, particularly myocardial and endocardial-layer, dominate ventricular activation dynamics and should be prioritized when personalizing cardiac digital twins. Accounting for CV uncertainty is essential for accurate prediction and therapy optimization.
A. Busatto, J.A. Bergquist, T. Tasdizen, B.A. Steinberg, R. Ranjan, R.S. MacLeod.
Predicting Ventricular Arrhythmia in Myocardial Ischemia Using Deep Learning, In Heart Rhythm O2, Elsevier, 2026.
Background Myocardial ischemia can trigger ventricular arrhythmias with life-threatening consequences. Current monitoring is largely reactive, limiting opportunities for preventive intervention. Objective To determine whether high-resolution epicardial electrograms contain predictive signatures that enable forecasting the timing of premature ventricular contractions (PVCs) during acute ischemia, and to quantify subject-specific data requirements for effective personalization. Methods We analyzed epicardial sock electrograms (247 electrodes, 1 kHz) from 21 porcine acute ischemia experiments comprising 2,252 spontaneous PVCs. Signals were segmented into overlapping sequences of 3, 5, or 7 consecutive non-PVC beats with a continuous target of time-to-next PVC. A 6-layer Long Short-Term Memory (LSTM) network (hidden size 128) with temporal attention was trained using mean absolute error (MAE). Performance was evaluated in (A) pooled 80/10/10 cross-validation and (B) leave-one-experiment-out testing with subject-specific fine-tuning using 10% or 15% of held-out data. Results In Paradigm A, MAE decreased with longer context (6.50 s for 3 beats, 5.97 s for 5 beats, 4.73 s for 7 beats) with excellent calibration (R2>0.996). In Paradigm B, increasing fine-tuning from 10% to 15% reduced mean MAE by 9.6–14.6 s and flattened error growth with prediction horizon, improving the fraction of predictions within 30–60 s windows. Conclusion Epicardial electrograms support accurate PVC time-to-event forecasting during acute ischemia, and modest subject-specific adaptation substantially improves generalization, motivating development of real-time predictive monitoring tools.
A. Cattaneo, M.K. Ballard, R.M. Kirby, V. Shankar.
JetSCI: A Hybrid JAX-PETSc Framework for Scalable Differentiable Simulation, Subtitled arXiv:2604.22087v1, 2026.
The rapid rise of scientific machine learning (SciML) has expanded the role of differentiable modeling, surrogate modeling, and data-driven constitutive laws in large-scale simulation. The JAX framework provides an attractive environment for these workflows through automatically differentiable programs, vectorization, GPU acceleration, and while enabling seamless learning of surrogate models. However, large-scale simulation still relies on mature HPC infrastructure. Libraries, such as PETSc, provide scalable MPI-based parallelism, robust linear and nonlinear solvers, and advanced preconditioning capabilities that remain difficult to reproduce in JAX-only workflows. We present JetSCI, a hybrid JAX-PETSc framework that unifies these complementary strengths. JetSCI uses JAX for GPU-parallel differentiable discretizations and PETSc for robust, scalable solution of the resulting systems on distributed-memory architectures, exposing multilevel parallelism through GPU acceleration within nodes and MPI parallelism across nodes. For finite element discretizations of heterogeneous micromechanics problems, JetSCI outperforms JAX-only implementations in efficiency and accuracy.
L. Chenarides, R. Ladislau, M. Parashar, S. Porter, J. Lane.
Data-usage descriptors as search metadata: the case of food security data and the National Data Platform (2015-2025), In Scientific Data, Vol. 13, No. 1150, Nature, 2026.
DOI: https://doi.org/10.21203/rs.3.rs-8569040/v1
Scientific data is a critical input into scientific research. Yet the research data landscape is constantly changing as new datasets emerge, others are retired, or some disappear altogether. Data-usage descriptors can substantially advance research productivity by reducing the time that researchers spend finding new and relevant datasets in their research field. This paper describes how to generate data usage descriptors by finding how datasets are used in publications and then linking the dataset information to the publication metadata. It also shows how usage descriptors can be used to find other related datasets and their usage. It concludes by arguing that the approach represents a critical piece of foundational infrastructure that could be deployed in repositories as part of a referenceable, navigable, and contextual data framework. This article contains a reproducible workflow for constructing data-usage descriptors, based on analyzing the full text of publications in the Dimensions database. The illustrative use case is research on food security. The illustrative repository is the National Data Platform.
C. Christenson, S. Viknesh, R.L. Judson-Torres, A. Arzani.
Data-driven system identification in cancer systems biology: A multiscale modeling approach to melanoma, In Computer Methods and Programs in Biomedicine, Vol. 285, Elsevier, 2026.
Background and Objective: Developing systems biology models of cancer is critical for improving understanding of complex biological interactions and advancing appropriate therapeutic strategies. However, constructing complex, multiscale models that are easily parameterized by the available data can be difficult. Advances in computational methods have provided data-driven frameworks utilizing machine learning approaches to build predictive systems biology models. These methods display a strong ability to fit data and predict cellular and molecular trajectories but can be limited by either poor mechanistic interpretations or require a large prior understanding of the system of interest. In many cases, such as cancer, the systems are highly complex, resulting in limited availability of prior knowledge and a strong need for mechanistic insight from the developed models. This work builds upon recent developments in system identification to construct fully mechanistic models directly from the available data, serving as a proof-of-concept study for multi-scale system identification with synthetic data. Methods: We utilize ADAM-SINDy (sparse identification of nonlinear dynamics with ADAM) to accomplish this objective, a recently developed differentiable optimization method for the identification of nonlinear, parameterized dynamical systems, with advancements to facilitate its use in systems biology settings. Specifically, we focus on the usage of prior biological knowledge and the subsequent finetuning through a two-stage identification process. We apply the framework in the context of melanoma (a dangerous skin cancer) and quantitative systems pharmacology, developing a benchmark multiscale model connecting sub-cellular signaling and cell scale dynamics in response to therapy. Results: The ADAM-SINDy framework is shown to be capable of identifying a ground truth system of equations directly from synthetically generated data. Coupling mechanisms between scales (subcellular and cellular) are also identifiable with this framework. At high temporal resolutions, the method achieves 100% reconstruction accuracy, dropping to 89.9% with a two times coarser resolution. Conclusions: The framework presented in this work is capable of constructing systems biology networks and mathematical models directly from data, while maintaining clear mechanistic interpretations. This system identification paradigm in systems biology can help remove the need for extensive biological insight into the network of interest a priori, which can be exceedingly difficult in the context of cancer.
L. Cicci, S. Qian, C. Rodero, M. Strocchi, C. Corrado, F. Campos, S. Malik, A. Lee, A. Qayyum, K. Gillette, J. Isbister, R. Sy, M. Lee, M. Noseda, R. Wilkinson, G. Plank, M. Bishop, S. Niederer.
Personalising cardiac electrophysiology models from CT and ECG for 3D activation imaging and tissue characterisation, Subtitled Research Square Preprint, 2026.
Background: Electrocardiographic imaging maps cardiac electrical activity
non-invasively but is restricted to the epicardium. Computational electrophysiol-
ogy models can predict 3D activation and tissue properties but require extensive
parameter calibration.
Methods: We introduce an unbiased workflow combining sensitivity analysis
with emulator-based Bayesian history matching to calibrate over 100 organ- and
tissue-scale parameters. The framework incorporates CT-scan images and 12-
lead ECGs with a multi-scale electrophysiology model to generate personalised
ventricular simulations.
Results: The framework was tested on seven subjects (four with synthetic and
three with clinical ECGs), with validation performed using high-density body
surface potentials from a 252-electrode vest for the clinical cases. Calibrated
models reproduced individual ECG morphologies and showed strong agreement
with independent measurements (Pearson’s correlation coefficient: 0.80 ± 0.04).
Conclusions: The study links non-invasive data with high-fidelity simulations to
estimate spatially-varying properties, supporting personalised cardiac modelling
for clinical use.
Page 1 of 154
