Skip to content
Fitila Labs

Research · 03 AI

Scientific machine learning

  • Livesurrogate and reduction library
  • Researchphysics-guaranteed surrogates

A full-physics simulation can take hours, and a design study, calibration or control loop may need thousands of runs. Surrogate models make those studies affordable, but only if they are accurate where the decision depends on them. We test each surrogate against the full model and record the range in which it holds.

Current work

  • Interchangeable surrogates. Neural operators, physics-residual models, Gaussian processes and polynomial-chaos expansions behind one predictive contract, so an optimizer, a solver or an uncertainty study can take a stand-in or the exact model interchangeably, and export it without a machine-learning framework at run time.
  • Consistency tests. The library tests that sensitivity indices computed from a model and from its surrogate agree within their error bars.
  • Reduced-order models. Proper orthogonal decomposition and discrete empirical interpolation for time-dependent problems, kept beside the full model they were trained on.
  • EnergyGPT. A language model specialized for the energy sector, fine-tuned from Llama 3.1 8B on a curated corpus of energy literature, in full-parameter and LoRA variants. Both outperform the base model on most energy question-answering tasks in our benchmarks, the LoRA variant at a fraction of the training cost.

A. Chebbi and B. Kolade, “Towards EnergyGPT: A Large Language Model Specialized for the Energy Sector,” arXiv preprint 2509.07177, 2025. Read the preprint ↗ (opens in a new tab)

Background

This work builds on our earlier research in generative modeling and manifold learning, including manifold-stabilized diffusion for image synthesis and synthetic data for scenarios that are rare or expensive to collect.

Open problems

  • Surrogates with physical guarantees. Operator inference and reduced bases trained on full-physics runs, with the range of validity computed and enforced.
  • Scientific foundation models. Models pretrained to predict any measured quantity from the others, then adapted to a commodity, region or question, with calibrated uncertainty on each prediction. We have not trained such a model yet.
  • Citations in domain language models. We are working on retrieval and citation for domain models, and on evaluations that check whether a cited passage supports the claim.