Sarang Patil

Sarang Patil

PhD Student, Department of Data Science, New Jersey Institute of Technology

Profile photo

I am a third-year PhD student in the Department of Data Science at the New Jersey Institute of Technology (NJIT), advised by Dr. Mengjia Xu and a member of the Xu Lab. My research explores how hyperbolic geometry, state-space models, and graph neural networks can be used to build efficient large language models and dynamic graph embeddings.

My most recent work is HyperForecast, a long-term time-series forecasting framework that operates in hyperbolic space. Earlier work includes the Hierarchical Mamba (HiM) framework (published in TMLR, 2026), which combines efficient Mamba2 state-space models with hyperbolic spaces to capture hierarchical relationships in language data, and a comparative study of dynamic graph embedding approaches using transformers and the Mamba architecture. In addition to these projects, I continue to develop hyperbolic models for other domains, exploring how curvature-aware embeddings can benefit a wide range of applications. I was also involved in surveying, and organizing the rapidly growing body of work on hyperbolic large language models, which forms the basis of my accepted survey paper in SIAM Review.

Beyond hyperbolic learning, I work on scientific machine learning: multiscale graph-wavelet compressed sensing for physical simulation data, and machine-learning forecasting of solar active-region emergence with collaborators.

Outside of research, I enjoy playing chess ♟️, hiking 🥾, exploring new cuisines 🍜🌍, and playing video games 🎮.

Education

New Jersey Institute of Technology

Ph.D. in Data Science (2024 â€“ Present)

University of Maryland Baltimore County

Master of Professional Studies in Data Science (2021 â€“ 2022)

Savitribai Phule Pune University, India

Bachelor of Engineering in Computer Engineering (2016 â€“ 2020)

Interests

  • Hyperbolic geometry and non-Euclidean representation learning
  • State-space models (SSMs) and efficient sequence modeling
  • Large language models and hierarchy-aware embeddings
  • Graph neural networks and dynamic graph embedding
  • Curvature-aware optimization and geometric deep learning
  • Time-series forecasting and scientific machine learning

Experience

Research Assistant, New Jersey Institute of Technology, NJ

September 2024 â€“ Present

Research Assistant, University of Maryland Baltimore County, MD

Jan 2022 â€“ Dec 2022

Data Science Intern, CoReCo Technologies, Pune, India

Aug 2019 â€“ Jun 2020

Project Intern, Aalborg University, Copenhagen, Denmark

Jan 2018 â€“ Feb 2018

Publications

Hyperbolic Representation Learning & LLMs

Hierarchy-aware embeddings, state-space models and large language models in hyperbolic space.

HiM diagram
Sarang Patil, Ashish Parmanand Pandey, Ioannis Koutis, Mengjia Xu
TMLRTransactions on Machine Learning Research, 2026
This work introduces the Hierarchical Mamba (HiM) model, which integrates efficient Mamba2 state-space models with hyperbolic representations (PoincarĂ© and Lorentz manifolds) to learn hierarchy-aware language embeddings. HiM projects Mamba2 outputs into hyperbolic space with learnable curvature and hyperbolic loss functions, capturing relational distances across levels of a hierarchy. Experiments on linguistic and medical datasets show that HiM outperforms Euclidean baselines and highlights the trade-offs between the PoincarĂ© and Lorentz variants. Source code is available online.
Hyperbolic LLM diagram
Sarang Patil, Zeyong Zhang, Yiran Huang, Tengfei Ma, Mengjia Xu
SIAM ReviewSIAM Review (accepted, to appear)
Large language models excel at many tasks but often fail to capture the non-Euclidean hierarchies present in real-world data. This survey paper reviews recent progress on Hyperbolic LLMs (HypLLMs) and has been accepted for publication in the SIAM Review. It categorizes existing models into four groups—models using exponential/logarithmic maps, hyperbolic fine-tuned models, fully hyperbolic models, and hyperbolic state-space models—and discusses applications across language, vision and multimodal domains. A companion repository of papers, code and datasets is maintained on GitHub.
HypEHR architecture
Yuyu Liu, Sarang Patil, Mengjia Xu, Tengfei Ma
ACL 2026Findings of the Association for Computational Linguistics: ACL 2026, pp. 10849–10862
EHR question answering usually relies on costly LLM pipelines that ignore the tree-like organization of clinical data. HypEHR is a compact model that embeds medical codes, visits and questions in hyperbolic (Lorentz) space, pretrained on next-visit diagnosis prediction with ICD-aligned hierarchy regularization. On MIMIC-IV EHR-QA benchmarks it approaches the performance of LLM-based methods with far fewer parameters.

Scientific Machine Learning

Learning-based compression and recovery of graph-structured scientific simulation data.

Graph Wavelet Compressed Sensing pipeline
Amirhossein Nouranizadeh*, Sarang Patil*, Alan John Varghese, Varsha Narayanan, Amit Chakraborty, Mengjia Xu (*equal contribution)
SDM 2026SIAM International Conference on Data Mining (SDM), 2026
Training scientific machine learning models such as neural operators requires large volumes of simulated data, which makes data preparation and storage expensive. This paper proposes Graph Wavelet Compressed Sensing (GWCS), a learning-based framework that compresses graph signals offline into sparse, interpretable wavelet-domain representations using the spectral graph wavelet transform, multilevel importance sampling, and a neural inverse graph wavelet transform for scale-aware recovery.

Machine Learning for Solar Physics

Datasets and deep-learning models for forecasting solar active-region emergence from NASA SDO/HMI observations.

Transformer pipeline for continuum intensity forecasting
Jonas Tirona, Sarang Patil, Spiridon Kasapis, Eren Dogan, John T. Stefan, Irina N. Kitiashvili, Alexander G. Kosovichev, Mengjia Xu
JGR: MLCJournal of Geophysical Research: Machine Learning and Computation, 3(4), e2025JH001207, 2026
A transformer-based model that detects faint precursor signals of solar active-region emergence in helioseismic acoustic-power and magnetic-flux data from NASA's Solar Dynamics Observatory, forecasting emergence on average about 9 hours in advance and outperforming a standard transformer and a previous benchmark. The work was featured in NJIT News.
SolARED data processing pipeline
Spiridon Kasapis, Eren Dogan, Irina N. Kitiashvili, Alexander G. Kosovichev, John T. Stefan, Jake D. Butler, Jonas Tirona, Sarang Patil, Mengjia Xu
Solar PhysicsSolar Physics, 301(7), 106, 2026
A machine-learning-ready dataset built from full-disk SDO/HMI observations that tracks 50 large active regions (2010–2023), capturing acoustic power, unsigned magnetic flux and continuum intensity before, during and after emergence, together with an interactive web portal for visualization, to support forecasting of solar active-region emergence.

Talks & Presentations

Upcoming

Past

Academic Service

Reviewer:

News

Contact

Email: sp3463@njit.edu
Address: New Jersey Institute of Technology, Newark, NJ 07102

View Larger Map