Xiaolong Ma

AI Researcher at Argonne National Laboratory

prof_pic.jpg

Argonne National Laboratory

Lemont, IL

Be professional, precise and sometimes be perfect.

Xiaolong Ma is currently a Postdoctoral Appointee in the Data Science and Learning (DSL) Division at Argonne National Laboratory, supervised by Dr. Rajkumar Kettimuthu. He received his Ph.D. in Computer Science from University of Nevada, Reno, advised by Prof. Feng Yan and Prof. Lei Yang. His research lies at the intersection of high-performance computing, distributed system and machine learning system, with a particular focus on building scalable, efficient, and reliable infrastructure for large-scale AI training and inference.

His recent work focuses on optimizing large language model inference on leadership-class computing systems, including Aurora and Polaris. His research explores adaptive and resource-efficient LLM serving, with topics such as prefill/decode disaggregation, dynamic batching and resource scaling, hierarchical KV-cache management, resilient and fault-tolerant inference, and the use of opportunistic or preemptible GPU resources. He is also interested in designing efficient and scalable multi-agent systems for scientific workflows, autonomous scientific discovery, and complex decision-making applications.

His broader research interests include distributed and parallel systems, cloud computing, storage systems, multi-agent systems, and AI for Science.

news

Sep 06, 2026 Our paper “OPSERVE: Opportunistic LLM Inference over Fragmented GPU Capacity in HPC Systems” was accepted to the AI on HPC Workshop (AIonHPC) at SC 2026.
Aug 14, 2026 New preprint: “Wrong but Useful: Trajectory Value Beyond Answer Correctness in Multi-Agent Messages”.
Aug 01, 2026 Our paper “BatchFlow: Shared Batch Reuse and Adaptive Scheduling for Multi-Job Training” was accepted to SC 2026 (19% acceptance rate).
May 13, 2026 New preprint: “PEML: Parameter-efficient Multi-Task Learning with Optimized Continuous Prompts”.
Apr 09, 2026 New preprint: “SOLAR: Communication-Efficient Model Adaptation via Subspace-Oriented Latent Adapter Reparametrization”.