Research

LLM Inference and Serving

I design resource-aware systems that improve latency, throughput, and Service Level Objective compliance for large language model inference. Current work studies heterogeneous batching, coordinated waiting and execution time, KV-cache management, and multi-resource utilization.

Distributed Machine Learning Systems

My research explores data and model parallelism, fault tolerance, elastic scheduling, and reinforcement-learning-based orchestration for distributed deep learning across heterogeneous devices.

Edge AI and Real-Time Inference

I develop techniques for fast and accurate DNN inference on low-cost edge platforms, including heterogeneous accelerator execution, adaptive model partitioning, and resource-aware task assignment.

Secure and Robust AI Systems

I have investigated adversarial attacks and defenses for time-series and autonomous-driving models, as well as real-time detection of memory denial-of-service attacks in cloud systems.

Selected Projects

SLO-Aware LLM Serving
Heterogeneous batching, KV-cache management, and coordinated scheduling
Flex
Fast, accurate DNN inference using heterogeneous accelerator execution on low-cost edges
Fault-Tolerant Distributed Deep Learning
Data/model parallel training and inference across unreliable edge networks
AnyOpt
Prediction and optimization of IP anycast performance using operational network data