Research
LLM Inference and Serving
I design resource-aware systems that improve latency, throughput, and Service Level Objective compliance for large language model inference. Current work studies heterogeneous batching, coordinated waiting and execution time, KV-cache management, and multi-resource utilization.
Distributed Machine Learning Systems
My research explores data and model parallelism, fault tolerance, elastic scheduling, and reinforcement-learning-based orchestration for distributed deep learning across heterogeneous devices.
Edge AI and Real-Time Inference
I develop techniques for fast and accurate DNN inference on low-cost edge platforms, including heterogeneous accelerator execution, adaptive model partitioning, and resource-aware task assignment.
Secure and Robust AI Systems
I have investigated adversarial attacks and defenses for time-series and autonomous-driving models, as well as real-time detection of memory denial-of-service attacks in cloud systems.