Few-Shot Computer Vision, Multimodal Foundation Models, 3D Reconstruction and Neural Scene Representation, Data-Efficient Medical Vision
I am a 4th year PhD candidate at Stony Brook University, under the mentorship of Dr. Zhaozheng Yin. My research focusses on few-shot learning, Multimodal Foundation Models, 3D Reconstruction and Neural Scene Representation, Data-Efficient Medical Vision.
Previously, I interned at Centific, where my work included egocentric video simulation dataset generation using Gaussian Splatting, reducing hallucination in video captioning using Cosmos3-Nano Reasoner and Qwen-Instruct vLLM, 3D face pose estimation for selfie videos, and Indian-language OCR for tables and graph understanding.
SUNY Stony Brook | 2023 - 2028 (Expected)
Jadavpur University | 2019 - 2023 | GPA: 9.29/10
Adapting vision-language and foundation models for data-efficient visual understanding.
Building 3D reconstruction pipelines from real-world egocentric data to generate simulated environments, reconstructed assets, and hand-object interactions.
Developing structure-aware AI systems for medical and scientific imaging for medical diagnosis.
ECCV 2026
WACV 2026
Under Review, ACCV 2026
Under Review, IEEE TIP
Journal of Data, Information and Management 2024
VLSID 2022
Handbook of Moth-Flame Optimization Algorithm, 2022
Diagnostics, 2022
Computers in Biology and Medicine, 2021
Applied Sciences, 2021
ISDA 2022