I'm an MSR student at Carnegie Mellon University, advised by Prof. Laszlo A. Jeni. I am broadly interested in facilitating general-purpose dexterous manipulation using visual understanding. Currently, I work on 3D/4D visual understanding and hand-object interaction.
I have been lucky to work with some wonderful people along the way. Before CMU, I spent two years at Samsung R&D Institute India as a Research Engineer, building inference pipelines for serving on-device vision and language models on Samsung S24 and S25 Series, and also worked on building a text-based multilingual (12+ languages) safety filter for its AI models. During undergrad, I worked at the Vision and AI Lab, IISc Bangalore with Prof. R. Venkatesh Babu and Varun Jampani on long-tailed image generation with generative models, focusing on GANs.
I graduated from BITS Pilani with a dual degree in Computer Science & Economics.
A central open problem in robotics is general-purpose manipulation: how can a robot understand and interact with objects it has never seen before, rather than learning a separate representation or policy for every object and task? I am interested in learning visual representations that capture geometry, articulation, affordances, and dynamics, and generalize across rigid, articulated, and deformable objects.
Humans and many animals can manipulate unfamiliar objects by inferring where to interact, what motions are possible, and how objects will respond to actions. I am interested in learning similar priors from videos of human-object interactions and use them to generalize robot manipulation to unseen objects and tasks.
My work focuses on 3D/4D reconstruction, object geometry, and hand-object interaction. I am also interested in self-supervised learning, generative models, and world models for learning predictive representations of objects and interactions from large-scale visual data.
👁️
Learning Manipulation from ObservationLearning reusable manipulation priors from human demonstrations and interaction videos.
🧩
Geometry & 4D InteractionRecovering object structure, motion, articulation, deformation, and hand-object contact from video.
🌎
World Models for Embodied AgentsLearning generative and predictive models of physical interaction from largely self-supervised visual experience.
Publications
ECCV 2026 JointHOI: Jointly Generating Contact Maps Enhances Hand Object Interaction Generation
M. Song, J. Cho, J. Kim, A. Bal, Kartik Sharma, Y. Yu, L.A. Jeni, J. Noh European Conference on Computer Vision (ECCV), 2026
A single-stage diffusion framework that jointly generates 3D hand-object motion and dynamic contact maps from text, with contact-guided sampling to reduce penetration and floating.
CVPR 2026 Workshop MAD-HOI: Masked Autoregressive Diffusion for Generating Articulated Hand Object Interactions from Text
A. Bal, Kartik Sharma, E. Lai, S. Tiwari, L. Dahiya, C. Chawla, L.A. Jeni 1st PhysHuman Workshop, IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2026
Unifies masked autoregression with flow-matching diffusion in a continuous latent space for text-to-HOI generation, enabling variable-length, composite, infilling, and end-of-motion prediction from a single objective.
CVPR 2026 FPSBench: A Benchmark for Video Understanding at High Frame Rates
R. Choudhury, J.S. Dandurand, K. Qiu, K.M. Bhat, Kartik Sharma, L. Dahiya, Y. Zhao, S. Kundu, C.H. Lin, K. Kitani, L.A. Jeni IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2026
A large-scale video QA benchmark to evaluate VLMs at high frame rates, introducing the minFPS metric.
CVPR 2023 NoisyTwins: Class-Consistent and Diverse Image Generation through StyleGANs
H. Rangwani, L. Bansal, Kartik Sharma, T. Karmali, V. Jampani, R. Venkatesh Babu IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023
IEEE Big Data 2022 A Generalized Multimodal Deep Learning Model for Early Crop Yield Prediction
A. Kaur, P. Goyal, Kartik Sharma, L. Sharma, N. Goyal IEEE International Conference on Big Data, 2022
Research Engineer· Samsung R&D Institute IndiaNov 2023 – May 2025
Built and deployed compact safety and grounding models for multilingual LLM/LVM systems on flagship devices. Achieved 95% accuracy across 12+ locales with 45% smaller models. Improved cross-lingual image grounding for better object localization in low-resource languages.
Data Scientist· PrivateBlokFeb 2023 – Nov 2023
Built PrivateBlok's MVP: a GPT-3.5-powered financial QA chatbot for 10K+ companies. Enhanced retrieval accuracy with a custom re-ranking algorithm and created a temporal knowledge graph for detailed financial insights.
Project Assistant· Video Analytics Lab, IISc BangaloreAug 2022 – Jan 2023
Improved long-tailed image generation using StyleGANs, achieving 19% better FID scores. Published work at CVPR 2023, setting a new state-of-the-art for long-tailed datasets.
Developed an object-action detection system with 82% precision and optimized transformers for large-scale action recognition.
Projects
CMU 16-831 · Spring 2026 RL for Articulated Object Manipulation in ManiSkill3 Kartik Sharma, Kshitiz, Soumojit Bhattacharya
Benchmarked PPO, SAC, and Model-Based RL for the OpenCabinetDrawer-v1 task. Proposed three modifications: ICM+PPO (70.1% success), Demonstration-Augmented SAC (63.3%), and BC warm-start + RL fine-tuning (58.0%).
CMU 10-799 · Spring 2026 Diffusion & Flow Matching Kartik Sharma
PyTorch implementation of DDPM, DDIM, Flow Matching, and Flow Map Matching for image generation on CelebA-64. Features Dual-Time U-Net, diagonal-annealed time-pair sampling, and JVP-based Lagrangian PDE loss via torch.func.