Dr. Pinaki Bhaskar
I take multimodal AI from research to Galaxy phones.
GenAI Research Scientist & Architect Samsung R&D Institute India – Bangalore 2015 – present
I lead the Advanced Technology Lab at Samsung R&D Institute India – Bangalore, and I still design the systems my teams build. My work is multimodal AI on phones: vision-language models, multimodal LLMs and knowledge graphs that have to fit in the pocket of a device rather than a data centre.
The project that best shows what we do is natural-language search in Samsung Gallery — a vision-language model that lets you find a photo by describing it. It shipped in the Galaxy S25 series after a three-site effort across Seoul, Suwon and Bangalore, and won Samsung Research’s Best Project award. Before that, my team took an on-device personalised knowledge graph from research into the Galaxy S21 series. Today I’m building multimodal LLMs and Large Action Models for agentic AI inside mobile apps.
I came to this from language. My Ph.D. at Jadavpur University (2014) was about getting a machine to read many documents and answer in a few good sentences, and our systems ranked top at CLEF QA4MRE three years running. A postdoctoral year at CNR in Pisa on biomedical question answering followed, and in 2015 I joined Samsung. Along the way: more than 35 publications, seven patents, and the Zinnov Awards 2022 Technical Role Model award.
Selected work
All case studies-
Natural-language search in Samsung Gallery
Find a photo by describing it.
A vision-language model built as one backbone for several Gallery features. Commercialised in the Galaxy S25 series.
Case study -
On‑device personalised knowledge graph
Personal context, built and kept on the device.
A user-centric knowledge graph from the phone’s own data, published at WACV 2024 and ICON, and shipped from the Galaxy S21 series.
Case study -
Multimodal LLMs and Large Action Models for mobile agents
From understanding the screen to acting on it.
Current research: an on-device multimodal LLM and action models that power agentic AI in mobile apps, with related generative-imaging patents.
Case study
Current research
- Multimodal large language models that run on‑device.
- Large Action Models (LAMs) that power agentic AI inside mobile apps.
- One vision-language backbone serving many Gallery features, instead of one model per feature.
Photographs
All photographsAway from work I photograph landscapes and cities. A few frames from the Himalaya, rural Bengal and Europe.
Let’s talk.
I enjoy conversations about multimodal AI, on-device intelligence, and building research teams that ship. Email is the fastest way to reach me.





