Portrait of Dr. Pinaki Bhaskar

Dr. Pinaki Bhaskar

I take multimodal AI from research to Galaxy phones.

GenAI Research Scientist & Architect Samsung R&D Institute India – Bangalore 2015 – present

I lead the Advanced Technology Lab at Samsung R&D Institute India – Bangalore, and I still design the systems my teams build. My work is multimodal AI on phones: vision-language models, multimodal LLMs and knowledge graphs that have to fit in the pocket of a device rather than a data centre.

The project that best shows what we do is natural-language search in Samsung Gallery — a vision-language model that lets you find a photo by describing it. It shipped in the Galaxy S25 series after a three-site effort across Seoul, Suwon and Bangalore, and won Samsung Research’s Best Project award. Before that, my team took an on-device personalised knowledge graph from research into the Galaxy S21 series. Today I’m building multimodal LLMs and Large Action Models for agentic AI inside mobile apps.

I came to this from language. My Ph.D. at Jadavpur University (2014) was about getting a machine to read many documents and answer in a few good sentences, and our systems ranked top at CLEF QA4MRE three years running. A postdoctoral year at CNR in Pisa on biomedical question answering followed, and in 2015 I joined Samsung. Along the way: more than 35 publications, seven patents, and the Zinnov Awards 2022 Technical Role Model award.

Selected work

All case studies
  • Certificate of Appreciation: New Valley Project of the Year (R&D), Gallery Search

    Natural-language search in Samsung Gallery

    Find a photo by describing it.

    A vision-language model built as one backbone for several Gallery features. Commercialised in the Galaxy S25 series.

    Case study
  • First page of the paper POP-VQA: Privacy Preserving, On-Device, Personalized Visual Question Answering (WACV 2024)

    On‑device personalised knowledge graph

    Personal context, built and kept on the device.

    A user-centric knowledge graph from the phone’s own data, published at WACV 2024 and ICON, and shipped from the Galaxy S21 series.

    Case study
  • System diagram from US patent application 2026/0099969: adding an entity of interest to a captured image

    Multimodal LLMs and Large Action Models for mobile agents

    From understanding the screen to acting on it.

    Current research: an on-device multimodal LLM and action models that power agentic AI in mobile apps, with related generative-imaging patents.

    Case study

Current research

  1. Multimodal large language models that run on‑device.
  2. Large Action Models (LAMs) that power agentic AI inside mobile apps.
  3. One vision-language backbone serving many Gallery features, instead of one model per feature.

Photographs

All photographs

Away from work I photograph landscapes and cities. A few frames from the Himalaya, rural Bengal and Europe.

  • Snow-dusted Himalayan peaks wrapped in cloud beneath a deep blue sky streaked with cirrus.
    Rohtang Pass, Himachal Pradesh
  • Low sun above a small wooded island, its reflection a bright path across the still reservoir.
    Mukutmanipur, West Bengal
  • The iron lattice of the Eiffel Tower seen from directly below at night, lit gold against a black sky.
    Under the Eiffel Tower, Paris
  • A dark mountain road curving through fresh snow towards a valley filled with storm cloud.
    Rohtang Pass, Himachal Pradesh
  • A line of birds flying low between two wooded headlands over water turned gold by the evening sun.
    Mukutmanipur, West Bengal
  • A rainbow arcing through dark storm clouds above sunlit trees and moored boats on the Seine.
    Rainbow over the Seine, Paris

Let’s talk.

I enjoy conversations about multimodal AI, on-device intelligence, and building research teams that ship. Email is the fastest way to reach me.

pinaki.bhaskar@gmail.com