Multimodal
AI that handles more than one kind of input or output, such as text, images, audio and video.
In everyday terms
You can show a multimodal AI a photo of your fridge and ask what to cook, or speak to it and hear it reply.
For professionals
Models with encoders (or shared tokenisation) for several modalities, mapped into a common representation space.
Think of it like…
A colleague who can read, look at pictures and listen, not just read emails.
You've already seen it
Uploading screenshots to ChatGPT, live voice mode, describing an image for accessibility.
Myth vs reality
Myth: Multimodal AI "sees" like we do.
Reality: Images are converted into numbers it processes. It can miss things humans find obvious.
Quick check
Which is a multimodal task?
- Spell-check a sentence
- Describe what's in a photo
- Translate French to English
- Count words
Show answer
Describe what's in a photo: Image in, text out = two modalities.