Also known as: multimodal model
In plain English
A multimodal AI can work with different kinds of input. You can show it a photo and ask a question about it, or talk to it and have it talk back.
In practice
Multimodal models open up new uses: reading scanned forms and invoices, checking photos from field inspections, or analysing charts in reports. Test accuracy on your own documents, since image understanding is still less reliable than text.
Under the hood
Multimodal models map different data types into a shared representation, for example by encoding images into tokens the language model can attend to. Some are trained natively on mixed data; others connect separate encoders to a language model.
Example
"The multimodal model reads the photo of the meter and logs the reading."