# Multimodal AI

> Multimodal AI refers to models that understand and generate more than one type of data — for example text, images, audio and video together. It enables use cases like describing an image, reading a document scan, or answering questions about a diagram.

*Source: https://www.lazlosoftwaresolution.com/glossary/multimodal-ai*

Multimodal models unlock document understanding, visual inspection, accessibility and richer assistants. In enterprise software they power workflows such as extracting data from scanned forms or letting users ask questions about charts and screenshots.
