Glossary
Definition
Models that accept and combine more than one input type — text, images, audio, video. Multimodal assistants can look at a screenshot while discussing it.
Browse related tools →Glossary
Models that accept and combine more than one input type — text, images, audio, video. Multimodal assistants can look at a screenshot while discussing it.
Browse related tools →