Vision-Language Model (VLM) Glossary Definition

Term Vision-Language Model (VLM)
Definition

A vision-language model is an artificial-intelligence model that connects visual information with natural-language input or output.

Context & Usage

Teams use VLMs for image captioning, visual question answering, document understanding, screenshot interpretation, and multimodal assistants. This is an emerging term, so teams should verify how a specific source, platform, or standard uses it before treating it as a settled category.

Additional Resource: Authoritative source.

Related concepts include Multimodal Model, Computer Vision and Natural Language Processing.

Categories Software Development and Programming, Artificial Intelligence