MOSS-VL-Realtime

Multimodal vision-language model for realtime streaming video understanding. Upload a video (or an image) and ask any question — the model perceives the video frame-by-frame and streams its answer as the stream is observed.

Model Card | GitHub

Examples