MOSS-VL-Realtime
Multimodal vision-language model for realtime streaming video understanding.
Upload a video (or an image) and ask any question — the model perceives the
video frame-by-frame and streams its answer as the stream is observed.
Model Card |
GitHub