South Korea3.63B parametersImage + video16K context

Explore HyperCLOVA X SEED Vision Instruct 3B

Use NAVER's lightweight Korean-first vision-language model for text, images, video, visual questions, charts, diagrams, documents, and OCR-aware analysis.

Model

naver-hyperclovax/HyperCLOVAX-SEED-Vision-Instruct-3B

Input

Configure your request

The processor conversation array. Content supports text, image, and video. Images accept image, url, path, or base64 sources plus filename, ocr, lens_keywords, and lens_local_keywords. Videos accept video, url, or path plus filename, speech_to_text, and Lens metadata.