Vismayam-V01: Open Interior Intelligence
Today, we are introducing Vismayam-V01 — our specialized zero-shot interior object detection model. Designed explicitly for interior environments, Vismayam-V01 leverages open-vocabulary feature parsing to accurately localize and bound interior elements like furniture, architectural tiles, and sanitary fixtures directly from raw room imagery without requiring hardcoded training labels or dataset fine-tuning.
Vismayam
Vision Detection Benchmarks
Zero-shot open-vocabulary object detection and localization on interior imagery — hover a model to focus
Note: Results are averaged over multiple interior datasets and public benchmarks. Higher is better for all metrics.
Vision detection benchmarks across interior object localization tasks.
Early Access
Vismayam-V01 is being prepared for its first public access, with a primary focus on architects and interior designers working with real-world spaces, materials, furniture, tiles, sanitary fixtures, and architectural elements. We are opening early access for professionals, teams, developers, researchers, and businesses interested in exploring zero-shot interior object detection and turning complex room imagery into localized, searchable visual objects for applications such as design analysis, material discovery, product identification, and visual workflows.
Early users will have the opportunity to test Vismayam-V01 on their own interior images and help shape future versions of the model.
Join the early access program to evaluate Vismayam-V01 on your own interior imagery.
Contact us for early accessAn Open Interior Vision Model
Vismayam-V01 is our specialized open-vocabulary vision model built specifically for interior object detection and localization. It is designed to move beyond fixed detection classes, enabling architects, interior designers, and visual intelligence systems to locate and extract objects such as furniture, tiles, sanitary fixtures, and architectural elements directly from real-world interior imagery using natural-language queries.
Vismayam-V01 is built on a native-resolution vision-language architecture and Parallel Box Decoding (PBD), two key design choices that improve fine-grained spatial understanding while making object localization more efficient. The model combines high-resolution visual feature extraction, language-conditioned feature alignment, and structured box prediction, allowing complete bounding boxes to be decoded in parallel rather than generating coordinates sequentially. We further specialize these capabilities for interior imagery, focusing on tiles, furniture, sanitary fixtures, architectural elements, and other fine-grained objects, enabling the model to convert natural-language queries into precise spatial detections.
Open VLM Model Size Evolution
Parameter count over time — hover a line or card to focus
At a Glance (Latest Available Scale)
Note: Parameter counts are approximate. MoE = Mixture of Experts. Updated August 2026.
From general vision-language models to interior-specialized spatial grounding.
Detection
Vismayam-V01 is designed for long-horizon visual detection in complex interior scenes. Rather than relying on a fixed set of predefined classes, it accepts natural-language queries and uses them to localize objects across furniture, architectural elements, tiles, sanitary fixtures, lighting, and decorative components. Its parallel box decoding approach allows multiple objects to be localized within a scene while maintaining a structured spatial representation.
We evaluate Vismayam-V01 on real-world interior imagery containing dense object layouts, repeated elements, small fixtures, overlapping objects, and visually similar materials. Each model is evaluated using the same images and detection queries, with performance measured across localization accuracy, object-count accuracy, open-vocabulary query matching, and inference efficiency.
The evaluation below examines how different vision systems translate natural-language descriptions into precise object localization. Vismayam-V01 is specifically optimized for the challenges encountered by architects and interior designers, where finding all relevant objects in a scene is often more important than recognizing only the most obvious objects.
| Model | Instances detected | Accuracy |
|---|---|---|
| Vismayam-V01 | 72 | 91% |
| YOLO-World | 32 | 84% |
| Gemini | 45 | 68% |
| OpenAI | 20 | 42% |
Latest Milestone — Extended Detection Test
In our most recent Vismayam-V01 evaluation pass, we re-tested the model on denser interior scenes with overlapping furniture, repeated fixtures, and fine-grained architectural elements. This milestone run reflects improved instance recall and more reliable localization across the full scene.
This milestone reflects the latest Vismayam-V01 test pass and is separate from the multi-model comparison above.
Architecture and Infrastructure
Vismayam-V01 is built around a compact vision-language architecture optimized for open-vocabulary spatial localization. A visual encoder extracts high-resolution features from interior imagery, while a lightweight language backbone processes natural-language object queries and projects them into a shared multimodal representation space. This allows the model to connect visual regions with semantic descriptions without depending on a fixed detection vocabulary.
At the detection stage, Vismayam-V01 uses Parallel Box Decoding (PBD) to predict complete bounding-box representations in parallel rather than generating coordinate tokens sequentially. The detection head converts the fused visual-language representation into structured spatial predictions, allowing multiple objects to be localized within the same scene while preserving the correspondence between object queries and their bounding boxes. This design is particularly useful for dense interior scenes containing repeated objects, overlapping furniture, small fixtures, and architectural elements.
The underlying language component is a small Qwen-based backbone fine-tuned specifically for interior object localization. Rather than optimizing the model for general-purpose visual conversation, training concentrates on spatial grounding and the visual vocabulary encountered in architecture and interior design, including furniture, tiles, sanitary fixtures, lighting, surfaces, cabinetry, windows, doors, and decorative elements. The training pipeline progressively adapts the model from multimodal representation learning toward language-conditioned grounding and dense multi-object localization.
Vismayam-V01 is designed to remain efficient while retaining the flexibility of an open-vocabulary detector. Its inference pipeline accepts an image and natural-language query, produces localized object predictions, and can immediately convert those predictions into individual object crops for downstream visual search, embedding generation, catalog retrieval, and multimodal RAG. This separation between visual perception, localization, and downstream intelligence allows the same detection layer to scale across different interior applications without requiring a new detection head for every product or object category.
Our infrastructure is designed around practical deployment rather than model size alone. We evaluate inference latency, memory consumption, multi-object throughput, localization quality, and object-count accuracy alongside conventional detection metrics. This allows Vismayam-V01 to remain useful in real-world interior workflows where a model must not only identify an object, but reliably locate all relevant instances within a complex scene.
Further technical details on the model architecture, training methodology, PBD implementation, datasets, evaluation protocol, and deployment configuration will be provided in the upcoming Vismayam-V01 Technical Report.
Availability
Vismayam-V01 is currently available through an early access program for architects, interior designers, developers, and research teams. Participants can submit interior imagery, run open-vocabulary detection queries, and explore downstream workflows for material discovery, product identification, and design analysis.
To request access or discuss integration with your workflow:
- Share sample interior imagery and your use case
- Evaluate zero-shot detection on furniture, tiles, fixtures, and architectural elements
- Provide feedback to help shape future model releases
Frequently asked questions
Common questions about Vismayam-V01 for architects, interior designers, and research partners.
What is Vismayam-V01?
Vismayam-V01 is HALLOHOM AI's proprietary interior vision model for zero-shot open-vocabulary object detection. It localizes furniture, tiles, sanitary fixtures, and architectural elements from natural-language queries in room photos.
What is Parallel Box Decoding (PBD)?
PBD predicts complete bounding boxes in parallel instead of generating coordinates one token at a time. This improves multi-object localization in dense interior scenes with overlapping furniture and repeated elements.
How can I access Vismayam-V01?
Vismayam-V01 is available through early access for architects, designers, and qualified partners. Request access on the contact page with sample imagery and your use case.
Is Vismayam-V01 open source?
No. Vismayam-V01 is proprietary. Model weights and training data are not publicly released. Early-access partners receive an evaluation summary and deployment guidance.