INTERNALOur own product

VIT Intercore: a private AI assistant you address by voice and by hand

Our private assistant with a 3D memory planet: it wakes to an open palm, answers aloud, searches its memory graph and shares a workbench for assembling parts by hand. We show it in a one-minute trailer filmed from the real interface.

Voice-and-gesture AI assistant · 2026

Amber holographic planet with a memory graph of glowing nodes and links unfolded around it: the VIT Intercore interface
trailer length
1:04
memory nodes and links in the filmed build
387 / 656
fragments of light that form the planet
76,000
hand-training drills on the workbench
6

Videos

VIT Intercore: the trailer1 min 4 s, with sound. Filmed frame by frame from the real interface; the music is an original score written for the film. The hands are scripted tracks, spoken commands appear as on-screen text, and the assistant speaks in its real synthesized voice.

Poster of the VIT Intercore trailer: an amber holographic planet with a memory graph of glowing nodes and links unfolded around it

Task

Build the assistant for our own daily work, and make it something you address the way you would a person: by name and by hand. Instead of windows and menus the screen holds one object, a planet made of 76,000 fragments of light. When it wakes, its memory unfolds around it as a graph: people, the company, systems, skills and music. VIT Intercore is VITON13’s own internal product.

What we did

We built the core scene, the hand control, the voice and a second scene, the workbench. The hand works as a pointer: an open palm wakes the core, a pinch grabs and turns the view, two pinches zoom, and a short pinch presses a node. The microphone waits for the word “Vit”, and the assistant answers aloud.

The workbench has real depth: parts, wires, a soldering iron and a bin. Smart glasses and a watch are assembled step by step, by hand or by voice, and six drills train the hand: grab and throw, press, sort, stack, zoom and turn. To show the product we made a trailer of 1 minute 4 seconds; its music is an original score written for the film.

How it works

The camera follows both hands. Hand tracking runs locally with MediaPipe, and no camera frame is stored. Voice goes through a single brain: it understands the request, calls tools (open a section, find a node, play music, assemble a kit) and speaks the answer. When music is playing and the name is heard, the track steps back so the assistant can listen and reply.

The planet, the graph, the table and every part are procedural: drawn in code with three.js, with no 3D model files. The trailer is filmed frame by frame from the real interface.

What exists now

The assistant is private, so what we show is its trailer. The film runs 1 minute 4 seconds and goes through the wake-up, the memory, the gestures, a voice search, music, the workbench and hand training. The build we filmed on 7 October 2026 holds 387 memory nodes and 656 links, and its planet is made of 76,000 fragments of light.

The gallery shows screens of the same interface: waking with an open palm, zooming with two hands, the card of a node after a voice search, the glasses being assembled and the finished watch.

What this does not prove

This is our own internal product, not client work. The app is private: there is no public demo and no download, so a reader cannot try it; the film and the screens are all we show.

In the film the hands are scripted tracks sent through the same pointer the camera drives, and spoken commands appear as on-screen text; the assistant’s lines are its real synthesized voice. Personal records in memory were replaced with demo data before filming. The case claims no clients, revenue, user numbers or benchmarks, and the film does not measure how accurately the assistant recognises live hands or speech.

How to check a web studio’s portfolio

Services this case shows

Describe the task on the service page; a written scope, timeline and price come back within one working day.

Discuss a similar project

More cases

All cases