以三個相似度空間探索油麻地的生成影像:Video Embedding、Pixel Information,以及 Text Embedding。
每個空間都可以用 3D、2D Grid 或 2D Scatter 閱讀;Text Embedding 另有 BEAMS 分類視圖。
點選節點進入焦點模式:左側放大焦點影片,右側排列四個最相關影片。按 Esc 或點空白處離開。
使用滑鼠探索,或切換 Orbit 自動旋轉模式。
Space first — a space defines what “nearby” means; a layout only changes how that space is projected on screen.
Video Embedding — clips are positioned from CLIP video coordinates. Nearby clips look or feel semantically similar.
Pixel Information — clips are positioned from ten standardized still-frame ImagePlot features. Nearby clips share visual surface properties such as brightness, saturation, hue, edges, and contrast; this is not CLIP similarity.
Text Embedding — keywords live in a separate BGE text-embedding space. Matching-video thumbnails stay in the HUD.
Layouts — 3D preserves the selected 3D coordinates; 2D Grid packs items into readable cells using the selected space’s 2D projection; 2D Scatter shows that projection directly.
BEAMS — Belief, Everyday life, Arts, Memories, Subjective experience: five category grids available only in Text Embedding.
77 ImagePlot dimensions — the full feature sidecar is an analysis artifact. Pixel HUD shows only five readable fields: brightness median, saturation median, mean hue, edge density, and local contrast.
Similarity — white links use the active space’s similarity graph above the Related threshold. Pixel links are feature-space similarity, not semantic similarity.
Provenance — Source is AI or Human. Pipeline is the AI image→video chain when present.
Focus — tap a node to enlarge it on the left and show up to four related videos from top to bottom on the right. Connector length and white similarity numbers communicate closeness. Tap the center, empty space, or Escape to leave; tap a neighbor to transfer focus.