Unknown
Content creator highlighting spatio-temporal grounding feature for tracking subjects
How media typically covers Rachel Kim
Directly quoted in these articles
ByteDance's Vidi2 multimodal video model outperforms GPT-5 and Gemini 3 Pro on video understanding benchmarks, enabling temporal retrieval, spatio-temporal grounding, and intelligent video editing for videos up to 30 minutes long.
“Content creator highlighting spatio-temporal grounding feature for tracking subjects”