What a machine has to do to read a recipe video the way a cook does…
What a machine has to do to read a recipe video the way a cook does — screenshot every three seconds, read the captions, the comments and the voiceover, and tell instructions from song lyrics.
The clearest spec in the project for multimodal video understanding: seven capabilities an ensemble has to run to produce a shoppable list.
Transcribed from the asset above and checked against the record. Figures are reproduced as printed, including where the original is itself a claim rather than an audited number.
Teaching computers to watch TikToks like a human.
Video-to-image: Watch video. Take 1 screenshot every 3 seconds. Understand what is in each image. Are we making beans or just eating them?
Captions: Read text on screen. Read captions.
Comments: Read some comments to understand relevancy and context.
Date posted: How relevant, how recent?
Voiceover: Listen to sound. Is it music or voice over or both? Understand the voice over instructions, not the song lyrics.
- Source
- I built an AI that watches TikTok and shops Instacart →
- File
- /images/notion/8713aad7b627c628.png
- Description
- Dark slide headed “Teaching computers to watch TikToks like a human”, with capability bubbles around a recipe video