What a machine has to do to read a recipe video the way a cook does…

What a machine has to do to read a recipe video the way a cook does — screenshot every three seconds, read the captions, the comments and the voiceover, and tell instructions from song lyrics.

Dark slide headed “Teaching computers to watch TikToks like a human”, with capability bubbles around a recipe video
What this documents

The clearest spec in the project for multimodal video understanding: seven capabilities an ensemble has to run to produce a shoppable list.

Text in this image

Transcribed from the asset above and checked against the record. Figures are reproduced as printed, including where the original is itself a claim rather than an audited number.

Teaching computers to watch TikToks like a human.

Video-to-image: Watch video. Take 1 screenshot every 3 seconds. Understand what is in each image. Are we making beans or just eating them?

Captions: Read text on screen. Read captions.

Comments: Read some comments to understand relevancy and context.

Date posted: How relevant, how recent?

Voiceover: Listen to sound. Is it music or voice over or both? Understand the voice over instructions, not the song lyrics.

Provenance
Source
I built an AI that watches TikTok and shops Instacart →
File
/images/notion/8713aad7b627c628.png
Description
Dark slide headed “Teaching computers to watch TikToks like a human”, with capability bubbles around a recipe video