Recipe Video Matching for Accurate Cooking Step Segmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods fail to effectively generate comprehensive recipe videos that accurately match cooking steps with corresponding video segments, leading to incomplete or confusing cooking instructions.

Innovation Solution

A method and device that utilize machine-learned models to correlate text feature information from cooking recipes with image feature information from cooking videos, enabling precise matching of recipe steps with video segments.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional methods are used to create recipe videos, then manual editing and selection of video segments is required, but this leads to incomplete or confusing cooking instructions and fails to accurately match cooking steps with corresponding video segments

Engineering Contradiction:
Improvematching accuracy between recipe steps and video segmentsVSAvoidsystem complexity for video processing
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent introduces a machine learned matching model as an intermediary between text feature information and image feature information. This model automatically correlates recipe step descriptions with corresponding video segments, eliminating manual editing while achieving accurate matching through learned patterns from training data

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent replaces manual mechanical editing processes with automated machine learning-based matching. Instead of manually selecting and editing video segments to match recipe steps, the system uses trained models to automatically correlate text and image features, substituting human effort with computational intelligence

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Loss of information

If comprehensive cooking videos are used to ensure all cooking steps are covered, then the video includes unnecessary content such as non-recipe actions, but this reduces user understanding and creates confusion

Engineering Contradiction:
Improvecompleteness of cooking instructionsVSAvoidconfusion from unnecessary content
Core Design Contradiction:
Loss of informationVSObject-generated harmful factors

Solution Approach 1:

The patent extracts only the relevant portions of cooking videos that correspond to actual recipe steps. By using the machine learned matching model to identify and extract video segments that correlate with recipe step text features, the system removes unnecessary content such as non-recipe actions while preserving all essential cooking instructions

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent uses partial action by selecting only the specific video segments needed for each recipe step rather than using entire cooking videos. This partial selection approach ensures completeness of instructions while eliminating excessive unnecessary content through precise matching

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentEP4730823A1Method and apparatus for providing recipe image
Publication Date: 2026.04.22 SAMSUNG ELECTRONICS CO LTD
  • EP4730823A1 patent drawingFigure 1
  • EP4730823A1 patent drawingFigure 2
  • EP4730823A1 patent drawingFigure 3

AI summary

A device and method for providing a recipe video are provided, the method including obtaining text feature information corresponding to each step of a cooking recipe, the cooking recipe comprising a plurality of steps corresponding to a plurality of coking actions; obtaining image feature information corresponding to each frame of a plurality of frames constituting a cooking video; matching each step of the cooking recipe with a respective section of the cooking video based on a correlation between the obtained text feature information and the obtained image feature information obtained using a machine learned matching model, wherein the respective section of the cooking video corresponds to a respective step among the plurality of steps; and generating the recipe video by using the respective section of the cooking video matched with each step of the cooking recipe.