Recipe Video Matching for Accurate Cooking Step Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods fail to effectively generate comprehensive recipe videos that accurately match cooking steps with corresponding video segments, leading to incomplete or confusing cooking instructions.
Innovation Solution
A method and device that utilize machine-learned models to correlate text feature information from cooking recipes with image feature information from cooking videos, enabling precise matching of recipe steps with video segments.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional methods are used to create recipe videos, then manual editing and selection of video segments is required, but this leads to incomplete or confusing cooking instructions and fails to accurately match cooking steps with corresponding video segments
Solution Approach 1:
The patent introduces a machine learned matching model as an intermediary between text feature information and image feature information. This model automatically correlates recipe step descriptions with corresponding video segments, eliminating manual editing while achieving accurate matching through learned patterns from training data
Solution Approach 2:
The patent replaces manual mechanical editing processes with automated machine learning-based matching. Instead of manually selecting and editing video segments to match recipe steps, the system uses trained models to automatically correlate text and image features, substituting human effort with computational intelligence
2Loss of information
If comprehensive cooking videos are used to ensure all cooking steps are covered, then the video includes unnecessary content such as non-recipe actions, but this reduces user understanding and creates confusion
Solution Approach 1:
The patent extracts only the relevant portions of cooking videos that correspond to actual recipe steps. By using the machine learned matching model to identify and extract video segments that correlate with recipe step text features, the system removes unnecessary content such as non-recipe actions while preserving all essential cooking instructions
Solution Approach 2:
The patent uses partial action by selecting only the specific video segments needed for each recipe step rather than using entire cooking videos. This partial selection approach ensures completeness of instructions while eliminating excessive unnecessary content through precise matching
Data Source
Figure 1
Figure 2
Figure 3
AI summary
A device and method for providing a recipe video are provided, the method including obtaining text feature information corresponding to each step of a cooking recipe, the cooking recipe comprising a plurality of steps corresponding to a plurality of coking actions; obtaining image feature information corresponding to each frame of a plurality of frames constituting a cooking video; matching each step of the cooking recipe with a respective section of the cooking video based on a correlation between the obtained text feature information and the obtained image feature information obtained using a machine learned matching model, wherein the respective section of the cooking video corresponds to a respective step among the plurality of steps; and generating the recipe video by using the respective section of the cooking video matched with each step of the cooking recipe.