Cooking video generation method and cooking video generation system

By identifying abnormal segments in the cooking process and extracting replacement segments from the historical database for fusion processing, the problem of low video quality caused by cooking failures or recording interruptions has been solved, thus improving the integrity and watchability of the video.

CN121486519APending Publication Date: 2026-02-06HANGZHOU ROBAM APPLIANCES CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511653309.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-12
Publication Date
2026-02-06

AI Technical Summary

Technical Problem

Existing technologies often result in low-quality video footage due to user errors or network interruptions during cooking, leading to cooking failures or recording interruptions. They also cannot effectively handle abnormal segments, resulting in low utilization of video footage.

Method used

By acquiring real-time video data and status parameters during the cooking process, using preset video analysis methods to identify abnormal segments, and extracting alternative segments from historical cooking databases for fusion processing, a complete cooking video is generated.

Benefits of technology

It has achieved integrity restoration and optimization of cooking videos, improved the quality of video materials and user experience, and enhanced the coherence and watchability of the videos.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121486519A_ABST
    Figure CN121486519A_ABST
Patent Text Reader

Abstract

The invention provides a cooking video generation method and a cooking video generation system. The cooking video generation method comprises the following steps: acquiring real-time video data in a cooking process; identifying abnormal segments in the real-time video data based on a preset video analysis method; based on the abnormal fragment, extracting at least one substitute fragment corresponding to the abnormal fragment from a preset historical cooking database; the preset historical cooking database comprises video data of successful cooking; and carrying out fusion processing on the replacement fragment and the real-time video data to generate a target cooking video. In the mode, through intelligent identification and replacement of the abnormal segments in the cooking video, low-quality materials generated due to cooking failure or recording interruption can be repaired and optimized into a high-quality video with visual coherence and complete content, so that the integrity of the cooking video is ensured, and the user experience is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of smart home appliance technology, and in particular to a cooking video generation method and a cooking video generation system. Background Technology

[0002] With the development of smart homes, smart kitchen appliances with built-in cameras are becoming increasingly popular, allowing users to easily record the cooking process and upload it to a cloud platform for sharing. This feature meets users' needs for recording and sharing their culinary creations.

[0003] Existing technologies typically employ a simple direct recording and uploading model: capturing the entire cooking video using the device's built-in camera and then directly uploading it to the cloud for user playback or sharing. However, in actual cooking, user errors (such as accidentally opening the oven door) or network interruptions often lead to cooking failures or recording interruptions. These unexpected situations result in failed videos containing burnt or undercooked content, or multiple incomplete video segments. Existing systems lack the ability to effectively process these abnormal segments, failing to intelligently identify and compensate for them. This results in low-quality video footage that is difficult to meet users' sharing needs and has low material utilization. Summary of the Invention

[0004] In view of this, the purpose of this application is to provide a cooking video generation method and a cooking video generation system, which can repair and optimize low-quality and incomplete video materials caused by cooking failure or recording interruption, thereby ensuring the integrity of cooking videos and improving user experience.

[0005] In a first aspect, the present invention provides a method for generating cooking videos, comprising: Acquire real-time video data during the cooking process.

[0006] Based on a pre-defined video analysis method, abnormal segments in real-time video data are identified.

[0007] Based on the abnormal segment, at least one alternative segment corresponding to the abnormal segment is extracted from a preset historical cooking database; the preset historical cooking database includes video data of successful cooking.

[0008] The replacement footage is fused with real-time video data to generate the target cooking video.

[0009] In an optional implementation, prior to the step of acquiring real-time video data during the cooking process, the method further includes: The system captures cooking videos from inside the cooking equipment and adds timestamp information to each video.

[0010] The cooking status parameters of the cooking equipment are collected synchronously; wherein, the cooking status parameters include at least one of sensor parameters and sound signals.

[0011] The cooking video is associated with and saved with the cooking status parameters to generate real-time video data.

[0012] In an optional implementation, before the step of identifying abnormal segments in real-time video data based on a preset video analysis method, the method further includes: Get the current cooking method.

[0013] The sampling frequency is determined based on the pre-defined correspondence between cooking methods and sampling frequencies.

[0014] Based on the acquisition frequency, at least one keyframe is extracted from the real-time video data.

[0015] In an optional implementation, when the preset video analysis method is a video feature analysis method, the step of identifying abnormal segments in real-time video data includes: Extract the food feature data of the target food in each keyframe; the food feature data includes color features and texture features.

[0016] When the proportion of pixels belonging to the preset abnormal color gamut in the color features exceeds the preset color range threshold, or when the texture features meet the preset abnormal state, the video time segment corresponding to the key frame is determined to be an abnormal segment.

[0017] In an optional implementation, when the preset video analysis method is a comparative analysis method, the step of identifying abnormal segments in real-time video data includes: Retrieve recipe identifiers from real-time video data.

[0018] Based on recipe identifiers, corresponding standard video data is extracted from a pre-set historical cooking database.

[0019] Calculate the deviation between each keyframe in the real-time video data and the standard keyframe in the standard video data.

[0020] Determine whether each deviation value is greater than a preset deviation threshold.

[0021] If the deviation value is greater than the preset deviation threshold, the video time segment corresponding to the key frame of the deviation value is determined to be an abnormal segment.

[0022] In an optional implementation, when the preset video analysis method is a multimodal data analysis method, the step of identifying abnormal segments in real-time video data includes: Extract the audio signal corresponding to the real-time video data; the audio signal includes cooking sounds and user feedback sounds.

[0023] Identify whether the audio signal contains preset abnormal information.

[0024] When a preset abnormal information is detected in the audio signal, the corresponding video time segment is determined to be an abnormal segment.

[0025] In an optional implementation, before the step of identifying abnormal segments in real-time video data based on a preset video analysis method, the method further includes: Acquire sensor parameters corresponding to real-time video data and identify user motion trajectories in the real-time video data.

[0026] When sensor parameters exceed preset sensor parameter thresholds, or when user movement trajectories match preset abnormal movement trajectories, it is determined that there is an abnormal event in the real-time video data, and the video time period corresponding to the abnormal event is recorded in order to identify abnormal segments within the video time period.

[0027] In an optional implementation, based on the anomalous segment, at least one alternative segment corresponding to the anomalous segment is extracted from a preset historical cooking database, including: Obtain the metadata corresponding to the abnormal fragment; the metadata includes at least one of the recipe identifier, ingredient type, and cooking method.

[0028] Retrieve target cooking video data that matches the abnormal segment from the preset historical cooking database; the metadata of the target cooking video data is the same as the metadata of the abnormal segment.

[0029] Extract historical cooking segments from the target cooking video data that are at the same cooking stage as the abnormal segments, and use them as replacement segments.

[0030] In an optional implementation, the step of fusing the alternative segments with real-time video data to generate the target cooking video includes: Align the boundary regions of the replacement fragments with the boundary regions of the abnormal fragments in time.

[0031] The optical flow frame interpolation algorithm generates transition frames between boundary regions and adjacent frames, and then merges the transition frames with replacement segments to generate the video frames to be inserted.

[0032] The video frames to be inserted are synthesized into the real-time video data based on a preset pixel fusion algorithm to obtain the target cooking video.

[0033] In a second aspect, the present invention provides a cooking video generation system for performing the cooking video generation method of any of the foregoing embodiments; the system includes: a cooking device and a cloud platform connected in communication; the cooking device is equipped with a camera and at least one sensor.

[0034] This application provides a cooking video generation method and system. By acquiring real-time video data during the cooking process and analyzing it in conjunction with cooking state parameters and audio signals, abnormal segments in the cooking process can be accurately identified. Based on a historical cooking database, successful segments corresponding to the abnormal segments are extracted and replaced, avoiding video defects caused by burnt ingredients, undercooked food, or operational errors, thereby improving the coherence and watchability of the target video. Through time alignment, optical flow frame interpolation, and pixel-level synthesis in the fusion processing, seamless connection of video segments can be achieved, enhancing the natural transition effect of the video and thus improving the user viewing experience and the sharing value of the video.

[0035] Other features and advantages of this application will be set forth in the following description and will be apparent in part from the description or may be learned by practicing the application.

[0036] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description

[0037] To more clearly illustrate the technical solutions in the specific embodiments of this application or the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0038] Figure 1 This is a schematic diagram of a cooking video generation system provided in an embodiment of this application; Figure 2 A schematic diagram of the cooking equipment provided in the embodiments of this application; Figure 3 This is a flowchart of a cooking video generation method provided in an embodiment of this application.

[0039] Icons: 1-Cooking equipment; 2-Cloud platform; 11-Camera device; 12-Sensor. Detailed Implementation

[0040] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0041] To facilitate understanding of this embodiment, the embodiments of this application will be described in detail below.

[0042] This application provides a cooking video generation system, referring to... Figure 1 The system includes: a cooking device 1 and a cloud platform 2 connected by communication; the cooking device 1 is equipped with a camera device 11 and at least one sensor 12. The cooking video generation system is used to execute the cooking video generation method.

[0043] Here, cooking device 1 can be a smart oven, smart integrated stove, or other kitchen appliance with network communication capabilities. (Refer to...) Figure 2 The cooking equipment 1 is equipped with hardware for data acquisition and transmission, mainly including: Camera device 11: For example, a wide-angle camera installed inside the cavity of the cooking appliance 1, used to capture video data of the cooking process in real time.

[0044] At least one sensor 12 is used to synchronously acquire status parameters during the cooking process. The sensor 12 may include a temperature sensor 12 for monitoring the cavity temperature, a humidity sensor 12 for monitoring the amount of steam, and a microphone for capturing cooking sounds, etc.

[0045] Communication module: such as Wi-Fi (wireless local area network) module, which is responsible for packaging the real-time video data and cooking status parameters collected by camera device 11 and sensor 12, and sending them to cloud platform 2 in real time or in segments.

[0046] Cloud platform 2 can be a single server or a server cluster, with multiple functional modules deployed on its hardware to implement the methods of this application. These functional modules can be software programs, software code segments, or logical units, mainly including: Data receiving module: used to receive real-time video data and cooking status parameters from cooking device 1.

[0047] Analysis and Recognition Module: This module performs video analysis to identify anomalous segments. It can integrate one or more pre-trained AI models, such as convolutional neural networks for visual feature analysis or the YOLOv5 model (a deep learning model for object detection), to identify anomalous segments like burnt or immature segments by analyzing the color and texture of keyframes in the video or comparing them to a database. The module may also include an audio analysis unit to analyze abnormal sounds in the audio signal.

[0048] Historical cooking database: used to store a large amount of successful cooking video data with metadata such as recipe representation, ingredient type, etc.

[0049] Segment Extraction and Fusion Module: After the analysis and identification module identifies abnormal segments, the segment extraction and fusion module is responsible for retrieving and extracting successful video segments at the same cooking stage from the historical cooking database based on the metadata of the abnormal segments as replacement segments. Then, it uses techniques such as motion interpolation and pixel-level fusion to seamlessly synthesize the replacement segments with the original video data.

[0050] The video optimization and storage module is responsible for post-processing the target cooking video generated after fusion processing. This module may include an adaptive keyframe acquisition unit, such as one based on a Long Short-Term Memory (LSTM) network model, which dynamically adjusts the keyframe acquisition frequency according to the cooking method. The module may also include a compression encoding unit, which compresses the video using the H.265 (High Efficiency Video Coding) standard, adds metadata such as keyframe positions, and stores it in a cloud database for users to access and share at any time.

[0051] In this embodiment, after the user starts the cooking program, the camera and sensors inside the cooking equipment begin to work, and send the collected real-time video data and cooking status parameters to the cloud platform via the communication module. After receiving the data, the cloud platform's data receiving module hands it over to the analysis and recognition module for real-time analysis of the video content to determine if any abnormal segments exist. If an abnormal segment is identified, the segment extraction and fusion module retrieves suitable successful segments from the historical cooking database for replacement and fusion. Finally, the video optimization and storage module generates a complete target cooking video.

[0052] Based on the above embodiments, this application provides a method for generating cooking videos, referring to... Figure 3 The general flow of the cooking video generation method provided in this application embodiment is as follows: Step S101: Acquire real-time video data during the cooking process.

[0053] In a preferred embodiment, a camera device installed inside the smart cooking device captures video streams in real time, while sensors on the device synchronously collect key parameters (such as temperature and sound signals) during the cooking process. These video and parameter data are timestamped to form real-time video data, which is then uploaded to a cloud platform for processing via a network module. To improve subsequent analysis efficiency, a capture frequency can be dynamically adjusted on the cloud or device based on the characteristics of the current cooking method (such as stir-frying or stewing), and keyframes can be extracted from the real-time video stream at this frequency for subsequent recognition steps.

[0054] In other feasible embodiments, the source of real-time video data is not limited to built-in camera devices, but can also be a user-held mobile device (such as a mobile phone), a stand-alone camera mounted in the kitchen, or a pre-recorded cooking video uploaded by the user. Cooking status parameters are not limited to temperature and sound, but can also include stove power level data, humidity data within the cavity, weight change data from the smart scale, or user voice commands. Data processing is not limited to extracting keyframes; where computing resources permit, the entire video stream can also be analyzed.

[0055] Step S102: Based on a preset video analysis method, identify abnormal segments in the real-time video data.

[0056] Here, "abnormal fragments" refers to fragments that are unsuitable for sharing due to cooking failures or interruptions.

[0057] The preset video analysis method may include at least one of the following.

[0058] 1. Visual feature analysis methods are used to directly analyze video image content. For example, deep learning models such as convolutional neural networks are used to analyze keyframes in videos. Abnormal segments can be identified by recognizing abnormal colors such as charred black spots from burning, the original color of uncooked meat, or abnormal textures such as hardness or cracking on the surface of food.

[0059] 2. Multimodal data analysis methods are used to combine data other than video for comprehensive judgment. For example, by analyzing synchronously acquired audio signals, abnormal sounds such as the rapid cracking sound when water boils away can be identified; or, natural language processing technology can be used to analyze user reviews of past videos in sharing communities to extract keywords related to failure, such as "burnt" and "half-cooked," to assist in judgment.

[0060] 3. Comparative analysis method: Identifying anomalies by comparing with successful cases. For example, keyframes of the current cooking video are matched with keyframes of a large number of successful recipes in a historical cooking database. When the calculated deviation exceeds a preset threshold, it is judged as a failure, thus identifying the abnormal segment.

[0061] 4. Process parameter monitoring methods: Risk warnings are issued by analyzing sensor data or user behavior. For example, when parameters such as temperature or duration deviate significantly from the normal range of the standard recipe, or when users are found to be frequently opening and closing doors or suddenly adjusting the heat, the corresponding time period can be marked as a high-risk area for focused analysis to identify abnormal segments.

[0062] In other feasible embodiments, the identification method may also include using more advanced machine learning models for anomaly detection. Alternatively, instead of comparing with successful cases, it may compare with a database of failed cases, and if the similarity is too high, it is determined to be an anomaly. It may also be based on real-time user input; for example, if a user presses a "mark as failed" button when they discover a cooking failure, the system marks a video segment before and after that point in time as an anomaly.

[0063] Step S103: Based on the abnormal segment, extract at least one alternative segment corresponding to the abnormal segment from the preset historical cooking database; the preset historical cooking database includes video data of successful cooking.

[0064] In a preferred embodiment, metadata corresponding to the anomalous segment is obtained, such as recipe representation, ingredient type, and cooking method. This metadata is then used to search a historical cooking database for historical videos marked as successfully cooked and possessing the same or similar metadata. From the retrieved successful videos, video segments at the same cooking stage as the current anomalous segment (e.g., from the browning stage to the doneness stage) are extracted as replacement segments.

[0065] In other feasible embodiments, the source and extraction method of the replacement fragment can be more diverse. For example, instead of relying on metadata, content-based image retrieval technology can be used to analyze the last frame before the occurrence of the anomalous fragment and search the database for the most visually similar successful fragment as a connector. Alternatively, when multiple matching successful fragments are retrieved, the best one can be selected as the replacement fragment based on its popularity, number of likes, or user ratings. Furthermore, the replacement fragment does not necessarily have to come from real user cooking videos; it can be a pre-produced professional instructional video clip, or even a computer-generated animation clip depicting an ideal cooking effect.

[0066] Step S104: The replacement segment is fused with real-time video data to generate the target cooking video.

[0067] In a preferred embodiment, the boundaries between the replacement segment and the anomalous segment in the original video are precisely time-aligned. Then, motion interpolation is used to intelligently generate several smooth transition frames between the last frame of the original video and the first frame of the replacement segment, and between the last frame of the replacement segment and the next frame of the original video. Finally, pixel-level fusion is used to make slight adjustments and overlays of the replacement segment and the generated transition frames at the pixel level, ensuring that their hue and brightness are consistent with the context of the original video, achieving seamless compositing.

[0068] In other feasible embodiments, the fusion processing technique can also be from other video processing fields. For example, more advanced gradient domain fusion or Poisson fusion algorithms can be used to achieve more natural edge transitions. Generative adversarial network-based video inpainting techniques can also be employed, allowing AI models to draw the most natural transition effects. Furthermore, in addition to pursuing seamless transitions, stylized transition effects such as fade-in / fade-out and smooth wipes can be provided according to user preferences to artistically replace segments.

[0069] In one embodiment, before step S101, the method further includes the following steps S201-S203.

[0070] Step S201: Collect cooking videos from inside the cooking equipment and add timestamp information to each cooking video.

[0071] Here, a camera device, such as a high-temperature resistant wide-angle camera, is pre-installed inside the cooking equipment's cavity. When the user starts the cooking program, the camera device begins working in real time, continuously capturing the dynamic changes of the food inside the cavity during the cooking process, forming a video stream. Each frame or segment of the video stream is precisely timestamped.

[0072] Step S202: Synchronously collect cooking status parameters of the cooking equipment; wherein, the cooking status parameters include at least one of sensor parameters and sound signals.

[0073] Here, while the camera captures video, at least one sensor installed inside the cooking equipment also starts working simultaneously to collect status parameters during the cooking process. These cooking status parameters may include, but are not limited to: Sensor parameters: For example, the real-time temperature of the cavity obtained by a temperature sensor, or the amount of steam obtained by a humidity sensor.

[0074] Sound signals: For example, sound signals generated during cooking can be captured via a built-in microphone. Sound signals can reflect the state of the food at a specific stage, such as the "sizzling" sound during normal cooking or the "rapid popping" sound when the moisture evaporates.

[0075] Step S203: Associate and save the cooking video with the cooking status parameters to generate real-time video data.

[0076] Here, timestamps are used as a unified time reference to associate and bind video frames or video clips captured at a specific moment with cooking status parameters such as temperature, steam volume, and sound signals captured at the same moment. This comprehensive data packet, containing time-synchronized video information and various status parameters, constitutes real-time video data. Real-time video data is then sent to the cloud platform via the cooking equipment's built-in communication module (such as a Wi-Fi module).

[0077] In one embodiment, after step S102, the method further includes the following steps S301-S303.

[0078] Step S301: Obtain the current cooking method.

[0079] Here, the current cooking method is obtained through multiple means. For example, when a user selects a preset cooking program (such as roasted chicken wings or stir-fried vegetables) on the interface of a smart cooking device, the cooking method information, such as roasting or stir-frying, can be directly obtained. In other embodiments, the corresponding cooking method can also be determined by analyzing metadata such as the recipe identifier or dish name entered by the user.

[0080] Step S302: Determine the sampling frequency based on the preset correspondence between cooking methods and sampling frequencies.

[0081] Here, the correspondence between cooking methods and acquisition frequencies can be preset or learned through machine learning models (such as long short-term memory network models). The correspondence between cooking methods and acquisition frequencies is used to map different cooking methods to the optimal acquisition frequency.

[0082] For cooking methods where the visual characteristics of food, such as shape and color, change slowly and gradually during the cooking process, such as roasting, stewing, and baking, a lower acquisition frequency (i.e., a longer time interval) will be determined, such as acquiring one frame every 5 or 10 seconds.

[0083] For cooking methods where visual features change drastically and rapidly during the cooking process, such as stir-frying, deep-frying, and grilling, the system will determine a higher acquisition frequency (i.e., a shorter time interval), such as acquiring one frame per second, to ensure that any instantaneous key changes can be captured.

[0084] Step S303: Based on the acquisition frequency, extract at least one keyframe from the real-time video data.

[0085] Here, the real-time video data stream is processed according to the acquisition frequency. For example, if the acquisition frequency is determined to be one frame every 5 seconds, then single-frame images will be extracted at the 5th, 10th, 15th, and so on time points in the video stream. These extracted single-frame images, containing temporal information, constitute the keyframe set for subsequent analysis. Compared to the original video data stream containing dozens of frames per second, the data volume is greatly reduced, but the core information of the cooking process at key time points is preserved.

[0086] In one embodiment, when the preset video analysis method is a video feature analysis method, step S102, the step of identifying abnormal segments in real-time video data, includes the following steps S401-S402.

[0087] Step S401: Extract the food feature data of the target food in each keyframe; the food feature data includes color features and texture features.

[0088] Here, one or more keyframes are extracted, and the target food area in the keyframe image is analyzed to extract the corresponding visual feature data. The visual feature data includes color features and texture features.

[0089] Step S402: When the proportion of pixels belonging to the preset abnormal color gamut in the color features exceeds the preset color range threshold, or when the texture features meet the preset abnormal state, the video time segment corresponding to the key frame is determined to be an abnormal segment.

[0090] Here, for color features, the system monitors the color gamut distribution of the food area. For example, the system can predefine an abnormal color gamut model, which includes dark brown or carbonized black spots in a charred state, as well as the original pink or light color in an uncooked state (such as meat). When the system detects that the proportion of pixels in a keyframe that conform to this abnormal color gamut model exceeds a preset threshold (such as 20%), it determines that the keyframe is abnormal.

[0091] For texture features, the system analyzes the texture patterns on the food surface. For example, the system can identify the hard and brittle texture that burnt areas typically exhibit, or the overly smooth or stiff surface state that uncooked food may retain. When these textures that match preset abnormal states are identified, the keyframe is determined to be abnormal.

[0092] Finally, when a keyframe is determined to be abnormal, the system identifies the time period corresponding to that keyframe in the original video (e.g., the interval of a few seconds before and after the keyframe) as an abnormal segment.

[0093] In one embodiment, when the preset video analysis method is a comparative analysis method, step S102, the step of identifying abnormal segments in real-time video data, includes the following steps S501-S505.

[0094] Step S501: Obtain the recipe identifier from the real-time video data.

[0095] Here, we retrieve the metadata for the current cooking task, such as the recipe identifier (ID) selected by the user.

[0096] Step S502: Based on the recipe identifier, extract the corresponding standard video data from the preset historical cooking database.

[0097] Here, based on the recipe identifier, one or more videos marked as successful cooking with the same recipe identifier are extracted from the historical cooking database as a standard reference.

[0098] Step S503: Calculate the deviation between each keyframe in the real-time video data and the standard keyframe in the standard video data.

[0099] Here, similarity matching is performed. Keyframes at a specific point in time are extracted from the current real-time video, and standard keyframes at the same cooking stage are found from a standard reference video. An image comparison algorithm is used to calculate the deviation value between these two keyframes. The deviation value can be a score that integrates multiple dimensions such as color, texture, and contour.

[0100] Step S504: Determine whether each deviation value is greater than a preset deviation threshold.

[0101] Here, the preset deviation threshold is set in advance according to the actual situation, and can be 15%. The calculated deviation value is compared with the preset deviation threshold.

[0102] Step S505: If the deviation value is greater than the preset deviation threshold, determine that the video time segment corresponding to the key frame of the deviation value is an abnormal segment.

[0103] Here, if the deviation value is greater than the preset deviation threshold, it means that the current cooking state is too different from the successful case. The system will then determine that the key frame is abnormal and identify the corresponding time period as the abnormal segment.

[0104] If the deviation value is less than or equal to the preset deviation threshold, the video is normal.

[0105] In one embodiment, when the preset video analysis method is a multimodal data analysis method, step S102, the step of identifying abnormal segments in real-time video data, includes the following steps S601-S603.

[0106] Step S601: Extract the audio signal corresponding to the real-time video data; the audio signal includes cooking sounds and user feedback sounds.

[0107] Here, audio signals are extracted in sync with real-time video data. Audio analysis algorithms are then used to identify characteristic sounds during the cooking process.

[0108] Step S602: Identify whether the audio signal contains preset abnormal information.

[0109] Here, the preset abnormal information is pre-set according to the actual situation. For example, the system can distinguish between a normal, uniform "sizzling" sound and a "rapid cracking sound" caused by water drying out or food exploding.

[0110] Step S603: When the audio signal is found to contain preset abnormal information, the video time segment corresponding to the audio signal is determined to be an abnormal segment.

[0111] Here, when the system detects the latter type of pre-defined abnormal sound pattern, it identifies the video segment corresponding to the time period in which the sound appears as an abnormal segment.

[0112] In one feasible embodiment, the system can also incorporate natural language processing techniques to mine user reviews or feedback text from historical videos. For example, when the system detects that multiple users frequently use negative keywords such as "burnt" or "half-cooked" in their comments on a particular historical video, the system can mark that historical video as a sample containing abnormal segments for use in optimizing future recognition models. In some real-time interactive scenarios, if a user inputs commands such as "failed" via voice or text, the system can also mark the current time period as an abnormal segment in real time.

[0113] This application embodiment can also monitor objective parameters and user behavior during the cooking process to pre-screen high-risk time periods that may pose a risk of cooking failure, thereby enabling more targeted analysis and improving overall identification efficiency and accuracy. In one embodiment, the method further includes the following steps S701-S702.

[0114] Step S701: Obtain the sensor parameters corresponding to the real-time video data and identify the user's motion trajectory in the real-time video data.

[0115] Here, the system extracts objective data generated by device sensors (such as temperature sensors and timers) from cooking status parameters that are synchronized with real-time video data. For example, the system acquires real-time temperature readings inside the cavity and the runtime since the start of the cooking program.

[0116] By performing preliminary image analysis on the video stream or by monitoring the device's operation logs, the system can identify the user's specific actions. For example, the system can track the frequency with which a chef flips ingredients, or monitor actions such as adjusting the stove's heat or opening and closing the oven door.

[0117] Step S702: When the sensor parameters exceed the preset sensor parameter threshold, or the user's action trajectory matches the preset abnormal action trajectory, it is determined that there is an abnormal event in the real-time video data, and the video time period corresponding to the abnormal event is recorded in order to identify abnormal segments within the video time period.

[0118] Here, the system determines whether the sensor parameters exceed preset safety or standard thresholds. For example, if the oil temperature exceeds 200°C, or the actual cooking time has not reached 80% of the time specified in the standard recipe, the system determines that there is an anomaly.

[0119] The system will determine whether a user's actions match preset abnormal behavior patterns. For example, in a cooking process that requires stable heating, if the system detects unusual actions such as frequent stirring or suddenly reducing the heat, it will determine that an anomaly has occurred.

[0120] Once an anomaly is confirmed, the system does not immediately designate that time period as the final anomaly segment. Instead, it marks and records it as a high-risk or pending investigation period. For example, the system can initiate visual feature analysis or comparative analysis only within this marked time period to further pinpoint and confirm the final anomaly segment within that high-risk period, thereby improving response speed and operational efficiency.

[0121] In one embodiment, step S103 includes the following steps S801-S803.

[0122] Step S801: Obtain the metadata corresponding to the abnormal fragment; the metadata includes at least one of the recipe identifier, ingredient type, and cooking method.

[0123] Here, the contextual information associated with the anomalous fragment is parsed, i.e., metadata. Metadata includes at least one of the following: a recipe identifier used to uniquely identify the recipe, ingredient type (such as meat, vegetables), and cooking method (such as frying, boiling, roasting).

[0124] Step S802: Retrieve target cooking video data that matches the abnormal segment from the preset historical cooking database; the metadata of the target cooking video data is the same as the metadata of the abnormal segment.

[0125] Here, the system uses metadata as search keywords to query a database containing a large number of historical cooking videos. All videos in the database are marked as successfully cooked and also have detailed metadata tags. The system searches for target cooking videos with metadata identical or similar to the current anomalous segment. For example, if the metadata of the anomalous segment is "Recipe ID: 101, Cooking Method: Bake", the system will search the database for all successful videos that match these two conditions.

[0126] Step S803: Extract historical cooking segments from the target cooking video data that are at the same cooking stage as the abnormal segments, and use them as replacement segments.

[0127] Here, a segment that perfectly corresponds to the abnormal segment in the cooking process is extracted from the target cooking video data. The system analyzes the stage of the abnormal segment in the entire cooking process (e.g., the browning stage of chicken wings or 5 minutes before taking them out of the oven). Then, the system locates the exact same cooking stage in the retrieved target cooking video data and extracts it as the final replacement segment.

[0128] In one embodiment, step S104 includes the following steps S901-S903.

[0129] Step S901: Time-align the boundary region of the replacement fragment with the boundary region of the abnormal fragment.

[0130] Here, the boundaries of the replacement clip are aligned with the boundaries of the aberrant clip in the original video on the timeline.

[0131] Step S902: Based on the optical flow frame interpolation algorithm, a transition frame is generated between the boundary region and the adjacent frame, and the transition frame is fused with the replacement segment to generate the video frame to be inserted.

[0132] Here, to avoid abrupt cuts between the last frame of the original video and the first frame of the replacement segment, the system employs motion interpolation techniques, such as optical flow interpolation, to analyze the motion vectors of objects in the image frames on both sides of the boundary. Based on this, it calculates and generates several new transition frames that can depict smooth motion trajectories. These newly generated transition frames, together with the replacement segment itself, constitute the video frame to be inserted.

[0133] Step S903: Based on a preset pixel fusion algorithm, the video frames to be inserted are synthesized into the real-time video data to obtain the target cooking video.

[0134] Here, after smoothing the motion trajectory, a pixel fusion algorithm, such as alpha blending, is used to process the video frames to be inserted. Alpha blending can fine-tune the pixel values ​​of the inserted video frames to ensure consistency with the original video context in terms of brightness, color saturation, etc., achieving seamless pixel-level compositing. In a preferred embodiment, post-processing techniques such as Gaussian filtering can be further employed to smooth the stitching edges, or transition animations can be added, making the final target cooking video visually more perfect.

[0135] The cooking video generation method provided in this application can intelligently identify and replace abnormal segments in cooking videos, thereby repairing and optimizing low-quality materials caused by cooking failures or recording interruptions into high-quality videos with visual coherence and complete content, thus improving the generation rate of successful videos and enhancing the user experience.

[0136] The computer program product provided in this application includes a computer-readable storage medium storing program code. The instructions included in the program code can be used to execute the methods in the preceding method embodiments. For specific implementation details, please refer to the method embodiments, which will not be repeated here.

[0137] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the system and apparatus described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0138] Furthermore, in the description of the embodiments of this application, unless otherwise expressly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in this application based on the specific circumstances.

[0139] If a function is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0140] In the description of this application, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are used only for the convenience of describing this application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on this application. Furthermore, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.

[0141] Finally, it should be noted that the above embodiments are merely specific implementations of this application, used to illustrate the technical solutions of this application, and not to limit them. The protection scope of this application is not limited thereto. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments within the scope of the technology disclosed in this application, or make equivalent substitutions for some of the technical features. Such modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be covered within the protection scope of this application.

Claims

1. A method for generating cooking videos, characterized in that, include: Acquire real-time video data during the cooking process; Based on a preset video analysis method, abnormal segments in the real-time video data are identified; Based on the abnormal segment, at least one alternative segment corresponding to the abnormal segment is extracted from a preset historical cooking database; the preset historical cooking database includes video data of successful cooking. The replacement segment is fused with the real-time video data to generate the target cooking video.

2. The cooking video generation method according to claim 1, characterized in that, Prior to the step of acquiring real-time video data during the cooking process, the method further includes: The cooking video inside the cooking equipment is captured, and a timestamp is added to each of the cooking videos; The cooking status parameters of the cooking equipment are collected synchronously; wherein, the cooking status parameters include at least one of sensor parameters and sound signals; The cooking video is associated with and saved along with the cooking status parameters to generate the real-time video data.

3. The cooking video generation method according to claim 2, characterized in that, Before the step of identifying abnormal segments in the real-time video data based on a preset video analysis method, the method further includes: Get the current cooking method; The sampling frequency is determined based on the pre-defined correspondence between cooking methods and sampling frequencies; Based on the acquisition frequency, at least one keyframe is extracted from the real-time video data.

4. The cooking video generation method according to claim 3, characterized in that, When the preset video analysis method is a video feature analysis method, the step of identifying abnormal segments in the real-time video data includes: Extract the food feature data of the target food in each keyframe; the food feature data includes color features and texture features; When the proportion of pixels belonging to a preset abnormal color gamut in the color feature exceeds a preset color range threshold, or when the texture feature meets a preset abnormal state, the video time segment corresponding to the key frame is determined to be the abnormal segment.

5. The cooking video generation method according to claim 3, characterized in that, When the preset video analysis method is a comparative analysis method, the step of identifying abnormal segments in the real-time video data includes: Obtain the recipe identifier from the real-time video data; Based on the recipe identifier, extract the corresponding standard video data from the preset historical cooking database; Calculate the deviation value between each keyframe in the real-time video data and the standard keyframe in the standard video data; Determine whether each deviation value is greater than a preset deviation threshold; If the deviation value is greater than the preset deviation threshold, the video time segment corresponding to the key frame corresponding to the deviation value is determined to be the abnormal segment.

6. The cooking video generation method according to claim 1, characterized in that, When the preset video analysis method is a multimodal data analysis method, the step of identifying abnormal segments in the real-time video data includes: Extract the audio signal corresponding to the real-time video data; the audio signal includes cooking sounds and user feedback sounds. Identify whether the audio signal contains preset abnormal information; When the preset abnormal information is detected in the audio signal, the video time segment corresponding to the audio signal is determined to be the abnormal segment.

7. The cooking video generation method according to claim 1, characterized in that, Before the step of identifying abnormal segments in the real-time video data based on a preset video analysis method, the method further includes: Obtain the sensor parameters corresponding to the real-time video data, and identify the user's motion trajectory in the real-time video data; When the sensor parameters exceed a preset sensor parameter threshold, or when the user's action trajectory matches a preset abnormal action trajectory, it is determined that there is an abnormal event in the real-time video data, and the video time period corresponding to the abnormal event is recorded, so as to identify the abnormal segment within the video time period.

8. The cooking video generation method according to claim 1, characterized in that, Based on the anomalous segment, at least one alternative segment corresponding to the anomalous segment is extracted from a preset historical cooking database, including: Obtain the metadata corresponding to the abnormal fragment; the metadata includes at least one of recipe identifier, ingredient type, and cooking method; Retrieve target cooking video data that matches the abnormal segment from a preset historical cooking database; the metadata of the target cooking video data is the same as the metadata of the abnormal segment. Historical cooking segments that are at the same cooking stage as the abnormal segment are extracted from the target cooking video data and used as replacement segments.

9. The cooking video generation method according to claim 1, characterized in that, The step of fusing the replacement segment with the real-time video data to generate the target cooking video includes: The boundary region of the replacement fragment is time-aligned with the boundary region of the abnormal fragment. Based on the optical flow frame interpolation algorithm, a transition frame is generated between the boundary region and the adjacent frame, and the transition frame is fused with the replacement segment to generate the video frame to be inserted; The video frames to be inserted are synthesized into the real-time video data based on a preset pixel fusion algorithm to obtain the target cooking video.

10. A cooking video generation system, characterized in that, The system is used to perform the cooking video generation method according to any one of claims 1-9; the system includes: a cooking device and a cloud platform connected in communication; the cooking device is equipped with a camera and at least one sensor.