A cooking teaching video processing method, device and equipment and storage medium

By recognizing and matching cooking action information during the cooking process, the system automatically loops cooking demonstration video clips, solving the problem of users manually looping the videos and improving the user experience of cooking tutorial videos.

CN122160562APending Publication Date: 2026-06-05NINGBO FOTILE KITCHEN WARE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
NINGBO FOTILE KITCHEN WARE CO LTD
Filing Date
2026-01-09
Publication Date
2026-06-05

Smart Images

  • Figure CN122160562A_ABST
    Figure CN122160562A_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of kitchen appliances, and particularly relates to a cooking teaching video processing method and device, equipment and a storage medium, the method comprises the following steps: in response to a cooking instruction carrying cooking video information, determining cooking demonstration video data corresponding to the cooking video information, the cooking demonstration video data comprises a plurality of demonstration cooking action segments; in the cooking process of a user, obtaining cooking image data of the cooking process; performing action feature recognition on the cooking image data to obtain cooking action information; matching and analyzing the cooking action information with each demonstration cooking action segment of the cooking demonstration video data to obtain a target demonstration cooking action segment; and controlling a cooking device to play the target demonstration cooking action segment in a loop until the cooking action information is updated. Through the above method, the target demonstration cooking action segment can be positioned in real time according to the current cooking behavior of the user, and the user experience is significantly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of kitchen appliance technology, and in particular to a method, apparatus, device, and storage medium for processing cooking instructional videos. Background Technology

[0002] Currently, with the improvement of living standards, people's pursuit of delicious food is also increasing, and more people are willing to experience the joy of cooking at home.

[0003] When people cook at home, they often need to refer to recipe videos. Recipe videos are usually played all at once from beginning to end, and since the cooking process is relatively slow, users have to manually loop the cooking segments themselves, resulting in a poor user experience. Summary of the Invention

[0004] To address the aforementioned technical problems, this invention provides a method, apparatus, device, and storage medium for processing cooking instructional videos.

[0005] Firstly, a method for processing cooking instructional videos is provided, applied to cooking equipment, the method comprising: In response to a cooking instruction carrying cooking video information, cooking demonstration video data corresponding to the cooking video information is determined, wherein the cooking demonstration video data includes multiple demonstration cooking action segments; During the user's cooking process, acquire cooking image data of the cooking process; The cooking image data is subjected to motion feature recognition to obtain cooking motion information, which represents the current cooking behavior. The cooking action information is matched and analyzed with each demonstration cooking action segment of the cooking demonstration video data to obtain the target demonstration cooking action segment; The cooking device is controlled to play the target demonstration cooking action clip in a loop until the cooking action information is updated.

[0006] In a possible implementation, before determining the cooking demonstration video data corresponding to the cooking video information in response to a cooking instruction carrying cooking video information, the method further includes: Feature extraction is performed on the cooking video information to obtain multiple ingredient feature information, multiple time feature information, and multiple demonstration cooking action information; Based on the matching analysis of multiple ingredient feature information, multiple time feature information and multiple demonstration cooking action information, multiple demonstration cooking action segments are obtained. Each demonstration cooking action segment is used to indicate the ingredients and cooking actions used within a specific time period.

[0007] In a possible implementation, the step of matching and analyzing the cooking action information with each demonstration cooking action segment of the cooking demonstration video data to obtain the target demonstration cooking action segment includes: Obtain a sample cooking action segment that matches the cooking action information; If the demonstration cooking action segment appears only once in the cooking demonstration video data, the demonstration cooking action segment is determined to be the target demonstration cooking action segment.

[0008] In a possible implementation, the step of matching and analyzing the cooking action information with each demonstration cooking action segment of the cooking demonstration video data to obtain the target cooking action segment includes: If the demonstration cooking action segment appears more than once in the cooking demonstration video data, the preceding cooking action information of the cooking action information is obtained, and the preceding cooking action information represents the cooking action preceding the current cooking action. Obtain multiple preceding demonstration cooking action segments that match the cooking action information, wherein the preceding demonstration cooking action segment represents the previous demonstration cooking action segment that matches the current cooking behavior; The preceding cooking action information is matched and analyzed with multiple preceding demonstration cooking action segments to obtain the matching analysis results; If the matching analysis result indicates that there is only one preceding demonstration cooking action segment that matches the preceding cooking action information, the target preceding demonstration cooking action segment is determined, and the demonstration cooking action segment corresponding to the target preceding demonstration cooking action segment is determined as the target demonstration cooking action segment.

[0009] In a possible implementation, the method further includes: If the matching analysis result indicates that there are multiple preceding demonstration cooking action segments that match the preceding cooking action information, multiple target preceding demonstration cooking action segments are identified. Obtain predicted demonstration cooking action segments of the plurality of target preceding demonstration cooking action segments, wherein the predicted demonstration cooking action segments represent all demonstration cooking action segments preceding the target preceding demonstration cooking action segments; Obtain predicted cooking action information from the preceding cooking action information, wherein the predicted cooking action information represents all cooking behaviors prior to the preceding cooking action information; The predicted demonstration cooking action segment is matched and analyzed with the predicted cooking action information until the predicted demonstration cooking action segment matches the predicted cooking action information. The target preceding cooking action information corresponding to the predicted demonstration cooking action segment is determined, and the demonstration cooking action segment corresponding to the target preceding cooking action information is determined as the target demonstration cooking action segment.

[0010] In a possible implementation, the cooking device includes a projection device and a three-dimensional projection area. The projection device is used to project onto the three-dimensional projection area of ​​the cooking device. Before controlling the cooking device to loop the demonstration cooking action clip, the method further includes: The target demonstration cooking action segment is projected onto the three-dimensional projection area through the projection device, and the cooking device is controlled to play the target demonstration cooking action segment in a loop until the cooking action information is updated.

[0011] Secondly, a cooking instruction video processing device is provided, applied to cooking equipment, the device comprising: A cooking video determination module is used to determine cooking demonstration video data corresponding to the cooking video information in response to a cooking instruction carrying cooking video information. The cooking demonstration video data includes multiple demonstration cooking action segments. The acquisition module is used to acquire cooking image data of the cooking process during the user's cooking process; The recognition module is used to perform motion feature recognition on the cooking image data to obtain cooking motion information, which represents the current cooking behavior; The matching analysis module is used to match and analyze the cooking action information with each demonstration cooking action segment of the cooking demonstration video data to obtain the target demonstration cooking action segment. The control module is used to control the cooking equipment to play the target demonstration cooking action clip in a loop until the cooking action information is updated.

[0012] Thirdly, a cooking instruction video processing device is provided. The device includes a processor and a memory. The memory stores at least one instruction or at least one program. The at least one instruction or the at least one program is loaded and executed by the processor to implement any of the cooking instruction video processing methods described in the above embodiments.

[0013] Fourthly, a computer-readable storage medium is provided, wherein at least one instruction or at least one program is stored therein, the at least one instruction or the at least one program being loaded and executed by a processor in accordance with any of the cooking instruction video processing methods described in the above embodiments.

[0014] Fifthly, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the cooking video instruction video processing method as described in any of the above embodiments.

[0015] By adopting the above technical solution, the present invention has the following beneficial effects: This invention, in response to a cooking command carrying cooking video information, determines the corresponding cooking demonstration video data, which includes multiple demonstration cooking action segments. During the user's cooking process, cooking image data is acquired, and motion feature recognition is performed on the cooking image data to obtain cooking action information. This cooking action information represents the current cooking behavior. The cooking action information is matched and analyzed with each demonstration cooking action segment in the cooking demonstration video data to obtain a target demonstration cooking action segment. The cooking device is then controlled to loop the target demonstration cooking action segment until an update to the cooking action information is detected. By matching and analyzing the user's current cooking behavior with each demonstration cooking action segment to obtain the target demonstration cooking action segment, and looping the target demonstration cooking action segment until an update to the user's current cooking behavior is detected, this invention can locate the target demonstration cooking action segment that matches the user's current cooking behavior in real time and loop the target demonstration cooking action segment, significantly improving the user experience. Attached Figure Description

[0016] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention, and the same reference numerals usually represent the same parts. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0017] Figure 1 A flowchart illustrating a cooking instruction video processing method provided in this embodiment of the invention; Figure 2 This is a schematic diagram of the structure of a cooking instruction video processing device provided in an embodiment of the present invention.

[0018] Figure 3 This is a hardware structure block diagram of an electronic device for a cooking instructional video processing method provided in an embodiment of the present invention. Detailed Implementation

[0019] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.

[0020] The term "an embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the invention. In the description of the invention, it should be understood that the terms "upper," "lower," "top," "bottom," etc., indicating orientation or positional relationships based on the orientation or positional relationships shown in the accompanying drawings, are only for the convenience of describing the invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the invention. Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined with "first" and "second" may explicitly or implicitly include one or more of that feature. Moreover, the terms "first," "second," etc., are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of the invention described herein can be used in real time in orders other than those illustrated or described herein.

[0021] Reference Appendix Figure 1 This embodiment provides a method for processing cooking instructional videos, applied to cooking equipment, the method including: S1: In response to a cooking instruction carrying cooking video information, determine the cooking demonstration video data corresponding to the cooking video information. The cooking demonstration video data includes multiple demonstration cooking action segments. Specifically, cooking video information can refer to any recipe video, such as braised pork ribs with potatoes, scrambled eggs with tomatoes, fried rice with eggs, etc., and this application does not make a specific limitation in this regard. Cooking demonstration video data can refer to dividing the recipe video of braised pork ribs with potatoes into different short videos according to the operation sequence.

[0022] In some embodiments, the cooking video information may include, but is not limited to, 2D recipe videos, 3D recipe videos, etc.

[0023] In some embodiments, the cooking equipment may include, but is not limited to, a range hood.

[0024] In some embodiments, the cooking device includes a controller, a cloud platform, and a database. The database is used to collect cooking video information and send the determined cooking video information to the cloud platform. The cloud platform is used to acquire cooking action information and match and analyze the cooking action information with each demonstration cooking action segment of the cooking demonstration video data. The controller is able to respond to the cooking command carrying the cooking video information and control the operation of the database and the cloud platform.

[0025] In some embodiments, the controller, cloud platform, and database can be connected via wired or wireless means. In some embodiments, users can activate cooking instructions carrying cooking video information via the LED (Light Emitting Diode) screen on the cooking device, or they can control the cooking device by selecting cooking video information and activating cooking instructions carrying cooking video information via a mobile application.

[0026] In some embodiments, before determining the cooking demonstration video data corresponding to the cooking video information in response to a cooking instruction carrying cooking video information, the method further includes: S11: Extract features from cooking video information to obtain multiple ingredient feature information, multiple time feature information, and multiple demonstration cooking action information; S12: Based on multiple ingredient feature information, multiple time feature information and multiple demonstration cooking action information, a matching analysis is performed to obtain multiple demonstration cooking action segments. Each demonstration cooking action segment is used to indicate the ingredients and cooking actions used within a specific time period.

[0027] Specifically, the cloud platform receives the cooking video information and extracts features from the cooking video information to generate corresponding cooking demonstration video data.

[0028] In this embodiment, the ingredient characteristic information may include, but is not limited to: ingredient type, ingredient quantity, ingredient freshness, seasoning type, seasoning quantity, etc.; for example, the ingredient type may be rice, eggs, green onions, etc., the ingredient quantity may be one serving of rice, 1.2g of green onions or 1 egg, etc., and the seasoning type may include 0.2g of salt, 0.3g of pepper powder, 0.4g of chili powder, etc.

[0029] In this embodiment, the demonstrated cooking action information includes, but is not limited to, cooking behaviors such as pouring, stir-frying, adding, deep-frying, simmering, and pan-frying in the cooking video; In this embodiment, the time feature information is the duration of the cooking behavior in the cooking video.

[0030] For example, if a user wants to make fried rice, the ingredient feature information in the cooking video information includes oil, eggs, water, scallions, etc., the demonstration cooking action information includes cooking behaviors such as pouring and stirring, and the time feature information includes the duration of pouring for 3 seconds, the duration of stirring for 3 seconds, etc., thus obtaining demonstration cooking action segments including a video segment of pouring oil for 3 seconds, a video segment of stirring eggs for 3 seconds, etc.

[0031] By extracting features from cooking video information, demonstration cooking action segments are obtained, which include ingredients and cooking actions. By using this method to decompose the cooking video information into multiple demonstration cooking action segments, it is possible to more accurately match the user's current cooking behavior.

[0032] S2: During the user's cooking process, acquire cooking image data of the cooking process; In some embodiments, the cooking device includes a behavior recognition component for acquiring cooking image data of the cooking process.

[0033] In some embodiments, the behavior recognition component includes an image acquisition device for real-time monitoring and storage of cooking image data during the cooking process. For example, the image acquisition device may be a camera. Specifically, a camera is installed outside the cooking equipment or in another visible area to capture the user's cooking behavior during the cooking process.

[0034] S3: Perform motion feature recognition on cooking image data to obtain cooking motion information, which represents the current cooking behavior; In this embodiment, cooking action information includes, but is not limited to, cooking behaviors such as pouring, stir-frying, adding, deep-frying, simmering, and pan-frying in the cooking video; In some embodiments, the behavior recognition component includes a recognition sensor for identifying cooking behaviors in cooking image data during the cooking process, performing motion feature recognition to obtain cooking motion information, and sending the cooking motion information to a cloud platform. The recognition sensor may include, but is not limited to, infrared sensors, millimeter-wave radar, etc. Specifically, the image acquisition device sends cooking image data to the recognition sensor, the recognition sensor performs motion feature recognition to obtain the user's current cooking behavior, and sends the cooking behavior to the cloud platform.

[0035] Specifically, the behavior recognition component can be integrated into cooking equipment.

[0036] In some embodiments, the method further includes: performing ingredient feature recognition on cooking image data to obtain cooking ingredient feature information, wherein the cooking ingredient feature information represents the ingredients used in the current cooking behavior.

[0037] Specifically, the characteristics of cooking ingredients may include, but are not limited to: ingredient type, ingredient quantity, and ingredient freshness.

[0038] Specifically, the recognition sensor can also be used to identify the feature information of cooking ingredients in cooking image data during the cooking process. Specifically, the image acquisition device sends cooking image data to the recognition sensor, which identifies the features of the ingredients used in the user's current cooking activity and sends this information to the cloud platform.

[0039] In some embodiments, the behavior recognition components, cooking equipment, and cloud platform can be connected via wired or wireless means.

[0040] S4: Match and analyze the cooking action information with each demonstration cooking action segment of the cooking demonstration video data to obtain the target demonstration cooking action segment; In some embodiments, the cooking action information is matched and analyzed with various demonstration cooking action segments of the cooking demonstration video data, including: S401: Obtain a sample cooking action segment that matches the cooking action information; S402: If the demonstration cooking action segment appears only once in the cooking demonstration video data, the demonstration cooking action segment is determined as the target demonstration cooking action segment. When there is only one demonstration cooking action segment that matches the cooking action information, that demonstration cooking action segment is directly determined as the target demonstration cooking action segment, greatly simplifying the matching process.

[0041] Specifically, the cloud platform matches and analyzes the user's current cooking behavior with various demonstration cooking action segments to obtain the target demonstration cooking action segment.

[0042] In this embodiment, for example, if the obtained user's cooking action information is "pouring", the demonstration cooking action segments are arranged from front to back as "pouring", "stir-frying", "deep-frying" and "pan-frying". It can be seen that there is only one "pouring" action in the demonstration cooking action segment, that is, the demonstration cooking action segment is located as the target demonstration cooking action segment.

[0043] In some embodiments, the cooking action information is matched and analyzed with various demonstration cooking action segments of the cooking demonstration video data, including: S403: If the demonstration cooking action segment appears more than once in the cooking demonstration video data, obtain the previous cooking action information of the cooking action information. The previous cooking action information represents the cooking action before the current cooking action. S404: Obtain multiple preceding demonstration cooking action segments that match the cooking action information. The preceding demonstration cooking action segment represents the previous demonstration cooking action segment that matches the current cooking action. S405: Match and analyze the information of the preceding cooking actions with multiple preceding demonstration cooking action segments to obtain the matching analysis results; S406: If the matching analysis result indicates that there is only one preceding demonstration cooking action segment that matches the preceding cooking action information, determine the target preceding demonstration cooking action segment, and determine the demonstration cooking action segment corresponding to the target preceding demonstration cooking action segment as the target demonstration cooking action segment.

[0044] When multiple demonstration cooking action segments match the cooking action information, the preceding demonstration cooking action segment and the preceding cooking action information are searched and matched to obtain the target demonstration cooking action segment. This avoids the problem of not being able to accurately locate the target demonstration cooking action segment due to the existence of multiple demonstration cooking action segments matching the cooking action information.

[0045] In this embodiment, for example, if the obtained user's cooking action information is stir-frying 1 and pouring 2, and the demonstration cooking action segments are in the order of stir-frying 1, pouring 2, stir-frying 3, deep-frying 4, pouring 5, and pan-frying 6, it can be seen that there are two pouring actions in the demonstration cooking action segments. The preceding demonstration cooking action segments for the pouring action are stir-frying and deep-frying. At this time, the user's preceding cooking action information is stir-frying 1, so the target preceding demonstration cooking action segment is determined to be the stir-frying 1 action, that is, pouring 2 is determined to be the target demonstration cooking action segment.

[0046] In some embodiments, the method further includes: S407: If the matching analysis results indicate that there are multiple preceding demonstration cooking action segments that match the preceding cooking action information, identify multiple target preceding demonstration cooking action segments. S408: Obtain the predicted demonstration cooking action segments of multiple target preceding demonstration cooking action segments, the predicted demonstration cooking action segments representing all demonstration cooking action segments preceding the target preceding demonstration cooking action segments; S409: Obtain the predicted cooking action information based on the preceding cooking action information. The predicted cooking action information represents all cooking behaviors prior to the preceding cooking action information. S410: Perform matching analysis on the predicted demonstration cooking action segment and the predicted cooking action information until the predicted demonstration cooking action segment and the predicted cooking action information match, determine the target preceding cooking action information corresponding to the predicted demonstration cooking action segment, and determine the demonstration cooking action segment corresponding to the target preceding cooking action information as the target demonstration cooking action segment.

[0047] When multiple preceding demonstration cooking action segments match the preceding cooking action information, the target demonstration cooking action segment cannot be accurately located. Therefore, by matching and analyzing the predicted demonstration cooking action segments with the predicted cooking action information until the predicted demonstration cooking action segments match the predicted cooking action information, the target demonstration cooking action segment can be determined, thereby solving the above problem.

[0048] In this embodiment, for example, if the obtained user's cooking action information is in the order of simmering 1, frying 2, pouring 3, and the user's current cooking behavior is pouring 3, and the demonstration cooking action segments are in the order of simmering 1, frying 2, pouring 3, stir-frying 4, frying 5, pouring 6, and pan-frying 7, then there are two preceding demonstration cooking action segments that match the preceding cooking action information frying 2. The obtained user's predicted cooking action information is simmering 1, the first predicted demonstration cooking action segment is simmering 1, and the second predicted demonstration cooking action segment is simmering 1, frying 2, pouring 3, and stir-frying 4. Therefore, the first predicted demonstration cooking action segment matches the predicted cooking action information, that is, the target preceding cooking action information is frying 2, and the demonstration cooking action segment corresponding to the target preceding cooking action information frying 2 is pouring 3. This pouring 3 is the target demonstration cooking action segment.

[0049] In some embodiments, the cooking action information is matched and analyzed with various demonstration cooking action segments of the cooking demonstration video data, including: S411: If the demonstration cooking action segment appears more than once in the cooking demonstration video data, obtain the cooking ingredient feature information, which represents the ingredients used in the current cooking behavior. S412: Obtain multiple ingredient feature information from multiple demonstration cooking action segments that match the cooking action information, where each ingredient feature information represents the ingredient used in the demonstration cooking action segment; S413: Match and analyze the current ingredient feature information with multiple ingredient feature information to obtain the matching analysis results; S414: If the matching analysis results indicate that only one ingredient feature matches the current ingredient feature information, determine the target ingredient feature information and identify the corresponding demonstration cooking action segment as the target demonstration cooking action segment. By combining cooking ingredients and cooking behavior to match demonstration cooking action segments, the demonstration cooking action segment that the user wants to play can be located more accurately.

[0050] In this embodiment, for example, if the obtained user's cooking action information is 1. Stir-frying eggs and 2. Pour rice in. The demonstration cooking action segments are in the following order from front to back: 1. Stir-frying eggs and 2. Pour rice in. Fried eggs and 4. Pour scallions in. It can be seen that there are two pouring actions in the demonstration cooking action segments. The current cooking behavior is pouring. The current ingredient feature information is rice. Then the ingredient feature information in the first demonstration cooking action segment is rice. The ingredient feature information in the second demonstration cooking action segment is scallions. The target ingredient feature information is determined to be rice. The demonstration cooking action segment corresponding to the target ingredient feature information is determined to be the second demonstration cooking action segment.

[0051] In some embodiments, the cooking action information is matched and analyzed with various demonstration cooking action segments of the cooking demonstration video data, including: If the matching analysis results indicate that there are multiple ingredient feature information that match the current ingredient feature information, obtain the preceding cooking action information of the cooking action information. The preceding cooking action information represents the cooking action before the current cooking action. Acquire multiple preceding demonstration cooking action segments that match the cooking action information. The preceding demonstration cooking action segment represents the previous demonstration cooking action segment that matches the current cooking action. The information on the preceding cooking actions is matched and analyzed with multiple preceding demonstration cooking action segments to obtain the matching analysis results; If the matching analysis result indicates only one, determine the target preceding demonstration cooking action segment, and determine the demonstration cooking action segment corresponding to the target preceding demonstration cooking action segment as the target demonstration cooking action segment.

[0052] In this embodiment, if the obtained user's cooking action information is scrambled eggs 1 and pouring rice 2, and the demonstration cooking action segments are scrambled eggs 1, pouring rice 2, frying eggs 3, and pouring rice 4 in the order from front to back, it can be seen that there are two pouring actions in the demonstration cooking action segments. The current cooking behavior is pouring, and the current ingredient feature information is rice. There are multiple ingredient feature information that match the current ingredient feature information. The first demonstration cooking action segment is pouring rice 2, and the second demonstration cooking action segment is pouring rice 4. Therefore, it is determined that the cooking behavior preceding the current cooking behavior of pouring rice 2 is scrambled eggs 1. The preceding demonstration cooking action segment of the first demonstration cooking action segment is scrambled eggs 1, and the preceding demonstration cooking action segment of the second demonstration cooking action segment is frying eggs 3. It can be seen that the preceding demonstration cooking action segment that matches the preceding cooking action information scrambled eggs 1 is the preceding demonstration cooking action segment of the first demonstration cooking action segment, that is, the target demonstration cooking action segment is the first demonstration cooking action segment.

[0053] S5: Control the cooking equipment to play the target demonstration cooking action clips in a loop until the cooking action information is updated.

[0054] In some embodiments, controlling the cooking device to loop a segment of a target demonstration cooking action until an update to the cooking action information is detected includes: S51: Obtain updated cooking action information, which represents the current cooking behavior; S52: Match and analyze the updated cooking action information with each demonstration cooking action segment in the cooking demonstration video data to obtain the updated target demonstration cooking action segment. S53: Control the cooking equipment to play the updated target demonstration cooking action clip in a loop until the updated cooking action information is detected.

[0055] In this embodiment, when the cooking action information is updated—that is, when the user's current cooking behavior changes—the behavior recognition component uploads the changed cooking behavior to the cloud platform. The cloud platform then performs matching analysis on the behavior recognition component and the demonstration cooking action clip again to obtain a new target demonstration cooking action clip, and controls the cooking device to loop the new target demonstration cooking action clip. In this way, the cooking video can change according to the user's cooking behavior, greatly improving the user experience and eliminating the need for the user to manually loop the cooking video.

[0056] In some embodiments, the cooking device includes a projection device and a three-dimensional projection area. The projection device is used to project onto the three-dimensional projection area of ​​the cooking device. Before controlling the cooking device to loop a demonstration cooking action clip, the method further includes: The target demonstration cooking action clip is projected onto a three-dimensional projection area using a projection device. The cooking equipment is then controlled to loop the target demonstration cooking action clip until the cooking action information is updated. Since existing recipe videos are generally two-dimensional, lacking intuitiveness and a sense of immersive instruction, this invention sets the target demonstration cooking action clip as a three-dimensional video and projects it onto a three-dimensional projection area, providing users with a more immersive and engaging learning experience.

[0057] Specifically, the three-dimensional projection area can be located in any area that the user can see during the cooking process, or it can be set on the cooking equipment.

[0058] Specifically, the projection device can connect to the cloud platform wirelessly or via wired connection. The cloud platform sends a video clip of the target demonstration cooking action to the projection device, which then projects the video onto a three-dimensional projection area.

[0059] This application embodiment also provides a cooking instruction video processing device 100, such as... Figure 2As shown, Figure 2 This illustration shows a structural diagram of a cooking instruction video processing device 100 provided in an embodiment of this application. The device may include the following modules: The cooking video determination module 101 is used to determine the cooking demonstration video data corresponding to the cooking video information in response to a cooking instruction carrying cooking video information. The cooking demonstration video data includes multiple demonstration cooking action segments. The feature extraction module is used to extract features from cooking video information to obtain multiple ingredient feature information, multiple time feature information, and multiple demonstration cooking action information; The cooking demonstration video data generation module is used to perform matching analysis based on multiple ingredient feature information, multiple time feature information, and multiple demonstration cooking action information to obtain multiple demonstration cooking action segments. Each demonstration cooking action segment is used to indicate the ingredients and cooking actions used within a specific time period.

[0060] The acquisition module 102 is used to acquire cooking image data of the cooking process during the user's cooking process; The recognition module 103 is used to perform motion feature recognition on cooking image data to obtain cooking motion information, which represents the current cooking behavior. The ingredient feature recognition module is used to identify the ingredient features of cooking image data and obtain cooking ingredient feature information, which represents the ingredients used in the current cooking behavior.

[0061] The matching analysis module 104 is used to match and analyze the cooking action information with each demonstration cooking action segment of the cooking demonstration video data to obtain the target demonstration cooking action segment. The matching analysis module includes a first acquisition unit, a first matching analysis unit, a second acquisition unit, a third acquisition unit, a second matching analysis unit, a first determination unit, a second determination unit, a fourth acquisition unit, a fifth acquisition unit, a third matching analysis unit, a sixth acquisition unit, a seventh acquisition unit, a fourth matching analysis unit, a third determination unit, an eighth acquisition unit, and a ninth acquisition unit.

[0062] The first acquisition unit is used to acquire demonstration cooking action segments that match the cooking action information; The first matching analysis unit is used to determine the demonstration cooking action segment as the target demonstration cooking action segment if the demonstration cooking action segment appears only once in the cooking demonstration video data.

[0063] The second acquisition unit is used to acquire the preceding cooking action information if the demonstration cooking action segment appears more than once in the cooking demonstration video data. The preceding cooking action information represents the cooking action preceding the current cooking action. The third acquisition unit is used to acquire multiple preceding demonstration cooking action segments that match the cooking action information. The preceding demonstration cooking action segment represents the previous demonstration cooking action segment that matches the current cooking behavior. The second matching analysis unit is used to match and analyze the information of the preceding cooking actions with multiple preceding demonstration cooking action segments to obtain the matching analysis results. The first determining unit is used to determine the target pre-prepared demonstration cooking action segment if the matching analysis result indicates that there is only one pre-prepared demonstration cooking action segment that matches the pre-prepared cooking action information, and to determine the demonstration cooking action segment corresponding to the target pre-prepared demonstration cooking action segment as the target demonstration cooking action segment. The second determining unit is used to determine multiple target pre-prepared demonstration cooking action segments if the matching analysis results indicate that there are multiple pre-prepared demonstration cooking action segments that match the pre-prepared cooking action information. The fourth acquisition unit is used to acquire predicted demonstration cooking action segments of multiple target preceding demonstration cooking action segments. The predicted demonstration cooking action segments represent all demonstration cooking action segments before the target preceding demonstration cooking action segments. The fifth acquisition unit is used to acquire the predicted cooking action information based on the preceding cooking action information. The predicted cooking action information represents all cooking behaviors prior to the preceding cooking action information. The third matching analysis unit is used to perform matching analysis on the predicted demonstration cooking action segment and the predicted cooking action information until the predicted demonstration cooking action segment and the predicted cooking action information match, determine the target preceding cooking action information corresponding to the predicted demonstration cooking action segment, and determine the demonstration cooking action segment corresponding to the target preceding cooking action information as the target demonstration cooking action segment.

[0064] The sixth acquisition unit is used to acquire cooking ingredient feature information if the demonstration cooking action segment appears more than once in the cooking demonstration video data. The cooking ingredient feature information represents the ingredients used in the current cooking behavior. The seventh acquisition unit is used to acquire multiple ingredient feature information from multiple demonstration cooking action segments that match the cooking action information. Each ingredient feature information represents the ingredient used in the demonstration cooking action segment. The fourth matching analysis unit is used to match and analyze the current ingredient feature information with multiple ingredient feature information to obtain the matching analysis results; The third determining unit is used to determine the target ingredient feature information if the matching analysis result indicates that there is only one ingredient feature information that matches the current ingredient feature information, and to determine the demonstration cooking action segment corresponding to the target ingredient feature information as the target demonstration cooking action segment.

[0065] The eighth acquisition unit is used to acquire the preceding cooking action information if the matching analysis result indicates that there are multiple ingredients feature information that match the current ingredients feature information. The preceding cooking action information represents the previous cooking action of the current cooking action. The ninth acquisition unit is used to acquire multiple preceding demonstration cooking action segments that match the cooking action information. The preceding demonstration cooking action segment represents the previous demonstration cooking action segment that matches the current cooking behavior. The control module 105 is used to control the cooking equipment to play the target demonstration cooking action clips in a loop until the cooking action information is detected and updated.

[0066] The projection module is used to project a segment of the target demonstration cooking action onto a three-dimensional projection area using a projection device.

[0067] It should be noted that the above-described device embodiments and method embodiments are based on the same implementation methods.

[0068] This application provides a cooking instruction video processing device, which can be a terminal or a server, including a processor and a memory. The memory stores at least one instruction or at least one program, which is loaded and executed by the processor to implement the cooking instruction video processing method provided in the above method embodiments.

[0069] Memory can be used to store software programs and modules. The processor executes these stored software programs and modules to perform various functional applications and image object recognition. Memory can primarily include a program storage area and a data storage area. The program storage area stores the operating system, application programs required for the functions, etc.; the data storage area stores data created based on device usage, etc. Furthermore, memory can include high-speed random access memory (RAM) and non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, memory can also include a memory controller to provide the processor with access to the memory.

[0070] The methods and embodiments provided in this application can be executed in electronic devices such as mobile terminals, computer terminals, servers, or similar computing devices. Figure 3 This is a hardware structure block diagram of an electronic device for a cooking instructional video processing method provided in an embodiment of this application. For example... Figure 3As shown, the electronic device 500 can vary significantly due to differences in configuration or performance. It may include one or more central processing units (CPUs) 510 (CPUs 510 may include, but are not limited to, microprocessors such as MCUs or programmable logic devices such as FPGAs), a memory 530 for storing data, and one or more storage media 520 (e.g., one or more mass storage devices) for storing application programs 523 or data 522. The memory 530 and storage media 520 may be temporary or persistent storage. The program stored in the storage media 520 may include one or more modules, each module may include a series of instruction operations on the electronic device. Furthermore, the CPU 510 may be configured to communicate with the storage media 520 and execute the series of instruction operations in the storage media 520 on the electronic device 500. Electronic device 500 may also include one or more power supplies 560, one or more wired or wireless network interfaces 550, one or more input / output interfaces 540, and / or one or more operating systems 521, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, etc.

[0071] The input / output interface 540 can be used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by the communication provider of the electronic device 500. In one example, the input / output interface 540 includes a network interface controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the input / output interface 540 may be a radio frequency (RF) module used for wireless communication with the Internet.

[0072] Those skilled in the art will understand that Figure 3 The structure shown is for illustrative purposes only and does not limit the structure of the electronic device described above. For example, the electronic device 500 may also include... Figure 3 The more or fewer components shown, or having the same Figure 3 The different configurations shown.

[0073] Embodiments of this application also provide a computer-readable storage medium, which can be disposed in an electronic device to store at least one instruction or at least one program related to implementing a cooking instruction video processing method in the method embodiment. The at least one instruction or the at least one program is loaded and executed by the processor to implement the cooking instruction video processing method provided in the above method embodiment.

[0074] Optionally, in this embodiment, the storage medium may be located at at least one of the multiple network servers in a computer network. Optionally, in this embodiment, the storage medium may include, but is not limited to, various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.

[0075] According to one aspect of this application, a computer program product or computer program is provided, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the methods provided in the various alternative implementations described above.

[0076] It should be noted that the order of the above embodiments of the present invention is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. Furthermore, the above description focuses on specific embodiments, while other embodiments fall within the scope of the appended claims. In some cases, the actions or steps described in the claims can be performed in a different order than those shown in the embodiments and still achieve the desired results. Additionally, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0077] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the apparatus embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.

[0078] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.

[0079] The above description is merely a preferred embodiment of the present invention and should not be construed as limiting the scope of the invention. Therefore, any equivalent variations made in accordance with the claims of the present invention are still within the scope of the present invention.

Claims

1. A method for processing cooking instructional videos, applied to cooking equipment, characterized in that, The method includes: In response to a cooking instruction carrying cooking video information, cooking demonstration video data corresponding to the cooking video information is determined, wherein the cooking demonstration video data includes multiple demonstration cooking action segments; During the user's cooking process, acquire cooking image data of the cooking process; The cooking image data is subjected to motion feature recognition to obtain cooking motion information, which represents the current cooking behavior. The cooking action information is matched and analyzed with each demonstration cooking action segment of the cooking demonstration video data to obtain the target demonstration cooking action segment; The cooking device is controlled to play the target demonstration cooking action clip in a loop until the cooking action information is updated.

2. The cooking instruction video processing method according to claim 1, characterized in that, Before determining the cooking demonstration video data corresponding to the cooking video information in response to a cooking instruction carrying cooking video information, the method further includes: Feature extraction is performed on the cooking video information to obtain multiple ingredient feature information, multiple time feature information, and multiple demonstration cooking action information; Based on the matching analysis of multiple ingredient feature information, multiple time feature information and multiple demonstration cooking action information, multiple demonstration cooking action segments are obtained. Each demonstration cooking action segment is used to indicate the ingredients and cooking actions used within a specific time period.

3. The cooking instruction video processing method according to claim 1, characterized in that, The step of matching and analyzing the cooking action information with each demonstration cooking action segment of the cooking demonstration video data to obtain the target demonstration cooking action segment includes: Obtain a sample cooking action segment that matches the cooking action information; If the demonstration cooking action segment appears only once in the cooking demonstration video data, the demonstration cooking action segment is determined to be the target demonstration cooking action segment.

4. The cooking instruction video processing method according to claim 3, characterized in that, The step of matching and analyzing the cooking action information with each demonstration cooking action segment of the cooking demonstration video data to obtain the target cooking action segment includes: If the demonstration cooking action segment appears more than once in the cooking demonstration video data, the preceding cooking action information of the cooking action information is obtained, and the preceding cooking action information represents the cooking action preceding the current cooking action. Obtain multiple preceding demonstration cooking action segments that match the cooking action information, wherein the preceding demonstration cooking action segment represents the previous demonstration cooking action segment that matches the current cooking behavior; The preceding cooking action information is matched and analyzed with multiple preceding demonstration cooking action segments to obtain the matching analysis results; If the matching analysis result indicates that there is only one preceding demonstration cooking action segment that matches the preceding cooking action information, the target preceding demonstration cooking action segment is determined, and the demonstration cooking action segment corresponding to the target preceding demonstration cooking action segment is determined as the target demonstration cooking action segment.

5. The cooking instruction video processing method according to claim 4, characterized in that, The method further includes: If the matching analysis result indicates that there are multiple preceding demonstration cooking action segments that match the preceding cooking action information, multiple target preceding demonstration cooking action segments are identified. Obtain predicted demonstration cooking action segments of the plurality of target preceding demonstration cooking action segments, wherein the predicted demonstration cooking action segments represent all demonstration cooking action segments preceding the target preceding demonstration cooking action segments; Obtain predicted cooking action information from the preceding cooking action information, wherein the predicted cooking action information represents all cooking behaviors prior to the preceding cooking action information; The predicted demonstration cooking action segment is matched and analyzed with the predicted cooking action information until the predicted demonstration cooking action segment matches the predicted cooking action information. The target preceding cooking action information corresponding to the predicted demonstration cooking action segment is determined, and the demonstration cooking action segment corresponding to the target preceding cooking action information is determined as the target demonstration cooking action segment.

6. The cooking instruction video processing method according to any one of claims 1-5, characterized in that, The cooking equipment includes a projection device and a three-dimensional projection area. The projection device is used to project onto the three-dimensional projection area of ​​the cooking equipment. Before controlling the cooking equipment to loop the demonstration cooking action clip, the method further includes: The target demonstration cooking action segment is projected onto the three-dimensional projection area through the projection device, and the cooking device is controlled to play the target demonstration cooking action segment in a loop until the cooking action information is updated.

7. A cooking instruction video processing device, applied to cooking equipment, characterized in that, The device includes: A cooking video determination module is used to determine cooking demonstration video data corresponding to the cooking video information in response to a cooking instruction carrying cooking video information. The cooking demonstration video data includes multiple demonstration cooking action segments. The acquisition module is used to acquire cooking image data of the cooking process during the user's cooking process; The recognition module is used to perform motion feature recognition on the cooking image data to obtain cooking motion information, which represents the current cooking behavior; The matching analysis module is used to match and analyze the cooking action information with each demonstration cooking action segment of the cooking demonstration video data to obtain the target demonstration cooking action segment. The control module is used to control the cooking equipment to play the target demonstration cooking action clip in a loop until the cooking action information is updated.

8. A cooking instruction video processing device, characterized in that, The device includes a processor and a memory, the memory storing at least one instruction or at least one program, the at least one instruction or the at least one program being loaded and executed by the processor to implement the cooking instruction video processing method as described in any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, The storage medium stores at least one instruction or at least one program, which is loaded and executed by a processor to implement the cooking instruction video processing method as described in any one of claims 1 to 6.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the cooking instructional video processing method as described in any one of claims 1 to 6.