Recipe recommendation method and device, intelligent device and storage medium
By recognizing and processing text and image recipe information to generate voice recipes, and identifying cooking ingredients and object locations in real time, the system recommends cooking steps, solving the problem that existing systems cannot effectively present visual information and improving the convenience and safety of cooking.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- NINGBO FOTILE KITCHEN WARE CO LTD
- Filing Date
- 2026-02-10
- Publication Date
- 2026-05-26
AI Technical Summary
Existing cooking systems cannot effectively present the visual information in recipes, leading to inconvenience in operation, especially for visually impaired people, which poses safety risks. Furthermore, existing voice recommendation systems cannot provide real-time, accurate, and context-aware voice guidance.
By performing image and text recognition processing on the text and image recipe information, voice recipe information is generated, and the location of target cooking materials and objects in the cooking area is identified in real time, and cooking steps are recommended based on the location relationship.
It improves the smoothness and safety of the cooking process, enhances the real-time nature and scene awareness of cooking instructions, avoids misoperation and safety risks, and improves the user experience.
Smart Images

Figure CN122087136A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of cooking equipment control technology, and in particular to a recipe recommendation method, device, smart device, and storage medium. Background Technology
[0002] With the continuous advancement of artificial intelligence technology, intelligent cooking assistance systems are gradually entering people's daily lives. Currently, most systems rely primarily on recipes in the form of pictures and text. Users need to read the text and images on the screen or in paper documents to obtain cooking instructions and then perform the operations accordingly. However, this interaction method has obvious limitations: when users' hands are covered in oil, water, or flour, making it inconvenient to touch the device, frequently checking the recipe can seriously affect the smoothness of operation and even cause cooking to be interrupted.
[0003] Existing audio recipe recommendation systems are mostly based on audio tapes and audio e-books. While these alleviate the problems of manual operation to some extent, they still have shortcomings such as unintuitive information delivery and difficulty in quickly locating key steps. Furthermore, these solutions cannot effectively present the visual information in the recipe—such as the initial state of ingredients, their changes in shape during processing, and the specific location of cooking materials—limiting ease of use. Especially for visually impaired individuals, existing systems cannot provide real-time, accurate, and context-aware voice guidance, making it difficult for them to complete the cooking process independently and exposing them to higher safety risks. For example, when using knives or operating open flames or high-temperature stoves, the lack of spatial positioning and status feedback greatly increases the risk of accidental contact, burns, or cuts, seriously affecting the user experience and personal safety. Summary of the Invention
[0004] This application provides a recipe recommendation method, apparatus, smart device, and storage medium to at least address the problem of how to improve user convenience in cooking in related technologies. The technical solution of this application is as follows: According to a first aspect of the embodiments of this application, a recipe recommendation method is provided, including: In response to a target voice command, the system selects the target voice recipe information corresponding to the target voice command from multiple pre-stored voice recipe information; the multiple pre-stored voice recipe information is obtained by image and text recognition processing of multiple image and text recipe information. Based on the target voice recipe information, the location information of the target cooking materials and the target object in the cooking area is identified and processed to obtain the location relationship information between the target cooking materials and the target object; Based on location relationship information and target voice recipe information, cooking steps corresponding to target cooking ingredients are recommended to the target audience.
[0005] According to a second aspect of the embodiments of this application, a recipe recommendation device is provided, comprising: The filtering module is used to filter out the target voice recipe information corresponding to the target voice command from a plurality of pre-stored voice recipe information in response to the target voice command; the plurality of pre-stored voice recipe information is obtained by performing image and text recognition processing on a plurality of image and text recipe information; The positional relationship determination module is used to identify and process the positional information of the target cooking materials and the target object in the cooking area according to the target voice recipe information, so as to obtain the positional relationship information between the target cooking materials and the target object; The cooking step recommendation module is used to recommend cooking step information corresponding to the target cooking ingredients to the target object based on the location relationship information and the target voice recipe information.
[0006] According to a third aspect of the embodiments of this application, a smart device is provided, including: a range hood, a refrigerator, a dishwasher, and smart glasses.
[0007] According to a fourth aspect of the present application, an electronic device is provided, comprising: a processor; a memory for storing processor-executable instructions; wherein the processor is configured to execute the instructions to implement the method as described in any one of the first aspects above.
[0008] According to a fifth aspect of the present application, a computer-readable storage medium is provided, wherein when instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to perform any of the methods described in the first aspect of the present application. According to a sixth aspect of the embodiments of this application, a computer program product is provided, including computer instructions that, when executed by a processor, cause a computer to perform the method described in any one of the first aspects of the embodiments of this application.
[0009] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this application.
[0010] The technical solutions provided by the embodiments of this application bring at least the following beneficial effects: By responding to a target voice command, the system filters out the target voice recipe information corresponding to the target voice command from multiple pre-stored voice recipe information. This eliminates the need for users to frequently touch the screen or flip through paper recipes during the cooking process, avoiding inconvenience and hygiene issues caused by oil, water stains, or other contaminants on their hands, and improving the smoothness and ease of operation of the cooking process. Based on the target voice recipe information, the location information of the target cooking materials and the target object in the cooking area is identified and processed to obtain the location relationship information between the target cooking materials and the target object. This allows for real-time perception of the cooking environment and enhances the real-time nature and scene perception capabilities of cooking guidance. Based on the location relationship information and the target voice recipe information, the cooking steps corresponding to the target cooking ingredients are recommended to the target object. This can effectively prevent users from making mistakes in cooking steps and mixing up seasonings, thereby improving cooking accuracy. At the same time, it can also focus on the safety of using pots and knives during the cooking process, avoiding risks such as burns or cuts to users, thus improving cooking safety and further enhancing the user experience.
[0011] Other features and aspects of this application will become clear from the following detailed description of exemplary embodiments with reference to the accompanying drawings. Attached Figure Description
[0012] To more clearly illustrate the technical solutions and advantages in the embodiments or prior art of this specification, the drawings used in the description of the embodiments or prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0013] Figure 1 This is a flowchart illustrating a recipe recommendation method according to an exemplary embodiment.
[0014] Figure 2 This is a flowchart illustrating a recipe information conversion method according to an exemplary embodiment.
[0015] Figure 3 This is a block diagram of a recipe device according to an exemplary embodiment.
[0016] Figure 4 This is a block diagram of an electronic device for recipe recommendation according to an exemplary embodiment. Figure 1 .
[0017] Figure 5 This is a block diagram of an electronic device for recipe recommendation according to an exemplary embodiment. Figure 2 . Detailed Implementation
[0018] To enable those skilled in the art to better understand the technical solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments in the specification, and not all of the embodiments. Based on the embodiments in this specification, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0019] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or server that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or devices.
[0020] Various exemplary embodiments, features, and aspects of the present invention will now be described in detail with reference to the accompanying drawings. The same reference numerals in the drawings denote elements that have the same or similar functions. Although various aspects of the embodiments are shown in the drawings, they are not necessarily drawn to scale unless specifically indicated otherwise.
[0021] The term "exemplary" as used herein means "serving as an example, embodiment, or illustration." Any embodiment illustrated herein as "exemplary" is not necessarily to be construed as superior to or better than other embodiments. The term "and / or" in this document is merely a description of the relationship between related objects, indicating that three relationships may exist, for example, A and / or B, which can represent: A alone, A and B simultaneously, and B alone. Furthermore, the term "at least one" in this document means any combination of at least two of any one or more of a plurality, for example, including at least one of A, B, and C, which can represent including any one or more elements selected from the set consisting of A, B, and C.
[0022] Unless otherwise specified, the directions in this article should be understood as follows: the direction closer to the user is forward, and the direction farther from the user is backward.
[0023] Furthermore, to better illustrate the present invention, numerous specific details are set forth in the following detailed embodiments. Those skilled in the art will understand that the present invention can be practiced without certain specific details. In some instances, methods, means, elements, and circuits well known to those skilled in the art have not been described in detail in order to highlight the spirit of the invention.
[0024] It should be noted that the following diagram shows one possible sequence of steps, and it is not strictly necessary to follow this order. Some steps can be executed in parallel without interdependence.
[0025] Before introducing the method embodiments provided by the present invention, a brief introduction will be given on the application scenarios, related terms or nouns that may be involved in the method embodiments of the present invention, so as to facilitate the understanding of those skilled in the art.
[0026] To avoid inconvenience and hygiene issues caused by users frequently touching electronic screens such as mobile phones, tablets, and range hood displays during cooking, or by having oil, water, or other contaminants on their hands when flipping through paper recipes, and because existing voice-recommended recipe systems are simply based on audio tapes or audio e-books and cannot present the visual effects of the recipes, thus reducing the user experience, this application provides a recipe recommendation method that can perceive the cooking process in real time, thereby improving user safety, convenience, and overall user experience.
[0027] Figure 1 This is a flowchart illustrating a recipe recommendation method according to an exemplary embodiment. For example... Figure 1 As shown, it may include the following steps.
[0028] In step S101, in response to the target voice command, the target voice recipe information corresponding to the target voice command is selected from a plurality of pre-stored voice recipe information.
[0029] In the embodiments of this specification, the target voice command can refer to a voice command issued by a target object that is related to recipe lookup and cooking operations. The target object can be a user who is cooking, preparing to cook, or looking up a recipe; this application does not limit this. Voice recipe information can refer to information played to the user using voice as a medium. Voice recipe information is obtained by performing image and text recognition processing on multiple image and text recipe information.
[0030] The target voice recipe information can refer to the information corresponding to the cooking or recipe query that the target object is about to perform. For example, if the target voice command is to make scrambled eggs with tomatoes, then the target voice recipe information can be the voice recipe steps corresponding to scrambled eggs with tomatoes.
[0031] In one possible implementation, before step S101, multiple text and image recipe information is acquired; text and image recognition processing is performed on the multiple text and image recipe information to obtain multiple voice recipe information, and the multiple voice recipe information is stored.
[0032] Among them, multiple illustrated recipe information can refer to information that uses pictures and locations as carriers, including pictures of dishes, lists of ingredients, proportions of seasonings, and specific seasoning data, such as 5 grams of light soy sauce and 5 grams of salt. This application does not limit this.
[0033] The recipe recommendation method of this application can be applied to smart devices, such as range hoods, smart glasses, dishwashers, etc., and this application does not limit it.
[0034] The intelligent devices may include high-definition cameras, supplementary lighting systems, and various detection sensors, etc., and this application does not limit the specific type of these devices. High-definition cameras can be used to recognize image and text information in graphic recipes. They can also be used to scan ingredients and cooking materials to better provide cooking services to the target audience, and this application does not limit the type of camera used. High-definition cameras can include various types. For example, an RGB camera can be used to capture red, green, and blue primary color information to generate a color image. In the recipe recommendation method of this application, it can be used to recognize graphic recipe information and the color and shape of ingredients, such as tomatoes and green vegetables, and this application does not limit the type of camera used. Depth cameras can not only acquire color images but also the distance from each point in the scene to the camera, thereby generating a three-dimensional point cloud map. Infrared cameras can be used for night vision or scenes with low lighting, and this application does not limit the specific type of high-definition camera used. Supplementary lighting systems can provide a stable and controllable light source environment when the high-definition camera is acquiring images, to overcome recognition interference caused by insufficient or changing ambient light during the recognition process, and to ensure that the acquired image and video data are clear and accurate. Multiple detection sensors can be used to detect the position of recipes, ingredients, and the positional relationship between cookware and target objects.
[0035] Figure 2 This is a flowchart illustrating a recipe information conversion method according to an exemplary embodiment. For example... Figure 2 As shown, distortion correction and image enhancement processing are performed on each image information in multiple image and text recipe information to obtain the cooking materials corresponding to each of the multiple image and text recipe information; text recognition processing and semantic correction processing are performed on multiple text regions in multiple image and text recipe information to obtain the text recipe information corresponding to each of the multiple text regions; and voice recipe information is determined based on the cooking materials and text recipe information.
[0036] The image information can refer to the pictures in the illustrated recipe information, including images of ingredients, such as cucumbers, and the shape of the cucumbers, such as shredded or diced. This application does not limit this. Cooking materials can include the materials needed for cooking, such as recipe ingredient information, pots, stoves, and spatulas or ladles. Recipe ingredient information includes all ingredients in the target speech recipe information, and may include seasonings and specific ingredients for cooking. This application does not limit this. The text area can refer to the area corresponding to the text in the illustrated recipe information. The text recipe information can refer to the recipe corresponding to the recognized text portion.
[0037] In one possible implementation, geometric distortion correction is performed on multiple image information to obtain multiple image correction information; image enhancement processing is performed on the multiple image correction information according to the histogram equalization algorithm to obtain multiple image enhancement information; the multiple image enhancement information is input into an image parsing model for image parsing processing to obtain cooking materials.
[0038] In the embodiments of this specification, image correction information can refer to recipe image data after image correction. Histogram equalization algorithm can refer to an algorithm that enhances global contrast by redistributing the brightness values of pixels. Image enhancement information can refer to image data with enhanced contrast. The image parsing model is based on a deep learning object detection and recognition model, such as a faster region-based convolutional neural network (Faster R-CNN), etc., and this application does not limit it to this.
[0039] For example, geometric distortion correction processing is performed on the image information in the illustrated recipe information, that is, the original image information captured by a high-definition camera is corrected. Since high-definition cameras may cause straight lines at the image edges to become curved, for example, straight lines may become barrel-shaped or pincushion-shaped distortions, this application does not limit this. Geometric distortion correction processing can use Zhang's calibration algorithm, which is a flexible and convenient camera calibration method. It only requires taking several photos of a planar calibration board from different angles to calculate the camera's internal and external parameters, which can then be used to correct distortion. This application does not limit the algorithm for geometric distortion correction processing; preferably, Zhang's calibration algorithm is used to perform geometric distortion correction processing on the image information in the illustrated recipe information. This results in multiple image correction information after distortion correction processing. The error of the multiple image correction information after geometric distortion correction processing can be less than or equal to 0.5 pixels.
[0040] Furthermore, image enhancement processing is applied to the multiple image correction information output in the previous step. This is done by redistributing the pixel intensity values of the image, such as widening the histogram, though this application does not limit the specifics. This is particularly useful for recipe images that were taken in poor lighting, are too dark or too bright, or have low contrast. After image enhancement processing, multiple image enhancement information is obtained, which can accurately reveal the edges, textures, colors, and types of cooking ingredients in the image.
[0041] Finally, the multiple image enhancement information output from the previous step is input into an image parsing model, which can be a deep learning-based object detection and recognition model. This image parsing model has been trained using multiple historical image enhancement information. It can then identify the position of each object in the image enhancement information, as well as the objects within each bounding box. The image enhancement information is then input into the image parsing model for image parsing processing, outputting cooking ingredients, such as tomatoes, eggs, a wok, and a spatula; this application does not limit the specific ingredients.
[0042] In one possible implementation, text detection processing is performed on the image and text recipe information according to the differentiable binarization algorithm to determine multiple text regions; text recognition processing is performed on the text in the text regions to determine multiple initial text information; semantic correction processing is performed on the multiple initial text information to obtain the text recipe information corresponding to each of the multiple text regions.
[0043] In the embodiments of this specification, the differentiable binarization algorithm can be used to solve the problem of inaccurate text edge detection in traditional text detection methods under complex backgrounds. Initial text information can refer to the information obtained after text recognition.
[0044] For example, a differentiable binarization algorithm is used to perform text detection processing on the illustrated recipe information to identify multiple text regions, i.e., the text portions of the printed recipe information, such as the textual steps for cooking ingredients, the specific quantities and weights of seasonings, etc. Then, text recognition processing is performed on the text regions to determine the initial text information; for example, directly recognizing the text regions in the illustrated recipe information to obtain all the text information in the recipe. Correspondingly, semantic correction processing is performed on the initial text information to obtain the text recipe information corresponding to each text region. For example, regarding the instruction to stir-fry over high heat until just cooked, visually impaired individuals or those who do not cook frequently might not be able to determine the exact cooking time. Further semantic correction processing can be performed, such as setting the stove to the highest heat level, for example, level three, and then continuing to stir-fry for 2 minutes after the stove is at level three. This modification can accurately inform users of the cooking time and heat level, improving the user experience.
[0045] After processing the image and text information as described above, the cooking ingredients corresponding to each of the multiple image and text recipe information and the text recipe information corresponding to each of the multiple text regions are obtained. Based on the recipe ingredient information, cooking ingredients and text recipe information corresponding to each recipe, the corresponding voice recipe information for each image and text recipe information is determined.
[0046] When a user is cooking and uses voice to activate a recipe, the recipe recommendation system responds to the target voice command and then filters out the target voice recipe from a pre-stored list of multiple voice recipes. For example, if a user says, "I want to make scrambled eggs with tomatoes," the recipe recommendation system immediately responds to that voice command and then filters out the target voice recipe from a pre-stored list of multiple voice recipes.
[0047] In step S103, based on the target voice recipe information, the location information of the target cooking materials and the target object in the cooking area is identified and processed to obtain the location relationship information between the target cooking materials and the target object.
[0048] In the embodiments of this specification, the location information of the target object may refer to the user's specific location. Target cooking materials include target cooking ingredients and target cooking tools. For example, in scrambled eggs with tomatoes, the target cooking ingredients are tomatoes, eggs, and salt; the target cooking tools are bowls, chopsticks, woks, stoves, and spatulas, but this application does not limit these.
[0049] Locational relationship information can refer to the corresponding positions of the target cooking material and the target object. For example, if the target voice recipe is for scrambled eggs with tomatoes, then the tomato is located 30 centimeters to the left of the target object.
[0050] In one possible implementation, semantic analysis is performed on the target voice recipe information to determine the target cooking materials; the target cooking materials in the cooking area are identified to determine the cooking location information; the location information of the target object is obtained; and the cooking location information and the location information are spatially matched to determine the location relationship information.
[0051] In the embodiments of this specification, cooking location information may refer to the location information of the target cooking material relative to the high-definition camera.
[0052] For example, the target cooking ingredients are parsed from the target language recipe information. For instance, the target cooking ingredients used in each step are found from the cooking step information; for example, if the current step is to add 5 grams of salt, the target cooking ingredient is extracted as salt. Further, a high-definition camera is used to scan the cooking area, which can refer to areas such as the kitchen; this application is not limited to this. The target cooking ingredients in the cooking area are scanned to determine the cooking location information corresponding to the target cooking ingredients. For example, the location corresponding to salt is found, i.e., coordinates (x1, y1, z1).
[0053] Next, the location information of the target object, i.e., the user issuing the target language command, can be obtained, with coordinates (x2, y2, z2). By calculating the Euclidean distance and movement trend between the target object and the target cooking material, that is, by spatially matching the cooking location information and the location information, the Euclidean distance and movement trend between the target object and the target cooking material are determined. Furthermore, based on the Euclidean distance and movement trend, the positional relationship information is determined, for example, the salt shaker is 30 cm southeast of the left hand. This allows the system to accurately perceive the user's actions and position, thereby predicting the user's next operation, avoiding user errors, and improving the user experience. In particular, when visually impaired users are cooking, it can improve their cooking safety, prevent them from encountering dangers during the cooking process, improve the ease of operation for visually impaired users, and enhance the intelligence of recipe recommendations.
[0054] In step S105, based on the location relationship information and the target voice recipe information, cooking step information corresponding to the target cooking material is recommended to the target object.
[0055] In the embodiments of this specification, cooking step information may refer to cooking operation guidance steps recommended to the target object in the form of voice.
[0056] In one possible implementation, the high-definition camera can perceive positional information in real time. If it detects that the target object has not yet begun any operation, it can verbally play the vegetable-washing step from the cooking steps, such as making scrambled eggs with tomatoes, to the target object. At this point, the voice prompt might say, "Wash the tomatoes on your right, facing due east, and take three eggs from the refrigerator behind you." This application does not limit the content of the voice prompt. After the high-definition camera detects that the target object has completed this step, it can jump to the next step and continue the playback process.
[0057] In one possible implementation, in response to a target jump instruction, the target cooking step corresponding to the target jump instruction is located from the cooking step information, and the target cooking step information is recommended to the target object.
[0058] In the embodiments of this specification, the target jump instruction can refer to the step jump instruction issued by the target object. For example, if the target object is currently in the first step of cooking, but wants to know how much seasoning to add, it may issue a voice command to "jump directly to the third step". This application does not limit this.
[0059] The target cooking step can refer to the step information corresponding to the jump.
[0060] For example, a target user wants to make scrambled eggs with tomatoes, but the first step has already been completed, such as beating the eggs, washing and chopping the tomatoes. The target user then wants to proceed directly to the scrambled egg process and issues a "jump to the third step of scrambled eggs with tomatoes" instruction. The recipe recommendation system receives this jump instruction and immediately locates the third step of scrambled eggs with tomatoes, i.e., the target cooking step, for example, "adjust the stove heat to level two, pour in 30 ml of oil after 10 seconds, and then pour in the beaten eggs and stir-fry after another 10 seconds." This application does not limit this step.
[0061] Figure 3 This is a block diagram illustrating a recipe device according to an exemplary embodiment. (Refer to...) Figure 3 The device may include: The filtering module 301 is used to filter out the target voice recipe information corresponding to the target voice command from a plurality of pre-stored voice recipe information in response to the target voice command; the plurality of pre-stored voice recipe information is obtained by performing image and text recognition processing on a plurality of image and text recipe information. The position relationship determination module 303 is used to identify and process the position information of the target cooking materials and the target object in the cooking area according to the target voice recipe information, so as to obtain the position relationship information between the target cooking materials and the target object; In one possible implementation, the positional relationship determination unit 303 includes: A target cooking ingredient determination unit is used to perform semantic analysis on the target speech recipe information to determine the target cooking ingredients. The cooking location information determination unit is used to identify the target cooking material in the cooking area and determine the cooking location information; A location information acquisition unit is used to acquire the location information of the target object; The positional relationship information determination unit is used to perform spatial matching processing on the cooking position information and the position information to determine the positional relationship information.
[0062] The cooking step recommendation module 305 is used to recommend cooking step information corresponding to the target cooking material to the target object based on the location relationship information and the target voice recipe information.
[0063] In one possible implementation, prior to the filtering module 301, the following is included: The illustrated recipe information acquisition unit is used to acquire multiple illustrated recipe information. The voice recipe information storage unit is used to perform image and text recognition processing on the multiple image and text recipe information to obtain the multiple voice recipe information and store the multiple voice recipe information.
[0064] In one possible implementation, the voice recipe information storage unit includes: The cooking material determination unit is used to perform distortion correction and image enhancement processing on each image information in the plurality of graphic recipe information to obtain the cooking materials corresponding to each of the plurality of graphic recipe information. The text recipe information determination unit is used to perform text recognition processing and semantic correction processing on multiple text regions in the multiple text and image recipe information to obtain the text recipe information corresponding to each of the multiple text regions. The voice recipe information determination unit is used to determine the voice recipe information based on the cooking ingredients and the text recipe information.
[0065] In one possible implementation, the cooking material determining unit includes: An image correction unit is used to perform geometric distortion correction processing on the multiple image information to obtain multiple image correction information; The image enhancement unit is used to perform image enhancement processing on the plurality of image correction information according to the histogram equalization algorithm to obtain the plurality of image enhancement information; An image analysis unit is used to input the multiple image enhancement information into the image analysis model for image analysis processing to obtain the cooking materials.
[0066] In one possible implementation, the text recipe information determining unit includes: The text region determination unit is used to perform text detection processing on the graphic recipe information according to the differentiable binarization algorithm to determine the multiple text regions; An initial recognition unit is used to perform text recognition processing on the text in the text region and determine multiple initial text information; The text semantic correction unit is used to perform semantic correction processing on the multiple initial text information to obtain the text recipe information corresponding to each of the multiple text regions.
[0067] In one possible implementation, the recipe recommendation device also includes: The target jump unit is used to respond to the target jump instruction, locate the target cooking step corresponding to the target jump instruction from the cooking step information, and recommend the target cooking step information to the target object.
[0068] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.
[0069] Figure 4 This is a block diagram illustrating an electronic device for recipe recommendation according to an exemplary embodiment. The electronic device may be a terminal, and its internal structure diagram may be as follows: Figure 4 As shown, the electronic device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage medium. The network interface is used to communicate with external terminals via a network connection. When the computer program is executed by the processor, it implements a recipe recommendation method. The display screen can be an LCD screen or an e-ink screen. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad mounted on the device's casing, or an external keyboard, touchpad, or mouse. Those skilled in the art will understand that Figure 4 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the electronic device to which the present application is applied. The specific electronic device may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.
[0070] Figure 5 This is a block diagram illustrating an electronic device for recipe recommendation according to an exemplary embodiment. The electronic device may be a server, and its internal structure diagram may be as follows: Figure 5 As shown, this electronic device includes a processor, memory, and a network interface connected via a system bus. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage medium. The network interface is used to communicate with external terminals via a network connection. When the computer program is executed by the processor, it implements a recipe recommendation method. Those skilled in the art will understand that Figure 5 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the electronic device to which the present application is applied. The specific electronic device may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements. In an exemplary embodiment, an electronic device is also provided, including: a processor; and a memory for storing processor-executable instructions; wherein the processor is configured to execute the instructions to implement the recipe recommendation method as described in the embodiments of this application.
[0071] In an exemplary embodiment, a computer-readable storage medium is also provided, which, when executed by a processor of an electronic device, enables the electronic device to perform the recipe recommendation method of the present application embodiments. The computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, or optical data storage device, etc.
[0072] In an exemplary embodiment, a computer program product containing instructions is also provided, which, when run on a computer, causes the computer to execute the recipe recommendation method in the embodiments of this application.
[0073] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and RAMbus dynamic RAM (RDRAM), etc.
[0074] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this application are indicated by the following claims.
[0075] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.
Claims
1. A recipe recommendation method, characterized in that, include: In response to a target voice command, the target voice recipe information corresponding to the target voice command is selected from a plurality of pre-stored voice recipe information; The pre-stored multiple voice recipe information is obtained by performing image and text recognition processing on multiple image and text recipe information; Based on the target voice recipe information, the location information of the target cooking materials and the target object in the cooking area is identified and processed to obtain the location relationship information between the target cooking materials and the target object. Based on the location relationship information and the target voice recipe information, cooking steps corresponding to the target cooking ingredients are recommended to the target object.
2. The recipe recommendation method according to claim 1, characterized in that, Before the step of filtering out the target voice recipe information corresponding to the target voice command from a plurality of pre-stored voice recipe information in response to the target voice command, the method includes: Get multiple recipe information with pictures and text; The multiple text and image recipe information is processed by image and text recognition to obtain the multiple voice recipe information, and the multiple voice recipe information is stored.
3. The recipe recommendation method according to claim 2, characterized in that, The step of performing image and text recognition processing on the multiple image and text recipe information to obtain the multiple voice recipe information includes: Distortion correction and image enhancement processing are performed on each image in the multiple illustrated recipe information to obtain the cooking materials corresponding to each of the multiple illustrated recipe information; Text recognition and semantic correction are performed on multiple text regions in the multiple text and image recipe information to obtain the text recipe information corresponding to each of the multiple text regions. The voice recipe information is determined based on the cooking ingredients and the text recipe information.
4. The recipe recommendation method according to claim 3, characterized in that, The process of performing distortion correction and image enhancement on each image in the plurality of illustrated recipe information to obtain the cooking ingredients corresponding to each of the plurality of illustrated recipe information includes: Geometric distortion correction processing is performed on the multiple image information to obtain multiple image correction information; Based on the histogram equalization algorithm, image enhancement processing is performed on the plurality of image correction information to obtain the plurality of image enhancement information; The multiple image enhancement information is input into the image parsing model for image parsing processing to obtain the cooking materials.
5. The recipe recommendation method according to claim 3, characterized in that, The step of performing text recognition and semantic correction processing on multiple text regions in the multiple text and image recipe information to obtain the text recipe information corresponding to each of the multiple text regions includes: Based on the differentiable binarization algorithm, text detection processing is performed on the graphic recipe information to determine the multiple text regions; The text in the text region is subjected to text recognition processing to determine multiple initial text information; Semantic correction processing is performed on the multiple initial text information to obtain the text recipe information corresponding to each of the multiple text regions.
6. The recipe recommendation method according to claim 1, characterized in that, The step of identifying and processing the location information of the target cooking materials and the target object in the cooking area based on the target voice recipe information to obtain the location relationship information between the target cooking materials and the target object includes: Semantic analysis is performed on the target speech recipe information to determine the target cooking ingredients; The target cooking materials in the cooking area are identified to determine the cooking location information; Obtain the location information of the target object; The cooking location information and the location information are spatially matched to determine the location relationship information.
7. The recipe recommendation method according to claim 1, characterized in that, The method further includes: In response to a target jump command, the target cooking step corresponding to the target jump command is located from the cooking step information, and the target cooking step information is recommended to the target object.
8. A recipe recommendation device, characterized in that, include: A filtering module is used to filter out the target voice recipe information corresponding to the target voice command from a plurality of pre-stored voice recipe information in response to the target voice command; The pre-stored multiple voice recipe information is obtained by performing image and text recognition processing on multiple image and text recipe information; The positional relationship determination module is used to identify and process the positional information of the target cooking materials and the target object in the cooking area according to the target voice recipe information, so as to obtain the positional relationship information between the target cooking materials and the target object; The cooking step recommendation module is used to recommend cooking step information corresponding to the target cooking ingredients to the target object based on the location relationship information and the target voice recipe information.
9. A smart device, characterized in that, The recipe recommendation method as described in any one of claims 1-7 is used, wherein the smart device is one of a range hood, refrigerator, dishwasher, or smart glasses.
10. A computer-readable storage medium, characterized in that, When the instructions in the computer-readable storage medium are executed by the processor of the electronic device, the electronic device is able to perform the recipe recommendation method as described in any one of claims 1 to 7.