A food material input method and apparatus
By integrating image recognition and text recognition on shopping list images, the problem of cumbersome refrigerator food entry process and high misidentification rate has been solved, achieving simple and efficient food management.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HISENSE GRP HLDG CO LTD
- Filing Date
- 2021-11-30
- Publication Date
- 2026-06-05
AI Technical Summary
In existing technologies, the process of entering food information into refrigerators is cumbersome, complicated, and has a high rate of misidentification, making it difficult to manage food information conveniently.
By acquiring images of shopping lists, image recognition is performed, and image and text information are identified and merged to determine ingredient information. This information is then displayed on the mobile interface for user confirmation before being entered into the ingredient database.
It enables one-click batch entry of food information, simplifies the operation process, improves entry efficiency and accuracy, and makes it easier for users to manage refrigerator food.
Smart Images

Figure CN116206296B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of smart home appliance technology, and in particular to a method and device for inputting food ingredients. Background Technology
[0002] With the continuous development of smart home appliances, food management based on machine vision has become a development trend in the home appliance industry. For example, machine vision can be used to input food information into refrigerators.
[0003] Currently, the main methods for entering food items into refrigerators are manual entry via an app or refrigerator camera recognition. However, the first method, manual entry via an app, is cumbersome, requiring the input of a large amount of food information and is complex. Furthermore, due to the rapid development of online shopping and the proliferation of various platforms, there are often issues with these platforms not being able to interoperate with the smart refrigerator's food management system. This results in purchased food items not being able to be easily added to the refrigerator's food management system, making it difficult to truly help users conveniently manage their refrigerator's food. The second method, using the refrigerator camera for recognition, is often limited by issues such as food packaging, obstructions, and lighting, leading to low recognition rates and similarly failing to truly help users manage their food.
[0004] Therefore, there is a need to provide a more efficient food ingredient entry solution to address the problems of cumbersome entry processes and high misidentification rates. Summary of the Invention
[0005] This invention provides a method and device for inputting food ingredients, which solves the problems of cumbersome food ingredient input process and high misidentification rate.
[0006] In a first aspect, an embodiment of the present invention provides a method for inputting food ingredients, comprising:
[0007] Obtain an image of a shopping list to be entered; perform image recognition based on the shopping list image to obtain at least one product information included in the shopping list image; display the ingredient information included in the at least one product information on a first interface, the ingredient information including one or more of ingredient name, ingredient quantity and ingredient price; and enter the ingredient information into the ingredient database in response to the user's input operation triggered by the ingredient information.
[0008] Using the above method, image recognition is performed on the shopping list image to determine the ingredient information that needs to be entered into the ingredient database and presented on the mobile interface for the user to confirm. This achieves one-click batch entry of ingredient information, which is simple to operate and effectively improves the efficiency and accuracy of ingredient entry, making it more convenient for users to manage refrigerator ingredients.
[0009] In one possible design, image recognition is performed based on the shopping list image to obtain information about at least one product included in the shopping list image, including:
[0010] The shopping list image is divided into image regions and text regions; at least one image recognition result corresponding to the image region and at least one text recognition result corresponding to the text region are identified; the at least one image recognition result and the at least one text recognition result are fused to obtain at least one product information included in the shopping list.
[0011] The above method provides a way to determine product information based on a shopping list image. For example, the shopping list image can be divided into image areas and text areas. Then, at least one image recognition result corresponding to the image area and at least one text recognition result corresponding to the text area can be identified respectively. The image recognition results and text recognition results can be fused to obtain at least one product information. By combining images and text to identify product information, the content in the shopping list image is fully utilized, and the accuracy of recognition is effectively improved.
[0012] In one possible design, the image recognition result includes the coordinate information of the image, the product type represented by the image, and the confidence level of the image recognition result; the text recognition result includes the coordinate information of the text box corresponding to the text, the confidence level of the text recognition result, and the text content corresponding to the text, wherein the text content includes one or more of the following: product name, product quantity, and product price.
[0013] In the above manner, the image recognition result in this application embodiment may also include the coordinate information of the image, the product name represented by the image, and the confidence level of the image recognition result. The text recognition result may also include the coordinate information of the text box corresponding to the text, the confidence level of the text recognition result, and the text content corresponding to the text. Thus, more content can be obtained based on the image recognition result and the text recognition result, making it more applicable.
[0014] In one possible design, the method further includes:
[0015] Based on the coordinate information of each image recognition result and each text recognition result, determine the image recognition result and text recognition result representing the same product; based on the image recognition result and text recognition result corresponding to each product, determine the output result for each product.
[0016] By using the coordinate information of each image recognition result as described above, the image recognition result and text recognition result corresponding to the same product can be determined. Then, based on the image recognition result and text recognition result of the same product, the output result of the corresponding product can be determined, which can effectively improve the accuracy of the product output result.
[0017] In one possible design, the output result for each product is determined based on the image recognition result and text recognition result corresponding to each product, including:
[0018] Determine whether the image recognition result and the text recognition result of the same product are consistent; if so, the output result of the product includes the image recognition result and the text recognition result.
[0019] If not, when the confidence level of the image recognition result is higher than the confidence level of the text recognition result, the output result of the product is one or more of the product type in the image recognition result and the product name, product quantity, and product price included in the text content of the text recognition result; or, when the confidence level of the text recognition result is higher than the confidence level of the image recognition result, the output result of the product is the text recognition result; or, when both the confidence level of the text recognition result and the confidence level of the image recognition result are less than a preset confidence level threshold, the output result of the product is unknown.
[0020] In this embodiment of the application, when outputting product recognition results, the final output result of the product can be determined jointly based on the confidence level of the image recognition result and the confidence level of the text recognition result corresponding to the product, effectively ensuring the accuracy of the product output result.
[0021] In one possible design, the first interface displays ingredient information included in the at least one product information, including:
[0022] The information of ingredients that need to be placed in the refrigerator is determined from the information of at least one product; the storage location of each ingredient in the refrigerator is determined according to the degree of matching between the storage conditions provided by each storage location in the refrigerator and the storage conditions required by each ingredient in the information of ingredients; the storage location of each ingredient in the refrigerator is displayed on the first interface.
[0023] In this embodiment of the application, after determining the information of the ingredients that need to be placed in the refrigerator, the storage location of each ingredient in the information of the ingredients in the refrigerator is determined according to the degree of matching between the storage conditions provided by each storage location in the refrigerator and the storage conditions required by each ingredient in the information of the ingredients. The storage location of each ingredient in the information of the ingredients in the refrigerator is also displayed on the display interface, which can effectively improve the intelligence of food storage management.
[0024] In one possible design, before entering the ingredient information into the ingredient database after responding to a user's input operation, the method further includes:
[0025] In response to a user's instruction to modify the ingredient information displayed on the first interface, the ingredient information displayed on the first interface is modified according to the modification instruction.
[0026] Through the above method, this application embodiment allows users to further modify the ingredient information to be entered before entering the ingredient database, making it more intelligent and convenient.
[0027] Secondly, an embodiment of the present invention provides a method for inputting food ingredients, including:
[0028] The system acquires an image of a shopping list to be entered; receives ingredient information, which is determined from at least one product information obtained through image recognition based on the shopping list image; displays the ingredient information on a first interface, which includes one or more of the following: ingredient name, ingredient quantity, ingredient price, and the storage location of the ingredient in the refrigerator; and enters the ingredient information into the ingredient database in response to a user's input operation triggered by the ingredient information.
[0029] Using the above method, image recognition is performed on the shopping list image to determine the ingredient information that needs to be entered into the ingredient database and presented on the mobile interface for the user to confirm. This achieves one-click batch entry of ingredient information, which is simple to operate and effectively improves the efficiency and accuracy of ingredient entry, making it more convenient for users to manage refrigerator ingredients.
[0030] In one possible design, the ingredient information is displayed on the first interface, including:
[0031] Based on the degree of matching between the storage conditions provided by each storage location in the refrigerator and the storage conditions required by each ingredient in the ingredient information, the storage location of each ingredient in the ingredient information in the refrigerator is determined; the storage location of each ingredient in the ingredient information in the refrigerator is displayed on the first interface.
[0032] In this embodiment of the application, after determining the information of the ingredients that need to be placed in the refrigerator, the storage location of each ingredient in the information of the ingredients in the refrigerator is determined according to the degree of matching between the storage conditions provided by each storage location in the refrigerator and the storage conditions required by each ingredient in the information of the ingredients. The storage location of each ingredient in the information of the ingredients in the refrigerator is also displayed on the display interface, which can effectively improve the intelligence of food storage management.
[0033] In one possible design, before entering the ingredient information into the ingredient database after responding to a user's input operation, the method further includes:
[0034] In response to a user's instruction to modify the ingredient information displayed on the first interface, the ingredient information displayed on the first interface is modified according to the modification instruction.
[0035] Through the above method, this application embodiment allows users to further modify the ingredient information to be entered before entering the ingredient database, making it more intelligent and convenient.
[0036] Thirdly, an embodiment of the present invention provides a method for inputting food ingredients, including:
[0037] The system receives a shopping list image sent by a terminal device; performs image recognition based on the shopping list image to obtain at least one product information included in the shopping list image; sends the ingredient information included in the at least one product information to the terminal device, so that the terminal device displays the ingredient information on a first display interface, and enters the ingredient information into an ingredient database in response to a user's input operation triggered by the ingredient information; wherein the ingredient information includes one or more of the following: ingredient name, ingredient quantity, ingredient price, and the storage location of the ingredient in the refrigerator.
[0038] Using the above method, image recognition is performed on the shopping list image to determine the ingredient information that needs to be entered into the ingredient database and presented on the mobile interface for the user to confirm. This achieves one-click batch entry of ingredient information, which is simple to operate and effectively improves the efficiency and accuracy of ingredient entry, making it more convenient for users to manage refrigerator ingredients.
[0039] In one possible design, image recognition is performed based on the shopping list image to obtain information about at least one product included in the shopping list image, including:
[0040] The shopping list image is divided into image areas and text areas;
[0041] Identify at least one image recognition result corresponding to the image region and at least one text recognition result corresponding to the text region;
[0042] The at least one image recognition result is fused with the at least one text recognition result to obtain at least one product information included in the shopping list.
[0043] The above method provides a way to determine product information based on a shopping list image. For example, the shopping list image can be divided into image areas and text areas. Then, at least one image recognition result corresponding to the image area and at least one text recognition result corresponding to the text area can be identified respectively. The image recognition results and text recognition results can be fused to obtain at least one product information. By combining images and text to identify product information, the content in the shopping list image is fully utilized, and the accuracy of recognition is effectively improved.
[0044] In one possible design, the image recognition result includes the coordinate information of the image, the product type represented by the image, and the confidence level of the image recognition result;
[0045] The text recognition result includes the coordinate information of the text box corresponding to the text, the confidence level of the text recognition result, and the text content corresponding to the text, which includes one or more of the following: product name, product quantity, and product price.
[0046] In the above manner, the image recognition result in this application embodiment may also include the coordinate information of the image, the product name represented by the image, and the confidence level of the image recognition result. The text recognition result may also include the coordinate information of the text box corresponding to the text, the confidence level of the text recognition result, and the text content corresponding to the text. Thus, more content can be obtained based on the image recognition result and the text recognition result, making it more applicable.
[0047] In one possible design, the method further includes:
[0048] Based on the coordinate information of each image recognition result and the coordinate information of each text recognition result, determine the image recognition result and text recognition result representing the same product;
[0049] Based on the image recognition and text recognition results for each product, the output result for each product is determined.
[0050] By using the coordinate information of each image recognition result as described above, the image recognition result and text recognition result corresponding to the same product can be determined. Then, based on the image recognition result and text recognition result of the same product, the output result of the corresponding product can be determined, which can effectively improve the accuracy of the product output result.
[0051] In one possible design, the output result for each product is determined based on the image recognition result and text recognition result corresponding to each product, including:
[0052] Determine whether the image recognition results and text recognition results for the same product are consistent;
[0053] If so, the output of the product includes the image recognition result and the text recognition result;
[0054] If not, when the confidence level of the image recognition result is higher than the confidence level of the text recognition result, the output result of the product is one or more of the product type in the image recognition result and the product name, product quantity, and product price included in the text content of the text recognition result; or, when the confidence level of the text recognition result is higher than the confidence level of the image recognition result, the output result of the product is the text recognition result; or, when both the confidence level of the text recognition result and the confidence level of the image recognition result are less than a preset confidence level threshold, the output result of the product is unknown.
[0055] In this embodiment of the application, when outputting product recognition results, the final output result of the product can be determined jointly based on the confidence level of the image recognition result and the confidence level of the text recognition result corresponding to the product, effectively ensuring the accuracy of the product output result.
[0056] Fourthly, embodiments of the present invention provide an intelligent refrigerator, comprising:
[0057] An image acquisition device, used to acquire images of the shopping list to be entered;
[0058] A display, the display being used to display a first interface;
[0059] The processor is configured to perform image recognition based on the shopping list image to obtain at least one product information included in the shopping list image; display food information included in the at least one product information on a first interface, the food information including one or more of the following: food name, food quantity, food price, and food storage location in the refrigerator; and, in response to a user's input operation triggered by the food information, input the food information into a food database.
[0060] In one possible design, the processor is specifically configured to divide the shopping list image into image regions and text regions; identify at least one image recognition result corresponding to the image regions and at least one text recognition result corresponding to the text regions; and fuse the at least one image recognition result with the at least one text recognition result to obtain at least one product information included in the shopping list.
[0061] In one possible design, the image recognition result includes the coordinate information of the image, the product type represented by the image, and the confidence level of the image recognition result;
[0062] The text recognition result includes the coordinate information of the text box corresponding to the text, the confidence level of the text recognition result, and the text content corresponding to the text, which includes one or more of the following: product name, product quantity, and product price.
[0063] In one possible design, the processor is further configured to determine the image recognition result and text recognition result representing the same product based on the coordinate information of each image recognition result and the coordinate information of each text recognition result; and to determine the output result of each product based on the image recognition result and text recognition result corresponding to each product.
[0064] In one possible design, the processor is specifically used to determine whether the image recognition result and the text recognition result of the same product are consistent;
[0065] If so, the output of the product includes the image recognition result and the text recognition result;
[0066] If not, when the confidence level of the image recognition result is higher than the confidence level of the text recognition result, the output result of the product is one or more of the product type in the image recognition result and the product name, product quantity, and product price included in the text content of the text recognition result; or, when the confidence level of the text recognition result is higher than the confidence level of the image recognition result, the output result of the product is the text recognition result; or, when both the confidence level of the text recognition result and the confidence level of the image recognition result are less than a preset confidence level threshold, the output result of the product is unknown.
[0067] In one possible design, the processor is configured to determine, from the at least one product information, information about the food items that need to be placed in the refrigerator;
[0068] Based on the degree of matching between the storage conditions provided by each storage location in the refrigerator and the storage conditions required by each ingredient in the ingredient information, the storage location of each ingredient in the ingredient information in the refrigerator is determined.
[0069] The first interface displays the storage location of each ingredient in the refrigerator from the ingredient information.
[0070] In one possible design, the processor is also configured to respond to a user's instruction to modify the food information displayed on the first interface;
[0071] The ingredient information displayed on the first interface is modified according to the modification instruction.
[0072] Fifthly, embodiments of the present invention provide an intelligent refrigerator, comprising:
[0073] An image acquisition device, used to acquire an image of the shopping list to be entered;
[0074] A display, the display being used to display a first interface;
[0075] The processor is configured to display the food information on a first interface. The food information is determined from at least one product information obtained by image recognition based on the shopping list image. The food information includes one or more of the following: food name, food quantity, food price, and food storage location in the refrigerator. In response to a user's input operation on the food information, the processor inputs the food information into the food database.
[0076] In one possible design, the processor is configured to determine the storage location of each ingredient in the food information in the refrigerator based on the degree of matching between the storage conditions provided by each storage location in the refrigerator and the storage conditions required by each ingredient in the food information; and to display the storage location of each ingredient in the food information in the refrigerator on the first interface.
[0077] In one possible design, the processor is further configured to respond to a user's instruction to modify the food information displayed on the first interface; and to modify the food information displayed on the first interface according to the modification instruction.
[0078] Sixthly, embodiments of the present invention provide a food ingredient input device, comprising:
[0079] A transceiver, the transceiver being used to receive a shopping list image sent by a terminal device;
[0080] A processor, configured to perform image recognition based on the shopping list image to obtain information about at least one product included in the shopping list image;
[0081] The transceiver is also used to send the ingredient information included in the at least one product information to the terminal device, so that the terminal device displays the ingredient information on the first display interface, and enters the ingredient information into the ingredient database after responding to the user's input operation triggered by the ingredient information.
[0082] The ingredient information includes one or more of the following: ingredient name, ingredient quantity, ingredient price, and the storage location of the ingredient in the refrigerator.
[0083] In one possible design, the processor is specifically configured to divide the shopping list image into image regions and text regions; identify at least one image recognition result corresponding to the image regions and at least one text recognition result corresponding to the text regions; and fuse the at least one image recognition result with the at least one text recognition result to obtain at least one product information included in the shopping list.
[0084] In one possible design, the image recognition result includes the coordinate information of the image, the product type represented by the image, and the confidence level of the image recognition result;
[0085] The text recognition result includes the coordinate information of the text box corresponding to the text, the confidence level of the text recognition result, and the text content corresponding to the text, which includes one or more of the following: product name, product quantity, and product price.
[0086] In one possible design, the processor is further configured to determine the image recognition result and text recognition result representing the same product based on the coordinate information of each image recognition result and the coordinate information of each text recognition result; and to determine the output result of each product based on the image recognition result and text recognition result corresponding to each product.
[0087] In one possible design, the processor is specifically used to determine whether the image recognition result and the text recognition result of the same product are consistent;
[0088] If so, the output of the product includes the image recognition result and the text recognition result;
[0089] If not, when the confidence level of the image recognition result is higher than the confidence level of the text recognition result, the output result of the product is one or more of the product type in the image recognition result and the product name, product quantity, and product price included in the text content of the text recognition result; or, when the confidence level of the text recognition result is higher than the confidence level of the image recognition result, the output result of the product is the text recognition result; or, when both the confidence level of the text recognition result and the confidence level of the image recognition result are less than a preset confidence level threshold, the output result of the product is unknown.
[0090] In a seventh aspect, embodiments of the present invention also provide a computing device, comprising: a memory for storing a computer program; and a processor for calling the computer program stored in the memory and executing the method described in various possible designs of the first to third aspects according to the obtained program.
[0091] Eighthly, embodiments of the present invention also provide a computer-readable non-volatile storage medium including a computer-readable program that, when read and executed by a computer, causes the computer to perform the method described in various possible designs of the first to third aspects. Attached Figure Description
[0092] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0093] Figure 1a An exemplary schematic diagram of the smart refrigerator in the off state in an embodiment of this application is shown;
[0094] Figure 1b An exemplary schematic diagram of the smart refrigerator in the open state in an embodiment of this application is shown;
[0095] Figure 2 An exemplary system architecture diagram is shown in an embodiment of this application;
[0096] Figure 3 An exemplary schematic diagram of the functional structure of the controller of the smart refrigerator in an embodiment of this application is shown;
[0097] Figure 4 A flowchart of a method for inputting food ingredients provided in an embodiment of the present invention;
[0098] Figure 5This is a schematic diagram of the food ingredient input stage provided in an embodiment of the present invention;
[0099] Figure 6 A shopping list image diagram provided for an embodiment of the present invention;
[0100] Figure 7 This is a schematic diagram of the image region division of a shopping list provided in an embodiment of the present invention;
[0101] Figure 8 This is a schematic diagram of a food ingredient display interface provided in an embodiment of the present invention. Detailed Implementation
[0102] To enable those skilled in the art to better understand the technical solutions in this disclosure, the technical solutions in the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this disclosure, and not all embodiments. Based on the embodiments in this disclosure, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of this disclosure.
[0103] In the description of this disclosure, it should be understood that the terms “center,” “upper,” “lower,” “front,” “rear,” “left,” “right,” “vertical,” “horizontal,” “top,” “bottom,” “inner,” and “outer,” etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings, and are only for the convenience of describing this disclosure and simplifying the description, and are not intended to indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of this disclosure.
[0104] The terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this disclosure, unless otherwise stated, "a plurality of" means two or more.
[0105] In the description of this disclosure, it should be noted that, unless otherwise expressly specified and limited, the terms "installation," "connection," and "linkage" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection between two components. Those skilled in the art can understand the specific meaning of the above terms in this disclosure based on the specific circumstances.
[0106] To address the cumbersome and complex process of entering food information into refrigerators in existing technologies, this disclosure provides a smart refrigerator and a food information entry method to simplify the process and improve efficiency. It should be noted that the method provided in this embodiment is applicable not only to smart refrigerators but also to smart freezers and other similar devices.
[0107] Figure 1a and Figure 1b An exemplary embodiment of the present application provides the structure of a smart refrigerator.
[0108] like Figure 1a As shown, the smart refrigerator includes a cabinet 10, a refrigeration unit (not shown in the figure), and other accessories (such as lights and thermometers that can be installed inside the cabinet, not shown in the figure). The refrigeration system mainly consists of a compressor, condenser, evaporator, and capillary tube throttling device, forming a closed-loop system. The evaporator can be installed at the top inside the smart refrigerator, while the other components are installed at the back of the smart refrigerator.
[0109] The enclosure 10 is equipped with a door 20, and a display screen 50 may be further installed on the door 20. The display screen 50 is coupled to the controller (e.g., through a circuit connection).
[0110] A camera module 30 can also be installed on the refrigerator body 10. This camera module can capture images of the front area of the refrigerator body 10, so as to capture images of food that the user puts into the refrigerator or food that the user takes out of the refrigerator. Specifically, taking the plane where the refrigerator door is located as the first plane, the front area of the refrigerator body 10 includes at least an area extending a certain distance outward from the first plane. The camera module can capture images of this area, that is, it can capture images of the user's hand movements during the process of storing and retrieving food after opening the door 20, as well as images of the stored and retrieved food.
[0111] In some embodiments, the camera module 30 may be positioned on the upper part of the housing 10 near the door 20 so as to capture images of the front area of the housing 10.
[0112] like Figure 1b As shown, the smart refrigerator's body 10 may include multiple compartments (compared to compartments 50a to 50e in the figure) to facilitate users in classifying and storing different types of food. Some compartments are semi-open (compared to compartments 50a to 50c in the figure), while others are closed (compared to compartments 50d to 50e in the figure). In this embodiment, weight sensors (not shown in the figure) may also be installed in the compartments to detect the weight of the food in that compartment.
[0113] It should be noted that, Figure 1a and Figure 1bThe structure of the smart refrigerator shown is merely an example. This application does not limit the size of the smart refrigerator, the number of doors (e.g., a single door or multiple doors), or the number and type of other accessories. For instance, in some embodiments, the smart refrigerator is equipped with a Radio Frequency Identification (RFID) reader, which can be used to read RFID tags on food packaging to obtain information such as the type and quantity of the food. In other embodiments, the smart refrigerator also has a voice function, capable of recognizing input voice to obtain information such as the type and quantity of food input by the user via voice.
[0114] Figure 2 An exemplary network architecture diagram applicable to embodiments of this application is shown.
[0115] As shown in the figure, the smart refrigerator 101 is connected to the server 103 via network 102. The server 103 can also communicate with the user's mobile terminal 105 via mobile communication network 104. In some application scenarios, the smart terminal connects to the gateway 106 via a local area network, and the gateway 106 can connect to the server via the Internet, enabling communication between the smart refrigerator 101 and the server 103.
[0116] based on Figure 2 In some embodiments of the system architecture shown, the smart refrigerator 101 can plan the storage location of food items, that is, rationally plan the storage location of food items. In other embodiments, the smart refrigerator 101 can send information about the food items to be planned to the server 103, which will then rationally plan the storage location of the food items. Furthermore, the server 103 can also send the food storage location planning results to the smart refrigerator, so that the smart refrigerator can display the planning results on a display screen or output them through other means (such as through voice broadcast). The server 103 can also send the food storage location planning results to the user's mobile terminal 105 through the mobile communication network 104, so that the user can conveniently view the food storage location planning results through the mobile terminal 105.
[0117] Figure 3 An exemplary schematic diagram of a controller in a smart refrigerator is shown, which can realize food identification and input functions. As shown in the figure, the controller may include: an information acquisition module 301, a food determination module 302, and a location allocation module 303. Further, it may also include a display module 304 and a communication module 305.
[0118] The information acquisition module 301 is used to acquire a shopping list image to be entered, and perform image recognition based on the shopping list image to acquire at least one product information included in the shopping list image; wherein, the information acquisition module 301 can acquire the shopping list image to be entered through a camera device in the smart refrigerator; it can also receive the shopping list image transmitted by a terminal device, such as a mobile phone, through a communication module 305 in the smart refrigerator.
[0119] In addition, the information acquisition module 301 in this embodiment can also acquire information from the food database and / or the refrigerator information database.
[0120] The ingredient determination module 302 is used to determine the ingredient information to be entered from the at least one product information;
[0121] The location allocation module 303 is used to determine the storage location of each ingredient in the food information in the refrigerator based on the degree of matching between the storage conditions provided by each storage location in the refrigerator and the storage conditions required by each ingredient in the food information.
[0122] Display module 304 is used to display the ingredient information included in the at least one product information on the first interface. Furthermore, display module 304 can also display the storage location of each ingredient in the ingredient information within the refrigerator on the first interface.
[0123] The communication module 305 is used to communicate with terminal devices or servers.
[0124] Figure 4 The illustration shows a schematic diagram of the food ingredient entry process provided in an embodiment of this application. This process can be executed by a smart refrigerator or by a server. As shown in the figure, the process may include the following steps:
[0125] S401: Obtain the image of the shopping list to be entered.
[0126] In this application embodiment, there are multiple ways to obtain the shopping list image to be entered, and the specific methods are not limited to the following:
[0127] Method 1: The shopping list image to be entered obtained in this embodiment of the application can be obtained by the user taking a picture of the shopping list page through the terminal device, and the pictured shopping list image can be uploaded to the execution device, such as uploading it to a smart refrigerator or server.
[0128] Method 2: The shopping list image to be entered obtained in this embodiment can be obtained by the user through a screenshot of the shopping list page on the terminal device, and the screenshot of the shopping list image can be uploaded to the execution device, such as a smart refrigerator or a server.
[0129] Understandably, users can shop through online shopping software or platforms on their terminal devices. The online shopping software or platforms on the terminal devices will generate electronic shopping lists, and the terminal devices can capture images of the electronic shopping lists by taking screenshots and uploading them to the execution device.
[0130] It should be noted that if the electronic shopping list in the online shopping software or platform can be saved as an image, the terminal device can send the saved electronic shopping list image directly to the execution device without taking a screenshot.
[0131] Method 3: The shopping list image to be entered obtained in this embodiment of the application can be obtained by the user by taking a picture of the shopping list page through the execution device. For example, the user takes a picture of the shopping list through the camera device in the smart refrigerator to obtain the image of the shopping list.
[0132] It should be noted that the images of the shopping list obtained in this application embodiment are obtained with the user's authorization or by the user's own operation, and all comply with the requirements of relevant national laws and regulations.
[0133] S402: Based on the shopping list image, perform image recognition to obtain information about at least one product included in the shopping list image.
[0134] In this step, when performing image recognition based on the shopping list image, the ROI (Region of Interest) can be divided according to the layout pattern of the shopping list image, and recognition can be performed based on different ROI regions.
[0135] For example, suppose the order details page of the shopping list image adopts a left-right layout. For instance, the left side of the interface displays the image information of the product, and the right side displays the text information such as the product name, price, specifications, and quantity.
[0136] Therefore, in this embodiment, the ROI region can be divided according to the layout rules, dividing the shopping list image into image regions and text regions. In this embodiment, a lightweight deep learning method can be used to divide the image and text regions, but this is not a limitation.
[0137] In addition, in order to effectively improve the accuracy of image recognition in this embodiment, the shopping list image can be preprocessed such as filtering before image recognition to remove noise interference and improve image quality.
[0138] Furthermore, in this embodiment of the application, the image region and the text region can be identified to obtain at least one product information included in the shopping list image.
[0139] In this embodiment of the application, at least one image recognition result corresponding to the image region and at least one text recognition result corresponding to the text region can be identified; then, the at least one image recognition result and the at least one text recognition result are fused to obtain at least one product information included in the shopping list.
[0140] For example, in this embodiment of the application, the image recognition result and the text recognition result representing the same product can be determined based on the coordinate information of each image recognition result and the coordinate information of each text recognition result. Then, the output result of each product can be determined based on the image recognition result and the text recognition result corresponding to each product.
[0141] Optionally, in this embodiment, the image recognition result includes the coordinate information of the image, the product name represented by the image, and the confidence level of the image recognition result; the text recognition result includes the coordinate information of the text box corresponding to the text, the confidence level of the text recognition result, and the text content corresponding to the text, wherein the text content includes one or more of the following: product name, product quantity, and product price.
[0142] Furthermore, in order to effectively improve the accuracy of the product output results in this embodiment, when determining the output result of each product based on the image recognition result and text recognition result corresponding to each product, the confidence level of the image recognition result and the confidence level of the text recognition result corresponding to the product can also be combined to determine the output result of the product. Specifically, this is not limited to the following situations:
[0143] Case 1: If the image recognition result and the text recognition result of the product are consistent, then the output result of the product includes the image recognition result and the text recognition result.
[0144] Scenario 2: If the image recognition result and the text recognition result of the product are inconsistent, and the confidence level of the image recognition result is higher than that of the text recognition result and higher than a preset confidence level threshold, the output result of the product is the product type in the image recognition result, and one or more of the product name, product quantity and product price included in the text content of the text recognition result.
[0145] Case 3: If the image recognition result and the text recognition result of the product are inconsistent, and the confidence level of the text recognition result is higher than that of the image recognition result and higher than a preset confidence level threshold, then the output result of the product shall be the text recognition result.
[0146] Case 4: If the image recognition result and the text recognition result of the product are inconsistent, and the confidence scores of both the text recognition result and the image recognition result are less than the preset confidence score threshold, then the output result of the product is unknown.
[0147] S403: Display the ingredient information included in the at least one product information on the first interface.
[0148] Optionally, the ingredient information described in this application embodiment may include the type, quantity, and price of the ingredients. The quantity of ingredients can be represented by their weight, volume (e.g., a 1000ml box of ice cream), or other units, such as a bunch of green vegetables, or by the number of individual ingredients, such as 10 eggs.
[0149] In this embodiment of the application, at least one product identified based on the shopping list image may not all be food ingredients.
[0150] For example, a shopping list includes six items: one piece of clothing, one pair of pants, two heads of cabbage, one radish, one head of cabbage, and 500g of beef. Of these, the two heads of cabbage, one radish, one head of cabbage, and 500g of beef are considered food ingredients. Therefore, to improve the accuracy of food ingredient entry and effectively reduce the process of users manually deleting or adding information, this embodiment of the application can filter out food ingredients from at least one product identified based on the shopping list image.
[0151] Furthermore, in this embodiment of the application, the storage location of each ingredient in the ingredient information in the refrigerator can be determined based on the degree of matching between the storage conditions provided by each storage location in the refrigerator and the storage conditions required by each ingredient in the ingredient information. For example, pork belly is suitable for freezing, and baby bok choy is suitable for refrigeration. Thus, the storage location of each ingredient in the ingredient information in the refrigerator is displayed on the first interface, which enriches the display content and makes it more applicable.
[0152] Furthermore, in this embodiment of the application, after identifying the food information and the storage location of each food item in the refrigerator from at least one product information through an intelligent algorithm, this information is displayed on the front-end interface. This allows users to further select or modify the food items to be added to the refrigerator via voice or touch screen, thus achieving better interaction with the user.
[0153] For example, the ingredient information displayed on the first interface includes ingredient 1 to ingredient 5. In response to the user's modification instruction to delete ingredient 5 from the ingredient information displayed on the first interface, ingredient 5 is deleted from the first interface according to the modification instruction. After deletion, the ingredient information displayed on the first interface includes ingredient 1 to ingredient 4.
[0154] S404: In response to the user's input operation triggered by the ingredient information, the ingredient information is entered into the ingredient database.
[0155] Among them, see Figure 5 As shown in the embodiment of this application, the process of inputting ingredients can be divided into five stages: shooting stage, recognition stage, display stage, verification stage, and input stage.
[0156] In order to better understand the embodiments of this application, the implementation process of the embodiments of this application will be described in detail below based on different execution stages.
[0157] Phase 1: Filming Phase
[0158] Assuming that the shopping list image to be entered in this embodiment of the application is as follows: Figure 6 As shown, the order details page of the shopping list image adopts a left-right layout, with the left side of the interface displaying the product image information and the right side displaying the product name, price, specifications, quantity, and other text information.
[0159] After the terminal device obtains the shopping list image by taking a picture or screenshot, it uploads the shopping list image to the smart refrigerator.
[0160] Phase Two: Identification Phase
[0161] After acquiring the shopping list image, the smart refrigerator performs image recognition on the shopping list image.
[0162] like Figure 7 As shown, the shopping list image is divided into two parts: the image area of the products on the left and the text area of the products on the right.
[0163] Then, at least one image recognition result is obtained for the image region recognition, and at least one text recognition result is obtained for the text region recognition.
[0164] Specifically, for image regions, image recognition algorithms can be used to identify the product types in each image within the image region.
[0165] For example, according to the above Figure 7 The content shown identifies images 1-4 in the shopping list image, representing the categories of goods: clothing, fruit, vegetables, and fruits.
[0166] For image regions, neural networks can also be used for detection and recognition to identify the position coordinates of the detection box corresponding to each image in the image region.<x,y,w,h> And the corresponding confidence level. Where x and y in the position coordinates are the coordinates of the top-left corner of the detection box, and w and h are the width and height of the detection box. It should be noted that, in the embodiments of this application, the neural network used for recognition can be, but is not limited to, one or a derivative model of a network structure such as (deep) neural network, convolutional neural network, deep confidence network, deep stacked neural network, deep fusion network, deep recurrent neural network, deep recurrent neural network, deep Bayesian neural network, deep generative network, deep reinforcement learning, etc.
[0167] For example, according to the above Figure 6 The content shown below includes the following: the coordinates of the detection boxes corresponding to the four images in the identified shopping list image, and their confidence scores:
[0168] The coordinates of the detection box corresponding to image 1, which represents the product type as clothing, are <100, 100, 50, 50>, with a confidence level of 0.93.
[0169] The coordinates of the detection box corresponding to image 2, which indicates that the product type is fruit, are <100, 150, 50, 50>, with a confidence level of 0.82.
[0170] The coordinates of the detection box corresponding to image 3, which indicates that the product type is fruit, are <100, 200, 50, 50>, with a confidence level of 0.71.
[0171] The coordinates of the detection box corresponding to image 4, which indicates that the product type is vegetables, are <100, 250, 50, 50>, with a confidence level of 0.58.
[0172] For ease of representation, the product name, location coordinates, and corresponding confidence score of each image identified for the image region can be represented as {N, x, y, w, h, D}, where N represents the product name, x and y are the coordinates of the top-left corner of the detection box, w and h are the width and height of the detection box, and D is the confidence score of the image recognition result. For example, the recognition result of each image identified for the image recognition region can be represented as: {Short-sleeved shirt, 100, 100, 50, 50, 0.93}, {Banana, 100, 150, 50, 50, 0.82}, {Chicken cutlet, 100, 200, 50, 50, 0.71}, {Apple, 100, 250, 50, 50, 0.58}.
[0173] Specifically, for text areas, character recognition algorithms can be used to identify the text information in each text box within the text area. For example, based on the above... Figure 6The content shown identifies four text boxes in the shopping list image, and the text information corresponding to each text box is as follows:
[0174] The text information corresponding to text box 1 is: Product name is short sleeve, product quantity is 1, and product price is 156;
[0175] The text information corresponding to text box 2 is: Product type is vegetables, product name is banana, product quantity is 700g, and product price is 18.8.
[0176] The text information corresponding to text box 3 is: Product type is poultry, product name is steak, product quantity is 500g, and product price is 30.
[0177] The text information corresponding to text box 4 is: Product name is pear, product quantity is 450g, and product price is 16.7.
[0178] For text regions, neural networks can be used for detection and recognition to identify the position coordinates <[x1, y1], [x2, y2], [x3, y3], [x4, y4]> of each text box within the text region, along with the corresponding confidence level. Here, [x1, y1] represents the top-left corner coordinate of the text box, [x2, y2] represents the top-right corner coordinate, [x3, y3] represents the bottom-right corner coordinate, and [x4, y4] represents the bottom-left corner coordinate.
[0179] For example, according to the above Figure 6 The content shown below, along with the position coordinates and confidence levels of the four text boxes identified in the shopping list image, are as follows:
[0180] The position coordinates corresponding to text box 1 are: <[160, 90], [260, 90], [260, 60], [160, 60]>, with a confidence level of 0.98;
[0181] The position coordinates of text box 2 are: <[160, 140], [260, 140], [260, 110], [160, 110]>, with a confidence level of 0.81;
[0182] The position coordinates of text box 3 are: <[160, 190], [260, 190], [260, 160], [160, 160]>, with a confidence level of 0.76;
[0183] The position coordinates of text box 4 are: <[160, 240], [260, 240], [260, 210], [160, 210], with a confidence level of 0.61.
[0184] For ease of representation, the text information represented by each text box identified for the text region, the position coordinates of the text box, and the corresponding confidence score can be represented as {"text", [x1, y1], [x2, y2], [x3, y3], [x4, y4], D}, where text represents the text information in the text box, [x1, y1] are the coordinates of the upper left corner of the text box, [x2, y2] are the coordinates of the upper right corner of the text box, [x3, y3] are the coordinates of the lower right corner of the text box, [x4, y4] are the coordinates of the lower left corner of the text box, and D is the confidence score of the text recognition result.
[0185] For example, the recognition result of each character in the character recognition region can be represented as:
[0186] {"Short-sleeved shirt, women's, quantity 1, price 156", [160, 90], [260, 90], [260, 60], [160, 60], 0.98}, {"700g bananas, quantity 1, price 18.8", [160, 140], [260, 140], [260, 110], [160, 110], 0.81}, {"500g steak, quantity 1, price 30", [160, 190], [260, 190], [260, 160], [160, 160], 0.76}, {"450g pears, quantity 1, price 16.7", [160, 240], [260, 240], [260, 210], [160, 210], 0.61}.
[0187] Then, based on the coordinate information of each image recognition result and the coordinate information of each text recognition result, the image recognition result and text recognition result representing the same product are determined.
[0188] Optionally, in the embodiments of this application, the image recognition result and text recognition result representing the same product can be determined according to the following formula 1.
[0189]
[0190] According to Formula 1 above, it can be determined that image 1 and text box 1 represent the same product, for example, product 1; image 2 and text box 2 represent the same product, for example, product 2; image 3 and text box 3 represent the same product, for example, product 3; and image 4 and text box 4 represent the same product, for example, product 4.
[0191] Then, the image recognition results and text recognition results are merged, and the output result for each product is determined based on the image recognition results and text recognition results corresponding to each product.
[0192] In this embodiment of the application, in order to effectively improve the accuracy of the product output results, when determining the output result of each product based on the image recognition result and the text recognition result corresponding to each product, the confidence level of the image recognition result and the confidence level of the text recognition result corresponding to the product can also be combined to determine the output result of the product. Specifically, it is not limited to the following situations:
[0193] Case 1: If the image recognition result and the text recognition result of the product are consistent, then the output result of the product includes the image recognition result and the text recognition result.
[0194] For example, suppose the pre-set confidence threshold is 0.62, where the confidence of the image recognition result corresponding to product 1 is 0.93 and the confidence of the text recognition result corresponding to product 1 is 0.98.
[0195] The product category in the image recognition result is clothing, and the text content in the text recognition result includes the product name as short sleeves, the product quantity as 1, and the product price as 156.
[0196] Since the image recognition result and the text recognition result corresponding to product 1 are consistent, and the confidence levels of both the image recognition result and the text recognition result are higher than the preset confidence threshold of 0.62, the image recognition result and the text recognition result corresponding to product 1 can be used together as the output result of product 1.
[0197] For example, the output of product 1 is: product type is clothing, product name is short sleeve, product quantity is 1, and product price is 156.
[0198] Scenario 2: If the image recognition result and the text recognition result of the product are inconsistent, and the confidence level of the image recognition result is higher than that of the text recognition result and higher than a preset confidence level threshold, the output result of the product is the product type in the image recognition result, and one or more of the product name, product quantity and product price included in the text content of the text recognition result.
[0199] For example, assuming the pre-set confidence threshold is 0.62, the confidence of the image recognition result corresponding to product 2 is 0.82, and the confidence of the text recognition result corresponding to product 2 is 0.81.
[0200] The product category in the image recognition result is fruit, and the text content in the text recognition result includes the product category as vegetables, the product name as banana, the product quantity as 700g, and the product price as 18.8.
[0201] Since the image recognition result and text recognition result corresponding to product 2 are inconsistent, and the confidence level of the image recognition result is higher than that of the text recognition result and higher than the preset confidence level threshold of 0.62, then the product type in the image recognition result corresponding to product 2, and the product name, product quantity and product price included in the text content in the text recognition result can be used as the output result of product 2.
[0202] Understandably, although the text recognition result includes product types, since the product types in the text recognition result are different from the product types in the image recognition result, and the confidence level of the image recognition result is higher than that of the text recognition result, the product type corresponding to product 2 shall be based on the image recognition result.
[0203] For example, the output of product 2 is: product type is fruit, product name is banana 700g, product quantity is 1, and product price is 18.8.
[0204] Case 3: If the image recognition result and the text recognition result of the product are inconsistent, and the confidence level of the text recognition result is higher than that of the image recognition result and higher than a preset confidence level threshold, then the output result of the product shall be the text recognition result.
[0205] For example, assuming the pre-set confidence threshold is 0.62, the confidence of the image recognition result corresponding to product 3 is 0.71, and the confidence of the text recognition result corresponding to product 3 is 0.76.
[0206] The product category in the image recognition result is fruit, and the text content in the text recognition result includes the product category as poultry, the product name as steak, the product quantity as 500g, and the product price as 30.
[0207] Since the image recognition result and the text recognition result corresponding to product 3 are inconsistent, and the confidence level of the text recognition result is higher than that of the image recognition result and higher than the preset confidence level threshold of 0.62, the text recognition result corresponding to product 3 can be used as the output result of product 3.
[0208] For example, the output result of product 3 is: product type is poultry, product name is steak, product quantity is 500g, and product price is 30.
[0209] Case 4: If the image recognition result and the text recognition result of the product are inconsistent, and the confidence scores of both the text recognition result and the image recognition result are less than the preset confidence score threshold, then the output result of the product is unknown.
[0210] For example, assuming the pre-set confidence threshold is 0.62, the confidence of the image recognition result corresponding to product 4 is 0.58, and the confidence of the text recognition result corresponding to product 4 is 0.61.
[0211] Since the image recognition result and text recognition result corresponding to product 4 are both lower than the preset confidence threshold of 0.62, the output result of the product is unknown at this time.
[0212] For example, the output result of product 4 is unknown.
[0213] For a better understanding of the above four situations, please refer to Table 1 below for the specific analysis process.
[0214]
[0215] Table 1. Schematic diagram of output results determined based on image recognition results and character recognition results.
[0216] Through the above-mentioned recognition of the shopping list image, it can be determined that the shopping list image includes a total of 4 products: short-sleeved shirt, banana, steak, and unknown product 4.
[0217] Since not all products identified from the shopping list image in this embodiment are food ingredients, this embodiment can filter out food ingredients from the at least one product identified from the shopping list image in order to improve the accuracy of food ingredient entry and effectively reduce the process of users manually deleting and entering information.
[0218] For example, since the short-sleeved shirts in the above four products are not food ingredients, they can be removed. The final filtered food ingredients include bananas, steak, and unknown product 4.
[0219] Phase Three: Display Phase
[0220] In this embodiment of the application, the relevant information of the banana, steak and unknown product 4 can be displayed on the display interface of the smart refrigerator.
[0221] For example, such as Figure 8 As shown, the quantities of the bananas, steaks, and unknown product 4, the required storage conditions, and their respective storage locations in the refrigerator are displayed.
[0222] Phase Four: Verification Phase
[0223] Furthermore, the embodiments of this application can also allow users to further select or modify the ingredients to be added to the refrigerator via voice or touch screen, thus achieving better interaction with the user.
[0224] For example, the smart refrigerator receives a modification instruction triggered by the user for product 4, such as changing product 4 to a pear.
[0225] Phase 5: Data Entry Phase
[0226] The smart refrigerator, in response to a user's input operation triggered by the food information, inputs the food information into the food database.
[0227] For example, the relevant information about the bananas, steaks, and pears can be entered into the food database.
[0228] It should be noted that the above embodiments are merely examples of this application and do not constitute a limitation on the method of inputting ingredients in this application.
[0229] This application also provides a computer program product, including a computer program, which includes program instructions. When the program instructions are executed by an electronic device, the electronic device performs the food ingredient input method provided in the above embodiments.
[0230] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the above-described device and module can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0231] In the embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple modules or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces, apparatuses, or modules, and may be electrical, mechanical, or other forms.
[0232] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical modules; that is, they may be located in one place or distributed across multiple network modules. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0233] Furthermore, the functional modules in the various embodiments of this application can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium.
[0234] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product.
[0235] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a server or data center that integrates one or more available media. The available medium may be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state disk (SSD)).
[0236] The technical solutions provided in this application have been described in detail above. Specific examples have been used in this application to illustrate the principles and implementation methods of this application. The description of the above embodiments is only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A method for inputting food ingredients, characterized in that, The method includes: Obtain the image of the shopping list to be entered; The shopping list image is divided into image areas and text areas; The system identifies at least one image recognition result corresponding to the image region and at least one text recognition result corresponding to the text region; wherein, the image recognition result includes the coordinate information of the image, the product type represented by the image, and the confidence level of the image recognition result; the text recognition result includes the coordinate information of the text box corresponding to the text, the confidence level of the text recognition result, and the text content corresponding to the text, wherein the text content includes one or more of the following: product type, product name, product quantity, and product price. Based on the coordinate information of each image recognition result and each text recognition result, determine the image recognition result and text recognition result representing the same product; Based on the image recognition results and text recognition results corresponding to each product, the output result of each product is determined to obtain at least one product information included in the shopping list image; wherein, determining the output result of each product based on the image recognition results and text recognition results corresponding to each product includes: determining whether the image recognition result and text recognition result of the same product are consistent; if yes, the output result of the product includes the image recognition result and the text recognition result; if no, when it is determined that the confidence level of the image recognition result is higher than the confidence level of the text recognition result, the output result of the product is one or more of the product type in the image recognition result and the product name, product quantity, and product price included in the text content of the text recognition result; or, when it is determined that the confidence level of the text recognition result is higher than the confidence level of the image recognition result, the output result of the product is the text recognition result; or, when it is determined that the confidence level of both the text recognition result and the image recognition result is less than a preset confidence level threshold, the output result of the product is unknown; The first interface displays the ingredient information included in the at least one product information, which includes one or more of the following: ingredient name, ingredient quantity, ingredient price, and the storage location of the ingredient in the refrigerator. In response to the user's input operation triggered by the ingredient information, the ingredient information is entered into the ingredient database.
2. The method as described in claim 1, characterized in that, The display of ingredient information included in the at least one product information on the first interface includes: Determine the food information that needs to be placed in the refrigerator from the at least one product information; Based on the degree of matching between the storage conditions provided by each storage location in the refrigerator and the storage conditions required by each ingredient in the ingredient information, the storage location of each ingredient in the ingredient information in the refrigerator is determined. The first interface displays the ingredient information and the storage location of each ingredient in the refrigerator.
3. The method as described in claim 1, characterized in that, Before entering the food information into the food database after responding to the user's input operation on the food information, the method further includes: In response to the user's instruction to modify the ingredient information displayed on the first interface; The ingredient information displayed on the first interface is modified according to the modification instruction.
4. A method for inputting food ingredients, characterized in that, The method includes: Obtain the image of the shopping list to be entered; The process involves receiving ingredient information, which is determined from at least one product information obtained through image recognition based on the shopping list image. This process includes: identifying at least one image recognition result corresponding to the image region and at least one text recognition result corresponding to the text region; wherein the image recognition result includes the coordinate information of the image, the product type represented by the image, and the confidence level of the image recognition result; the text recognition result includes the coordinate information of the text box corresponding to the text, the confidence level of the text recognition result, and the text content corresponding to the text, wherein the text content includes one or more of the following: product type, product name, product quantity, and product price; determining the image recognition result and text recognition result representing the same product based on the coordinate information of each image recognition result and each text recognition result; and determining the output result for each product based on the image recognition result and text recognition result corresponding to each product to obtain at least one product information included in the shopping list image; wherein determining the output result for each product based on the image recognition result and text recognition result corresponding to each product includes: determining the image recognition result and text recognition result of the same product... If the text recognition results are consistent, then the product output includes both the image recognition result and the text recognition result; if not, when the confidence level of the image recognition result is higher than that of the text recognition result, the product output is one or more of the product type in the image recognition result and the product name, product quantity, and product price included in the text content of the text recognition result; or, when the confidence level of the text recognition result is higher than that of the image recognition result, the product output is the text recognition result; or, when both the confidence level of the text recognition result and the confidence level of the image recognition result are less than a preset confidence threshold, the product output is unknown. The first interface displays the ingredient information, which includes one or more of the following: ingredient name, ingredient quantity, ingredient price, and the storage location of the ingredient in the refrigerator. In response to the user's input operation triggered by the ingredient information, the ingredient information is entered into the ingredient database.
5. The method as described in claim 4, characterized in that, The display of the ingredient information on the first interface includes: Based on the degree of matching between the storage conditions provided by each storage location in the refrigerator and the storage conditions required by each ingredient in the ingredient information, the storage location of each ingredient in the ingredient information in the refrigerator is determined. The first interface displays the ingredient information and the storage location of each ingredient in the refrigerator.
6. The method as described in claim 4 or 5, characterized in that, Before entering the food information into the food database after responding to the user's input operation on the food information, the method further includes: In response to the user's instruction to modify the ingredient information displayed on the first interface; The ingredient information displayed on the first interface is modified according to the modification instruction.
7. A method for inputting food ingredients, characterized in that, The method includes: Receive shopping list images sent by the terminal device; Based on the shopping list image, image recognition is performed to obtain at least one product information included in the shopping list image; wherein, based on the shopping list image, image recognition is performed to obtain at least one product information included in the shopping list image, including: recognizing at least one image recognition result corresponding to the image region and at least one text recognition result corresponding to the text region; wherein, the image recognition result includes the coordinate information of the image, the product type represented by the image, and the confidence level of the image recognition result; the text recognition result includes the coordinate information of the text box corresponding to the text, the confidence level of the text recognition result, and the text content corresponding to the text, wherein the text content includes one or more of the following: product type, product name, product quantity, and product price; based on the coordinate information of each image recognition result and each text recognition result, the image recognition result and text recognition result representing the same product are determined; based on the image recognition result and text recognition result corresponding to each product, the output result of each product is determined to obtain at least one product information included in the shopping list image; wherein, the step of determining the output result of each product based on the image recognition result and text recognition result corresponding to each product includes: determining the image recognition result and text recognition result of the same product... If the text recognition results are consistent, then the product output includes both the image recognition result and the text recognition result; if not, when the confidence level of the image recognition result is higher than that of the text recognition result, the product output is one or more of the product type in the image recognition result and the product name, product quantity, and product price included in the text content of the text recognition result; or, when the confidence level of the text recognition result is higher than that of the image recognition result, the product output is the text recognition result; or, when both the confidence level of the text recognition result and the confidence level of the image recognition result are less than a preset confidence threshold, the product output is unknown. The ingredient information included in the at least one product information is sent to the terminal device so that the terminal device displays the ingredient information on the first display interface, and after responding to the user's input operation triggered by the ingredient information, the ingredient information is entered into the ingredient database. The ingredient information includes one or more of the following: ingredient name, ingredient quantity, ingredient price, and the storage location of the ingredient in the refrigerator.
8. A smart refrigerator, characterized in that, include: An image acquisition device, used to acquire images of the shopping list to be entered; A display, the display being used to display a first interface; A processor configured to perform the method as described in any one of claims 1-3.
9. A smart refrigerator, characterized in that, include: An image acquisition device, used to acquire an image of the shopping list to be entered; A display, the display being used to display a first interface; The processor is configured to perform the method as described in any one of claims 4-6.
10. A computing device, characterized in that, include: Memory, used to store program instructions; A processor is configured to invoke program instructions stored in the memory and execute the method as described in any one of claims 1-3; or execute the method as described in any one of claims 4-6; or execute the method as described in claim 7.