A food material recognition method, device, equipment and medium
By using a pre-trained food recognition model and multi-angle image acquisition, combined with hand detection, the problems of low food recognition accuracy and low automation were solved, achieving efficient food recognition even under occlusion and rapid storage conditions.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HISENSE GRP HLDG CO LTD
- Filing Date
- 2021-12-22
- Publication Date
- 2026-04-28
AI Technical Summary
Existing technologies for food identification have low accuracy and automation, especially when food is obscured or stored quickly, making accurate identification difficult and requiring manual input.
A pre-trained food identification model is used to acquire images to be identified through a main image acquisition device. If the confidence level is low, hand detection is performed to determine the target hand position. An auxiliary image acquisition device is used to acquire multi-angle images. By combining non-maximum suppression method and hand detection, the type of food can be identified.
It improves the accuracy and automation of food ingredient identification, reduces the need for manual data entry, and ensures accurate identification of food ingredient types in complex scenarios.
Smart Images

Figure CN116363642B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of smart refrigerator technology, and in particular to a method, device, equipment and medium for food identification. Background Technology
[0002] With the popularization of the smart home concept, household users' requirements for home appliances are constantly increasing. As an essential appliance, the refrigerator's basic functions can already meet the needs of most families. However, in order to gain a foothold in the high-end home appliance market, refrigerators need to improve their practicality while meeting users' higher-level needs through intelligent technology.
[0003] To meet the needs of households for precise food management, more and more manufacturers are choosing to combine smart refrigerators with cameras, resulting in a variety of food recognition solutions.
[0004] Static recognition solutions include traditional external refrigerator screen image acquisition devices for identifying food and internal image acquisition devices for food identification. However, while internal image acquisition devices can capture images of the refrigerator interior, their field of view is limited, especially when food is stacked on the same shelf, making it easy for food to be obstructed. Using traditional external refrigerator screen image acquisition devices requires the user to move the food within the image acquisition range; the automation level of image acquisition is low, and the accuracy of food identification and the quality of food type input are not ideal. Therefore, static recognition solutions cannot effectively identify food in actual user food storage scenarios, requiring manual input of food types.
[0005] Figure 1 A schematic diagram illustrating food covering in existing technology, such as... Figure 1 As shown, the right side of the figure ( Figure 1 The ingredients (on the left and right sides) are obscured; Figure 2 An illustration of another food covering technique provided by existing technology, such as... Figure 2 As shown, the left side of the figure ( Figure 2 The ingredients (left and right sides) in the middle are obscured.
[0006] The dynamic recognition solution uses a single image acquisition device mounted on the top of the refrigerator compartment to detect and identify food items during real-time storage, eliminating the need for users to manually input information via an app or voice. However, this method has a limited perspective when using only a single image acquisition device, and its accuracy remains low when food is stored too quickly or obscured by hands.
[0007] Figure 3 An illustration of an existing technology that shows food is stored too quickly, such as... Figure 3 As shown, Figure 3Because the ingredients are stored quickly, the images are somewhat blurry. Figure 4 An illustration of food being obscured, provided for reference in existing technology, such as... Figure 4 As shown, a large area of the food is obscured by the hand.
[0008] Furthermore, existing technologies propose a method to identify the types of food stored in a refrigerator by acquiring images of food from multiple angles. By establishing a mapping relationship between multiple images inside the refrigerator, when the type of food cannot be determined, the image with the smallest obstructed area among the multi-angle images is selected, and the image recognition result is used as the final identified food type. However, this method only uses the analysis of static images from different angles to determine the food type. When the refrigerator is saturated with food, it is difficult to penetrate obstructions to identify food behind them. Therefore, this method still requires manual input to supplement the food category information.
[0009] Therefore, improving the accuracy of food identification, reducing manual data entry, and increasing the automation level of food identification have become urgent technical problems to be solved. Summary of the Invention
[0010] This application provides a method, apparatus, device, and medium for food ingredient identification, which addresses the problems of low accuracy and automation in food ingredient identification in the prior art.
[0011] Firstly, this application provides a method for identifying food ingredients, the method comprising:
[0012] The first image to be identified is acquired by the main image acquisition device during the food storage and retrieval process. Based on the pre-trained food identification model, the first confidence level corresponding to the first food type in the first image to be identified is determined.
[0013] If the first confidence level is less than the preset confidence threshold, then hand detection is performed on the first image to be identified to obtain the target hand position in the first image to be identified. Based on the target hand position and the pre-saved correspondence between the hand position and each auxiliary image acquisition device, each target auxiliary image acquisition device corresponding to the target hand position is determined, wherein the hand position includes left hand, right hand, front hand and back hand.
[0014] Acquire each second image to be identified by each target auxiliary image acquisition device, identify each second food type in each second image to be identified based on the food identification model, and determine the identified target food type based on each second food type.
[0015] Furthermore, if the first confidence level is not less than a preset confidence threshold, the method further includes:
[0016] The first food ingredient type is identified as the target food ingredient type.
[0017] Furthermore, the first confidence level corresponding to the first food category in the first image to be identified, based on the pre-trained food identification model, includes:
[0018] The first image to be identified is input into a pre-trained food identification model. Based on the food identification model, each prediction box, each corresponding prediction food type, and each corresponding prediction confidence level of the food in the first image to be identified are determined. Based on each prediction box and each corresponding prediction confidence level, a non-maximum suppression method is used to determine the target prediction box. The prediction food type corresponding to the target prediction box is determined as the first food type output, and the prediction confidence level corresponding to the target prediction box is determined as the first confidence level output.
[0019] Furthermore, the method also includes:
[0020] Based on the food ingredient recognition model, the first area of the target prediction box is determined and output.
[0021] The step of performing hand detection on the first image to be identified to obtain the target hand position in the first image to be identified includes:
[0022] Hand detection is performed on the first image to be identified to determine the second area of the hand region in the first image to be identified, the coordinates of each bone node, and the corresponding preset number.
[0023] Determine whether the preset number corresponding to each bone node contains a first preset number and a second preset number;
[0024] If not, determine the ratio of the first area to the second area. If the ratio is not less than the first preset threshold, determine that the target hand position of the hand in the first image to be identified is a forward hand. If the ratio is less than the second preset threshold, determine that the target hand position of the hand in the first image to be identified is a reverse hand. Wherein the first preset threshold is greater than the second preset threshold.
[0025] If so, determine whether the first abscissa of the bone node corresponding to the first preset number is less than the second abscissa of the bone node corresponding to the second preset number. If so, determine that the target hand position of the hand in the first image to be identified is the left hand. If not, determine that the target hand position of the hand in the first image to be identified is the right hand.
[0026] Furthermore, determining the identified target ingredient type based on each of the second ingredient types includes:
[0027] If each of the second ingredient types is the same, then any second ingredient type is identified as the target ingredient type.
[0028] Furthermore, if each of the second food ingredients is not identical, the step of identifying each second food ingredient in each second image to be identified based on the food ingredient recognition model includes:
[0029] Based on the food ingredient recognition model, identify each type of second food ingredient and its corresponding second confidence level in each second image to be identified;
[0030] The step of determining the identified target ingredient type based on each of the second ingredient types includes:
[0031] Determine the target second confidence level among each second confidence level, and determine the second food type corresponding to the target second confidence level as the target food type.
[0032] Furthermore, the process of training the food ingredient recognition model includes:
[0033] Obtain any sample image from the sample set and the first label information corresponding to the sample image, wherein the first label information is used to identify the type of food contained in the sample image;
[0034] Using the original food identification model, each predicted bounding box, each predicted food type, and each predicted confidence level of the food in the sample image are determined. Based on each predicted bounding box and each predicted confidence level, a non-maximum suppression method is used to determine the target predicted bounding box. The predicted food type corresponding to the target predicted bounding box is determined as the second label information in the sample image.
[0035] Based on the first label information and the second label information, the parameter values of each parameter in the original food ingredient recognition model are adjusted.
[0036] Secondly, this application provides a food ingredient identification device, the device comprising:
[0037] The recognition module is used to acquire the first image to be recognized captured by the main image acquisition device during the food storage and retrieval process, and to identify the first confidence level corresponding to the first food type in the first image to be recognized based on the pre-trained food recognition model.
[0038] The determination module is used to perform hand detection on the first image to be identified if the first confidence level is less than a preset confidence threshold, to obtain the target hand position of the hand in the first image to be identified, and to determine each target auxiliary image acquisition device corresponding to the target hand position according to the target hand position and the pre-saved correspondence between the hand position and each auxiliary image acquisition device, wherein the hand position includes left hand, right hand, front hand and back hand;
[0039] The recognition module is further configured to acquire each second image to be recognized acquired by each target auxiliary image acquisition device, identify each second food type in each second image to be recognized based on the food recognition model, and determine the identified target food type based on each second food type.
[0040] Furthermore, the determining module is also used to determine the first food ingredient type as the identified target food ingredient type if the first confidence level is not less than a preset confidence threshold.
[0041] Furthermore, the recognition module is specifically used to input the first image to be recognized into a pre-trained food recognition model, and based on the food recognition model, determine each prediction box of the food in the first image to be recognized, each corresponding prediction food type, and each corresponding prediction confidence level. Based on each prediction box and each corresponding prediction confidence level, a non-maximum suppression method is used to determine a target prediction box, the prediction food type corresponding to the target prediction box is determined as the first food type output, and the prediction confidence level corresponding to the target prediction box is determined as the first confidence level output.
[0042] Furthermore, the recognition module is also used to determine the first area output of the target prediction box based on the food ingredient recognition model;
[0043] The determining module is specifically used to perform hand detection on the first image to be identified, determine the second area of the hand region in the first image to be identified, the coordinates of each bone node, and the corresponding preset number; determine whether the preset number corresponding to each bone node contains a first preset number and a second preset number; if not, determine the ratio of the first area to the second area; if the ratio is not less than a first preset threshold, determine that the target hand position of the hand in the first image to be identified is a right-handed hand; if the ratio is less than a second preset threshold, determine that the target hand position of the hand in the first image to be identified is a left-handed hand, wherein the first preset threshold is greater than the second preset threshold; if yes, determine whether the first abscissa of the bone node corresponding to the first preset number is less than the second abscissa of the bone node corresponding to the second preset number; if yes, determine that the target hand position of the hand in the first image to be identified is a left-handed hand; if not, determine that the target hand position of the hand in the first image to be identified is a right-handed hand.
[0044] Furthermore, the identification module is specifically used to identify any second ingredient type as the target ingredient type if all the second ingredient types are the same.
[0045] Furthermore, the recognition module is specifically used to identify each second food type and its corresponding second confidence level in each second image to be recognized based on the food recognition model; determine the largest target second confidence level among the second confidence levels, and determine the second food type corresponding to the target second confidence level as the target food type.
[0046] Furthermore, the device also includes:
[0047] The training module is used to acquire any sample image in the sample set and the first label information corresponding to the sample image, wherein the first label information is used to identify the types of food contained in the sample image; through the original food recognition model, it determines each prediction box of the food in the sample image, each corresponding predicted food type, and each corresponding prediction confidence; based on each prediction box and each corresponding prediction confidence, it uses a non-maximum suppression method to determine the target prediction box, and determines the predicted food type corresponding to the target prediction box as the second label information in the sample image; based on the first label information and the second label information, it adjusts the parameter values of each parameter in the original food recognition model.
[0048] Thirdly, this application provides an electronic device, which includes a processor and a memory, wherein the memory is used to store program instructions, and the processor is used to execute the computer program stored in the memory to implement the steps of any of the above-described methods for identifying food ingredients.
[0049] Fourthly, this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of any of the above-described methods for identifying food ingredients.
[0050] This application provides a method, apparatus, device, and medium for food ingredient recognition. Because this method is based on a pre-trained food ingredient recognition model, when the first confidence level corresponding to the first food ingredient type in the first image to be recognized acquired by the main image acquisition device is less than a preset confidence threshold, each second image to be recognized is acquired by each target auxiliary image acquisition device corresponding to the target hand position in the first image to be recognized. Each second food ingredient type in each second image to be recognized is then identified, and the identified target food ingredient type is determined based on each second food ingredient type. This improves the accuracy of food ingredient recognition and eliminates the need for manual input to supplement food ingredient types, thus increasing the automation level of food ingredient recognition. Attached Figure Description
[0051] To more clearly illustrate the technical solutions in this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0052] Figure 1 A schematic diagram illustrating a method of food covering using existing technology;
[0053] Figure 2 A schematic diagram illustrating another method of food covering provided by existing technology;
[0054] Figure 3 A diagram illustrating an existing technology where food is stored too quickly;
[0055] Figure 4 A schematic diagram illustrating an example of food being obscured using existing technology;
[0056] Figure 5 A schematic diagram illustrating the process of a food ingredient identification method provided in this application;
[0057] Figure 6 A schematic diagram illustrating how this application identifies each prediction box of food ingredients in a first image to be identified;
[0058] Figure 7 A schematic diagram illustrating another method for identifying each predicted bounding box of food ingredients in a first image to be identified, as provided in this application;
[0059] Figure 8 A method provided for this application from Figure 6 A schematic diagram of the target prediction box identified in the diagram;
[0060] Figure 9 A method provided for this application from Figure 7A schematic diagram of the target prediction box identified in the diagram;
[0061] Figure 10 A schematic diagram of a hand region in a first image to be identified, provided for this application;
[0062] Figure 11 A schematic diagram of a hand region in another first image to be identified provided in this application;
[0063] Figure 12 A schematic diagram of a hand region in another first image to be identified provided in this application;
[0064] Figure 13 A schematic diagram of a bone node provided in this application;
[0065] Figure 14 A schematic diagram illustrating the process of a food ingredient identification method provided in this application;
[0066] Figure 15 This is a schematic diagram of the structure of a food ingredient identification device provided in some embodiments of this application;
[0067] Figure 16 This is a schematic diagram of the structure of another food ingredient recognition device provided in some embodiments of this application;
[0068] Figure 17 A hardware structure diagram of a smart refrigerator provided for some embodiments of this application;
[0069] Figure 18 This is a schematic diagram of an electronic device structure provided for some embodiments of this application. Detailed Implementation
[0070] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0071] To improve the accuracy and automation of food ingredient identification, this application provides a food ingredient identification method, apparatus, equipment, and medium.
[0072] Figure 5 This application provides a schematic diagram of a method for identifying food ingredients, which includes the following steps:
[0073] S501: Acquire the first image to be identified by the main image acquisition device during the food storage and retrieval process, and identify the first confidence level corresponding to the first food type in the first image to be identified based on the pre-trained food identification model.
[0074] The food identification method provided in this application is applied to an electronic device, which can be a smart refrigerator or a server, wherein the server can be a local server or a cloud server.
[0075] If the electronic device is a smart refrigerator, it is equipped with a main image acquisition device located at the top of the refrigerator's crisper compartment. The smart refrigerator directly acquires the first image to be recognized from the main image acquisition device. If the electronic device is a server, the smart refrigerator's main image acquisition device, after acquiring the first image to be recognized, sends it to the server.
[0076] The food identification method provided in this application is dynamic identification. In dynamic identification, the first image to be identified is the image of the food in the user's hand that is captured by the image acquisition device during the process of storing and retrieving the food.
[0077] After the electronic device acquires the first image to be identified, it identifies the first image based on the pre-trained food identification model, and can obtain the first confidence level corresponding to the first food type of the input first image to be identified.
[0078] The recognition model can be either a feature extractor and classifier-based model, or a neural convolutional network model using deep learning methods. If it's a feature extractor and classifier-based model, the feature extractor can be any of the following: Histogram of Oriented Gradients (HOG), Local Binary Pattern (LBP), or Deformable Part Model (DPM). The classifier can be any of the following: Support Vector Machine (SVM), Adaboost, Decision Tree, Bayesian Network, or Neural Network. If the recognition model is a neural convolutional network model, it can be any of the following: Fast-CNN, Faster-CNN, YOLO, or YOLO 9000.
[0079] S502: If the first confidence level is less than the preset confidence threshold, then hand detection is performed on the first image to be identified to obtain the target hand position in the first image to be identified. Based on the target hand position and the pre-saved correspondence between the hand position and each auxiliary image acquisition device, each target auxiliary image acquisition device corresponding to the target hand position is determined, wherein the hand position includes left hand, right hand, front hand and back hand.
[0080] To determine whether the identified first food type is accurate, a preset confidence threshold is stored in advance. Based on the first confidence level and the preset confidence threshold, if the first confidence level is less than the preset confidence threshold, it indicates that the first food type identified based on the first image to be identified is inaccurate. This may be because the food was stored too quickly or the food was obscured by a hand. In order to accurately determine the target food type, the electronic device also needs to acquire a second image to be identified from other angles.
[0081] In order to acquire a second image to be identified from other angles, the electronic device pre-stores the correspondence between hand positions and each auxiliary image acquisition device. The hand positions include left hand, right hand, front hand, and back hand. A front hand means that the palm of the hand is facing up, and a back hand means that the back of the hand is facing up. The auxiliary image acquisition devices include image acquisition devices on both sides and bottom of the refrigerator compartment.
[0082] The electronic device performs hand detection on the first image to be identified. Specifically, it can use the existing human key point recognition algorithm OpenPose to detect the hand and obtain the target hand position in the first image to be identified. Based on the target hand position and the pre-saved correspondence between the hand position and each auxiliary image acquisition device, it determines each target auxiliary image acquisition device corresponding to the target hand position from the correspondence.
[0083] S503: Acquire each second image to be identified by each target auxiliary image acquisition device, identify each second food type in each second image to be identified based on the food identification model, and determine the identified target food type according to each second food type.
[0084] The electronic device acquires each second image to be identified acquired by each target auxiliary image acquisition device. The second image to be identified may be an image acquired at the time of the acquisition of the first image to be identified, or it may be an image acquired by each target auxiliary image acquisition device after each target auxiliary image acquisition device is determined. This application does not limit this.
[0085] After acquiring each second image to be identified, the electronic device identifies each second image based on a pre-trained food identification model. It can obtain each second food type corresponding to each input second image to be identified, and determine the final identified target food type based on the determined second food type.
[0086] Specifically, the target ingredient can be any one of the second ingredient categories, the ingredient category with the most repetitions, or other methods can be used to determine the target ingredient.
[0087] S504: If the first confidence level is not less than the preset confidence threshold, the first food type is determined as the identified target food type.
[0088] Based on the first confidence level and the preset confidence threshold, if the first confidence level is not less than the preset confidence threshold, it means that the first food type identified based on the first image to be identified is accurate, and therefore the first food type is directly determined as the target food type.
[0089] Because this application uses a pre-trained food identification model, when the confidence level of the first food type in the first image to be identified acquired by the main image acquisition device is less than a preset confidence threshold, each second image to be identified is acquired by each target auxiliary image acquisition device according to the target hand position in the first image to be identified. Each second food type in each second image to be identified is then identified, and the identified target food type is determined based on each second food type. This improves the accuracy of food identification and eliminates the need for manual input to supplement the food types, thus increasing the automation level of food identification.
[0090] In order to identify the first confidence level corresponding to the first food type in the first image to be identified, based on the above embodiments, in this application, the step of identifying the first confidence level corresponding to the first food type in the first image to be identified based on the pre-trained food identification model includes:
[0091] The first image to be identified is input into a pre-trained food identification model. Based on the food identification model, each prediction box, each corresponding prediction food type, and each corresponding prediction confidence level of the food in the first image to be identified are determined. Based on each prediction box and each corresponding prediction confidence level, a non-maximum suppression method is used to determine the target prediction box. The prediction food type corresponding to the target prediction box is determined as the first food type output, and the prediction confidence level corresponding to the target prediction box is determined as the first confidence level output.
[0092] In order to determine the first confidence level corresponding to the first food type in the first image to be identified, in this application, the first image to be identified is input into a pre-trained food identification model. Based on the food identification model, each prediction box of the food in the first image to be identified, each predicted food type corresponding to each prediction box, and each prediction confidence level are determined.
[0093] Figure 6 This application provides a schematic diagram of identifying each prediction box of food ingredients in a first image to be identified, as shown below. Figure 6 As shown, Figure 6 Each box in the image represents a prediction box for each food item. Figure 7 Another schematic diagram provided by this application for identifying each predicted bounding box of food ingredients in a first image to be identified, as shown below. Figure 7 As shown, Figure 7 Each box in the image represents a prediction box for each food item.
[0094] Based on each prediction box and its corresponding prediction confidence, a non-maximum suppression method is used to determine the target prediction box. The predicted food type corresponding to the target prediction box is determined as the first food type output, and the prediction confidence corresponding to the target prediction box is determined as the first confidence output.
[0095] Record each identified prediction box as B. c B c =[bbox c1 bbox c2 ,…,bbox cn Each predicted food category corresponding to each prediction box is Class = [c1, c2, ..., c]. n ], the prediction confidence S corresponding to each prediction box c =[s c1 ,s c2 ,…,s cn The highest confidence level S is determined based on each prediction confidence level. cmax , of which S cmax =max[s c1 ,s c2 ,…,s cn ], and set its corresponding prediction box as M. c M c =bbox i Calculate M c Other prediction boxes besides M c Overlap coverage IOU i If IOU i If the value is greater than the preset threshold n, then in the prediction box B cDelete the corresponding predicted bounding box. Repeat the above steps until the final target predicted bounding box is determined.
[0096] Figure 8 A method provided for this application from Figure 6 A schematic diagram of the target prediction box determined in the image, as shown below. Figure 8 As shown, Figure 8 The small box in the image is the target prediction box; Figure 9 A method provided for this application from Figure 7 A schematic diagram of the target prediction box determined in the image, as shown below. Figure 9 As shown, Figure 9 The small box in the image is the target prediction box.
[0097] To determine the target hand position in the first image to be identified, based on the above embodiments, the method in this application further includes:
[0098] Based on the food ingredient recognition model, the first area of the target prediction box is determined and output.
[0099] The step of performing hand detection on the first image to be identified to obtain the target hand position in the first image to be identified includes:
[0100] Hand detection is performed on the first image to be identified to determine the second area of the hand region in the first image to be identified, the coordinates of each bone node, and the corresponding preset number.
[0101] Determine whether the preset number corresponding to each bone node contains a first preset number and a second preset number;
[0102] If not, determine the ratio of the first area to the second area. If the ratio is not less than the first preset threshold, determine that the target hand position of the hand in the first image to be identified is a forward hand. If the ratio is less than the second preset threshold, determine that the target hand position of the hand in the first image to be identified is a reverse hand. Wherein the first preset threshold is greater than the second preset threshold.
[0103] If so, determine whether the first abscissa of the bone node corresponding to the first preset number is less than the second abscissa of the bone node corresponding to the second preset number. If so, determine that the target hand position of the hand in the first image to be identified is the left hand. If not, determine that the target hand position of the hand in the first image to be identified is the right hand.
[0104] In order to determine the target hand position in the first image to be identified, this application, based on the food recognition model, outputs the first area of the target prediction box after determining the target prediction box.
[0105] The electronic device performs hand detection on the first image to be recognized, determining the second area of the hand region in the first image to be recognized, the coordinates of each bone node, and its corresponding preset number; wherein the hand region includes 21 bone nodes, and the bone nodes can be divided into [point0, point2, ..., point...]. 20 ].
[0106] The first image to be identified in this application is a depth map. Based on the depth map, the three-dimensional coordinates of each bone node can be identified, and the three-dimensional coordinates can be converted into world coordinates in the world coordinate system. According to the position of each bone node in the hand area, it is possible to identify which bone node is on which finger and determine its corresponding preset number.
[0107] Figure 10 A schematic diagram of a hand region in a first image to be identified, as provided in this application, is shown below. Figure 10 As shown, Figure 10 The box in the middle represents the hand area. Figure 11 A schematic diagram of the hand region in another first image to be identified provided in this application, as shown below. Figure 11 As shown, Figure 11 The box in the middle represents the hand area. Figure 12 A schematic diagram of the hand region in another first image to be identified provided in this application, as shown below. Figure 12 As shown, Figure 12 The box in the middle represents the hand area.
[0108] Figure 13 A schematic diagram of a bone node provided in this application, such as Figure 13 As shown, Figure 13 Each dot in the diagram represents a bone node.
[0109] To determine whether the hand position is left-handed or right-handed, the electronic device pre-stores a first preset number and a second preset number. These two preset numbers are pre-set. When the first preset number is the number of the thumb tip, the second preset number can be the number of any other bone node. When the first preset number is the number of any bone node among the thumb nodes (i.e., any one of point2, point3, and point4), the second preset number is the number of any bone node among the index, middle, and ring fingers (i.e., point4). j ,j≥5.
[0110] Based on the preset number corresponding to each detected bone node, it is determined whether it contains both the first preset number and the second preset number. If it contains both, it can be determined whether the target hand position in the first image to be identified is the left hand or the right hand.
[0111] According to the first preset number point i The first abscissa x corresponding to the bone node i and the second preset number point j The second abscissa x corresponding to the bone node j , determine whether the first abscissa x i is less than the second abscissa x j . If x i <x j , it means that the target hand position of the hand is the left hand. If x i >x j , it means that the target hand position of the hand is the right hand.
[0112] If both are not included, it means that it is impossible to determine whether the target hand position of the hand in the first image to be recognized is the left hand position or the right hand position. Therefore, the electronic device determines whether the target hand position is the forward hand or the reverse hand. According to the first area bbox of the target prediction box foodi , the second area bbox of the hand area handj , determine the ratio rateArea of the first area to the second area, where rateArea = bbox foodi ÷bbox handj .
[0113] In order to determine whether the target hand position is the forward hand or the reverse hand, in this application, a first preset threshold a and a second preset threshold b are pre-saved. Determine whether the ratio rateArea is not less than the first preset threshold a. If rateArea≥a, it means that the target hand position is the forward hand. Determine whether the ratio rateArea is less than the second preset threshold b. If rateArea < b, it means that the target hand position is the reverse hand, where the first preset threshold a is greater than the second preset threshold b.
[0114] In order to determine the type of the target food ingredient recognized, based on the above embodiments, in this application, the determining the type of the target food ingredient recognized according to each second food ingredient type includes:
[0115] If each of the second food ingredient types is the same, any one of the second food ingredient types is determined as the type of the target food ingredient recognized.
[0116] In order to determine the type of the target food ingredient, in this application, if each of the second food ingredient types determined based on the second image to be recognized is the same, it means that the same kind of food ingredient is recognized. Therefore, any one of the second food ingredient types is determined as the type of the target food ingredient recognized.
[0117] In order to determine the type of the target food ingredient recognized, in this application, if each of the second food ingredient types is not the same, the recognizing each second food ingredient type in each second image to be recognized based on the food ingredient recognition model includes:
[0118] Based on the food ingredient recognition model, identify each type of second food ingredient and its corresponding second confidence level in each second image to be identified;
[0119] The step of determining the identified target ingredient type based on each of the second ingredient types includes:
[0120] Determine the target second confidence level among each second confidence level, and determine the second food type corresponding to the target second confidence level as the target food type.
[0121] In order to determine the target ingredient type, in this application, if each second ingredient type determined based on the second image to be identified is not uniform, that is, some of the second ingredient types are the same and some are different, or all of the second ingredient types are different, then based on the ingredient identification model, each second ingredient type in each second image to be identified and the corresponding second confidence level are identified.
[0122] Based on each second confidence level, the highest target second confidence level is determined, and the second ingredient type corresponding to the target second confidence level is determined and designated as the target ingredient type.
[0123] The process of determining the type of target food ingredient in this application is described below through a specific embodiment. In this application, the auxiliary image acquisition device Cam... position There are three types, namely the image acquisition device (cam) on the left side of the refrigerator compartment of the smart refrigerator. left The image acquisition device (cam) on the right side of the refrigerator compartment of the smart refrigerator right , and the image acquisition device at the bottom of the refrigerator compartment of the smart refrigerator (cam) bottom .
[0124] Table 1 shows the correspondence between a hand position and each auxiliary image acquisition device provided in this application.
[0125] As shown in Table 1:
[0126]
[0127] When the target hand position is determined to be a forehand, the corresponding auxiliary image acquisition device is the cam image acquisition device on the left side of the smart refrigerator. left Image acquisition device cam on the right side of the refrigerator compartment of the smart refrigerator right When the target hand position is determined to be the left hand, the auxiliary image acquisition device corresponding to the left hand is the cam image acquisition device on the right side of the refrigerator compartment. right and the image acquisition device at the bottom of the refrigerator compartment of the smart refrigerator (cam) bottomWhen the target hand position is determined to be right-handed, the auxiliary image acquisition device corresponding to the right hand is the cam image acquisition device on the left side of the refrigerator compartment of the smart refrigerator. left and the image acquisition device at the bottom of the refrigerator compartment of the smart refrigerator (cam) bottom When the target hand position is determined to be a backhand, the auxiliary image acquisition device corresponding to the backhand is the cam image acquisition device on the left side of the refrigerator compartment of the smart refrigerator. left The image acquisition device (cam) on the right side of the refrigerator compartment of the smart refrigerator right , and the image acquisition device at the bottom of the refrigerator compartment of the smart refrigerator (cam) bottom .
[0128] When the target hand position is determined to be a forehand, if the image acquisition device cam on the left side of the smart refrigerator... left Image acquisition device cam on the right side of the refrigerator compartment of the smart refrigerator right If the two second food categories identified in the two second images to be identified are the same food category, then that food category is determined as the target food category; if the two second food categories identified in the two second images to be identified are not the same food category, then the second food category with higher confidence level is determined as the target food category.
[0129] When the target hand position is determined to be the left hand, if the image acquisition device cam on the right side of the refrigerator compartment of the smart refrigerator... right and the image acquisition device at the bottom of the refrigerator compartment of the smart refrigerator (cam) bottom If the two second food categories identified in the two second images to be identified are the same food category, then that food category is determined as the target food category; if the two second food categories identified in the two second images to be identified are not the same food category, then the second food category with higher confidence level is determined as the target food category.
[0130] When the target hand position is determined to be right-handed, if the image acquisition device cam on the left side of the smart refrigerator's freezer compartment... left and the image acquisition device at the bottom of the refrigerator compartment of the smart refrigerator (cam) bottom If the two second food categories identified in the two second images to be identified are the same food category, then that food category is determined as the target food category; if the two second food categories identified in the two second images to be identified are not the same food category, then the second food category with higher confidence level is determined as the target food category.
[0131] When the target hand position is determined to be a reverse hand, the three-dimensional coordinates (Position) of the center point of the hand region in the first image to be identified are determined based on hand detection. hand And a pre-stored image acquisition device (cam) for the left side of the smart refrigerator's freezer compartment. leftThe three-dimensional coordinates of the smart refrigerator's right-side image acquisition device (cam) right Using three-dimensional coordinates, a target auxiliary image acquisition device is determined that is close to the center point of the hand area according to Euclidean distance. If the target auxiliary image acquisition device and the image acquisition device at the bottom of the smart refrigerator's freezer compartment are... bottom If the two second food categories identified in the two second images to be identified are the same food category, then that food category is determined as the target food category; if the two second food categories identified in the two second images to be identified are not the same food category, then the second food category with higher confidence level is determined as the target food category.
[0132] To train the food ingredient recognition model, based on the above embodiments, the process of training the food ingredient recognition model in this application includes:
[0133] Obtain any sample image from the sample set and the first label information corresponding to the sample image, wherein the first label information is used to identify the type of food contained in the sample image;
[0134] Using the original food identification model, each predicted bounding box, each predicted food type, and each predicted confidence level of the food in the sample image are determined. Based on each predicted bounding box and each predicted confidence level, a non-maximum suppression method is used to determine the target predicted bounding box. The predicted food type corresponding to the target predicted bounding box is determined as the second label information in the sample image.
[0135] Based on the first label information and the second label information, the parameter values of each parameter in the original food ingredient recognition model are adjusted.
[0136] In order to train the food identification model, this application stores a sample set for training. The sample images in the sample set include food images of each type of food during the picking and placing process. The first label information of the sample images in the sample set is manually pre-annotated, wherein the first label information is used to identify the type of food contained in the sample image.
[0137] In this application, after obtaining any sample image from the sample set and its first label information, the sample image is input into an original food ingredient recognition model. This model determines each predicted bounding box of the food ingredient in the sample image, the corresponding predicted food ingredient type, and the corresponding prediction confidence level. Based on each predicted bounding box and its corresponding prediction confidence level, a non-maximum suppression method is used to determine the target predicted bounding box. The predicted food ingredient type corresponding to the target predicted bounding box is output as the second label information of the sample image. The second label information identifies the food ingredient type of the sample image recognized by the original food ingredient recognition model.
[0138] After determining the second label information of the sample image based on the original food identification model, the original food identification model is trained based on the second label information and the first label information of the sample image to adjust the parameter values of various parameters of the original food identification model.
[0139] The above operation is performed on each sample image in the sample set used to train the original food ingredient recognition model. When a preset condition is met, a trained recognition model is obtained. This preset condition may be that the number of sample images in the sample set whose first label information matches their second label information after training with the original food ingredient recognition model is greater than a set number; or that the number of iterations for training the original food ingredient recognition model reaches a set maximum number of iterations, etc. Specifically, this application does not impose any limitations on this.
[0140] As one possible implementation, when training the original food ingredient recognition model, the sample images in the sample set can be divided into training sample images and test sample images. The original food ingredient recognition model is first trained based on the training sample images, and then the reliability of the trained food ingredient recognition model is tested based on the test sample images.
[0141] The following specific embodiment illustrates one of the food ingredient identification methods of this application. Figure 14 A schematic diagram illustrating the process of a food ingredient identification method provided in this application, such as... Figure 14 As shown, the method includes the following steps:
[0142] S1401: Acquire the first image to be recognized captured by the main image acquisition device during the food storage and retrieval process.
[0143] S1402: Based on the pre-trained food identification model, identify the first confidence level corresponding to the first food category in the first image to be identified.
[0144] S1403: Determine whether the first confidence level is less than the preset confidence threshold. If yes, proceed to S1404; otherwise, proceed to S1408.
[0145] S1404: Perform hand detection on the first image to be identified to obtain the target hand position in the first image to be identified.
[0146] S1405: Based on the target hand position and the pre-saved correspondence between the hand position and each auxiliary image acquisition device, determine each target auxiliary image acquisition device corresponding to the target hand position.
[0147] S1406: Control the main image acquisition device to turn off and each target auxiliary image acquisition device to turn on, and acquire each second image to be identified acquired by each target auxiliary image acquisition device.
[0148] S1407: Based on the food ingredient recognition model, identify each second food ingredient type in each second image to be identified, and determine the identified target food ingredient type according to each second food ingredient type, and proceed to S1409.
[0149] S1408: Control the main image acquisition device to shut down, and determine the first food type as the identified target food type.
[0150] S1409: Record the target ingredient type in the ingredient database.
[0151] Based on the above embodiments, Figure 15 This is a schematic diagram of the structure of a food ingredient identification device provided in some embodiments of this application. The device includes:
[0152] The recognition module 1501 is used to acquire the first image to be recognized acquired by the main image acquisition device during the food storage and retrieval process, and to identify the first confidence level corresponding to the first food type of the food in the first image to be recognized based on the pre-trained food recognition model.
[0153] The determination module 1502 is used to perform hand detection on the first image to be identified if the first confidence level is less than a preset confidence threshold, to obtain the target hand position of the hand in the first image to be identified, and to determine each target auxiliary image acquisition device corresponding to the target hand position according to the target hand position and the pre-saved correspondence between the hand position and each auxiliary image acquisition device, wherein the hand position includes left hand, right hand, front hand and back hand;
[0154] The recognition module 1501 is further configured to acquire each second image to be recognized acquired by each target auxiliary image acquisition device, identify each second food type in each second image to be recognized based on the food recognition model, and determine the identified target food type according to each second food type.
[0155] Furthermore, the determining module is also used to determine the first food ingredient type as the identified target food ingredient type if the first confidence level is not less than a preset confidence threshold.
[0156] Furthermore, the recognition module is specifically used to input the first image to be recognized into a pre-trained food recognition model, and based on the food recognition model, determine each prediction box of the food in the first image to be recognized, each corresponding prediction food type, and each corresponding prediction confidence level. Based on each prediction box and each corresponding prediction confidence level, a non-maximum suppression method is used to determine a target prediction box, the prediction food type corresponding to the target prediction box is determined as the first food type output, and the prediction confidence level corresponding to the target prediction box is determined as the first confidence level output.
[0157] Furthermore, the recognition module is also used to determine the first area output of the target prediction box based on the food ingredient recognition model;
[0158] The determining module is specifically used to perform hand detection on the first image to be identified, determine the second area of the hand region in the first image to be identified, the coordinates of each bone node, and the corresponding preset number; determine whether the preset number corresponding to each bone node contains a first preset number and a second preset number; if not, determine the ratio of the first area to the second area; if the ratio is not less than a first preset threshold, determine that the target hand position of the hand in the first image to be identified is a right-handed hand; if the ratio is less than a second preset threshold, determine that the target hand position of the hand in the first image to be identified is a left-handed hand, wherein the first preset threshold is greater than the second preset threshold; if yes, determine whether the first abscissa of the bone node corresponding to the first preset number is less than the second abscissa of the bone node corresponding to the second preset number; if yes, determine that the target hand position of the hand in the first image to be identified is a left-handed hand; if not, determine that the target hand position of the hand in the first image to be identified is a right-handed hand.
[0159] Furthermore, the identification module is specifically used to identify any second ingredient type as the target ingredient type if all the second ingredient types are the same.
[0160] Furthermore, the recognition module is specifically used to identify each second food type and its corresponding second confidence level in each second image to be recognized based on the food recognition model; determine the largest target second confidence level among the second confidence levels, and determine the second food type corresponding to the target second confidence level as the target food type.
[0161] Furthermore, the device also includes:
[0162] The training module is used to acquire any sample image in the sample set and the first label information corresponding to the sample image, wherein the first label information is used to identify the types of food contained in the sample image; through the original food recognition model, it determines each prediction box of the food in the sample image, each corresponding predicted food type, and each corresponding prediction confidence; based on each prediction box and each corresponding prediction confidence, it uses a non-maximum suppression method to determine the target prediction box, and determines the predicted food type corresponding to the target prediction box as the second label information in the sample image; based on the first label information and the second label information, it adjusts the parameter values of each parameter in the original food recognition model.
[0163] The following describes a specific embodiment of the food ingredient identification device of this application. Figure 16 This is a schematic diagram of the structure of another food ingredient recognition device provided in some embodiments of this application, such as... Figure 16 As shown, the food identification device includes a food detection and identification module 1601 and a main and auxiliary camera collaboration module 1602.
[0164] The food ingredient detection and recognition module 1601 is used to acquire a first image to be recognized acquired by the main image acquisition device during the food ingredient storage and retrieval process, and to identify the first confidence level corresponding to the first food ingredient type in the first image to be recognized based on a pre-trained food ingredient recognition model; to acquire each second image to be recognized acquired by each target auxiliary image acquisition device, and to identify each second food ingredient type in each second image to be recognized based on the food ingredient recognition model, and to determine the identified target food ingredient type according to each second food ingredient type, which is equivalent to the recognition module 1501 in the above embodiment.
[0165] The main and auxiliary camera collaboration module 1602 is used to perform hand detection on the first image to be identified if the first confidence level is less than a preset confidence threshold, to obtain the target hand position in the first image to be identified, and to determine each target auxiliary image acquisition device corresponding to the target hand position based on the target hand position and the pre-saved correspondence between the hand position and each auxiliary image acquisition device; equivalent to the determination module 1502 in the above embodiment.
[0166] The following describes a specific embodiment of the food ingredient identification device of this application. Figure 17 A hardware structure diagram of a smart refrigerator is provided for some embodiments of this application, such as... Figure 17 As shown, the smart refrigerator includes a main camera, a processor, an auxiliary camera, a physical protective cover, a fruit and vegetable section, a fresh food section, and a dry goods section.
[0167] The main camera is a binocular camera with higher resolution and frame rate than the auxiliary camera. It is located at the top of the refrigerator compartment of the smart refrigerator. As the smart refrigerator door opens, the physical protective cover of the main camera slides to the left or right and begins to capture the first image to be identified of the food, hands and position information during the storage and retrieval process. The first image to be identified is then sent to the processor, which is equivalent to the main image acquisition device in the above embodiment.
[0168] The auxiliary camera, located on both sides of the refrigerator compartment and below the fruit and vegetable section, is used to acquire a second image to be identified and send it to the processor, which is equivalent to the auxiliary image acquisition device in the above embodiment.
[0169] The processor, located at the top of the cold storage compartment, is used to receive signals from the door and video streams input from the main and auxiliary cameras, control the switching of the main and auxiliary cameras and the physical protective cover, and support the calculation of underlying algorithm logic and the execution of commands.
[0170] Figure 18 The present application provides a schematic diagram of an electronic device structure based on some embodiments. In addition to the above embodiments, the present application also provides an electronic device, including a processor 1801, a communication interface 1802, a memory 1803 and a communication bus 1804, wherein the processor 1801, the communication interface 1802 and the memory 1803 communicate with each other through the communication bus 1804.
[0171] The memory 1803 stores a computer program, which, when executed by the processor 1801, causes the processor 1801 to perform the following steps:
[0172] The first image to be identified is acquired by the main image acquisition device during the food storage and retrieval process. Based on the pre-trained food identification model, the first confidence level corresponding to the first food type in the first image to be identified is determined.
[0173] If the first confidence level is less than the preset confidence threshold, then hand detection is performed on the first image to be identified to obtain the target hand position in the first image to be identified. Based on the target hand position and the pre-saved correspondence between the hand position and each auxiliary image acquisition device, each target auxiliary image acquisition device corresponding to the target hand position is determined, wherein the hand position includes left hand, right hand, front hand and back hand.
[0174] Acquire each second image to be identified by each target auxiliary image acquisition device, identify each second food type in each second image to be identified based on the food identification model, and determine the identified target food type based on each second food type.
[0175] The communication bus mentioned in the above electronic devices can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into address bus, data bus, control bus, etc. For ease of illustration, only one thick line is used to represent it in the diagram, but this does not mean that there is only one bus or one type of bus.
[0176] Communication interface 1802 is used for communication between the above-mentioned electronic device and other devices.
[0177] The memory may include random access memory (RAM) or non-volatile memory (NVM), such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.
[0178] The processors mentioned above can be general-purpose processors, including central processing units, network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits, field-programmable gate arrays or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.
[0179] For the concepts, explanations, detailed descriptions, and other steps related to the technical solutions provided in this application, please refer to the descriptions of these contents in the foregoing methods or other embodiments, which will not be repeated here.
[0180] Based on the above embodiments, this application also provides a computer-readable storage medium storing a computer program, which is executed by a processor in the following steps:
[0181] The first image to be identified is acquired by the main image acquisition device during the food storage and retrieval process. Based on the pre-trained food identification model, the first confidence level corresponding to the first food type in the first image to be identified is determined.
[0182] If the first confidence level is less than the preset confidence threshold, then hand detection is performed on the first image to be identified to obtain the target hand position in the first image to be identified. Based on the target hand position and the pre-saved correspondence between the hand position and each auxiliary image acquisition device, each target auxiliary image acquisition device corresponding to the target hand position is determined, wherein the hand position includes left hand, right hand, front hand and back hand.
[0183] Acquire each second image to be identified by each target auxiliary image acquisition device, identify each second food type in each second image to be identified based on the food identification model, and determine the identified target food type based on each second food type.
[0184] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0185] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0186] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0187] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0188] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.
Claims
1. A method for identifying food ingredients, characterized in that, The method includes: The first image to be identified is acquired by the main image acquisition device during the food storage and retrieval process. Based on the pre-trained food identification model, the first confidence level corresponding to the first food type in the first image to be identified is determined. If the first confidence level is less than the preset confidence threshold, then hand detection is performed on the first image to be identified to obtain the target hand position in the first image to be identified. Based on the target hand position and the pre-saved correspondence between the hand position and each auxiliary image acquisition device, each target auxiliary image acquisition device corresponding to the target hand position is determined, wherein the hand position includes left hand, right hand, front hand and back hand. Acquire each second image to be identified acquired by each target auxiliary image acquisition device, identify each second food type in each second image to be identified based on the food identification model, and determine the identified target food type based on each second food type; The step of performing hand detection on the first image to be identified to obtain the target hand position in the first image to be identified includes: Hand detection is performed on the first image to be identified to determine the second area of the hand region in the first image to be identified, the coordinates of each bone node, and the corresponding preset number. Determine whether the preset number corresponding to each bone node contains a first preset number and a second preset number. When the first preset number is the fingertip number of the thumb, the second preset number is the number of any other bone node. When the first preset number is the number of any bone node among the thumb nodes, the second preset number is the number of any bone node among the index finger, middle finger, and ring finger. If not, determine the ratio of the first area to the second area. If the ratio is not less than the first preset threshold, determine that the target hand position of the hand in the first image to be identified is a forward hand. If the ratio is less than the second preset threshold, determine that the target hand position of the hand in the first image to be identified is a reverse hand. The first preset threshold is greater than the second preset threshold. The first area is the area of the target prediction box corresponding to the first food type. If so, determine whether the first abscissa of the bone node corresponding to the first preset number is less than the second abscissa of the bone node corresponding to the second preset number. If so, determine that the target hand position of the hand in the first image to be identified is the left hand. If not, determine that the target hand position of the hand in the first image to be identified is the right hand.
2. The method according to claim 1, characterized in that, If the first confidence level is not less than a preset confidence threshold, the method further includes: The first food ingredient type is identified as the target food ingredient type.
3. The method according to claim 1, characterized in that, The first confidence level corresponding to the first food category in the first image to be identified, based on the pre-trained food identification model, includes: The first image to be identified is input into a pre-trained food identification model. Based on the food identification model, each prediction box, each corresponding prediction food type, and each corresponding prediction confidence level of the food in the first image to be identified are determined. Based on each prediction box and each corresponding prediction confidence level, a non-maximum suppression method is used to determine the target prediction box. The prediction food type corresponding to the target prediction box is determined as the first food type output, and the prediction confidence level corresponding to the target prediction box is determined as the first confidence level output.
4. The method according to claim 1, characterized in that, The step of determining the identified target ingredient type based on each of the second ingredient types includes: If each of the second ingredient types is the same, then any second ingredient type is identified as the target ingredient type.
5. The method according to claim 4, characterized in that, If each of the second food ingredients is not identical, the step of identifying each second food ingredient in each second image to be identified based on the food ingredient recognition model includes: Based on the food ingredient recognition model, identify each type of second food ingredient and its corresponding second confidence level in each second image to be identified; The step of determining the identified target ingredient type based on each of the second ingredient types includes: Determine the target second confidence level among each second confidence level, and determine the second food type corresponding to the target second confidence level as the target food type.
6. The method according to claim 1, characterized in that, The process of training the food ingredient recognition model includes: Obtain any sample image from the sample set and the first label information corresponding to the sample image, wherein the first label information is used to identify the type of food contained in the sample image; Using the original food identification model, each predicted bounding box, each predicted food type, and each predicted confidence level of the food in the sample image are determined. Based on each predicted bounding box and each predicted confidence level, a non-maximum suppression method is used to determine the target predicted bounding box. The predicted food type corresponding to the target predicted bounding box is determined as the second label information in the sample image. Based on the first label information and the second label information, the parameter values of each parameter in the original food ingredient recognition model are adjusted.
7. A food ingredient identification device, characterized in that, The device includes: The recognition module is used to acquire the first image to be recognized captured by the main image acquisition device during the food storage and retrieval process, and to identify the first confidence level corresponding to the first food type in the first image to be recognized based on the pre-trained food recognition model. The determination module is used to perform hand detection on the first image to be identified if the first confidence level is less than a preset confidence threshold, to obtain the target hand position of the hand in the first image to be identified, and to determine each target auxiliary image acquisition device corresponding to the target hand position according to the target hand position and the pre-saved correspondence between the hand position and each auxiliary image acquisition device, wherein the hand position includes left hand, right hand, front hand and back hand; The recognition module is also used to acquire each second image to be recognized acquired by each target auxiliary image acquisition device, identify each second food type in each second image to be recognized based on the food recognition model, and determine the identified target food type according to each second food type. The determining module is specifically used to perform hand detection on the first image to be identified, determine the second area of the hand region in the first image to be identified, the coordinates of each bone node, and the corresponding preset number; determine whether the preset number corresponding to each bone node contains a first preset number and a second preset number, wherein, when the first preset number is the fingertip number of the thumb, the second preset number is the number of any other bone node; when the first preset number is the number of any bone node among the thumb nodes, the second preset number is the number of any bone node among the index, middle, and ring fingers; if not, determine the ratio of the first area to the second area, and if the ratio is not small... If the ratio is less than a first preset threshold, the target hand position in the first image to be identified is determined to be a right hand. If the ratio is less than a second preset threshold, the target hand position in the first image to be identified is determined to be a left hand. The first preset threshold is greater than the second preset threshold. The first area is the area of the target prediction box corresponding to the first food type. If the ratio is less than the second preset threshold, it is determined whether the first abscissa of the bone node corresponding to the first preset number is less than the second abscissa of the bone node corresponding to the second preset number. If the ratio is less than the second preset threshold, the target hand position in the first image to be identified is determined to be a left hand. If the ratio is less than the second preset threshold, the target hand position in the first image to be identified is determined to be a right hand.
8. An electronic device, characterized in that, include: The processor, communication interface, memory, and communication bus are connected, with the processor, communication interface, and memory communicating with each other via the communication bus. The memory stores a computer program that, when executed by the processor, causes the processor to perform the method according to any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, It stores a computer program executable by a processor, which, when run on the processor, causes the processor to perform the method according to any one of claims 1-6.
Citation Information
Patent Citations
Food material management device and method
CN108154078A
Deep-learning-based food material identification system and food material identification method
CN108846314A