A method for estimating the weight of a refrigerator, a server, and food ingredients.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-03-31
- Publication Date
- 2026-08-14
AI Technical Summary
[0005]本申请提供一种冰箱、服务器及食材重量估计方法,以解决食材重量识别效率较低的问题
[0025]本实施例中,接收图像采集设备采集的视觉信息,对视觉信息中目标食材进行识别,获得目标食材对应的目标食材类别及目标食材对应的目标食材像素尺寸,确定目标食材对应的目标距离表征值,目标距离表征值用于表征目标食材与图像采集设备之间的距离,基于目标食材类别、目标食材像素尺寸及目标距离表征值对目标食材的重量进行估计,获得目标食材的重量估计值。从而简化了确定重量的过程,确定重量的复杂度较低,耗时较少,从而可以提高食材重量识别效率。
Smart Images

Figure CN122566471A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of refrigerator technology, and in particular to a refrigerator, a server, and a method for estimating the weight of food. Background Technology
[0002] With the development of intelligent AI technology, smart refrigerators are becoming increasingly intelligent and have more and more functions, such as supporting the recognition of food weight.
[0003] Currently, refrigerators typically first identify the volume of food and then determine its weight based on that volume.
[0004] However, the process of recognizing the volume of food is relatively complex and time-consuming, resulting in low efficiency in recognizing the weight of food. Summary of the Invention
[0005] This application provides a refrigerator, a server, and a method for estimating the weight of food ingredients to solve the problem of low efficiency in identifying the weight of food ingredients.
[0006] In a first aspect, some embodiments provide a refrigerator, including: an image acquisition device for acquiring visual information; and a controller configured to: receive the visual information acquired by the image acquisition device; identify a target food item in the visual information to obtain a target food item category and a target food item pixel size corresponding to the target food item; determine a target distance representation value corresponding to the target food item, the target distance representation value being used to represent the distance between the target food item and the image acquisition device; and estimate the weight of the target food item based on the target food item category, the target food item pixel size, and the target distance representation value to obtain a weight estimate of the target food item.
[0007] In this embodiment, visual information acquired by an image acquisition device is received, and the target food ingredient in the visual information is identified to obtain the target food ingredient category and the corresponding pixel size. A target distance representation value is determined, which represents the distance between the target food ingredient and the image acquisition device. Based on the target food ingredient category, pixel size, and target distance representation value, the weight of the target food ingredient is estimated to obtain a weight estimate. This simplifies the weight determination process, reduces complexity and time consumption, and thus improves the efficiency of food ingredient weight recognition.
[0008] In some embodiments, when the controller executes the step of determining the target distance representation value corresponding to the target food ingredient, it is configured to: determine the target area identifier corresponding to the target storage area to which the target food ingredient belongs; determine the area height correspondence of the refrigerator, the area height correspondence including the distance representation values corresponding to multiple reference area identifiers in the refrigerator, the reference area identifiers corresponding one-to-one with the storage areas in the refrigerator, and the distance representation value corresponding to the reference area identifier reflecting the distance between the storage area corresponding to the reference area identifier and the image acquisition device; match the target area identifier with the reference area identifiers in the area height correspondence; and determine the distance representation value corresponding to the successfully matched reference area identifier as the target distance representation value corresponding to the target food ingredient.
[0009] In this embodiment, the region height correspondence includes distance representation values corresponding to multiple reference region identifiers inside the refrigerator. Each reference region identifier corresponds one-to-one with a storage region inside the refrigerator. By matching the target region identifier with the reference region identifier in the region height correspondence, the distance representation value corresponding to the storage region to which the target food belongs can be determined, and the target distance representation value corresponding to the target food can be obtained. The method of determining the target distance representation value corresponding to the target food has low complexity and high efficiency.
[0010] In some embodiments, when the controller performs the process of estimating the weight of the target ingredient based on the target ingredient category, the target ingredient pixel size, and the target distance representation value to obtain the weight estimate of the target ingredient, it is configured to: determine a distance-weight correspondence associated with the target ingredient category, the distance-weight correspondence including the correspondence between the distance representation value and the weight per unit area; match the target distance representation value with the distance representation value in the distance-weight correspondence; and determine the weight estimate of the target ingredient based on the weight per unit area corresponding to the successfully matched distance representation value and the ingredient pixel size.
[0011] In this embodiment, since the distance-weight correspondence includes the correspondence between the distance representation value and the weight per unit area, the weight per unit area corresponding to the target distance representation value can be determined. Then, based on the weight per unit area corresponding to the target distance representation value, the weight estimate of the target food can be determined. The process of calculating the weight estimate is simple and efficient, thereby improving the efficiency of weight estimation.
[0012] In some embodiments, the weight estimate is determined by a weight prediction model, which is obtained by training a prediction model to be trained using training samples corresponding to the sample food ingredients. The training samples include the true category of the sample food ingredients, the pixel size of the sample food ingredients, a sample distance representation value, and the true weight of the sample food ingredients. The pixel size of the sample food ingredients is the size of the sample food ingredients identified from a sample image taken when the sample food ingredients are located in a refrigerator. The sample distance representation value reflects the distance between the sample food ingredients and the image acquisition device when the sample image was taken.
[0013] In this embodiment, the weight is estimated by using a trained quality prediction model, which can accurately and quickly estimate the weight of the target food, improving the efficiency and accuracy of weight estimation.
[0014] In some embodiments, when the controller performs the step of identifying the target food ingredient in the visual information to obtain the target food ingredient category and the target food ingredient pixel size, it is configured to: determine the food image region of the target food ingredient from the target image in the visual information, wherein the target food ingredient is food taken from the refrigerator, the target image contains the target food ingredient, and the food image region contains the target food ingredient; perform food identification on the food image region of the target food ingredient to obtain the target food ingredient category and the target food ingredient pixel size.
[0015] In this embodiment, since the target food is taken from the refrigerator, the weight of the food taken from the refrigerator can be estimated.
[0016] In some embodiments, when the controller performs the step of determining the target distance representation value corresponding to the target food ingredient, it is configured to: perform region recognition on the target image in the visual information to obtain region images corresponding to multiple region identifiers, wherein the region identifiers represent storage areas in the refrigerator; determine the target region image to which the target food ingredient belongs from the region images corresponding to the multiple region identifiers; and use the distance representation value associated with the region identifier corresponding to the target region image as the target distance representation value corresponding to the target food ingredient.
[0017] In this embodiment, the target region image to which the target ingredient belongs is determined from the region images corresponding to multiple region identifiers. The distance representation value associated with the region identifier corresponding to the target region image is used as the target distance representation value corresponding to the target ingredient. This allows for rapid determination of the target distance representation value.
[0018] In some embodiments, the target food is food taken out by the user from the refrigerator, and the controller is further configured to: acquire unit nutrient information corresponding to the category of the target food; generate nutritional information of the target food based on the unit nutrient information and the weight estimate; the controller is further configured to: analyze the nutritional information of each target food identified within a target time period to obtain the nutritional information analysis result corresponding to the target time period; and generate food removal information corresponding to the target time period based on each target food identified within the target time period.
[0019] In this embodiment, the nutritional information of each target ingredient identified within the target time period is analyzed to obtain the nutritional information analysis results corresponding to the target time period. Based on each target ingredient identified within the target time period, the ingredient retrieval status information corresponding to the target time period is generated, providing a basis for providing feedback to the user on the ingredient retrieval status and helping to improve the user experience.
[0020] Secondly, some embodiments also provide a server configured to: receive visual information acquired by an image acquisition device; identify target food items in the visual information to obtain the target food item category and the target food item pixel size; determine a target distance representation value corresponding to the target food item, the target distance representation value being used to represent the distance between the target food item and the image acquisition device; and estimate the weight of the target food item based on the target food item category, the target food item pixel size, and the target distance representation value to obtain a weight estimate of the target food item.
[0021] In this embodiment, visual information acquired by an image acquisition device is received, and the target food ingredient in the visual information is identified to obtain the target food ingredient category and the corresponding pixel size. A target distance representation value is determined, which represents the distance between the target food ingredient and the image acquisition device. Based on the target food ingredient category, pixel size, and target distance representation value, the weight of the target food ingredient is estimated to obtain a weight estimate. This simplifies the weight determination process, reduces complexity and time consumption, and thus improves the efficiency of food ingredient weight recognition.
[0022] In some embodiments, the weight estimate is determined by a weight prediction model, and the server is further configured to: acquire training samples corresponding to the sample food, the training samples containing the true category of the sample food, the pixel size of the sample food, the sample distance representation value, and the true weight of the sample food, the pixel size of the sample food being the size of the sample food identified from a sample image taken with the sample food in a refrigerator, and the sample distance representation value reflecting the distance between the sample food and the image acquisition device when the sample image was taken; input the training samples into a prediction model to be trained for weight prediction to obtain the predicted weight of the sample food; and adjust the model parameters of the prediction model based on the difference between the true weight of the sample food and the predicted weight of the sample food to obtain a weight prediction model.
[0023] In this embodiment, since the training samples include the true category of the sample ingredients, the pixel size of the sample ingredients, the sample distance representation value, and the true weight of the sample ingredients, the pixel size of the sample ingredients is the size of the sample ingredients identified from the sample images, the sample images are taken when the sample ingredients are in a refrigerator, and the sample distance representation value is used to reflect the distance between the sample ingredients and the image acquisition device when the sample image is taken, the prediction model can be trained using the training samples, so that the prediction model can learn the ability to predict weight based on category, size, and distance representation value, thereby ensuring the reliability of the weight prediction model.
[0024] Thirdly, some embodiments also provide a method for estimating the weight of food ingredients, applied to a server or refrigerator. The method includes: receiving visual information acquired by an image acquisition device; identifying target food ingredients in the visual information to obtain the target food ingredient category and the target food ingredient pixel size; determining a target distance representation value corresponding to the target food ingredient, the target distance representation value being used to represent the distance between the target food ingredient and the image acquisition device; and estimating the weight of the target food ingredient based on the target food ingredient category, the target food ingredient pixel size, and the target distance representation value to obtain an estimated weight value of the target food ingredient.
[0025] In this embodiment, visual information acquired by an image acquisition device is received, and the target food ingredient in the visual information is identified to obtain the target food ingredient category and the corresponding pixel size. A target distance representation value is determined, which represents the distance between the target food ingredient and the image acquisition device. Based on the target food ingredient category, pixel size, and target distance representation value, the weight of the target food ingredient is estimated to obtain a weight estimate. This simplifies the weight determination process, reduces complexity and time consumption, and thus improves the efficiency of food ingredient weight recognition. Attached Figure Description
[0026] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0027] Figure 1 Schematic diagrams of refrigerator structures provided for some embodiments of this application;
[0028] Figure 2 Schematic diagrams of the refrigerator structure provided for other embodiments of this application;
[0029] Figure 3 A schematic diagram of the hardware configuration of the controller and its associated devices provided in some embodiments of this application;
[0030] Figure 4 A flowchart illustrating the food weight estimation method provided in some embodiments of this application;
[0031] Figure 5 A schematic diagram of images captured by an image acquisition device provided in some embodiments of this application;
[0032] Figure 6 A schematic diagram illustrating the distance-weight correspondence provided in some embodiments of this application;
[0033] Figure 7 A schematic diagram of the training sample set provided for some embodiments of this application;
[0034] Figure 8 Schematic diagram of unit nutrient information provided in some embodiments of this application;
[0035] Figure 9 A schematic diagram illustrating the nutritional information of the target food ingredients provided in some embodiments of this application;
[0036] Figure 10A flowchart illustrating a nutritional information identification method based on food weight estimation provided in some embodiments of this application;
[0037] Figure 11 A system architecture diagram of a nutritional information identification method based on food weight estimation provided in some embodiments of this application;
[0038] Figure 12 Flowcharts for food ingredient identification provided in some embodiments of this application;
[0039] Figure 13 A timing diagram illustrating the food weight estimation method provided in some embodiments of this application;
[0040] Figure 14 This is a structural block diagram of a food weight estimation device provided in some embodiments of this application.
[0041] Figure 15 This is an internal structural diagram of a computer device provided in some embodiments of this application. Detailed Implementation
[0042] The embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described below do not represent all embodiments consistent with this application. They are merely examples of systems and methods consistent with some aspects of this application as detailed in the claims.
[0043] It should be noted that the brief descriptions of terms in this application are only for the convenience of understanding the embodiments described below, and are not intended to limit the embodiments of this application. Unless otherwise stated, these terms should be understood in their ordinary and common meaning.
[0044] The terms "first," "second," "third," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar or related objects or entities, and do not necessarily imply a specific order or sequence, unless otherwise specified. It should be understood that such terms are interchangeable where appropriate.
[0045] The terms “comprising” and “having”, and any variations thereof, are intended to cover but not exclude inclusion, for example, a product or device that includes a range of components is not necessarily limited to all of the components that are clearly listed, but may include other components that are not clearly listed or that are inherent to such product or device.
[0046] The term "module" refers to any known or subsequently developed hardware, software, firmware, artificial intelligence, fuzzy logic, or combination of hardware and / or software code that is capable of performing the functions associated with that element.
[0047] Figure 1 This is a schematic diagram of the structure of a refrigerator 100 provided in an embodiment of this application. The refrigerator 100 in this embodiment has an approximately rectangular shape. The refrigerator includes a cabinet defining a storage space and one or more doors, such as door bodies 101, located at the opening of the cabinet. Each door body includes a door shell located outside the cabinet, a door inner liner located inside the cabinet, an upper end cover, a lower end cover, and an insulation layer located between the door shell, door inner liner, upper end cover, and lower end cover. Typically, the insulation layer is filled with foam material. The cabinet has chambers, including component storage chambers for placing refrigerator components, such as a compressor compartment, and storage chambers for storing food, etc. Of course, the refrigerator in this application can also have other shapes; this application does not limit the external structure of the refrigerator.
[0048] The refrigerator interior includes an image acquisition device 104 located in the storage compartment 102, drawer 103, and on top. Depending on its purpose, the storage compartment can be configured as a refrigerator compartment, freezer compartment, variable temperature compartment, vacuum drawer, or humidifier drawer, etc. The storage compartment can include multiple storage areas, for example, in... Figure 1 In this refrigerator, the storage compartment 102 is divided into five storage areas by four partitions: partition 105, partition 106, partition 107, and partition 108. The five storage areas are: a first storage area between partition 105 and the top of the storage compartment; a second storage area between partition 105 and partition 106; a third storage area between partition 106 and partition 107; a fourth storage area between partition 107 and partition 108; and a fifth storage area between partition 108 and the bottom of the storage compartment. The image acquisition device can be, but is not limited to, a camera. The refrigerator in this application can also be of other shapes, and this application does not limit the external structure of the refrigerator. Figure 2 As shown, a schematic diagram of another refrigerator structure is provided.
[0049] like Figure 3As shown, the refrigerator may also include at least one of the following: a controller 200, a display 201, a communication device, a power supply, a memory, and a user input interface. The controller is the intelligent core of the refrigerator, responsible for managing its operating status. It monitors the refrigerator's sensor data and makes adjustments based on the operating environment. Especially under fault or abnormal conditions, the controller can control the compressor's operation according to set logic. Both the controller and the compressor are located inside the refrigerator. The controller can be a microcontroller unit (MCU) or other types of controllers. The communication device is a component used to communicate with external devices or servers according to various communication protocols. The communication device may include a WiFi module, a Bluetooth module, and / or an Ethernet module. The power supply provides power to the various components of the controller under its control. The memory can store various operating programs, data, and applications under the controller's control, and can also store various control signal commands input by the user. The user can input user commands through a graphical user interface (GUI) displayed on the display, and the user input interface can receive user-input commands through the graphical user interface. The controller may include a central processing unit, RAM (Random Access Memory), ROM (Read-Only Memory), a video processor and / or a graphics processor, and may also include interfaces for input / output, such as a first interface and a second interface.
[0050] The refrigerator also includes an integrated main inverter and display board, which is located inside the refrigerator body. The controller can be housed on this integrated board. The main inverter and display board includes a controller, power filter circuit, rectifier components, voltage detection circuit, three-phase inverter circuit, drive circuit, current sampling circuit, memory, voltage analog-to-digital converter module, pulse width modulation signal output module, temperature analog-to-digital converter module, operational amplifier, key detection circuit, display drive circuit, display module, fan, drive circuit, and fan interface. The power filter circuit stabilizes the DC voltage using the energy storage and release characteristics of capacitors. The rectifier components convert AC to DC. The three-phase inverter circuit converts DC to three-phase AC, providing a suitable three-phase AC power supply for the compressor.
[0051] The voltage detection circuit primarily detects the bus voltage using voltage divider resistors and sends the detected voltage to the voltage-to-digital converter (ADC). The ADC converts the received voltage into a voltage signal, enabling the controller to acquire it. The current sampling circuit samples the DC bus current and sends the sampled current to an operational amplifier. The operational amplifier processes the sampled current and sends the processed current to the current-to-digital converter (ADC). The ADC converts the received current into a current signal, allowing the controller to acquire it.
[0052] The controller analyzes and processes the digital current signal to obtain a Pulse Width Modulation (PWM) signal for controlling the compressor's operation. This PWM signal is then sent to the drive circuit via a PWM signal output module. The drive circuit uses the PWM signal to control the output of the three-phase inverter circuit, thereby controlling the compressor's operating state. The memory stores information such as the refrigerator's settings; however, this embodiment does not specifically limit the information that the memory can store. The temperature analog-to-digital converter module converts the temperature collected by the temperature sensor into a temperature signal, enabling the controller to obtain the temperature signal from the temperature analog-to-digital converter module.
[0053] The button detection circuit monitors the button status in real time and adjusts the refrigerator's settings and control modes accordingly. The display driver circuit drives the display module to display the settings and mode information. It should be noted that the buttons are located on the refrigerator itself, allowing users to adjust settings such as temperature. The fan driver operates the refrigerator's fan via a fan interface. The controller receives information from the button detection circuit through an interface and transmits data to the display driver circuit and fan driver circuit via the same interface.
[0054] Based on this, in some embodiments, this application provides a refrigerator, which includes: an image acquisition device for acquiring visual information; and a controller configured to: receive the visual information acquired by the image acquisition device; identify target food items in the visual information to obtain the target food item category and the target food item pixel size; determine a target distance representation value corresponding to the target food item, the target distance representation value being used to represent the distance between the target food item and the image acquisition device; and estimate the weight of the target food item based on the target food item category, the target food item pixel size, and the target distance representation value to obtain an estimated weight value of the target food item.
[0055] In this embodiment, visual information acquired by an image acquisition device is received, and the target food ingredient in the visual information is identified to obtain the target food ingredient category and the corresponding pixel size. A target distance representation value is determined, which represents the distance between the target food ingredient and the image acquisition device. Based on the target food ingredient category, pixel size, and target distance representation value, the weight of the target food ingredient is estimated to obtain a weight estimate. This simplifies the weight determination process, reduces complexity and time consumption, and thus improves the efficiency of food ingredient weight recognition.
[0056] Based on this, in some embodiments, this application provides a method for estimating the weight of food ingredients, applied to a refrigerator or server, wherein the refrigerator includes, for example... Figure 4 As shown, the method includes:
[0057] Step 402: Receive visual information acquired by the image acquisition device.
[0058] The visual information can be images or videos. The image acquisition device is located on top of the refrigerator, and its field of view includes a portion of the refrigerator's interior and the front, resulting in an image that resembles a top-down view of the refrigerator. If the refrigerator includes drawers, the image acquisition device cannot capture the drawers when they are closed, as they are hidden inside the refrigerator; however, it can capture the drawers when they are open.
[0059] Step 404: Identify the target food in the visual information to obtain the target food category and the target food pixel size.
[0060] The visual information includes images, and the food pixel size refers to the size of the food within the image. The target food can be any food in the visual information; for example, it could be food taken out of the refrigerator or food being placed into the refrigerator. The target food pixel size refers to the pixel dimensions of the target food. The food pixel size includes width and height, both measured in pixels.
[0061] For example, a target food ingredient can be located from visual information, and its storage area in the refrigerator can be determined. The target food ingredient is food taken from the refrigerator. A food image region containing the target food ingredient can be determined from the target image in the visual information; the target image contains the target food ingredient, and the food image region contains the target food ingredient. Food recognition can be performed on the food image region to obtain the food category of the target food ingredient and its position within the food image region. The position of the target food ingredient within the food image region can be represented by a bounding box of the target food ingredient.
[0062] For example, a food category detection model can be used to identify the category of food. The controller can directly identify the food in the food image region, or the controller can send the food image region to the server, which will then identify the food and return the result to the refrigerator. For instance, a food category detection model can be deployed in the refrigerator and / or the server. The controller and / or the server can input the food image region into the food category detection model to identify the food category and obtain the food category of the target food and its position in the food image region. The food category detection model is a neural network model, which can be, but is not limited to, models such as YOLO (You Only Look Once) and SSD (Single Shot MultiBox Detector). Both YOLO and SSD are single-stage object detection algorithms.
[0063] Step 406: Determine the target distance representation value corresponding to the target ingredient. The target distance representation value is used to represent the distance between the target ingredient and the image acquisition device.
[0064] For example, the refrigerator or server may store the weight per unit area corresponding to the target distance representation value. Weight per unit area refers to the weight corresponding to a unit area. The pixel area of the target food can be determined based on its pixel dimensions, where pixel area = width × height. The product of the weight per unit area corresponding to the target distance representation value and the pixel area is used as the estimated weight of the target food. For example, the estimated weight = width × height × weight per unit area. The unit of weight per unit area can be grams per square pixel.
[0065] For example, a weight prediction model can be deployed in the refrigerator or server. The weight prediction model is used to predict the weight of food based on its category, pixel size, and distance representation. For instance, the controller can input the target food category, pixel size, and distance representation into the weight prediction model to predict the weight, obtaining the weight output by the model, and using this weight as an estimate of the target food's weight. Alternatively, the controller can send the target food category, pixel size, and distance representation to the server, which then inputs these parameters into the weight prediction model to predict the weight, obtaining the output weight and returning it to the refrigerator or controller.
[0066] Step 408: Estimate the weight of the target ingredient based on the target ingredient category, the target ingredient pixel size, and the target distance representation value to obtain the weight estimate of the target ingredient.
[0067] For example, the location of the target food item in the refrigerator can be determined, and a target distance representation value can be determined based on its location. Taking the target food item as... Figure 5 Taking an apple in a drawer as an example, we can determine that the target food category is apple, and that its pixel dimensions (w, width, and h, height) are, for example, 20 pixels and 25 pixels respectively, and that its position is in the right drawer. Based on the right drawer, we can determine the target distance representation value. Furthermore, based on the target food category, pixel dimensions, and target distance representation value, we can estimate the weight of the target food to obtain its estimated weight.
[0068] In this embodiment, visual information acquired by an image acquisition device is received, and the target food ingredient in the visual information is identified to obtain the target food ingredient category and the corresponding pixel size. A target distance representation value is determined, which represents the distance between the target food ingredient and the image acquisition device. Based on the target food ingredient category, pixel size, and target distance representation value, the weight of the target food ingredient is estimated to obtain a weight estimate. This simplifies the weight determination process, reduces complexity and time consumption, and thus improves the efficiency of food ingredient weight recognition.
[0069] In some embodiments, when the controller performs the task of determining the target distance representation value corresponding to the target food ingredient, it is configured to: determine the target area identifier corresponding to the target storage area to which the target food ingredient belongs; determine the area height correspondence of the refrigerator, the area height correspondence includes the distance representation values corresponding to multiple reference area identifiers in the refrigerator, the reference area identifiers correspond one-to-one with the storage areas in the refrigerator, and the distance representation value corresponding to the reference area identifier is used to reflect the distance between the storage area corresponding to the reference area identifier and the image acquisition device; match the target area identifier with the reference area identifier in the area height correspondence; and determine the distance representation value corresponding to the successfully matched reference area identifier as the target distance representation value corresponding to the target food ingredient.
[0070] The storage areas within the refrigerator can be, but are not limited to, drawers and shelves, and multiple drawers can exist within the refrigerator. The storage area corresponding to the reference area identifier refers to the storage area represented by the reference area identifier. The target storage area can be a drawer or a shelf. The distance representation value corresponding to a successfully matched reference area identifier is the distance representation value corresponding to the target area identifier. The target distance representation value corresponding to the target food item is the distance representation value corresponding to the target area identifier.
[0071] In this embodiment, the region height correspondence includes distance representation values corresponding to multiple reference region identifiers inside the refrigerator. Each reference region identifier corresponds one-to-one with a storage region inside the refrigerator. By matching the target region identifier with the reference region identifier in the region height correspondence, the distance representation value corresponding to the storage region to which the target food belongs can be determined, and the target distance representation value corresponding to the target food can be obtained. The method of determining the target distance representation value corresponding to the target food has low complexity and high efficiency.
[0072] In some embodiments, when the controller performs weight estimation of the target ingredient based on the target ingredient category, the target ingredient pixel size, and the target distance representation value to obtain the weight estimate of the target ingredient, it is configured to: determine the distance-weight correspondence associated with the target ingredient category, the distance-weight correspondence including the correspondence between the distance representation value and the weight per unit area; match the target distance representation value with the distance representation value in the distance-weight correspondence; and determine the weight estimate of the target ingredient based on the weight per unit area and the ingredient pixel size corresponding to the successfully matched distance representation value.
[0073] Among them, the weight per unit area corresponding to the distance characterization value is used to characterize the weight per unit area of the target food category in the image region corresponding to the distance characterization value, and the image region corresponding to the distance characterization value refers to the area occupied by the storage area corresponding to the distance characterization value in the image acquired by the image acquisition device.
[0074] The weight per unit area corresponding to the distance representation value can be obtained through experimental measurement. For example, the storage area corresponding to the distance representation value stores ingredients of the target ingredient category. Multiple images can be acquired by an image acquisition device, and the bounding box corresponding to the target ingredient category can be identified from each image. The pixel size of the ingredient is determined based on the size of the bounding box, and the imaging area of the ingredient is determined based on the pixel size of the ingredient. The ratio of the ingredient's weight to the imaging area is calculated to obtain the candidate weight per unit area corresponding to the ingredient. The average of the candidate weights per unit area corresponding to the target ingredient category in each image is taken as the weight per unit area corresponding to the distance representation value.
[0075] The refrigerator or server can store distance-weight relationships associated with multiple food categories. The target food category is one of these multiple food categories. For example... Figure 6 The diagram illustrates the correspondence between distance and weight, where D1~Dn represent n different reference distance values, and a1~an represent the weight per unit area corresponding to each of the n different reference distance values. The weight per unit area corresponding to a successfully matched distance value is the weight per unit area corresponding to the target distance value.
[0076] If the reference distance representation value corresponds to the area identifier, then the reference distance representation value represents the distance between the storage area corresponding to the area identifier and the image acquisition device. Therefore, the reference distance representation value is known when the refrigerator structure is determined. For example, the structure of the refrigerator from top to bottom consists of shelf areas, a single drawer, and a double drawer, with the double drawer containing a left drawer and a right drawer. The space of the single drawer is larger than that of the left and right drawers, and the areas of the left and right drawers can be the same or different. Thus, D1~Dn can be D1, D2, D3, and D4. D1 can be the reference distance representation value corresponding to the shelf area, D2 can represent the reference distance representation value corresponding to the single drawer, D3 can represent the reference distance representation value corresponding to the left drawer, and D4 can represent the reference distance representation value corresponding to the right drawer. D3 and D4 can be the same. Since there may be multiple overlapping shelf areas inside the refrigerator, in this application, each shelf area corresponds to a different reference distance representation value, or the multiple overlapping shelf areas are regarded as a whole (e.g., a shelf compartment), and this shelf compartment corresponds to a reference distance representation value. Each shelf compartment corresponds to a reference distance value, representing the distance between the shelf compartment and the image acquisition device. For example, it could be the distance between the center of the shelf compartment and the image acquisition device. Of course, for higher accuracy, each shelf area can have its own reference distance value.
[0077] For example, the imaging area of the target food can be determined based on its pixel size. The weight per unit area corresponding to the successfully matched distance representation value is multiplied by the imaging area of the target food, and the result is used as the estimated weight of the target food. For instance, if the imaging area is 500×600 and the weight per unit area corresponding to the target distance representation value is 0.0076 g / px², then the estimated weight = 500×600×0.0076. Here, g is the unit of weight "grams", and px² is the square pixel, the unit of imaging area. px is short for pixel.
[0078] In this embodiment, since the distance-weight correspondence includes the correspondence between the distance representation value and the weight per unit area, the weight per unit area corresponding to the target distance representation value can be determined. Then, based on the weight per unit area corresponding to the target distance representation value, the weight estimate of the target food can be determined. The process of calculating the weight estimate is simple and efficient, thereby improving the efficiency of weight estimation.
[0079] The weight estimation based on weight per unit area in this application has been experimentally verified. Taking an apple as an example, with an apple weighing 249 grams located 30 cm below the image acquisition device, an image is captured using the image acquisition device. The bounding box of the apple is drawn in the captured image, and the size of the bounding box is 181 pixels × 182 pixels. Since 249 / (181 × 182) = 0.0076, the weight per unit area is 0.0076 g / px². Then, with an apple weighing 368 grams positioned 30 cm below the image acquisition device, an image was captured. A bounding box of the apple was drawn within this image, with dimensions of 208 pixels × 205 pixels. The weight per unit area (0.0076 g / px²) was calculated as the product of 208 pixels × 205 pixels, yielding 322.304 g. Since 322.304 g is very close to 368 grams, it can be used as a weight estimate, proving that weight estimation based on weight per unit area is reliable. The bounding box dimensions are also the pixel dimensions of the food.
[0080] In some embodiments, when the controller performs the task of identifying target food ingredients in visual information and obtaining the target food ingredient category and the target food ingredient pixel size, it is configured to: determine the food image region of the target food ingredient from the target image in the visual information, wherein the target food ingredient is food taken from the refrigerator, the target image contains the target food ingredient, and the food image region contains the target food ingredient; perform food ingredient identification on the food image region of the target food ingredient to obtain the target food ingredient category and the target food ingredient pixel size.
[0081] For example, visual information can be used to identify a picking action, such as picking food from a refrigerator. The target image can be the image (i.e., a video frame) corresponding to the start time of the picking action.
[0082] For example, a first video frame corresponding to the start time of the picking action can be determined from the visual information, and a second video frame corresponding to the end time of the picking action can be determined from the visual information. Image difference calculation is performed on the first video frame and the second video frame to obtain a difference map. The location information of the changed area is determined from the difference map, and the image area at the location information of the changed area in the first video frame is taken as the food image area of the target food ingredient.
[0083] For example, a food image region can be input into a food category detection model, which can then identify the target food category and its pixel size. This food category detection model can be a target detection model. The input to the food category detection model is the food image region, and the output includes the food category and a bounding box. The food category output by the food category detection model can be used as the target food category, and the size of the output bounding box can be used as the target food pixel size.
[0084] For example, a food category detection model may include a feature extraction network, a feature fusion network, and a detection head. The feature extraction network extracts features from the input image, such as food image regions. This feature extraction network may contain multiple downsampling layers with different downsampling ratios, allowing it to output feature maps at different downsampling ratios. The feature fusion network fuses the feature maps output by the feature extraction network at different downsampling ratios, outputting a fused feature map. The detection head predicts the food category and bounding box based on the fused feature map. The feature extraction network may be, but is not limited to, ResNet (Residual Network). The feature fusion network may be, but is not limited to, FPN (Feature Pyramid Network). The detection head may include an RPN (Region Proposal Network), a RoI Align layer, and fully connected layers. The Region Proposal Network outputs RoIs (Regions of Interest). The RoI Alignment layer pools RoIs of different sizes to a fixed size. The fully connected layers output the category and bounding box.
[0085] In this embodiment, since the target food is taken from the refrigerator, the weight of the food taken from the refrigerator can be estimated.
[0086] In some embodiments, when the controller performs the task of determining the target distance representation value corresponding to the target food, it is configured to: perform region recognition on the target image in the visual information to obtain region images corresponding to multiple region identifiers, where each region identifier represents a storage area in the refrigerator; determine the target region image to which the target food belongs from the region images corresponding to the multiple region identifiers; and use the distance representation value associated with the region identifier corresponding to the target region image as the target distance representation value corresponding to the target food.
[0087] Region detection models can be used to identify regions in a target image. The input to a region detection model is an image, and its output is a segmentation map of the same size as the input image. The value of each pixel in the segmentation map represents the region identifier of the stored area to which that pixel belongs. The region detection model can be, but is not limited to, U-Net (U-shaped network). After obtaining the segmentation map, it can be divided according to the pixel values to obtain the region image corresponding to each region identifier.
[0088] For example, the food image region of the target food can be matched with the region image. If the food image region belongs to the region image, then the region image is determined as the target region image to which the target food belongs.
[0089] In this embodiment, the target region image to which the target ingredient belongs is determined from the region images corresponding to multiple region identifiers. The distance representation value associated with the region identifier corresponding to the target region image is used as the target distance representation value corresponding to the target ingredient. This allows for rapid determination of the target distance representation value.
[0090] In some embodiments, the weight estimate is determined by a weight prediction model, which is obtained by training the prediction model to be trained using training samples corresponding to the sample ingredients. The training samples include the true category of the sample ingredients, the pixel size of the sample ingredients, the sample distance representation value, and the true weight of the sample ingredients. The pixel size of the sample ingredients is the size of the sample ingredients identified from the sample image, which is taken when the sample ingredients are in a refrigerator. The sample distance representation value is used to reflect the distance between the sample ingredients and the image acquisition device when the sample image is taken.
[0091] There can be multiple training samples. These training samples can come from a training sample set, which can contain training samples corresponding to different categories of food ingredients. For example... Figure 7 The diagram shown illustrates the training sample set. It should be noted that... Figure 7 This is merely an illustration of a training sample set and does not represent the actual training sample set used when training the prediction model.
[0092] In this embodiment, the weight is estimated by using a trained quality prediction model, which can accurately and quickly estimate the weight of the target food, improving the efficiency and accuracy of weight estimation.
[0093] In some embodiments, the target food is food taken out of the refrigerator by the user. The controller is further configured to: obtain unit nutrient information corresponding to the category of the target food; generate nutritional information of the target food based on the unit nutrient information and weight estimate; the controller is further configured to: analyze the nutritional information of each target food identified within the target time period to obtain the nutritional information analysis result corresponding to the target time period; and generate food removal information corresponding to the target time period based on each target food identified within the target time period.
[0094] The nutrient information per unit weight refers to the content of nutrients per unit weight. The unit weight can be set according to actual needs, such as 50g or 100g. The nutrient information per unit weight can include the content of multiple nutrients individually per unit weight. Taking apples as an example... Figure 8 The diagram provided illustrates the nutritional information per unit. RAE stands for Retinol Activity Equivalent. The nutritional information per unit indicates that an apple (100g) contains 0.02mg of riboflavin, 0.02mg of thiamine, 50μg of carotene, 13.7g of carbohydrates, 1.7g of dietary fiber, and 0.2g of fat.
[0095] For example, the estimated weight can be multiplied by the content of each nutrient in the unit nutrient information to obtain the nutritional value of the target ingredient under each nutrient. This nutritional value is then used as the nutritional information of the target ingredient. Figure 9 The diagram shown illustrates the nutritional information of the target ingredient.
[0096] In this embodiment, the nutritional information of each target ingredient identified within the target time period is analyzed to obtain the nutritional information analysis results corresponding to the target time period. Based on each target ingredient identified within the target time period, the ingredient retrieval status information corresponding to the target time period is generated, providing a basis for providing feedback to the user on the ingredient retrieval status and helping to improve the user experience.
[0097] For example, such as Figure 10 As shown, a nutritional information identification method based on the food weight estimation method of this application is provided, including:
[0098] 1. When the user opens the refrigerator door, the camera on top of the refrigerator will start recording video or acquiring spectral images.
[0099] 2. The MCU controller uploads the acquired food images or spectral image information to the cloud via the network module.
[0100] In this context, "cloud" refers to servers.
[0101] 3. The cloud-based food ingredient recognition model identifies the uploaded food ingredient images, obtains the types of food ingredients picked up and placed by the user, the size information of the food ingredients in the image, and identifies the area to which the food ingredients belong in the cloud.
[0102] Among them, the cloud-based food identification model refers to the food category detection model on the server. For example, it can infer the relative height of the food in the vertical direction and the estimated size information of the image based on the background.
[0103] 4. Based on the food category (i.e. food type), region and pixel size information identified in step 3, estimate the weight of the food to be taken / placed, and calculate the weight (i.e. content) of the nutrients in the food based on the estimated weight and the food composition content table.
[0104] Among them, pixel size information refers to the pixel size of the food ingredient.
[0105] 5. Based on the user identity information when the user's interconnected device is bound to the refrigerator, the family member who will be handling the food retrieval and placement can be identified.
[0106] 6. Based on long-term statistics such as one day, one week, and one month, the system generates a user's refrigerator food usage report, i.e., a food consumption report.
[0107] 7. Based on the food consumption report and user health information from step 6, healthy eating advice can be provided.
[0108] like Figure 11 The diagram illustrates a system architecture for a nutritional information recognition method. The MCU (Microcontroller Unit) refers to the controller. The camera (e.g., a spectral camera) is an image acquisition device, such as a camera on the top of a refrigerator, used to acquire information about food entering and leaving the refrigerator. The network module uploads the acquired image data to the cloud. The cloud-based food recognition module estimates the type and size of the food based on the acquired image data (including spectral information). The nutritional component estimation module estimates the nutritional components and content of the food. Specifically, the camera on the top of the refrigerator (which can integrate a spectral camera) acquires video or spectral information of food entering and leaving the refrigerator. The acquired video or spectral information is then uploaded to the cloud for analysis of food type, size, or spectral information to analyze the nutritional components and content of the food. Guidance or reminders regarding the nutritional components of the food are then provided to the user. Health-connected devices can be, but are not limited to, mobile phones or wearable devices. Interactive devices can be, but are not limited to, mobile phones or wearable devices. The food recognition model can include models for identifying the type and / or shape of food, and can also include models for spectral recognition.
[0109] With the development of AI (Artificial Intelligence) technology, smart refrigerators are becoming increasingly intelligent, capable of automatically identifying food types and placement areas. However, they lack effective sensors for detecting the nutritional components of food, and users, who are particularly concerned about dietary management, lack data support for this crucial aspect. This embodiment addresses this by using a camera on the top of the refrigerator to capture the types and shapes of food entering and leaving the refrigerator, estimating food size, and then combining this with the nutritional composition of each food type to calculate the corresponding component content. This provides a solution for automatically detecting the nutritional components of food in the refrigerator, helping users track their family's nutritional intake and providing data for healthy eating.
[0110] In some embodiments, "identifying target ingredients in visual information" can be "identifying ingredients in a drawer in visual information," and the weight of each identified ingredient can be estimated separately. For example... Figure 12 As shown, a flowchart for food identification is provided, including: when the refrigerator door is detected to be open, the camera is activated to record video, and it is determined whether to "capture the drawer". If yes, it indicates that food in the drawer needs to be identified; otherwise, it indicates that food outside the drawer needs to be identified. If yes, then "determine the starting coordinates of the drawer area and the rectangles W and H", that is, determine the position information of the drawer area frame. The starting coordinates are the initial coordinates of the drawer area frame, for example (x, y), the rectangle W is the width of the drawer area frame, and H is the height of the drawer area frame. Here, the drawer area frame represents the position of the drawer in the captured image. After "determining the starting coordinates of the drawer area and the rectangles W and H", "calculate the camera register and initialize the camera", that is, use the position information of the drawer area frame to set the register of the output window in the camera so that the camera only outputs the image area within the drawer area frame. Then, "the camera starts video output", that is, outputs the image area located within the output window of the captured image (i.e., the image area within the drawer area frame), and then performs food identification on the image area within the output window.
[0111] In some embodiments, a server is provided, configured to: receive visual information acquired by an image acquisition device; identify target food ingredients in the visual information to obtain the target food ingredient category and the target food ingredient pixel size; determine a target distance representation value corresponding to the target food ingredient, the target distance representation value being used to represent the distance between the target food ingredient and the image acquisition device; and estimate the weight of the target food ingredient based on the target food ingredient category, the target food ingredient pixel size, and the target distance representation value to obtain a weight estimate of the target food ingredient.
[0112] In this embodiment, visual information acquired by an image acquisition device is received, and the target food ingredient in the visual information is identified to obtain the target food ingredient category and the corresponding pixel size. A target distance representation value is determined, which represents the distance between the target food ingredient and the image acquisition device. Based on the target food ingredient category, pixel size, and target distance representation value, the weight of the target food ingredient is estimated to obtain a weight estimate. This simplifies the weight determination process, reduces complexity and time consumption, and thus improves the efficiency of food ingredient weight recognition.
[0113] In some embodiments, the weight estimate is determined by a weight prediction model, and the server is further configured to: acquire training samples corresponding to the sample ingredients, the training samples containing the true category of the sample ingredients, the pixel size of the sample ingredients, the sample distance representation value, and the true weight of the sample ingredients, the pixel size of the sample ingredients being the size of the sample ingredients identified from the sample image, the sample image being taken with the sample ingredients in a refrigerator, and the sample distance representation value being used to reflect the distance between the sample ingredients and the image acquisition device when the sample image was taken; input the training samples into the prediction model to be trained for weight prediction to obtain the predicted weight of the sample ingredients; and adjust the model parameters of the prediction model based on the difference between the true weight of the sample ingredients and the predicted weight of the sample ingredients to obtain the weight prediction model.
[0114] The prediction model can be a neural network model, such as a deep neural network (DNN).
[0115] For example, a deep neural network can be built using Keras and used as a prediction model. The process of building a deep neural network includes: creating a neural network model with layers connected sequentially; defining the input to the input layer, which can include the true category of the sample food, the pixel size of the sample food, and the distance representation value between the sample and the sample food; and also, the true weight of the sample food. "Inputting training samples into the prediction model to be trained" can mean inputting the true category, pixel size, and distance representation value of the sample food in the training samples into the model to be trained, or it can mean inputting the true category, pixel size, distance representation value, and true weight of the sample food in the training samples into the model to be trained. Define a first hidden layer containing 64 neurons with the ReLU activation function. Define a second hidden layer containing 32 neurons with the ReLU activation function. Define a third hidden layer containing 16 neurons with the ReLU activation function. Define an output layer that outputs the predicted weight. All three hidden layers are fully connected layers. This embodiment is merely an example of a deep neural network. The parameters and structure of the deep neural network in this application are not limited to the structure and parameters provided in this embodiment.
[0116] For example, before training the prediction model, the model is compiled, the optimizer is set to Adam, the learning rate is 0.001, the loss function is set to a mean squared error (MSE) function, and the evaluation metric of the model is set to a mean absolute error (MAE) function.
[0117] For example, multiple training samples can exist, and these multiple training samples can be used to iteratively train the prediction model to obtain a weight prediction model. These multiple training samples can include training samples corresponding to different categories of food ingredients. Thus, the trained weight prediction model can be used to predict the weight of different categories of food ingredients. Taking evaluation as an example, the training samples can be: [406 (width), 368 (height), 10 (distance representation value)] → 254 (weight), [241, 209, 20] → 254, [193, 173, 25] → 254, [169, 144, 30] → 254, or [150, 128, 35] → 254. Here, the units for width and height are pixels.
[0118] In this embodiment, since the training samples include the true category of the sample ingredients, the pixel size of the sample ingredients, the sample distance representation value, and the true weight of the sample ingredients, the pixel size of the sample ingredients is the size of the sample ingredients identified from the sample images, the sample images are taken when the sample ingredients are in a refrigerator, and the sample distance representation value is used to reflect the distance between the sample ingredients and the image acquisition device when the sample image is taken, the prediction model can be trained using the training samples, so that the prediction model can learn the ability to predict weight based on category, size, and distance representation value, thereby ensuring the reliability of the weight prediction model.
[0119] In some embodiments, such as Figure 13 As shown, a time series diagram of a method for estimating the weight of food ingredients is provided, including:
[0120] 1. The controller receives visual information acquired by the image acquisition device.
[0121] 2. The controller sends visual information to the server.
[0122] 3. The server identifies the target food in the visual information and obtains the target food category and the target food pixel size.
[0123] 4. The server determines the target area identifier corresponding to the target storage area of the target food, determines the area height correspondence of the refrigerator, matches the target area identifier with the reference area identifier in the area height correspondence, and determines the distance representation value corresponding to the successfully matched reference area identifier as the target distance representation value corresponding to the target food.
[0124] 5. The server determines the distance-weight correspondence of the target food category, matches the target distance representation value with the distance representation value in the distance-weight correspondence, and determines the weight estimate of the target food based on the unit area weight and food pixel size corresponding to the successfully matched distance representation value.
[0125] 6. The server pushes information to the user terminal based on the weight estimate.
[0126] It should be understood that although the steps in the flowcharts of the above embodiments are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the above embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.
[0127] Based on the same inventive concept, this application also provides a food weight estimation device for implementing the above-mentioned food weight estimation method. The solution provided by this device is similar to the solution described in the above method, and the specific limitations can be found in the limitations of the food weight estimation method above, which will not be repeated here.
[0128] In some embodiments, such as Figure 14 As shown, a food weight estimation device is provided, including: an information receiving module 1402, a food identification module 1404, a distance determination module 1406, and a weight estimation module 1408, wherein:
[0129] The information receiving module 1402 is used to receive visual information acquired by the image acquisition device.
[0130] The food ingredient recognition module 1404 is used to identify target food ingredients in visual information and obtain the target food ingredient category and the target food ingredient pixel size.
[0131] The distance determination module 1406 is used to determine the target distance representation value corresponding to the target food ingredient. The target distance representation value is used to represent the distance between the target food ingredient and the image acquisition device.
[0132] The weight estimation module 1408 is used to estimate the weight of the target food based on the target food category, the target food pixel size, and the target distance representation value, and to obtain the weight estimate of the target food.
[0133] In some embodiments, a computer device is provided, which may be a controller. The computer device includes a processor, memory, input / output (I / O) interfaces, and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is connected to the system bus via the I / O interfaces. The processor of the computer device provides computational and control capabilities. The memory of the computer device includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The database of the computer device stores data involved in the food weight estimation method. The I / O interfaces of the computer device are used for exchanging information between the processor and external devices. The communication interface of the computer device is used for communicating with external terminals via a network connection. When the computer program is executed by the processor, it implements a food weight estimation method.
[0134] In one exemplary embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 15 As shown, the computer device includes a processor, memory, input / output (I / O) interfaces, and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The database stores at least some of the data involved in the food weight estimation method. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communicating with external terminals via a network connection. When the computer program is executed by the processor, it implements a food weight estimation method.
[0135] In some embodiments, a computer device is provided, including a memory and a processor. The memory stores a computer program, and the processor executes the computer program to perform the following steps: receiving visual information acquired by an image acquisition device; identifying a target food ingredient in the visual information to obtain the target food ingredient category and the target food ingredient pixel size; determining a target distance representation value corresponding to the target food ingredient, the target distance representation value being used to represent the distance between the target food ingredient and the image acquisition device; and estimating the weight of the target food ingredient based on the target food ingredient category, the target food ingredient pixel size, and the target distance representation value to obtain a weight estimate of the target food ingredient.
[0136] In some embodiments, this application also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, performs the following steps: receiving visual information acquired by an image acquisition device; identifying a target food ingredient in the visual information to obtain the target food ingredient category and the target food ingredient pixel size; determining a target distance representation value corresponding to the target food ingredient, the target distance representation value being used to represent the distance between the target food ingredient and the image acquisition device; and estimating the weight of the target food ingredient based on the target food ingredient category, the target food ingredient pixel size, and the target distance representation value to obtain a weight estimate of the target food ingredient.
[0137] In some embodiments, this application also provides a computer program product, including a computer program that, when executed by a processor, performs the following steps: receiving visual information acquired by an image acquisition device; identifying target food items in the visual information to obtain the target food item category and the target food item pixel size; determining a target distance representation value corresponding to the target food item, the target distance representation value being used to represent the distance between the target food item and the image acquisition device; and estimating the weight of the target food item based on the target food item category, the target food item pixel size, and the target distance representation value to obtain a weight estimate of the target food item.
[0138] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.
[0139] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.
[0140] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. A refrigerator, characterized in that, The refrigerator includes: Image acquisition equipment used to acquire visual information; The controller is configured as follows: Receive visual information acquired by the image acquisition device; The target food ingredient in the visual information is identified to obtain the target food ingredient category and the target food ingredient pixel size. Determine the target distance representation value corresponding to the target food ingredient, wherein the target distance representation value is used to represent the distance between the target food ingredient and the image acquisition device; The weight of the target ingredient is estimated based on the target ingredient category, the target ingredient pixel size, and the target distance representation value to obtain the weight estimate of the target ingredient.
2. The refrigerator according to claim 1, characterized in that, When the controller executes the process of determining the target distance representation value corresponding to the target ingredient, it is configured as follows: Determine the target area identifier corresponding to the target storage area to which the target ingredient belongs; The region height correspondence of the refrigerator is determined. The region height correspondence includes the distance characterization values corresponding to multiple reference region identifiers in the refrigerator. Each reference region identifier corresponds one-to-one with a storage area in the refrigerator. The distance characterization value corresponding to the reference region identifier is used to reflect the distance between the storage area corresponding to the reference region identifier and the image acquisition device. Match the target region identifier with the reference region identifier in the region height correspondence relationship; The distance representation value corresponding to the successfully matched reference area is determined as the target distance representation value corresponding to the target ingredient.
3. The refrigerator according to claim 2, characterized in that, When the controller performs the process of estimating the weight of the target ingredient based on the target ingredient category, the target ingredient pixel size, and the target distance representation value to obtain the estimated weight value of the target ingredient, it is configured to: Determine the distance-weight correspondence associated with the target food category, wherein the distance-weight correspondence includes the correspondence between the distance representation value and the weight per unit area; Match the target distance representation value with the distance representation value in the distance-weight correspondence; Based on the unit area weight corresponding to the successfully matched distance representation value and the pixel size of the food ingredient, the weight estimate of the target food ingredient is determined.
4. The refrigerator according to claim 1, characterized in that, The weight estimate is determined by a weight prediction model, which is obtained by training the prediction model to be trained using training samples corresponding to the sample ingredients. The training samples include the true category of the sample ingredients, the pixel size of the sample ingredients, the sample distance representation value, and the true weight of the sample ingredients; The pixel size of the sample food is the size of the sample food identified from a sample image, which was taken while the sample food was in the refrigerator; The sample distance characterization value is used to reflect the distance between the sample food and the image acquisition device when the sample image is captured.
5. The refrigerator according to any one of claims 1 to 4, characterized in that, When the controller performs the process of identifying the target food ingredient in the visual information and obtaining the target food ingredient category and the target food ingredient pixel size, it is configured as follows: The target food ingredient image region is determined from the target image in the visual information, wherein the target food ingredient is food taken from the refrigerator, the target image contains the target food ingredient, and the food ingredient image region contains the target food ingredient; The target ingredient image region is used for ingredient recognition to obtain the target ingredient category and the target ingredient pixel size.
6. The refrigerator according to any one of claims 1 to 4, characterized in that, When the controller executes the process of determining the target distance representation value corresponding to the target ingredient, it is configured as follows: The target image in the visual information is subjected to region recognition to obtain region images corresponding to multiple region identifiers, where each region identifier represents a storage area in the refrigerator. From the region images corresponding to the multiple region identifiers, determine the target region image to which the target ingredient belongs; The distance representation value associated with the region identifier corresponding to the target region image is used as the target distance representation value corresponding to the target ingredient.
7. The refrigerator according to any one of claims 1 to 4, characterized in that, The target food item is the food item taken by the user from the refrigerator, and the controller is further configured to: Obtain the unit nutrient information corresponding to the target food category; Based on the unit nutrient information and the estimated weight, the nutritional information of the target food ingredient is generated; The controller is also configured to be at least one of the following: The nutritional information of each target food ingredient identified within the target time period is analyzed to obtain the nutritional information analysis results corresponding to the target time period. Based on the target ingredients identified within the target time period, information on the retrieval status of the ingredients corresponding to the target time period is generated.
8. A server, characterized in that, The server is configured as follows: Receive visual information acquired by image acquisition devices; The target food ingredient in the visual information is identified to obtain the target food ingredient category and the target food ingredient pixel size. Determine the target distance representation value corresponding to the target food ingredient, wherein the target distance representation value is used to represent the distance between the target food ingredient and the image acquisition device; The weight of the target ingredient is estimated based on the target ingredient category, the target ingredient pixel size, and the target distance representation value to obtain the weight estimate of the target ingredient.
9. The server according to claim 8, characterized in that, The weight estimate is determined using a weight prediction model, and the server is further configured to: Obtain training samples corresponding to the sample ingredients. The training samples include the true category of the sample ingredients, the pixel size of the sample ingredients, the sample distance representation value, and the true weight of the sample ingredients. The pixel size of the sample ingredients is the size of the sample ingredients identified from the sample image. The sample image is taken when the sample ingredients are in a refrigerator. The sample distance representation value is used to reflect the distance between the sample ingredients and the image acquisition device when the sample image is taken. The training samples are input into the prediction model to be trained to predict the weight and obtain the predicted weight of the sample ingredients. Based on the difference between the actual weight of the sample food and the predicted weight of the sample food, the model parameters of the prediction model are adjusted to obtain a weight prediction model.
10. A method for estimating the weight of food ingredients, characterized in that, Applied to a server or refrigerator, the method includes: Receive visual information acquired by image acquisition devices; The target food ingredient in the visual information is identified to obtain the target food ingredient category and the target food ingredient pixel size. Determine the target distance representation value corresponding to the target food ingredient, wherein the target distance representation value is used to represent the distance between the target food ingredient and the image acquisition device; The weight of the target ingredient is estimated based on the target ingredient category, the target ingredient pixel size, and the target distance representation value to obtain the weight estimate of the target ingredient.