Refrigerator and refrigerator control method

CN122813458APending Publication Date: 2026-09-25HISENSE XINGHAI TECHNOLOGY (HANGZHOU) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610957897.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-29
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

另一些冰箱虽然具备食材指引功能,但是其通过在冰箱各个角落设置复杂昂贵的相机实时采集冰箱内部食材图像,以监控更新冰箱内部食材的存储状态

Benefits of technology

[0011]可以理解的是,上述第二方面至第五方面的有益效果可以参见上述第一方面中的相关描述,在此不再赘述。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122813458A_ABST
    Figure CN122813458A_ABST
Patent Text Reader

Abstract

The application provides a refrigerator and a refrigerator control method. The refrigerator comprises: an image acquisition device arranged at the top of the front side of the refrigerator and configured to take a picking image of a user picking food; and a controller configured to: acquire the picking image; determine an initial height coordinate of a target access position of a target food in the refrigerator based on depth information of the picking image, wherein the height coordinate is a coordinate in a height direction of the refrigerator; determine a calibrated height coordinate of the target access position according to the initial height coordinate and a set of reference height coordinates of each storage position in the refrigerator, wherein the set of reference height coordinates comprises a plurality of reference height coordinates corresponding to each storage position respectively, and the plurality of reference height coordinates correspond to each storage position respectively; determine storage position information corresponding to the target access position according to the calibrated height coordinate; and update food inventory data of the refrigerator based on category information of the target food and the storage position information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of refrigerator technology, and more particularly to a refrigerator and a refrigerator control method. Background Technology

[0002] As an indispensable home appliance in modern households, the refrigerator's main function is to store food in a low-temperature environment to extend its shelf life. With the improvement of people's living standards, refrigerator capacities have continued to increase, and the internal storage compartments have become increasingly refined. Users need to locate the food before retrieving it during daily use.

[0003] Traditional refrigerators lack food guide functions, requiring users to rely on memory or searching to locate items. Other refrigerators, while possessing food guide functions, rely on complex and expensive cameras embedded in various corners to continuously capture images of the food inside, monitoring its storage status. This approach is not only costly in terms of equipment but also involves enormous computational demands due to the simultaneous data collection from multiple cameras. Furthermore, in scenarios where food is heavily obstructed, the accuracy of identifying obscured items is less than ideal, resulting in poor reliability of inventory updates and food guide functionality. Summary of the Invention

[0004] This application provides a refrigerator, a refrigerator control method, a refrigerator control device, a computer-readable storage medium, and a computer program product, which can not only improve the reliability of refrigerator food inventory updates and help achieve accurate food guidance, but also reduce hardware costs and computing overhead.

[0005] In a first aspect, a refrigerator is provided, comprising: An image acquisition device is installed at the top front side inside the refrigerator and is configured to take images of the user taking out and putting in food from above. The images include the user's hands and the imaging area of ​​the target food being taken out and put in. The controller, which communicates with the image acquisition device, is configured as follows: Get the pick-and-place image; Based on the depth information of the pick-up and put-down images, the initial height coordinates of the target food in the refrigerator are determined, where the height coordinates are coordinates along the height direction of the refrigerator. Based on the initial height coordinates and the set of reference height coordinates for each storage location inside the refrigerator, the calibration height coordinates of the target access location are determined. The set of reference height coordinates includes multiple reference height coordinates corresponding to each storage location. Based on the calibration height coordinates, determine the storage location information corresponding to the target access location; Update the refrigerator's food inventory data based on the category and storage location information of the target food ingredients.

[0006] Secondly, a refrigerator control method is provided, applicable to refrigerators in any embodiment of this application. The refrigerator includes an image acquisition device disposed on the top front side inside the refrigerator, configured to capture images of a user taking or placing food items from above. The images include the user's hand and the imaging area of ​​the target food item being taken or placed. The method includes: Get the pick-and-place image; Based on the depth information of the pick-up and put-down images, the initial height coordinates of the target food in the refrigerator are determined, where the height coordinates are coordinates along the height direction of the refrigerator. Based on the initial height coordinates and the set of reference height coordinates for each storage location inside the refrigerator, the calibration height coordinates of the target access location are determined. The set of reference height coordinates includes multiple reference height coordinates corresponding to each storage location. Based on the calibration height coordinates, determine the storage location information corresponding to the target access location; Update the refrigerator's food inventory data based on the category and storage location information of the target food ingredients.

[0007] Thirdly, a refrigerator control device is provided for performing the steps of the refrigerator control method provided in the second aspect of the embodiments of this application.

[0008] Fourthly, a computer-readable storage medium is provided, which stores a computer program that, when run by a refrigerator control device, causes the refrigerator control device to perform the refrigerator control method of the second aspect.

[0009] Fifthly, a computer program product is provided, comprising: a computer program that, when run by a processor, causes the processor to execute the refrigerator control method of the second aspect.

[0010] The refrigerator provided in the first aspect of this application, by placing the image acquisition device on the top front side of the refrigerator and using a top-down shooting method, requires only a small number of cameras to cover the user's operation area for retrieving and placing food. Compared with the prior art, which deploys complex and expensive camera arrays in multiple locations inside the refrigerator, this significantly reduces hardware costs and deployment difficulty. Simultaneously, compared with the prior art, since the refrigerator controller in this application only needs to process image data acquired by a small number of cameras, the computational load is greatly reduced, which helps to reduce the controller's computational burden and power consumption, and improves the refrigerator's real-time response capability. Furthermore, in this application embodiment, the refrigerator's image acquisition device focuses on the target food being handled by the user. This operation-centric image acquisition strategy enables the refrigerator to accurately acquire visual information directly related to the food being handled, reducing interference from irrelevant background information, thereby improving the efficiency and accuracy of subsequent processing. Based on this, the controller first determines the initial height coordinates of the target food using depth information in the image, and then, combined with a pre-stored set of reference height coordinates for each storage location, corrects the initial height coordinates, quickly and accurately determining the storage location information corresponding to the target food's storage location within the refrigerator. Specifically, by using image depth information, the coordinates of food items along the height of the refrigerator can be accurately obtained. Then, by using pre-stored reference height coordinates for each storage compartment for constraint correction, the depth measurement value can be precisely assigned to the nearest storage compartment level. This dual positioning mechanism, combining depth measurement and prior constraints, can stably output accurate information about the storage compartment where the target food item is located, even in complex scenarios where food items in different storage compartments are obscured or where lighting is uneven. This greatly improves the robustness and accuracy of spatial positioning. Finally, by associating the food item category information with the storage compartment information, the food inventory data is updated, achieving automated and precise inventory management. Because the refrigerator can continuously and accurately know what food items are placed in which storage compartment or what food items are taken out of which storage compartment during operation, it can maintain accurate food inventory records by accumulating this information, helping to provide users with accurate food guidance. Users can understand the distribution of food items in the refrigerator without manually recording or searching, saving them time and effort in finding food. Therefore, the refrigerator in this application embodiment achieves accurate perception of the storage location of food, especially the storage location in the vertical direction, with lower cost and computational overhead, which can improve the reliability of food guidance and thus significantly improve the intelligence level of refrigerator food management and user experience.

[0011] It is understood that the beneficial effects of the second to fifth aspects mentioned above can be found in the relevant descriptions in the first aspect mentioned above, and will not be repeated here. Attached Figure Description

[0012] Figure 1 A schematic diagram of the structure of a refrigerator according to an embodiment of this application is shown; Figure 2 A schematic flowchart of a refrigerator control method according to an embodiment of this application is shown; Figure 3a One of the schematic diagrams illustrating the implementation principle of a refrigerator control method according to an embodiment of this application is shown; Figure 3b This is a second schematic diagram illustrating the implementation principle of a refrigerator control method according to an embodiment of this application; Figure 3c This is shown as a third schematic diagram illustrating the implementation principle of a refrigerator control method according to an embodiment of this application; Figure 3d This is shown as a fourth schematic diagram illustrating the implementation principle of a refrigerator control method according to an embodiment of this application; Figure 4 A flowchart illustrating a refrigerator control method according to another embodiment of this application is shown; Figure 5 A schematic diagram of the structure of a control device provided in one embodiment of this application is shown. Detailed Implementation

[0013] The technical solutions of the embodiments of this application will be described below with reference to the accompanying drawings. In the description of the embodiments of this application, unless otherwise stated, " / " means "or," for example, A / B can mean A or B; "and / or" in this text is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Furthermore, in the description of the embodiments of this application, "multiple" refers to two or more than two.

[0014] Hereinafter, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first," "second," or "third" may explicitly or implicitly include one or more of that feature.

[0015] Specific details, such as particular system architectures and techniques, are set forth for illustrative purposes and not for limitation, to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application can be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods are omitted to avoid unnecessary detail that could obscure the description of this application.

[0016] First, it should be noted that the collection, storage, use, processing, transmission, provision, disclosure, and application of user personal information involved in the embodiments of this application all comply with the provisions of relevant laws and regulations, have obtained the user's authorization or consent, have taken necessary confidentiality measures, and do not violate public order and good morals.

[0017] As mentioned earlier, existing refrigerators with food guidance functions are expensive to install, and the computational load from simultaneous data collection by multiple cameras is enormous. Furthermore, in scenarios where food is severely obscured, this solution's accuracy in identifying obscured food is not ideal, resulting in poor reliability of inventory updates and food guidance.

[0018] Furthermore, further research revealed that while existing refrigerator technologies incorporate cameras in multiple locations, they fundamentally rely on static image analysis to determine the location and category of food items. Due to the multi-layered design of the refrigerator's internal structure, there is a natural mutual occlusion phenomenon in the height direction between storage compartments. This, coupled with the occlusion caused by the stacking of food items, makes it impossible for multi-camera solutions to accurately capture a complete internal view of the refrigerator. Specifically, existing refrigerator control methods suffer from the following three technical deficiencies: First, a lack of height-direction positioning capability. Existing image analysis methods are based on two-dimensional planar information and lack the ability to accurately perceive the height of food items (i.e., the direction of the refrigerator shelves). When users retrieve or place food items, the system struggles to accurately identify the specific location of occluded food items based on static two-dimensional images, especially determining which shelf the food is placed on. This is because the occlusion effect prevents traditional image analysis methods from effectively distinguishing food items on different shelves, thus failing to accurately answer the user's question about the shelf location of the food. Second, a lack of dynamic perception of user behavior. Existing technologies only perform static analysis of images inside the refrigerator and cannot perceive the dynamic behavior of users retrieving or placing food items. For example, when a user opens the refrigerator door and takes out a bottle of milk, the refrigerator cannot determine from a static image whether the milk was taken away or moved, nor can it distinguish between different operation types such as "taken out," "added," and "consumed." This static analysis method results in a significant lag in updating food inventory data compared to actual operations, often requiring manual confirmation from the user or reliance on additional input devices. Third, multimodal information is not effectively integrated. While some existing technologies attempt to introduce weight sensors or voice interaction modules, these modules operate independently, lacking a unified fusion and reasoning framework. For example, a weight sensor can detect changes in the weight of a shelf but cannot determine which food caused the change; a voice interaction module can receive verbal commands but cannot verify whether the commands match the actual operation. This lack of effective integration of heterogeneous information limits the refrigerator's overall sensing capabilities, leading to high false positive and false negative rates.

[0019] In summary, existing refrigerators with food guidance functions suffer from problems such as high device cost, large computational load, poor recognition accuracy in obstructed scenarios, lack of dynamic behavior perception capabilities, and failure to effectively integrate multimodal information. These issues result in poor reliability of inventory updates and food guidance, leading to a poor user experience.

[0020] To at least partially solve the above-mentioned technical problems, embodiments of this application provide a refrigerator, a refrigerator control method, a refrigerator control device, a computer-readable storage medium, and a computer program product, which can improve the reliability of refrigerator food inventory updates, help achieve accurate food guidance, and reduce hardware costs and computing overhead.

[0021] First, this application provides a refrigerator control method. This control method can be applied to various types of refrigerators. Specifically, it can be applied to the refrigerator controller. The following describes the method in conjunction with... Figures 1 to 4 The refrigerator control method of the present application will be described in detail.

[0022] like Figure 1 As shown in the illustration, the refrigerator 100 provided in this embodiment includes an image acquisition device 110 and a controller 120. The image acquisition device 110 is disposed on the top front side of the refrigerator 100 and is configured to capture images of the user taking or placing food items from above. In this embodiment, the image acquisition device 110 can be various types of visual acquisition devices. In one example, the image acquisition device 110 can be an RGB-D camera, which can simultaneously acquire RGB color images and corresponding depth images, wherein the depth image records the distance information from each pixel in the scene to the camera. For example, an RGB-D camera can be placed at the center of the top front area of ​​the refrigerator, with the camera's line of sight perpendicular to the height direction of the refrigerator and its line of sight pointing towards the bottom of the refrigerator. In another example, the image acquisition device 110 can be a binocular stereo camera, calculating depth information through the parallax between the left and right cameras. In yet another example, the image acquisition device 110 can also be a combination of a regular color camera and an independent depth sensor, or a solution using a monocular camera combined with a deep learning depth estimation algorithm. This application does not limit the specific implementation of the image acquisition device, as long as it can acquire images containing visual and depth information.

[0023] In this embodiment, the image captured by the image acquisition device 110 when the user opens the refrigerator door to retrieve or place food items is an image taken by the user. This image includes the user's hand and the imaging area of ​​the target food item being retrieved or placed. The target food item can be the food item the user is retrieving or placing; for example, when the user reaches into the refrigerator to take out a bottle of milk, the milk is the target food item; or when the user puts a bag of vegetables into the refrigerator, the vegetables are the target food item. It should be noted that the image is not a panoramic scan of all food items inside the refrigerator, but rather focuses on the target area the user is operating on. Therefore, it can accurately capture visual information directly related to the food item being operated on, reducing interference from irrelevant backgrounds. The image can be a single frame image, or it can be one or more representative frames selected from a continuously acquired image sequence.

[0024] In this embodiment, the controller 120 is communicatively connected to the image acquisition device 110 and is configured to execute the refrigerator control method of any embodiment of this application. The controller 120 can be a variety of computing devices. In one example, the controller 120 can be a microcontroller unit (MCU), which has high integration and low power consumption, making it suitable for embedded applications. In another example, the controller 120 can be a digital signal processor (DSP), which excels at processing images and signals. As yet another implementation, the controller 120 can be a field-programmable gate array (FPGA), which enables hardware-level parallel acceleration and is suitable for real-time image processing. In yet another example, the controller 120 can be a system-on-a-chip (SoC), integrating the processor, memory, peripheral interfaces, etc., onto a single chip, balancing performance and integration. Furthermore, the controller 120 can also be a general-purpose processor (such as an ARM Cortex series or x86 architecture processor), a graphics processing unit (GPU), a neural network processor (NPU), or a combination of these processors.

[0025] For example, the refrigerator may also include a variety of input / output devices and sensors to further enhance its intelligence.

[0026] For example, a refrigerator may include a door status sensor, installed on the refrigerator door frame or door body, to detect door opening and closing events. The door status sensor can be a reed switch, Hall effect sensor, photoelectric sensor, or microswitch, etc. As another example, a refrigerator may include a weight sensor array, installed under each storage shelf, to detect the total weight of food stored in the corresponding compartment. The weight sensor array can consist of multiple pressure sensors or strain gauges, with one or more sensors placed under each compartment to achieve independent measurement of the weight of each shelf.

[0027] For example, the refrigerator may also include a microphone array. Exemplarily, the microphone array may also be mounted on the refrigerator panel or door for capturing user voice commands. The microphone array may consist of multiple microphones and support beamforming and sound source localization functions, enabling effective pickup of user voice signals in noisy environments.

[0028] For example, a refrigerator may include an output device for providing feedback information to the user. The output device may include a speaker for playing voice feedback; it may also include a display screen for displaying graphic and text information; or it may include other forms of output devices such as indicator lights or vibration motors.

[0029] In addition, the refrigerator may include a communication module for data interaction with cloud servers or other external devices. The communication module can support wireless communication protocols such as Wi-Fi, Bluetooth, Zigbee, and cellular networks (such as 4G / 5G), as well as wired communication methods such as Ethernet.

[0030] For example, the refrigerator may also include a storage device for storing data such as food inventory data, a set of reference height coordinates, a storage location mapping table, visual change feature sequences, and conversation history, as mentioned below. The storage device may be in the form of flash memory, EEPROM, SD card, solid-state drive (SSD), or cloud storage.

[0031] The aforementioned components are coordinated and controlled by a controller to achieve intelligent food management in the refrigerator. For example, the controller receives data from sensors such as an image acquisition device, a door status sensor, a weight sensor array, and a microphone array. After processing, it generates control commands, drives the output device to output feedback information, and interacts with the cloud server through a communication module, thereby realizing functions such as intelligent sensing, positioning, inventory management, and human-computer interaction of food.

[0032] like Figure 2 As shown, the refrigerator control method provided in this application embodiment specifically includes the following steps S210, S220, S230, S240 and S250.

[0033] Step S210: Obtain the pick-up and put-down images.

[0034] For example, combined Figure 3a and Figure 3b The image acquisition device (such as an RGB-D camera) is positioned at the top front of the refrigerator's interior (i.e., the area near the refrigerator door at the top of the refrigerator), with the camera's line of sight pointing downwards along the height of the refrigerator (i.e., parallel to the refrigerator's interior). Figure 3a The refrigerator shown in the image has the opposite z-axis direction), and the camera's field of view (along...) Figure 3aThe y-axis of the refrigerator coordinate system shown in the diagram corresponds to the area between the edge of each storage compartment and the edge of the refrigerator's inner cavity (i.e., the contact door). In this way, the camera can capture images of the user's hand movements as they extend their hands into and out of the refrigerator's storage compartments, while minimizing interference from other food items in the refrigerator's storage compartments.

[0035] It's understandable that if the user is holding food, the image captured by the camera will include both the food and the hand. If the user is not holding food, the image may only include the hand and not the food. Therefore, the image acquired in this step, showing the handling of food, includes both the hand and the food; in other words, it captures an image of the person holding the food.

[0036] In one example, the controller can directly acquire a single frame image from the image acquisition device and determine whether the single frame image contains the imaging area of ​​a hand and food; if so, the frame image is identified as the pick-up / placement image. In another example, the controller can select a representative frame image from a continuously acquired time-series image sequence as the pick-up / placement image according to preset rules. For example, it can select an image containing a hand and food captured at the moment when the hand is closest to the storage location, or select an image at the moment when the contact between the hand and food is most significant.

[0037] Step S220: Based on the depth information of the pick-up and put-down image, determine the initial height coordinates of the target food item's storage location in the refrigerator, where the height coordinates are coordinates along the height direction of the refrigerator.

[0038] In this embodiment, the depth information can be the distance value from each pixel recorded in the depth channel of the image to the image acquisition device. The initial height coordinates can be the coordinate values ​​of the food in the height direction of the refrigerator obtained directly through coordinate system transformation or other conversion methods using the depth information. Figure 3a As shown, the height direction of the refrigerator can be the vertical direction from the bottom to the top of the refrigerator cavity (corresponding to the z-axis direction in the figure), which is consistent with the stacking direction of each shelf.

[0039] In one example, the controller can first perform semantic segmentation on the color image in the pick-up and place image to obtain a food segmentation mask for the target food. Then, it calculates the centroid pixel coordinates of this mask, reads the depth value at these centroid pixel coordinates from the depth image, and finally, combines the camera intrinsic parameter matrix and the pre-calibrated camera pose matrix relative to the refrigerator. Through inverse perspective projection transformation and coordinate transformation, the centroid pixel coordinates are converted into 3D spatial coordinates in the refrigerator coordinate system, and the height component is extracted from these coordinates as the initial height coordinates. In another example, the controller can directly average the depth values ​​of all pixels in the target food area of ​​the pick-up and place image, and then combine this with camera parameters to convert it into initial height coordinates. This is suitable for scenarios where the food area is large and the depth values ​​are uniform. In yet another example, the controller can use an object detection network to directly regress the 3D bounding box of the target food, then extract the height component from the center point coordinates of the bounding box, and obtain the initial height coordinates after coordinate system transformation. This is suitable for scenarios requiring fast inference.

[0040] Step S230: Determine the calibration height coordinates of the target storage location based on the initial height coordinates and the set of reference height coordinates for each storage location inside the refrigerator. The set of reference height coordinates includes multiple reference height coordinates corresponding to each storage location.

[0041] In this embodiment, the reference height coordinate set can be a pre-stored set of standard height values ​​that correspond one-to-one with each shelf or storage level inside the refrigerator, such as 150mm for the first shelf and 450mm for the second shelf. The calibration height coordinates can be height coordinate values ​​belonging to a certain standard storage level obtained by matching and correcting the initial height coordinates with the reference height coordinate set.

[0042] In one example, the controller can use a nearest neighbor matching algorithm to calculate the difference between the initial height coordinates and each reference height coordinate, selecting the reference height coordinate with the smallest difference as the calibration height coordinate. In another example, the controller can use a constrained optimization algorithm to correct the initial height coordinates to the closest reference height coordinates within a preset error tolerance range (e.g., ±20mm). If the initial height coordinates fall within the middle region of two reference height coordinates, the selection can be based on historical statistical patterns or confidence weighting. In yet another example, the controller can compare the initial height coordinates with a preset height range corresponding to each reference height coordinate. When the initial height coordinates fall within a certain preset height range, the reference height coordinate corresponding to that range is determined as the calibration height coordinate.

[0043] Step S240: Determine the storage location information corresponding to the target storage location based on the calibration height coordinates.

[0044] In this embodiment, storage location information can refer to information used to identify the specific storage location of food in the refrigerator, which may include storage location level (such as "second shelf"), partition location (such as "left side", "right side", "front", "rear"), or more refined coordinate description (such as "third cell on the right side of the second shelf").

[0045] In one example, the controller can look up the storage location identifier (e.g., "second shelf" or "shelf 3") that corresponds to the calibration height coordinates in a pre-stored calibration height-storage location mapping table, using this identifier as the storage location information for the target access position. In another example, the controller can directly derive the storage location identifier based on a pre-defined correspondence between the calibration height coordinates and storage locations, through rule matching. For example, when the calibration height coordinate is 450mm, the corresponding storage location identifier is "middle shelf." In yet another example, if the refrigerator has left / right or front / back partitions on the same height level, the controller can also combine the coordinate components of the target food in the length and width directions to look up the target storage location identifier that matches both the height coordinates and the partition position in the storage location mapping table, thereby achieving more precise storage location positioning.

[0046] Step S250: Update the refrigerator's food inventory data based on the category information and storage location information of the target food.

[0047] In this embodiment, category information can refer to food type identifiers obtained through image recognition or user input, such as "milk," "apple," and "beef." Food inventory data can be a structured dataset stored locally in the refrigerator or in a cloud database, recording information such as the name, quantity, location, storage time, and shelf life of all food items in the refrigerator.

[0048] In one example, the controller can combine the target ingredient's category (e.g., "milk"), storage location identifier (e.g., "right side of the second shelf"), and current timestamp into an inventory record and write it to the local ingredient inventory database. In another example, the controller can first query the database to see if a record of the same category and location already exists. If it does, the quantity or weight is updated; otherwise, a new record is created. In yet another example, the controller can also record the type of operation (e.g., adding or removing) for later tracking and analysis of user behavior. The updated ingredient inventory data can be used to generate ingredient distribution maps, near-expiration reminders, and intelligent recommendations. Users can check the location and status of ingredients at any time through the refrigerator's display screen or voice interaction.

[0049] like Figure 3aAs shown, during the process of a user opening the refrigerator to retrieve an apple, an image acquisition device on the top front of the refrigerator captures an image of the user taking and placing the apple. After acquiring this image, the controller determines the initial height coordinates of the apple within the refrigerator cavity based on the depth information of the depth image in the image. Subsequently, the initial height coordinates are compared with the pre-stored standard heights of each shelf (e.g., 150mm, 450mm, 750mm, 1050mm). The initial height coordinates are corrected based on the comparison results to obtain calibrated height coordinates. Next, the storage location mapping table is consulted based on the calibrated height coordinates to find that the apple's corresponding storage location identifier is "12," corresponding to the second shelf on the upper part of the refrigerator. Finally, the apple's category information is associated with this storage location information, updating the food inventory database and recording "the user took 1 apple from the second shelf." The number of apples on the second shelf in the database is equal to the previous update quantity minus 1. Subsequently, the user can accurately know the storage location and inventory changes of the apples through the refrigerator display screen or voice query.

[0050] The refrigerator provided in the first aspect of this application, by placing the image acquisition device on the top front side of the refrigerator and using a top-down shooting method, requires only a small number of cameras to cover the user's operation area for retrieving and placing food. Compared with the prior art, which deploys complex and expensive camera arrays in multiple locations inside the refrigerator, this significantly reduces hardware costs and deployment difficulty. Simultaneously, compared with the prior art, since the refrigerator controller in this application only needs to process image data acquired by a small number of cameras, the computational load is greatly reduced, which helps to reduce the controller's computational burden and power consumption, and improves the refrigerator's real-time response capability. Furthermore, in this application embodiment, the refrigerator's image acquisition device focuses on the target food being handled by the user. This operation-centric image acquisition strategy enables the refrigerator to accurately acquire visual information directly related to the food being handled, reducing interference from irrelevant background information, thereby improving the efficiency and accuracy of subsequent processing. Based on this, the controller first determines the initial height coordinates of the target food using depth information in the image, and then, combined with a pre-stored set of reference height coordinates for each storage location, corrects the initial height coordinates, quickly and accurately determining the storage location information corresponding to the target food's storage location within the refrigerator. Specifically, by using image depth information, the coordinates of food items along the height of the refrigerator can be accurately obtained. Then, by using pre-stored reference height coordinates for each storage compartment for constraint correction, the depth measurement value can be precisely assigned to the nearest storage compartment level. This dual positioning mechanism, combining depth measurement and prior constraints, can stably output accurate information about the storage compartment where the target food item is located, even in complex scenarios where food items in different storage compartments are obscured or where lighting is uneven. This greatly improves the robustness and accuracy of spatial positioning. Finally, by associating the food item category information with the storage compartment information, the food inventory data is updated, achieving automated and precise inventory management. Because the refrigerator can continuously and accurately know what food items are placed in which storage compartment or what food items are taken out of which storage compartment during operation, it can maintain accurate food inventory records by accumulating this information, helping to provide users with accurate food guidance. Users can understand the distribution of food items in the refrigerator without manually recording or searching, saving them time and effort in finding food. Therefore, the refrigerator in this application embodiment achieves accurate perception of the storage location of food, especially the storage location in the vertical direction, with lower cost and computational overhead, which can improve the reliability of food guidance and thus significantly improve the intelligence level of refrigerator food management and user experience.

[0051] The refrigerator control method of this application embodiment, by placing the image acquisition device on the top front side of the refrigerator and using a top-down shooting method, requires only a small number of cameras to cover the user's operation area for retrieving and placing food. Compared with the existing technology that deploys complex and expensive camera arrays in multiple locations inside the refrigerator, this significantly reduces hardware costs and deployment difficulty. Simultaneously, compared with the prior art, the refrigerator control method of this application embodiment only needs to process image data acquired by a small number of cameras, greatly reducing the computational load, which helps to reduce the controller's computational burden and power consumption, and improves the refrigerator's real-time response capability. Furthermore, in this application embodiment, the refrigerator's image acquisition device focuses on the target food being handled by the user. This operation-centric image acquisition strategy allows the refrigerator to accurately acquire visual information directly related to the food being handled, reducing interference from irrelevant background information, thereby improving the efficiency and accuracy of subsequent processing. Based on this, the initial height coordinates of the target food are first determined using depth information in the image, and then corrected by combining the pre-stored set of reference height coordinates for each storage location, quickly and accurately determining the storage location information corresponding to the target food's storage location in the refrigerator. Specifically, by using image depth information, the coordinates of food items along the height of the refrigerator can be accurately obtained. Then, by using pre-stored reference height coordinates for each storage compartment for constraint correction, the depth measurement value can be precisely assigned to the nearest storage compartment level. This dual positioning mechanism, combining depth measurement and prior constraints, can stably output accurate information about the storage compartment where the target food item is located, even in complex scenarios where food items in different storage compartments are obscured or where lighting is uneven. This greatly improves the robustness and accuracy of spatial positioning. Finally, by associating the food item category information with the storage compartment information, the food inventory data is updated, achieving automated and precise inventory management. Because the refrigerator can continuously and accurately know what food items are placed in which storage compartment or what food items are taken out of which storage compartment during operation, it can maintain accurate food inventory records by accumulating this information, helping to provide users with accurate food guidance. Users can understand the distribution of food items in the refrigerator without manually recording or searching, saving them time and effort in finding food. Therefore, the refrigerator in this application embodiment achieves accurate perception of the storage location of food, especially the storage location in the vertical direction, with lower cost and computational overhead, which can improve the reliability of food guidance and thus significantly improve the intelligence level of refrigerator food management and user experience.

[0052] In one embodiment, step S230 determines the calibration height coordinates of the target access location based on the initial height coordinates and the set of reference height coordinates for each storage location inside the refrigerator, including the following steps: Step S231: Compare the initial height coordinates with multiple reference height coordinates; Step S232: Based on the comparison results, select a target reference height coordinate from the set of reference height coordinates as the calibration height coordinate of the target access position, wherein the difference between the target reference height coordinate and the initial height coordinate is not greater than the difference between other reference height coordinates and the initial height coordinate; Step S240 determines the storage location information corresponding to the target access location based on the calibration height coordinates, including the following steps: Step S241: Find the storage location identifier of the target storage location that has a mapping relationship with the calibration height coordinate from the pre-stored storage location mapping table, and use it as the storage location information corresponding to the target access location. The storage location mapping table represents the mapping relationship between multiple reference height coordinates and each storage location.

[0053] In this embodiment, the initial height coordinates can be the uncorrected coordinate values ​​of the food in the refrigerator's height direction, directly calculated from depth information. The reference height coordinate set can be a pre-stored set of standard height values ​​that correspond one-to-one with each shelf or storage level inside the refrigerator, such as 150mm for the first shelf, 450mm for the second shelf, etc. Multiple reference height coordinates can correspond to each storage location, with each reference height coordinate uniquely identifying a storage location level. The calibration height coordinates can be the height coordinate values ​​belonging to a specific standard storage location level, obtained by matching and correcting the initial height coordinates with the reference height coordinate set. The target reference height coordinate can be the reference height coordinate that is closest to the initial height coordinates.

[0054] In one example, in step S231, the controller can calculate the difference between the initial height coordinates and each reference height coordinate to obtain a set of difference sequences. In another example, in step S231, the controller can compare the initial height coordinates with each reference height coordinate in sequence to determine whether the initial height coordinates fall within the preset height range corresponding to each reference height coordinate.

[0055] In step S232, various suitable methods can be used to select a target reference height coordinate from the set of reference height coordinates based on the comparison results, which will serve as the calibration height coordinate for the target access position. In one example, the controller can select the reference height coordinate with the smallest difference from the difference sequence calculated in step S231 as the target reference height coordinate. In another example, the controller can set a preset error tolerance range (e.g., ±20mm) to correct the initial height coordinate to the closest reference height coordinate whose difference is within the error tolerance range. In yet another example, if the initial height coordinate falls in the middle region of two reference height coordinates and the difference is equal, the refrigerator controller can select one of them as the target reference height coordinate according to a preset priority rule (e.g., rounding up or rounding down).

[0056] In the above solution, by comparing the initial height coordinates with a preset set of reference height coordinates and selecting the target reference height coordinates with the smallest difference as the calibration height coordinates, the refrigerator controller can effectively reduce positioning deviations caused by image acquisition errors or inaccurate food placement, thereby calibrating the food position to the standard height closest to the actual storage location. Furthermore, by querying a pre-stored storage location mapping table, the calibrated height coordinates are directly mapped to specific storage location identifiers, achieving an accurate conversion from visual information to physical storage location information. This significantly improves the accuracy and reliability of the refrigerator's judgment of food storage and retrieval locations, providing a solid foundation for refined and automated food inventory management, reducing inventory data errors caused by inaccurate positioning, and enhancing the user experience.

[0057] In one embodiment, the image acquisition device includes an RGB-D camera, and the pick-up and put-down image includes a color image acquired by the RGB-D camera and a corresponding depth image. Step S220 determines the initial height coordinates of the target food item's storage location in the refrigerator based on the depth information of the pick-up and put-down image, including the following steps: Step S221: Determine the two-dimensional image coordinates of the centroid pixel of the target ingredient based on the ingredient segmentation mask of the target ingredient. The ingredient segmentation mask is obtained by semantic segmentation of the color image, and the centroid pixel is the center pixel of the pixel region corresponding to the ingredient segmentation mask. Step S222: Based on the two-dimensional image coordinates of the centroid pixel, read the depth pixel value of the corresponding pixel of the centroid pixel from the depth image; Step S223: Determine the three-dimensional image coordinates of the target access position in the image coordinate system based on the two-dimensional image coordinates of the centroid pixel and the depth pixel value. Step S224: Based on the camera intrinsic parameter matrix of the RGB-D camera and the pre-calibrated pose matrix of the RGB-D camera relative to the refrigerator, the three-dimensional image coordinates are converted into three-dimensional spatial coordinates in the refrigerator coordinate system. The three coordinate axes of the three-dimensional spatial coordinates correspond to the length direction, width direction and height direction of the refrigerator, respectively, and the height direction is parallel to the line of sight of the RGB-D camera. Step S225: Extract the height coordinate components from the three-dimensional spatial coordinates and use them as the initial height coordinates.

[0058] In this embodiment, the image acquisition device can be an RGB-D camera. An RGB-D camera is a device capable of simultaneously capturing color images (RGB information) and depth images (D information). It can employ structured light principles, such as the Intel RealSense series or Microsoft Kinect series, to calculate depth by projecting a known pattern and analyzing its deformation; or it can employ the time-of-flight (ToF) principle, acquiring depth information by measuring the time difference between light emission and reception. RGB-D cameras can provide rich visual and spatial information, laying the foundation for subsequent food identification and positioning.

[0059] In this embodiment, the images to be captured and placed may include a color image and a corresponding depth image captured by an RGB-D camera. The color image provides information such as the visual texture and color of the food, which helps to identify the type and appearance characteristics of the food; the depth image provides the distance information between each pixel in the image and the camera, i.e., the depth value. These two types of images are usually acquired synchronously by the RGB-D camera in a single acquisition operation, and are correlated at the pixel level or can be aligned through calibration.

[0060] In this embodiment, the food segmentation mask can be a binary mask output after processing the color image using a semantic segmentation model, where each pixel is labeled as belonging to or not belonging to the target food. Semantic segmentation is an image processing technique that can classify each pixel in an image into a predefined category, thereby identifying the precise boundaries of different objects in the image. For example, a pre-trained deep learning model (such as U-Net, DeepLab series, etc.) can be trained on common food items in a refrigerator to recognize and segment various food items in an image.

[0061] For example, the centroid pixel can be the geometric center point of the pixel region corresponding to the food segmentation mask, and its coordinates can be obtained by calculating the average of the coordinates of all pixels within the mask region. The two-dimensional image coordinates can be the horizontal and vertical coordinates of the centroid pixel on the image plane, usually in pixels.

[0062] In step S221, various suitable methods can be used to determine the two-dimensional image coordinates of the centroid pixel. In one example, the controller can perform semantic segmentation on the color image of the pick-up and drop-down image to obtain a pixel-level mask of the target ingredient, and then calculate the average horizontal and vertical coordinates of all pixels within the mask region as the two-dimensional image coordinates of the centroid pixel. In another example, the controller uses connected component analysis to extract the largest connected region of the target ingredient and calculates the centroid coordinates of that region to eliminate the influence of sporadic noise. In yet another example, if the mask region of the target ingredient includes multiple discontinuous regions, the controller can select the region with the largest area to calculate the centroid coordinates.

[0063] In this embodiment, the depth pixel value can be the distance value stored in the corresponding pixel position in the depth image, representing the distance from that point to the light-sensitive plane of the RGB-D camera, usually expressed in millimeters or normalized units. The depth image and the color image are pre-aligned in pixel coordinates, so the same coordinate position corresponds to the same spatial point.

[0064] In step S222, in one example, the controller can directly use the two-dimensional image coordinates of the centroid pixel as an index to read the pixel value at the corresponding position in the depth image as the depth value. In another example, if the depth image contains holes or noise, the controller can take a small neighborhood (such as a 3×3 window) centered on the centroid pixel and calculate the mean or median of the effective depth pixel values ​​within that neighborhood as the depth value. In yet another example, the controller can first interpolate and fill the depth image before reading the depth value at the centroid pixel.

[0065] In this embodiment, the three-dimensional image coordinates of the image coordinate system can be represented as three-dimensional coordinates with the camera's optical center as the origin and the image plane as the reference, for example, denoted by (u,v,d), where u and v are pixel coordinates and d is the depth value. It can be understood that the three-dimensional image coordinates do not yet take into account the camera's intrinsic and extrinsic parameters and require further transformation to obtain the true position in physical space.

[0066] In step S223, various suitable methods can be used to determine the three-dimensional image coordinates of the centroid pixel of the target food ingredient. For example, the controller can directly combine the two-dimensional image coordinates (u,v) of the centroid pixel with the depth value d into a triple (u,v,d) as the three-dimensional image coordinates.

[0067] In this embodiment, the camera intrinsic parameter matrix can be a 3×3 matrix describing the camera's optical characteristics, including parameters such as focal length and principal point coordinates, used to convert pixel coordinates into normalized coordinates in the camera coordinate system. The pose matrix can be a 4×4 homogeneous transformation matrix describing the rotation and translation relationship between the camera coordinate system and the refrigerator coordinate system, obtained through pre-calibration. The refrigerator coordinate system can be a three-dimensional Cartesian coordinate system fixed on the refrigerator's inner cavity, with its origin set at a fixed position inside the refrigerator cavity (such as the lower left rear wall), and coordinate axes along the length, width, and height directions of the refrigerator, respectively.

[0068] In step S224, various suitable methods can be used for coordinate transformation. In one example, the controller first converts the 3D image coordinates (u, v, d) of the target food texture's core pixel into 3D coordinates (x_c, y_c, z_c) in the camera coordinate system using the inverse of the camera intrinsic matrix, and then uses the pose matrix to transform these coordinates into 3D spatial coordinates (x_w, y_w, z_w) in the refrigerator coordinate system. In another example, the controller directly uses the joint transformation matrix (the inverse of the product of the intrinsic and extrinsic matrices) to complete the mapping from pixel coordinates to world coordinates in one step. In yet another example, if camera distortion exists, the controller can first correct the pixel coordinate distortion before performing coordinate system transformation.

[0069] In this embodiment, the coordinate components in the height direction can be the components corresponding to the height axis in the refrigerator coordinate system, typically the components of the Y-axis or Z-axis (e.g., ...). Figure 3a The z-axis shown in the figure depends on the definition of the coordinate system. The initial height coordinates can be the original coordinates of the food in the height direction, directly calculated from the depth information without diaphragm constraint correction.

[0070] In step S225, various suitable methods can be used to extract the height coordinates. In one example, the controller directly takes the value of the corresponding height axis in the three-dimensional spatial coordinates as the initial height coordinates. In another example, if the height axis of the refrigerator coordinate system is the Y-axis, the controller extracts the y_w component of the three-dimensional coordinates (x_w, y_w, z_w) as the initial height coordinates.

[0071] Continue to refer to Figure 3aAs a user opens the refrigerator to retrieve an apple, an RGB-D camera located at the top front of the refrigerator captures an image of the user taking and placing the apple. After acquiring this image, the controller proceeds with the height coordinate extraction process. First, before step S221, the controller performs semantic segmentation on the color image captured by the RGB-D camera, generating a food segmentation mask corresponding to the apple. In step S221, the geometric center of this mask region is further calculated to determine the two-dimensional coordinates of the apple's centroid pixel on the image plane. Next, in step S222, the controller uses these two-dimensional coordinates as an index to accurately read the depth pixel value corresponding to the centroid pixel from the synchronously acquired depth image. This value represents the vertical distance between the apple's current position and the camera. Subsequently, in step S223, the controller fuses the two-dimensional image coordinates with the depth value to construct the three-dimensional image coordinates of the apple's centroid in the camera image coordinate system. Then, in step S224, the controller calls a pre-stored camera intrinsic parameter matrix and a pre-calibrated camera pose transformation matrix relative to the refrigerator body to convert the three-dimensional image coordinates into physical three-dimensional spatial coordinates in the refrigerator coordinate system. Finally, in step S225, the controller extracts the coordinate components corresponding to the refrigerator's height direction from the converted three-dimensional spatial coordinates and uses them as the initial height coordinates of the apple's current location. In this way, by deeply fusing visual and spatial information, the apple in the two-dimensional image is transformed into a three-dimensional coordinate point with precise physical height, providing accurate raw data support for subsequent constraint calibration of these coordinates based on the reference height of the refrigerator's internal shelves.

[0072] In the above scheme, an RGB-D camera can simultaneously acquire color and depth images, providing comprehensive data for food identification and localization. By performing semantic segmentation on the color image, the region of the target food can be accurately identified, and the two-dimensional image coordinates of its centroid pixel can be further determined, thus reducing depth measurement errors caused by irregular food shapes or partial occlusion. Subsequently, the depth value of the centroid pixel is read from the depth image, and combined with the camera intrinsic parameter matrix and a pre-calibrated camera pose matrix, the three-dimensional position of the food in the camera coordinate system is accurately transformed to the refrigerator coordinate system. By setting the height direction of the refrigerator coordinate system parallel to the line of sight of the RGB-D camera, the height coordinate component can be directly extracted from the transformed three-dimensional spatial coordinates as the initial height coordinates of the target food. This improves the accuracy and reliability of determining the initial height coordinates of the target food in the refrigerator, significantly enhancing the accuracy of food storage and retrieval location positioning.

[0073] In one embodiment, the refrigerator control method of this application further includes the following steps: Step S2401: Based on the coordinate components of the length and width directions in the three-dimensional spatial coordinates, determine the target partition position of the target access position on the same storage surface of the refrigerator, wherein the storage surface is perpendicular to the height direction. Step S240 determines the storage location information corresponding to the target access location based on the calibration height coordinates, including the following steps: Step S242: Find the storage location identifier of the target storage location that has a mapping relationship with both the calibration height coordinate and the target partition location from the pre-stored storage location mapping table. Use this identifier as the storage location information corresponding to the target access location. The storage location mapping table represents the mapping relationship between each storage location and multiple reference height coordinates and multiple partition locations.

[0074] It is understandable that, considering that the same horizontal plane inside the refrigerator may correspond to multiple storage locations, step S2401 utilizes the horizontal position information of the target food inside the refrigerator to further refine its specific area on a particular storage surface, thereby accurately locating the storage location of the target food. Specifically, the controller can pre-store the two-dimensional planar layout information of each storage surface (e.g., each shelf or drawer of the refrigerator), which divides the storage surface into multiple predefined partitions. After obtaining the length and width coordinate components in the three-dimensional spatial coordinates of the target food, the controller can compare these coordinates with the predefined partition boundaries to determine which partition the target food falls into, i.e., the target partition location. Alternatively, this can be achieved through a machine learning model. For example, a model can be pre-trained that takes the two-dimensional coordinates of the target food on the storage surface as input and outputs the target partition location to which it belongs. Training data can include food storage examples with labeled partition information.

[0075] In step S242, the storage location mapping table represents the mapping relationship between each storage location and multiple reference height coordinates and multiple partition locations. For example, the storage location mapping table can be a two-dimensional lookup table or database, where each row or entry contains a unique storage location identifier, along with associated reference height coordinates and partition location information. After the controller obtains the calibration height coordinates and target partition location of the target food ingredient's target storage location, it can perform a joint query in the storage location mapping table to find a unique storage location identifier that matches both conditions. Furthermore, the storage location mapping table can also be implemented using programming logic, such as a series of conditional statements. In this embodiment, the controller can first determine the storage surface based on the calibration height coordinates, and then further determine the specific storage location identifier on that storage surface based on the target partition location.

[0076] In the above scheme, the depth information of the images captured by the image acquisition device is first used, and combined with camera intrinsic parameters and pose matrix, the two-dimensional image coordinates and depth pixel values ​​of the target food are converted into three-dimensional spatial coordinates in the refrigerator coordinate system. These three-dimensional spatial coordinates contain precise positional information of the target food in the length, width, and height directions within the refrigerator. Based on this, the coordinate components in the length and width directions of these three-dimensional spatial coordinates are further used to determine the target partition position of the target food on the same storage surface within the refrigerator. This means that, in addition to the height information in the vertical direction, this embodiment also considers the specific area division in the horizontal direction of the refrigerator. Subsequently, the controller no longer determines the storage location information solely based on the calibration height coordinates, but instead searches a pre-stored storage location mapping table for the storage location identifier that has a mapping relationship with both the calibration height coordinates and the target partition position. This storage location mapping table is designed to represent the mapping relationship between each storage location and multiple reference height coordinates and multiple partition positions, thereby achieving refined management of the storage locations inside the refrigerator. In this way, even if multiple logical partitions exist on the storage surface at the same height, the refrigerator can accurately identify the specific storage location of the food, greatly improving the accuracy and precision of food inventory data updates. This dual positioning mechanism, which combines height and horizontal partitioning information, enables the refrigerator to manage its internal space more intelligently, providing users with more accurate food storage and retrieval records and management services.

[0077] In one embodiment, the refrigerator further includes a door status sensor configured to detect door opening and closing events. The door status sensor can be implemented in various suitable ways, including but not limited to a magnetic door switch, a Hall effect sensor, or a limit switch. For example, the door status sensor can be installed at the junction of the refrigerator door and the refrigerator body, and can output an open level signal when the refrigerator door is opened and a closed level signal when it is closed, allowing the refrigerator controller to identify changes in the door's state.

[0078] The refrigerator control method of this application embodiment further includes the following steps S201 to S207.

[0079] Step S201: Based on the door opening event and door closing event detected by the door status sensor, the time axis is divided into a steady-state segment and a variable segment, wherein the variable segment is the time period from the detection of the door opening event to the detection of the door closing event.

[0080] In this embodiment, the steady-state segment can refer to the time period when the refrigerator door is closed and there is no external interference inside the refrigerator; the changing segment can refer to the time period when the refrigerator door is open and the user may be taking or putting away food. A door opening event can be a signal detected by the door status sensor that the door changes from a closed state to an open state; a door closing event can be a signal detected by the door status sensor that the door changes from an open state to a closed state.

[0081] In one example, the controller can monitor the output signal of the door status sensor in real time. When it detects a signal level change from low to high (or from high to low, depending on the sensor type), it records it as a door opening event and marks the current time as the start time of the change segment. When it detects a reverse signal transition, it records it as a door closing event and marks the current time as the end time of the change segment. Furthermore, the controller can employ a software debouncing algorithm to continuously sample the door status signal, confirming an event only after the signal has stabilized for a preset time (e.g., 50 milliseconds) to avoid false triggers. For example, the controller can also record the timestamps of all change segments within a day in a log for subsequent analysis of user behavior patterns.

[0082] In step S202, during the change segment, the image acquisition device is controlled to acquire images of the refrigerator at a preset frequency to obtain multiple visual images.

[0083] In this embodiment, the visual image can refer to a sequence of image frames containing RGB information and depth information continuously acquired by the image acquisition device within the changing segment. The preset frequency can be a pre-set number of frames acquired per second, such as 15 frames / second, 30 frames / second, or 60 frames / second. For example, the image acquisition frequency in the changing segment can be higher than the image acquisition frequency in the steady-state segment. For example, no images may be acquired or images may be acquired at a lower frequency in the steady-state segment. This allows for accurate capture of the user's picking and placing operations during the changing segment while reducing unnecessary computing power consumption and storage occupation, thereby reducing the overall energy consumption of the refrigerator.

[0084] In one example, the controller can activate the image acquisition device after detecting a door opening event, continuously acquiring images at a frequency of 30 frames per second until a door closing event is detected, at which point acquisition stops. In another example, the controller can dynamically adjust the acquisition frequency based on the user's historical behavior, for example, increasing the frequency during peak hours and decreasing the frequency during low-frequency hours to conserve storage and computing resources. In yet another example, the controller can trigger high-frequency acquisition only when a hand is detected entering the frame, and decrease the acquisition frequency after the hand leaves the frame to reduce the generation of invalid data.

[0085] Step S203: For each visual image, input the visual image into the semantic segmentation model to obtain the hand segmentation mask of the user's hand, the food segmentation mask of each food item in the visual image, and the food category corresponding to the food segmentation mask.

[0086] In this embodiment, the semantic segmentation model can be a deep learning-based image segmentation network, such as VisionTransformer, U-Net, or DeepLab, capable of assigning a category label to each pixel in the image. The hand segmentation mask can be a binary mask that identifies only the pixel region of the user's hand. The food segmentation mask can be a binary mask that identifies the pixel region of each food item; each food item instance can have an independent mask. The food category can refer to the type of food, such as "apple," "milk," or "egg."

[0087] In step S203, various suitable methods can be employed for semantic segmentation. In one example, the controller inputs each visual image into a pre-trained Vision Transformer multi-task network, which simultaneously outputs a hand segmentation mask, a food segmentation mask, and a class confidence vector for each food instance. In another example, the controller first uses a lightweight object detection network (such as YOLO) to detect the bounding boxes of the hand and food, and then performs refined semantic segmentation on each bounding box region to improve processing speed. In yet another example, the controller uses an instance segmentation network (such as Mask R-CNN) to directly output the mask and class of each food instance, without post-processing.

[0088] Step S204: Based on the relative positional relationship between the hand segmentation mask and each food segmentation mask, determine the food contact state corresponding to the visual image. The food contact state includes an uncontacted state indicating that the user's hand is not in contact with any food, and a contact state indicating that the user's hand is in contact with any food.

[0089] In this embodiment, the relative positional relationship can refer to the degree of overlap or distance between the hand mask and the food mask in the image space. The food contact state can refer to the determination result of whether the user's hand has made physical contact with a certain food in the current frame.

[0090] In step S204, various suitable methods can be used to determine the food contact state. In one example, the controller calculates the intersection-over-union (IoU) ratio between the hand segmentation mask and each food segmentation mask. When the IoU is greater than a preset threshold (e.g., 0.3), the food is determined to be in contact; if all IoUs are less than the threshold, it is determined to be in a non-contact state. In another example, the controller calculates the minimum Euclidean distance between the boundary of the hand mask and the boundary of the food mask. When the distance is less than a preset value (e.g., 10 pixels), it is determined to be in contact. In yet another example, the controller combines depth information to determine whether the difference between the average depth of the hand mask area and the average depth of the food mask area is less than a preset threshold to confirm whether they are on the same depth plane, thereby more accurately determining whether physical contact has occurred.

[0091] Step S205: Based on the difference in hand segmentation masks of adjacent visual images in multiple visual images, determine the hand movement direction corresponding to the adjacent visual images. The hand movement direction includes the insertion direction indicating the hand reaching into the refrigerator and the withdrawal direction indicating the hand withdrawing from the refrigerator.

[0092] In this embodiment, the difference in the hand segmentation mask can refer to the amount of change in the position, shape, or area of ​​the hand mask between adjacent frames. The direction of hand movement can refer to the trend of hand movement within the refrigerator cavity; the direction of insertion indicates that the hand moves deeper into the refrigerator, and the direction of withdrawal indicates that the hand moves outward from the refrigerator.

[0093] In step S205, various suitable methods can be used to determine the direction of hand movement. In one example, the controller calculates the centroid displacement vector of the hand mask in adjacent frames. When the centroid moves along the depth direction of the refrigerator (e.g., the positive Z-axis), it is determined to be the insertion direction; when it moves in the opposite direction, it is determined to be the withdrawal direction. In another example, the controller calculates the area change rate of the hand mask. When the area increases, it is determined to be the insertion direction (the hand is closer to the camera, and the image becomes larger); when the area decreases, it is determined to be the withdrawal direction. In yet another example, the controller combines the centroid displacement and the area change rate for a comprehensive judgment. For example, when the centroid moves into the refrigerator and the area increases, it is determined to be the insertion direction.

[0094] Step S206: Based on the hand movement state corresponding to each adjacent visual image in multiple visual images, the food contact state corresponding to each visual image, and the food category of the contacted food, a visual change feature sequence of user hand-food interaction within the change segment is generated. The visual change feature sequence is a sequence of multiple event elements. Each event element corresponds to a continuous unidirectional hand movement event. The value of the event element includes the hand movement direction and timestamp of the hand movement event, the food contact state corresponding to each visual image within the timestamp, and the food category of the contacted food.

[0095] In this embodiment, the visual change feature sequence can be a structured data sequence that abstracts the interaction process between the hand and the food within the entire change segment. Event elements can be the basic units in the sequence, with each event element representing a continuous unidirectional hand movement (such as a continuous extension or withdrawal motion). The timestamp can be the start or end time of this hand movement segment.

[0096] In step S206, various suitable methods can be used to generate the visual change feature sequence. In one example, the controller traverses all keyframes, merges consecutive frames with the same hand movement direction into one event element, and records the direction, start and end timestamps, food contact state of each frame within the segment, and the category of the contacted food. In another example, the controller creates new event elements only when the hand movement direction changes or the food contact state changes, to reduce the sequence length. Furthermore, the controller can attach statistical information to each event element, such as the frequency of contact with food within the segment and the hand movement speed, to enrich the feature representation.

[0097] In a specific example, combined Figure 3b and Figure 3c When a user opens the refrigerator door, reaches their hand inside, touches an apple, and then removes their hand from the refrigerator while holding the apple, the resulting visual change feature sequence will contain three event elements: the first event element corresponds to the direction of reaching in, with a timestamp representing the time from when the door opens until the hand touches the apple (during which time the food is in a non-contact state); the second event element corresponds to the direction of leaving, with a timestamp representing the time from when the hand touches the apple until leaving the refrigerator door (during which time the food is in a contact state, and the contacted food type is apple). This visual change feature sequence can be represented as: [(reaching in, t0-t1, non-contact, none), (leaving out, t1-t2, contact, apple)]. Through this structured feature extraction, the entire process of the user's retrieval and placement operation can be clearly identified, providing an accurate basis for subsequent determination of the type of retrieval and placement action.

[0098] Step S207: Based on the visual change feature sequence, determine the user's handling status of ingredients within the change segment. The handling status includes the following: adding ingredients, taking out ingredients, storing and retrieving ingredients (meaning adding one ingredient and taking out another), consuming ingredients, and no change in ingredients.

[0099] In this embodiment, the processing state can refer to the type of operation performed by the user on the food during this door opening. The "Add Food" state indicates that the user puts the food into the refrigerator; the "Remove Food" state indicates that the user takes the food out of the refrigerator; the "Store and Retrieve Food" state indicates that the user puts in one type of food and takes out another type of food at the same time; the "Consume Food" state indicates that the user takes out part of the food, uses it, and then puts it back (such as a half-cut watermelon); and the "No Change in Food" state indicates that the user does not perform any actual operation on the food (such as just viewing it).

[0100] In step S207, various suitable methods can be used to determine the disposal state. In one example, the controller uses a rule engine to match patterns in the visual change feature sequence. If the sequence contains two event elements with values ​​of (extending direction, contact state) and (withdrawing direction, non-contact state) respectively, it is determined to be the state of adding food; if the values ​​are (extending direction, non-contact state) and (withdrawing direction, contact state) respectively, it is determined to be the state of removing food. In another example, the controller inputs the visual change feature sequence into a trained classifier (such as a random forest or MLP), outputs the probability distribution of each state, and selects the state with the highest probability as the final judgment result. In yet another example, the controller can combine detection information from other sensors for joint judgment. For example, it can combine the weight change features of each storage compartment in the refrigerator during the change period and the speech semantic features of the received user voice to perform multimodal fusion judgment to improve accuracy in complex scenarios.

[0101] Step S210 acquires the pick-up and drop-off image, including the following steps: Step S211: When the processing state is adding ingredients, taking out ingredients, storing ingredients, or consuming ingredients, select one or more visual images from multiple visual images as the picking and placing images according to the processing state.

[0102] In this embodiment, the image for picking up and placing food can refer to the representative image selected from multiple visual images acquired within the change segment that best reflects the type and location of the target food item. For example, a keyframe can be selected when the user is about to pick up the food after reaching into the refrigerator (e.g., when the hand is last captured by the RGB-D camera in the direction of reaching in, while the user is taking out the food), or when the user is about to remove the food from the storage compartment (e.g., when the hand is just captured by the RGB-D camera in the direction of removing the food, while the user is adding the food).

[0103] In step S211, various suitable methods can be used to select the pick-up and place-up images. For example, pick-up and place-up image filtering rules corresponding to different handling states can be preset, and then in this step, target keyframes that meet the rules can be selected as pick-up and place-up images through rule matching. Alternatively, a deep learning-based keyframe detection network can be used to directly locate the frame containing the target food ingredient and with the clearest interaction action from the sequence of changing segment images, and automatically output it as the pick-up and place-up image. Yet another example is selecting the most recent frame image when the food ingredient's contact state has just changed as the pick-up and place-up image.

[0104] Step S250 updates the refrigerator's food inventory data based on the target food's category and storage location information, including the following steps: Step S251: Update the refrigerator's food inventory data based on the target food's category information, storage location information, and the user's handling status of the target food.

[0105] For example, when the disposal status is "adding ingredients", the controller can add a new record in the database, including the ingredient category, storage location identifier, current timestamp, and default quantity; when the disposal status is "retrieving ingredients", the controller marks the corresponding ingredient record as retrieved and records the retrieval time; when the disposal status is "consuming ingredients", the controller can update the weight field of the corresponding ingredient, subtract the consumption amount, and retain other information; when the disposal status is "storing or retrieving ingredients", the controller performs update operations for both adding and retrieving ingredients simultaneously.

[0106] Through the aforementioned technical solution, the refrigerator can dynamically sense the user's interaction with food and accurately determine the user's intentions based on visual change feature sequences, such as adding, removing, storing, consuming, or simply viewing food. This refined understanding of user behavior allows the refrigerator to intelligently select the image that best reflects the actual operation and accurately update the food inventory data based on the processing status. Compared to methods that rely solely on a single image or simple door opening / closing events for judgment, this solution significantly improves the accuracy and reliability of food inventory data, effectively reducing inventory errors caused by misjudging user intentions, thus achieving more intelligent and user-friendly refrigerator food management.

[0107] In one embodiment, step S207 determines the user's handling status of the food within the change segment based on the visual change feature sequence, including at least one of the following steps S207a to S207e.

[0108] In step S207a, if the number of event elements in the visual change feature sequence is 2, and the value of the first event element corresponds to the insertion direction and contact state, and the value of the second event element corresponds to the withdrawal direction and non-contact state, then the handling state is determined to be the addition of ingredients state.

[0109] In this embodiment, the number of event elements can refer to the total number of event elements contained in the visual change feature sequence. The direction of reaching in can refer to the trend of the hand moving towards the inside of the refrigerator. The contact state can characterize the state of the hand contacting the food within the time period corresponding to the current event element. Specifically, it can be determined that the hand is in contact with the food based on whether the hand mask overlaps with a certain food mask or the distance is less than a distance threshold. The direction of withdrawal can refer to the trend of the hand moving towards the outside of the refrigerator. The non-contact state can refer to the state of the hand not contacting the food within the time period corresponding to the current event element. Specifically, it can be determined that the hand is not in contact with the food based on whether the hand mask overlaps with any food mask or the distance is not less than a distance threshold. The food addition state can refer to the user's operation type of placing food into the refrigerator.

[0110] For example, the controller can parse the visual change feature sequence, determine whether the number of event elements is 2, and verify one by one whether the hand movement direction of the first event element is the reaching direction and whether the food contact state is the contact state; and whether the hand movement direction of the second event element is the withdrawing direction and whether the food contact state is the non-contact state. If all matches, it is determined to be a food addition state. In addition, the controller can also confirm the category of food contact in the first event element before making the determination to ensure that there is valid food category information, so as to reduce false detections. In another example, the controller can also combine the weight change characteristics of the food in the refrigerator storage space during the change segment for verification. If the weight of the food in the storage space increases, the determination can be further confirmed.

[0111] like Figure 3b As shown, when a user's hand holding an apple enters the camera's field of view, by analyzing the visual images continuously captured by the RGB-D camera, it is possible to capture the hand touching the apple and simultaneously moving towards the storage space (i.e., along...). Figure 3a (As shown in the image, moving in the opposite direction of the y-axis) until the hand and apple are out of the camera's view, the user places the apple into the storage compartment, and then the hand is withdrawn from inside the storage compartment towards the outside of the refrigerator. Figure 3c As shown, the RGB-D camera can capture a hand that is not holding an apple entering the field of view and gradually moving towards the outside of the refrigerator (i.e., along...). Figure 3a The hand moves in the positive direction of the y-axis (as shown in the diagram). Thus, during this time, the generated visual change feature sequence contains two event elements: in the first event element, the hand movement direction is the insertion direction, the food contact state is "contact state," and the contacted food type is apple; in the second event element, the hand movement direction is the withdrawal direction, and the food contact state is "not contact state." Therefore, based on the logic of step S207a, the controller can quickly and accurately determine that this operation is an addition of food.

[0112] In step S207b, if the number of event elements in the visual change feature sequence is 2, and the value of the first event element corresponds to the insertion direction and the non-contact state, and the value of the second event element corresponds to the withdrawal direction and the contact state, then the handling state is determined to be the food removal state.

[0113] It is understandable that the core of step S207b lies in recognizing the visual feature sequence when a hand moves an item from inside the refrigerator to the outside. Similar to the logic of determining the state of added food in step S207a, this step can also accurately determine the state of the removed food based on the number and value of event elements in the visual change feature sequence. Figure 3d As shown, when a user's empty hand enters the camera's field of view from outside the refrigerator, the RGB-D camera continuously captures visual images, detecting the hand moving from the empty hand towards the storage compartment until it leaves the camera's view. Subsequently, after the user grabs the desired apple, their hand simultaneously withdraws from the storage compartment towards the outside of the refrigerator. The RGB-D camera can capture the hand entering the field of view and gradually moving towards the outside of the refrigerator. Thus, during this period, the generated visual change feature sequence contains two event elements: in the first event element, the hand's movement direction is the insertion direction, and the food contact state is the non-contact state; in the second event element, the hand's movement direction is the withdrawal direction, the food contact state is the contact state, and the food type is apple. Therefore, based on the logic of step S207b, the controller can quickly and accurately determine that this operation is a food removal state.

[0114] In step S207c, if the number of event elements in the visual change feature sequence is 4, and the value of the first event element corresponds to the insertion direction and the non-contact state, the value of the second event element corresponds to the withdrawal direction and the contact state, the value of the third event element corresponds to the insertion direction and the contact state, and the value of the fourth event element corresponds to the withdrawal direction and the non-contact state, and the food categories of the contacted food in the second and third event elements are the same, then the disposal state is determined to be the food consumption state.

[0115] In a specific application scenario, a user opens the refrigerator door, takes a whole watermelon from a storage compartment, cuts it in half on the counter behind the refrigerator, puts the remaining half back in the refrigerator, and closes the door. In this example, the user first reaches into the storage compartment with their empty hand. The controller can then extract the first event element from the image captured by the camera, with values ​​including: direction of entry and non-contact state. Next, the user grabs the whole watermelon from inside the compartment and removes it. The controller can then extract the second event element from the image captured by the camera, with values ​​including: direction of withdrawal, contact state, and food type: watermelon. Subsequently, the user cuts the watermelon in half on the counter outside the refrigerator, takes out one half, and then reaches into the refrigerator again with the remaining half in their hand. The controller can then extract the third event element from the image captured by the camera, with values ​​including: direction of entry, contact state, and food type: watermelon. Finally, the user puts the remaining half of the watermelon back in the storage compartment and withdraws it empty-handed. The controller can then extract the fourth event element from the image captured by the camera, with values ​​including: direction of withdrawal and non-contact state. During this process, although the watermelon's appearance and size changed significantly (from whole to half), the controller accurately identified the second and third encounters with the same "watermelon" category by recognizing the food type. Therefore, it could classify this series of actions as a "removal-partial consumption-return" consumption behavior, rather than simply removal or addition. Based on this four-event loop pattern, the controller could determine the disposal state as a food consumption state.

[0116] In step S207d, if the number of event elements in the visual change feature sequence is 2, and the value of the first event element corresponds to the insertion direction and contact state, and the value of the second event element corresponds to the withdrawal direction and contact state, and the food categories of the contacted food corresponding to the first event element and the second event element are different, then the handling state is determined to be the food storage and retrieval state.

[0117] In a specific application scenario, a user opens the refrigerator door, places a bag of apples into a storage compartment, and then randomly takes a banana from the compartment. In this example, after opening the refrigerator door, the user first reaches their hand, holding the bag of apples, into the storage compartment. The controller can then extract the first event element from the image captured by the camera, with values ​​including: direction of insertion, contact state, and the food type being apples. Subsequently, after placing the bag of apples into the compartment, the user doesn't withdraw their hand empty-handed but instead grabs a banana and takes it out. The second event element can then be extracted from the image captured by the camera, with values ​​including: direction of withdrawal, contact state, and the food type being bananas. During this process, the controller determines that the first and second event elements involve different food types (apples and bananas), thus accurately identifying this series of actions as a storage and retrieval behavior of "simultaneously adding one food item and removing another," rather than a simple addition or removal. Based on these two event patterns (first touching and then withdrawing from the contact, and the types of food touched are different), the controller can determine the food handling status during the door opening period as the food storage and retrieval status.

[0118] In step S207e, if the values ​​of all event elements in the visual change feature sequence correspond to the untouched state, the handling state is determined to be the state where the food has not changed.

[0119] In a specific application scenario, a user opens the refrigerator door, reaches inside to rummage through or arrange food, but does not actually take or put anything in, and then closes the door. It's understandable that in this example, the user's hand may have reached in and out one or more times after opening the refrigerator door, but it never actually came into contact with any food during this process. Throughout this process, the controller extracts all event elements from the images captured by the camera, and all of these events show a "no contact" state; that is, no event element records any contact between the hand and food. Based on this, the controller can determine that the user did not perform any substantial operation on the food during this door-opening period, and the status is "no change in food status."

[0120] Through the above technical solution, the refrigerator can efficiently and accurately identify various complex user operations inside the refrigerator based on the combination patterns of event elements in the visual change feature sequence. These operations include adding food, removing food, consuming food, storing food, and food remaining unchanged. This judgment logic reduces ambiguity in user behavior and significantly improves the accuracy and reliability of handling status recognition.

[0121] In one embodiment, step S211 selects one or more visual images from a plurality of visual images as pick-up and drop images according to the processing state, including the following steps: Step S2111: When the processing state is adding ingredients, taking out ingredients, storing ingredients, or consuming ingredients, the visual image with the same timestamp as the latest timestamp of the first target event element and / or the visual image with the same timestamp as the earliest timestamp of the second target event element is used as the pick-up and put-down image. The first target event element and the second target event element are located in the visual change feature sequence. The value of the first target event element corresponds to the insertion direction and contact state, and the value of the second target event element corresponds to the withdrawal direction and contact state.

[0122] In this embodiment, the first target event element can refer to an event element in the visual change feature sequence where the hand movement direction is the insertion direction and the food contact state is the contact state, corresponding to the path of the hand carrying the food into the refrigerator. The second target event element can refer to an event element where the hand movement direction is the withdrawal direction and the food contact state is the contact state, typically corresponding to the event of the hand carrying the food moving out of the refrigerator. The latest timestamp can refer to the acquisition time of the last keyframe within the time period corresponding to the event element. The earliest timestamp can refer to the acquisition time of the first keyframe within the time period corresponding to the event element.

[0123] In this embodiment, the controller can use the visual image corresponding to the latest timestamp of the first target event element (reaching direction and contact state) as the pick-up / placement image. This latest timestamp corresponds to the critical moment when the user's hand, grasping the food, reaches into the refrigerator, is about to reach the target storage location, and is about to release it. At this moment, the food is exactly above or in front of the target storage location, and its spatial position is closest to the target storage location. Therefore, the depth value of the food texture pixels extracted from this image can most accurately represent the actual depth and height of the target storage location, thus providing the most reliable raw data for subsequent coordinate correction and storage location binding. For example, when the handling state is adding food, the controller can use the visual image corresponding to the latest timestamp of the first event element (reaching direction and contact state) in the visual change feature sequence as the pick-up / placement image.

[0124] In this embodiment, the controller can use the visual image corresponding to the earliest timestamp of the second target event element (withdrawal direction and contact state) as the pick-up / placement image. This earliest timestamp corresponds to the moment when the user's hand has just taken the food from the storage location and returned to the edge of the storage location. At this time, the centroid height of the food is close to the height of the storage location. Selecting the image at this moment also ensures that the depth value corresponding to the centroid pixel of the food accurately reflects the spatial coordinates of the target storage location, reducing positioning errors caused by the positional shift of the food after it is lifted. For example, when the handling state is the food removal state, the controller can use the visual image corresponding to the earliest timestamp of the second event element (withdrawal direction and contact state) in the visual change feature sequence as the pick-up / placement image.

[0125] For example, when the processing state is the food consumption state or the food storage state, since the bidirectional flow of food is involved, the controller simultaneously selects the visual image corresponding to the latest timestamp of the first target event element and the visual image corresponding to the earliest timestamp of the second target event element, and records the time when the food is closest to the storage location before it is taken out and after it is put back, so as to fully reflect the positional relationship between the food and the storage location.

[0126] In the above scheme, based on the latest timestamp of the first target event element and the earliest timestamp of the second target event in the visual change feature sequence during door opening, the visual image closest to the target food item and the target storage location can be quickly and accurately selected as the retrieval image. This ensures that the centroid pixels extracted from the depth information of subsequent retrieval images can accurately represent the actual coordinates of the target storage location. This image selection strategy based on action thresholds effectively reduces the interference of factors such as hand occlusion, motion blur, and food displacement on depth measurement, making subsequent coordinate correction and storage location binding more accurate and reliable. This significantly improves the accuracy of food spatial positioning during dynamic food storage and retrieval, and enhances the reliability of automated refrigerator food inventory management.

[0127] In one embodiment, the refrigerator further includes a weight sensor array and / or a microphone array, wherein the weight sensor array is disposed below the load-bearing surface of each storage compartment and is configured to detect the total weight of the food items carried in the corresponding storage compartment; the microphone array is configured to collect the user's voice commands. The refrigerator control method of this application embodiment further includes the following steps S208 and / or S209: Step S208: Determine the weight change characteristics of each storage location based on the difference in the total weight of each storage location detected by the weight sensor array at the beginning and end of the change segment. Step S209: Perform speech recognition and semantic encoding on the speech commands within the changing segment to obtain the user's speech semantic features; Step S207 determines the user's handling status of the food within the change segment based on the visual change feature sequence, including the following steps: Step S2071: Input the visual change feature sequence, weight change feature and / or speech semantic feature into the disposal state transition probability model based on multimodal fusion, and output the probability distribution of the user's disposal state of food within the change segment, which belongs to each of the following states: adding food, taking out food, consuming food and no change in food. Step S2072: Determine the disposal status corresponding to the maximum probability value in the probability distribution as the user's disposal status for the ingredients.

[0128] In this embodiment, the weight sensor array can be a collection of sensors for measuring the weight of objects. It can consist of multiple independent weighing sensors (e.g., resistance strain gauge sensors, piezoelectric sensors, capacitive sensors, etc.), which are strategically arranged below the load-bearing surfaces of each storage compartment inside the refrigerator. Each sensor or sensor group is responsible for monitoring the total weight of the food in its assigned compartment. By acquiring data from these sensors in real time or periodically, changes in the weight of the food in the compartment can be accurately detected, thus providing objective physical evidence for judging the user's handling of the food.

[0129] A microphone array is an acoustic sensor system composed of multiple microphone units arranged in a specific geometry. Its main function is to collect sound signals from the environment, especially user voice commands. By processing the sound signals collected by multiple microphones (e.g., beamforming, noise suppression, sound source localization), microphone arrays can effectively improve the accuracy and anti-interference capability of speech recognition. They can be configured in various ways, such as linear arrays, circular arrays, or planar arrays, to adapt to different acoustic environments and application requirements.

[0130] Weight change characteristics can be the change in the total weight of food stored in each compartment of the refrigerator over a specific time period. For example, the weight change characteristics of each compartment can be obtained by comparing the difference in total weight detected at the beginning and end of the change period. For instance, if the weight of a compartment increases, it indicates that food is more likely to have been added; if the weight decreases, it may indicate that food is more likely to have been removed or consumed. This characteristic provides a physical and quantitative basis for judging the user's behavior in handling food.

[0131] Speech-semantic features refer to the information with clear meaning extracted from user speech commands through speech recognition and semantic encoding. Speech recognition technology converts the user's speech into text, while semantic encoding further analyzes the meaning and intent of the text. For example, if a user says "take an apple," after processing, we can obtain semantic information such as "take out" and "apple." This feature can directly reflect the user's clear intent and provide high-level semantic clues for judging the action status.

[0132] A multimodal fusion probabilistic model for handling state transitions can be an intelligent model capable of integrating information from different modalities (such as vision, weight, and speech) and outputting a probability distribution of user handling states. This model can be built based on machine learning or deep learning techniques. For example, it can employ sequence models such as Recurrent Neural Networks (RNNs), Long Short-Term Memory Networks (LSTMs), and Transformers to process visual change feature sequences, while simultaneously using weight change features and speech semantic features as additional input features. By learning the complex relationship between different modal features and user handling states, the model can comprehensively judge user behavior and provide the probability of each handling state in the form of a probability distribution, thereby improving the robustness and accuracy of the judgment.

[0133] In a specific example, the disposal state transition probability model can be expressed using the following formula: P(State|S_vis, S_wgt, S_txt) = sof ; Where State∈{State1,State2,State3,State4,State5}.

[0134] Wherein, S_vis represents the encoded visual change feature sequence, S_wgt represents the encoded weight change feature, S_txt represents the encoded speech semantic feature, MLP represents the multilayer perceptron network, State1 represents the state of adding ingredients, State2 represents the state of taking out ingredients, State3 represents the state of storing and retrieving ingredients, State4 represents the state of consuming ingredients, and State5 represents the state of no change in ingredients.

[0135] In step S2071, the features of the three different modalities can be concatenated and input into a multilayer perceptron network for processing. Finally, the model's predicted probability distribution of the current handling state is output through the softmax function.

[0136] For example, in step S2072, the controller can use the following formula to select the disposal state with the highest probability value from the probability distribution output in step S2071 as the final judgment result: State* = argmax_{State} P(State | S_vis, S_wgt, S_txt); Where State* represents the final optimal disposal state, and argmax represents the operation of taking the state with the highest probability among all candidate disposal states. For example, if the probability distribution of the model output is {State1: 0.75, State2: 0.08, State3: 0.1, State4: 0.05, State5: 0.02}, then the controller can determine the state corresponding to State1 (adding ingredients) as the final disposal state.

[0137] In the above solution, by introducing a weight sensor array and / or a microphone array, and combining it with a multimodal fusion model, the refrigerator can comprehensively analyze multi-source information such as vision, weight, and voice, thereby significantly improving the accuracy and robustness of judging the user's food handling behavior. This allows the refrigerator's food inventory management system to more accurately reflect actual inventory changes, reduce misjudgments, improve user experience, and provide a more reliable data foundation for subsequent intelligent recommendation and management functions. Especially in complex scenarios such as obstructed visual information, insufficient light, or unclear user operations, the supplementary weight and voice information can provide key auxiliary judgment criteria, improving the accuracy of recognizing the user's true behavioral intentions.

[0138] In one embodiment, the refrigerator control method provided in this application further includes the following steps: Step S261: Based on the detection signal from the door status sensor, establish a finite state machine for the door status of the refrigerator. The state set of the finite state machine includes the closed state, the open state, the open state, and the closed state. Step S262: When the state of the finite state machine transitions from the open state to the open state, a context freeze request is sent to the cloud server to pause the reasoning process of the large language model for the current dialogue and cache the current dialogue history. Step S263: When the state of the finite state machine transitions from the closed state to the closed state, the updated food inventory data along with the visual change feature sequence is uploaded to the cloud server, so that the cloud server can use the food inventory data and the visual change feature sequence as context summaries to inject into the prompt words of the large language model and resume the reasoning process of the large language model.

[0139] In this embodiment, the detection signal of the door status sensor refers to the electrical signal output by the door status sensor, which indicates the current physical state of the refrigerator door. This detection signal can be generated in various ways. For example, it can be generated by a mechanical switch: when the refrigerator door is closed, the switch is pressed, outputting a low-level signal; when the refrigerator door is opened, the switch is released, outputting a high-level signal. Alternatively, a Hall effect sensor can be used to generate a corresponding electrical signal by detecting changes in the magnetic field of a magnet on the refrigerator door.

[0140] In this embodiment, the finite state machine for the refrigerator door state is an abstract model used to describe the state transitions of the refrigerator door. This finite state machine includes predefined discrete states, such as closed, open, open, and closed states, as well as rules for transitioning from one state to another under specific conditions. This finite state machine can be implemented through software programming; for example, a state management module can run inside the controller, which determines and updates the current door state based on the real-time signal from the door state sensor and a preset time threshold. Alternatively, it can be implemented through hardware logic circuits, using flip-flops and gate circuits to construct the state transition logic.

[0141] In this embodiment, the closed state indicates that the refrigerator door is completely closed and in a stable state; the partially open state indicates that the refrigerator door is transitioning from the closed state to the open state; the open state indicates that the refrigerator door is completely open and in a stable state; and the partially closed state indicates that the refrigerator door is transitioning from the open state to the closed state. These states can be determined based on the signal change rate and duration of the door status sensor, as well as a comparison with a preset threshold.

[0142] In this embodiment, the context freeze request can be a specific instruction sent by the refrigerator controller to the cloud server, aimed at pausing the ongoing dialogue reasoning process of the large language model and saving the current dialogue history and state. This request can be sent via standard network communication protocols (such as HTTP / HTTPS, MQTT, etc.) and may include information such as user identifier, dialogue session identifier, and freeze operation instructions.

[0143] Large language models are natural language processing models built using deep learning techniques, capable of understanding, generating, and processing human language. Specifically, they can be pre-trained models based on the Transformer architecture, such as the GPT series and BERT series, which acquire powerful language understanding and generation capabilities through training on large-scale corpora.

[0144] The reasoning process of the current dialogue can be a computational process in which a large language model performs semantic analysis, intent recognition, knowledge retrieval, and generates corresponding responses based on user input and historical dialogue context. Caching the current dialogue history refers to temporarily storing all input, output, and intermediate state data generated by the large language model in the current session so that it can be quickly retrieved when needed. This can be stored in the memory of a cloud server, a distributed caching system (such as Redis), or a persistent database.

[0145] The updated food inventory data refers to structured data generated by the controller after the user completes the food retrieval and placement operations, based on visual information acquired by the image acquisition device. This data reflects the latest information on the type, quantity, and location of food items inside the refrigerator. The visual change feature sequence refers to a series of event elements generated by the image acquisition device and controller during the user's interaction with the food, describing the direction of hand movement, the state of contact with the food, and the type of food.

[0146] A contextual summary can be a concise text or structured data entry that extracts and integrates updated food inventory data and visual change feature sequences, containing key information. This summary can be generated using pre-defined template filling, keyword extraction algorithms, or a small language model. Injecting this contextual summary into the large language model's input means using it as part of the model's input to guide the model in considering this latest information during subsequent reasoning. This can be done by appending the summary before user input, as a system command, or by passing it to the large language model via specific API parameters. Resuming the large language model's reasoning process means that after the frozen state is lifted, the large language model can continue reasoning from previously cached dialogue history and states.

[0147] In the above scheme, a finite state machine for door states is introduced to achieve precise perception and management of the refrigerator door's open / closed state. When the door state sensor detects that the refrigerator door has transitioned from an open state to an open state, the controller immediately sends a context freeze request to the cloud server, thereby pausing the large language model's reasoning process for the current dialogue and caching its dialogue history. This mechanism effectively avoids inaccurate or irrelevant responses from the large language model due to its inability to perceive changes in the physical world in real time during user food handling operations. When the user completes the food handling and the door state sensor detects that the refrigerator door has transitioned from a closed state to a closed state, the controller uploads the updated food inventory data and visual change feature sequence to the cloud server. The cloud server uses this information as a context summary, injects it into the large language model's prompts, and resumes its reasoning process. In this way, when the large language model resumes reasoning, it can make judgments and responses based on the latest and most accurate food inventory information and user interaction behavior, ensuring that the intelligent assistant's dialogue context is highly synchronized with the actual physical state of the refrigerator. This, combined with the aforementioned solution of detecting user actions of taking food out of and putting it in the refrigerator and updating food inventory data through image acquisition devices and controllers, enables the refrigerator to not only accurately sense changes in food, but also intelligently manage interactions with cloud-based large language models, thereby providing a smoother, smarter, and more realistic user experience.

[0148] In one embodiment, after step S250, the refrigerator control method of this application embodiment further includes the following steps: Step S271: Based on the updated food inventory information, determine the feedback information related to the user's pick-up and drop-off of the target food. Step S272: Output feedback information; The feedback information includes at least one of the following: Event broadcast information, which includes at least one of the following: the time of retrieval and placement of the target food, the type of food, the quantity retrieved and placed, and the target storage location; Inventory status information, which includes the remaining quantity of the target food item in the refrigerator; Intelligent recommendation information, which includes at least one of the following: recipe recommendations, purchasing suggestions, or preservation suggestions related to the target ingredient; Alarm notification information, including near-expiration notification information indicating that the target food ingredient is close to its expiration date and / or storage notification information indicating that the target food ingredient is stored in an improper location.

[0149] For example, in step S271, the refrigerator controller can directly determine the feedback information related to the user's retrieval and placement of the target food based on the updated food inventory information. Alternatively, the updated food inventory information can be uploaded to a cloud-based big data model, which will then generate corresponding feedback information based on preset generation rules, combined with the current food inventory and historical interaction records, and then send it back to the refrigerator controller for output.

[0150] For example, in step S272, the generated feedback information can be presented in a user-perceptible manner, enabling the user to obtain this information in a timely manner. For instance, it can be displayed or broadcast through output devices such as the refrigerator's built-in display screen and speakers, or it can be achieved by pushing messages or displaying interfaces through a smart terminal (such as a mobile application) connected to the refrigerator.

[0151] In this embodiment, the event broadcast information can be real-time or near real-time feedback information that informs the user of the specific details of the food retrieval and placement event. It can include the timestamp of the retrieval and placement operation, the name of the food involved, the operation type (e.g., placement or removal), quantity changes, and the storage location identifier of the food. For example, the quantity retrieved or placed can be estimated based on changes in the volume or weight of the food.

[0152] In this embodiment, the inventory status information can be feedback information reflecting the current inventory status of a specific ingredient or the overall food inventory. This information can display the remaining quantity of a specific ingredient in the refrigerator, or display the inventory status of all ingredients in the form of an overview list. To improve intuitiveness, the inventory status information can also be displayed in a graphical interface, such as a progress bar or icons, which can intuitively present the inventory level.

[0153] In this embodiment, intelligent recommendation information can provide personalized suggestions based on the user's food inventory, preferences, or external data. Specifically, recipe recommendations can suggest suitable dishes based on existing ingredients in the refrigerator, combined with the user's historical cooking habits or popular recipes. Purchase suggestions can remind the user to buy ingredients in a timely manner based on their consumption rate, near-expiration status, or user-set inventory thresholds. Preservation suggestions can provide optimal storage methods or expiration date reminders based on the type of food, storage location, or current environmental parameters.

[0154] In this embodiment, alarm notifications can be warning feedback messages issued in specific abnormal situations or situations requiring user attention. Near-expiration notifications can be issued based on the food's shelf-life data, reminding the user before the food expires. Storage notifications can be based on the food's suitable storage conditions (such as temperature, humidity, and protection from light) and its actual storage location, determining whether it has been stored improperly and providing suggestions.

[0155] In the above solution, after updating the food inventory data, feedback information related to the user's food handling behavior can be generated and output immediately. This allows the user to instantly understand the results of their operation, the latest inventory status of the food, and receive personalized usage suggestions or warnings. This instant, multi-dimensional feedback mechanism not only enhances the user's perception and trust in the smart refrigerator's functions but also transforms the refrigerator from a passive recorder into a proactive intelligent assistant, greatly improving the user experience and the refrigerator's practical value.

[0156] The following is combined Figure 4 The implementation flow of a refrigerator control method according to another embodiment of this application will be described.

[0157] For example, the refrigerator may include an RGB-D camera disposed on the top front side of the interior cavity, a weight sensor array deployed under each storage shelf, a door status sensor for detecting door opening and closing events, a microphone array for collecting user voice, a speaker or display screen for outputting responses to user interaction, and a controller communicatively connected to the above components.

[0158] like Figure 4As shown, after the refrigerator is powered on and initializes, the refrigerator controller can control the door status sensor to continuously monitor the refrigerator door status. If no door opening event is detected, monitoring continues. If a door opening event is detected, the controller marks the start of the transition segment and switches the door status, while simultaneously sending a context freeze request to the cloud. If the refrigerator is currently in a dialogue with the user, the cloud server responds to the request by pausing the inference process of the cloud-based large language model for the current dialogue and caching the current dialogue history. Furthermore, after the transition segment begins, the refrigerator controller can control the RGB-D camera to capture a high-frequency sequence of visual images of the refrigerator, facilitating subsequent analysis of hand movement direction and food contact status. The controller can also control the microphone array to capture user voice, perform speech conversion, semantic analysis, and semantic feature encoding to obtain user voice semantic features. Additionally, at least at the start of the transition segment, the controller can control the weight sensor array to capture the initial weight of food in each storage compartment of the refrigerator. When the controller detects a door closing event, it marks the end of the transition segment and switches the door status. It can also control the weight sensor array to capture the final weight of food in each storage compartment of the refrigerator at the end of the transition segment. In this way, the controller can extract the weight change features of each storage location within the change segment, at least based on the difference between the initial and final weights. Furthermore, at the end of the change segment, the controller can extract keyframes from the visual image sequence acquired during the change segment (such as the RGB image sequence captured by a camera during door opening), perform semantic segmentation on each keyframe, and obtain hand segmentation masks, food segmentation masks, and food category labels. Then, based on the semantic segmentation results and temporal features of each keyframe, the controller analyzes the user's hand movement direction and food contact state during the change segment, and generates a visual change feature sequence within the change segment. For example, the visual change feature sequence is a sequence of multiple event elements, each event element corresponding to a continuous unidirectional hand movement event. The value of each event element includes the hand movement direction and timestamp of the hand movement event, the food contact state corresponding to each visual image within that timestamp, and the food category of the contacting food.

[0159] Next, the multimodal fusion judgment stage begins. The controller inputs the generated visual change feature sequence, the extracted weight change features, and the encoded speech semantic features into the multimodal fusion-based disposal state transition probability model. The model outputs the probability distribution of the user's disposal state (such as adding food, taking out food, consuming food, storing food, or no change in food) during this door opening period, and selects the disposal state corresponding to the highest probability as the final judgment result.

[0160] After determining the user's handling status of the food during the door opening period, the controller can perform operations such as storage location positioning, data updating, and feedback processing based on the handling status. Specifically, first, it can be determined whether the handling status is that the food has not changed. If so, only the various data generated during the door opening and closing can be stored, the food inventory database remains unchanged, and the system can directly jump to perform data upload and dialogue recovery related operations. If not, it means that the food inside the refrigerator has changed during this door opening period, and the corresponding keyframes for retrieval and placement can be selected as retrieval and placement images based on the handling status. For example, if the food is added, the latest frame of the hand-food contact state captured during the hand movement direction is the retrieval and placement image (including RGB image and corresponding depth image). Based on the food segmentation mask obtained by semantic segmentation of the RGB image in the retrieval and placement image, the pixel coordinates (x0, y0) corresponding to the target food texture center are located. Then, the depth value z0 corresponding to the target food texture center pixel is extracted from the depth image to obtain the three-dimensional image coordinates (x0, y0, z0) of the target food texture center. Next, the coordinates (x0, y0, z0) are transformed to the refrigerator coordinate system to obtain the new coordinates (x1, y1, z1) of the target food's center. Then, using the height constraint of the refrigerator shelf, the height coordinates of the target food's center are corrected to the closest refrigerator shelf reference height coordinate z1. Based on the corrected height coordinate z1, the corresponding storage location identifier is matched to accurately determine the target storage location for the user to retrieve or place the target food. A binding relationship is established between the target food category and the target storage location. Then, the controller can update the inventory database by combining the food handling status, target food category, and corresponding target storage location. After the update is completed, the controller packages and uploads all kinds of data generated during the door opening and closing to the cloud and requests the resumption of large language model reasoning and dialogue interaction. In addition, the cloud can generate various types of user feedback information based on the updated data and send it to the refrigerator controller. The refrigerator controller can output feedback information through speakers and displays, in both voice and text formats.

[0161] The above solution enables automatic sensing of users' food handling behavior, precise binding of food spatial location, automatic updating of food inventory, and human-computer interaction based on natural language, significantly improving the refrigerator's intelligence level and user experience.

[0162] It should be understood that the above examples are provided to help those skilled in the art understand the embodiments of this application, and are not intended to limit the embodiments of this application to the specific values ​​or scenarios illustrated. Those skilled in the art can obviously make various equivalent modifications or changes based on the above examples, and such modifications or changes also fall within the scope of the embodiments of this application.

[0163] The above text combined Figures 1 to 4 The refrigerator control method of the present application embodiments is described in detail below, in conjunction with Figure 5 This document describes in detail the device embodiments of this application. It should be understood that the control device in the embodiments of this application can execute various control methods described in the foregoing embodiments of this application. That is, the specific working processes of the various products described below can be referred to the corresponding processes in the foregoing method embodiments.

[0164] This application also provides a control device for executing the various steps of the refrigerator control method provided in any embodiment of this application.

[0165] Figure 5 Schematic diagrams of the control device in some embodiments of this application are shown. For example... Figure 5 As shown, the refrigerator control device 500 of this application embodiment includes: The acquisition module 510 is used to acquire images of the user picking up and placing food, wherein the images include the user's hand and the imaging area of ​​the target food being picked up and placed. The initial height determination module 520 is used to determine the initial height coordinates of the target food in the refrigerator based on the depth information of the pick-up and put-down image, wherein the height coordinates are coordinates along the height direction of the refrigerator. The height calibration module 530 is used to determine the calibration height coordinates of the target access position based on the initial height coordinates and the set of reference height coordinates of each storage position in the refrigerator. The set of reference height coordinates includes multiple reference height coordinates corresponding to each storage position. The storage location module 540 is used to determine the storage location information corresponding to the target storage location based on the calibration height coordinates. The data update module 550 is used to update the refrigerator's food inventory data based on the category and storage location information of the target food.

[0166] Each unit module of the refrigerator control device 500 can execute the corresponding steps in the above method embodiment, so the details of each unit module will not be elaborated here. Please refer to the description of the corresponding steps above for details.

[0167] It should be noted that the refrigerator control device 500 described above is embodied in the form of a functional unit. The term "unit" here can be implemented in software and / or hardware, without specific limitations.

[0168] For example, a “unit” can be a software program, a hardware circuit, or a combination of both that implements the above functions. The hardware circuit may include an application-specific integrated circuit (ASIC), electronic circuitry, a processor (e.g., a shared processor, a proprietary processor, or a group processor) and memory for executing one or more software or firmware programs, combined logic circuitry, and / or other suitable components that support the described functions.

[0169] Therefore, the units of the various examples described in the embodiments of this application can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the embodiments of this application.

[0170] This application also provides a refrigerator. For example... Figure 1 As shown, the refrigerator 100 provided in this application embodiment includes: The image acquisition device 110 is located on the top front side inside the refrigerator 100 and is configured to take images of the user taking out and putting in food from above. The images include the user's hands and the imaging area of ​​the target food being taken out and put in. The controller 120 is communicatively connected to the image acquisition device 110 and is configured as follows: Get the pick-and-place image; Based on the depth information of the pick-up and put-down images, the initial height coordinates of the target food in the refrigerator 100 are determined, where the height coordinates are the coordinates along the height direction of the refrigerator 100. Based on the initial height coordinates and the set of reference height coordinates for each storage location within the refrigerator 100, the calibration height coordinates for the target access location are determined. The set of reference height coordinates includes multiple reference height coordinates corresponding to each storage location. Based on the calibration height coordinates, determine the storage location information corresponding to the target access location; Based on the category and storage location information of the target ingredients, update the food inventory data of refrigerator 100.

[0171] In one implementation, the controller determines the calibrated height coordinates of the target access location based on the initial height coordinates and the set of reference height coordinates for each storage location within the refrigerator, and is configured as follows: Compare the initial altitude coordinates with multiple reference altitude coordinates; Based on the comparison results, a target reference height coordinate is selected from the set of reference height coordinates as the calibration height coordinate of the target access position, wherein the difference between the target reference height coordinate and the initial height coordinate is not greater than the difference between other reference height coordinates and the initial height coordinate; The controller determines the storage location information corresponding to the target access position based on the calibration height coordinates, and is configured as follows: The storage location identifier of the target storage location that has a mapping relationship with the calibration height coordinate is found from the pre-stored storage location mapping table. This identifier serves as the storage location information corresponding to the target access location. The storage location mapping table represents the mapping relationship between multiple reference height coordinates and each storage location.

[0172] In one embodiment, the image acquisition device includes an RGB-D camera, and the pick-up and put-down image includes a color image acquired by the RGB-D camera and a corresponding depth image. The controller, based on the depth information of the pick-up and put-down image, determines the initial height coordinates of the target food item's location in the refrigerator, and is configured as follows: Based on the food segmentation mask of the target food, determine the two-dimensional image coordinates of the centroid pixel of the target food. The food segmentation mask is obtained by semantic segmentation of the color image, and the centroid pixel is the center pixel of the pixel region corresponding to the food segmentation mask. Based on the two-dimensional image coordinates of the centroid pixel, read the depth pixel value of the corresponding pixel of the centroid pixel from the depth image; Based on the two-dimensional image coordinates of the centroid pixel and the depth pixel value, determine the three-dimensional image coordinates of the target access location in the image coordinate system. Based on the camera intrinsic parameter matrix of the RGB-D camera and the pre-calibrated pose matrix of the RGB-D camera relative to the refrigerator, the three-dimensional image coordinates are converted into three-dimensional spatial coordinates in the refrigerator coordinate system. The three coordinate axes of the three-dimensional spatial coordinates correspond to the length direction, width direction and height direction of the refrigerator, respectively, and the height direction is parallel to the line of sight of the RGB-D camera. Extract the height coordinate components from the three-dimensional spatial coordinates to serve as the initial height coordinates.

[0173] In one implementation, the controller is further configured to: Based on the coordinate components of the length and width directions in three-dimensional space, the target access location is determined to be the target partition location on the same storage surface of the refrigerator, where the storage surface is perpendicular to the height direction; The controller determines the storage location information corresponding to the target access position based on the calibration height coordinates, and is configured as follows: The storage location identifier is found in the pre-stored storage location mapping table. The storage location identifier is used as the storage location information corresponding to the target access location. The storage location mapping table represents the mapping relationship between each storage location and multiple reference height coordinates and multiple partition locations.

[0174] In one embodiment, the refrigerator further includes: The door status sensor is configured to detect door opening and door closing events of the refrigerator door; The controller also communicates with the door status sensor and is configured to: Based on the door opening and closing events detected by the door status sensor, the time axis is divided into a steady-state segment and a variable segment, where the variable segment is the time period from the detection of the door opening event to the detection of the door closing event; During the changing segment, the image acquisition device is controlled to acquire images of the refrigerator at a preset frequency, resulting in multiple visual images; For each visual image, input the visual image into the semantic segmentation model to obtain the hand segmentation mask of the user's hand, the food segmentation mask of each food item in the visual image, and the food category corresponding to the food segmentation mask; Based on the relative positional relationship between the hand segmentation mask and each food segmentation mask, the food contact state corresponding to the visual image is determined. The food contact state includes an uncontacted state indicating that the user's hand is not in contact with any food, and a contact state indicating that the user's hand is in contact with any food. Based on the difference in hand segmentation masks between adjacent visual images, the hand movement direction corresponding to the adjacent visual images is determined. The hand movement direction includes the direction of reaching into the refrigerator and the direction of withdrawing from the refrigerator. Based on the hand movement state corresponding to each adjacent visual image in multiple visual images, the food contact state corresponding to each visual image, and the food category of the contacted food, a visual change feature sequence of user hand-food interaction within a change segment is generated. The visual change feature sequence is a sequence of multiple event elements. Each event element corresponds to a continuous unidirectional hand movement event. The value of the event element includes the hand movement direction and timestamp of the hand movement event, the food contact state corresponding to each visual image within the timestamp, and the food category of the contacted food. Based on the visual change feature sequence, determine the user's handling status of food items within the change segment. The handling status includes the following: adding food item status, taking out food item status, storing and retrieving food item status (adding one food item and taking out another food item), consuming food item status, and food item status with no change. The controller acquires the pick-up and drop-down images and is configured as follows: When the processing state is adding ingredients, taking out ingredients, storing ingredients, or consuming ingredients, one or more visual images are selected from multiple visual images as the picking and placing images according to the processing state. The controller updates the refrigerator's food inventory data based on the target food's category and storage location information, and is configured as follows: The refrigerator's food inventory data is updated based on the target food's category information, storage location information, and the user's handling status of the target food.

[0175] In one implementation, the controller determines the user's handling status of the food within a change segment based on a sequence of visual change features, and is configured as follows: Perform at least one of the following decision operations: In the visual change feature sequence, the number of event elements is 2, and the value of the first event element corresponds to the insertion direction and contact state, while the value of the second event element corresponds to the withdrawal direction and non-contact state. In this case, the handling state is determined to be the addition of ingredients. In the visual change feature sequence, the number of event elements is 2, and the value of the first event element corresponds to the insertion direction and the non-contact state, while the value of the second event element corresponds to the withdrawal direction and the contact state. In this case, the handling state is determined to be the state of taking out the food. In the visual change feature sequence, the number of event elements is 4, and the value of the first event element corresponds to the direction of insertion and the non-contact state, the value of the second event element corresponds to the direction of withdrawal and the contact state, the value of the third event element corresponds to the direction of insertion and the contact state, the value of the fourth event element corresponds to the direction of withdrawal and the non-contact state, and if the food categories of the food in contact with the second and third event elements are the same, the handling state is determined to be the food consumption state. In the visual change feature sequence, the number of event elements is 2, and the value of the first event element corresponds to the insertion direction and contact state, the value of the second event element corresponds to the withdrawal direction and contact state, and the food categories of the contacted food corresponding to the first event element and the second event element are different, the handling state is determined to be the food storage and retrieval state. If the values ​​of all event elements in the visual change feature sequence correspond to the untouched state, the handling state is determined to be the state where the food has not changed. and / or The controller selects one or more visual images from multiple images as pick-up and drop-off images based on the processing status, and is configured as follows: When the processing state is adding ingredients, taking out ingredients, storing ingredients, or consuming ingredients, the visual image with the same timestamp as the latest timestamp of the first target event element and / or the visual image with the same timestamp as the earliest timestamp of the second target event element is used as the pick-up and put-down image. The first target event element and the second target event element are located in the visual change feature sequence. The value of the first target event element corresponds to the insertion direction and contact state, and the value of the second target event corresponds to the withdrawal direction and contact state.

[0176] In one embodiment, the refrigerator further includes a weight sensor array and / or a microphone array, wherein the weight sensor array is disposed below the load-bearing surface of each storage compartment and is configured to detect the total weight of the food items carried in the corresponding storage compartment; the microphone array is configured to collect the user's voice commands. The controller also communicates with a weight sensor array and / or a microphone array and is configured to: Based on the difference in the total weight of each storage location detected by the weight sensor array at the beginning and end of the change segment, the weight change characteristics of each storage location are determined. and / or Speech recognition and semantic encoding are performed on voice commands within the changing segments to obtain the user's speech semantic features; The controller determines the user's handling status of the food within the change segment based on the visual change feature sequence, and is configured as follows: Input the visual change feature sequence, weight change feature and / or speech semantic feature into the disposal state transition probability model based on multimodal fusion, and output the probability distribution of the user's disposal state of food within the change segment, which belongs to each of the following states: adding food, taking out food, consuming food and no change in food. The disposal status corresponding to the highest probability value in the probability distribution is determined as the user's disposal status for the ingredients.

[0177] In one implementation, the controller is further configured to: Based on the detection signal from the door status sensor, a finite state machine for the refrigerator door status is established. The state set of the finite state machine includes the closed state, the open state, the open state, and the closed state. When the state of the finite state machine transitions from the open state to the open state, a context freeze request is sent to the cloud server to pause the reasoning process of the large language model for the current dialogue and cache the current dialogue history. When the finite state machine transitions from the closed state to the closed state, the updated food inventory data along with the visual change feature sequence is uploaded to the cloud server. This allows the cloud server to use the food inventory data and visual change feature sequence as a context summary, inject it into the prompt words of the large language model, and resume the reasoning process of the large language model.

[0178] In one implementation, the controller is further configured to: Based on the updated food inventory information, determine the feedback information related to the user's pick-up and drop-off of the target food; Output feedback information; The feedback information includes at least one of the following: Event broadcast information, which includes at least one of the following: the time of retrieval and placement of the target food, the type of food, the quantity retrieved and placed, and the target storage location; Inventory status information, which includes the remaining quantity of the target food item in the refrigerator; Intelligent recommendation information, which includes at least one of the following: recipe recommendations, purchasing suggestions, or preservation suggestions related to the target ingredient; Alarm notification information, including near-expiration notification information indicating that the target food ingredient is close to its expiration date and / or storage notification information indicating that the target food ingredient is stored in an improper location.

[0179] Figure 1 This is merely an example of a refrigerator and does not constitute a limitation on refrigerators. A refrigerator may include more or fewer parts than shown, or combine certain parts, or use different parts.

[0180] In this embodiment, the refrigerator controller may include at least one processor, the refrigerator may also include a memory and a computer program stored in the memory and capable of running on at least one processor, and the processor executes the computer program to implement the steps of the above control method.

[0181] This application also provides a computer program product that, when executed by a processor, implements the refrigerator control method of any of the method embodiments of this application.

[0182] The computer program product can be stored in memory, for example, it is a program. The program is eventually converted into an executable object file that can be executed by the processor after processes such as preprocessing, compilation, assembly and linking.

[0183] This application also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a computer, implements the refrigerator control method of any of the method embodiments of this application. The computer program may be a high-level language program or an executable object program.

[0184] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A refrigerator, characterized in that, include: An image acquisition device is installed at the top front side inside the refrigerator and is configured to take images of the user taking and placing food, wherein the images include the user's hand and the imaging area of ​​the target food being taken and placed. The controller is communicatively connected to the image acquisition device and is configured to: Acquire the pick-and-place image; Based on the depth information of the retrieval and placement image, the initial height coordinates of the target food item's target storage and retrieval location in the refrigerator are determined, wherein the height coordinates are coordinates along the height direction of the refrigerator; Based on the initial height coordinates and the set of reference height coordinates for each storage location inside the refrigerator, the calibration height coordinates of the target access location are determined. The set of reference height coordinates includes multiple reference height coordinates corresponding to each storage location. Based on the calibration height coordinates, determine the storage location information corresponding to the target access location; Based on the category information of the target food ingredient and the storage location information, update the food inventory data of the refrigerator.

2. The refrigerator according to claim 1, characterized in that, The controller, based on the initial height coordinates and the set of reference height coordinates for each storage location within the refrigerator, determines the calibration height coordinates of the target access position, and is configured as follows: The initial height coordinates are compared with the plurality of reference height coordinates; Based on the comparison results, a target reference height coordinate is selected from the set of reference height coordinates as the calibration height coordinate of the target access position, wherein the difference between the target reference height coordinate and the initial height coordinate is not greater than the difference between other reference height coordinates and the initial height coordinate; The controller, based on the calibration height coordinates, determines the storage location information corresponding to the target access location, and is configured as follows: The storage location identifier of the target storage location that has a mapping relationship with the calibration height coordinate is found from the pre-stored storage location mapping table, and is used as the storage location information corresponding to the target access location. The storage location mapping table represents the mapping relationship between the multiple reference height coordinates and each storage location.

3. The refrigerator according to claim 1, characterized in that, The image acquisition device includes an RGB-D camera, and the retrieval image includes a color image acquired by the RGB-D camera and a corresponding depth image. The controller, based on the depth information of the retrieval image, determines the initial height coordinates of the target food item's location in the refrigerator, and is configured as follows: Based on the food segmentation mask of the target food, the two-dimensional image coordinates of the centroid pixel of the target food are determined, wherein the food segmentation mask is obtained by semantic segmentation of the color image, and the centroid pixel is the center pixel of the pixel region corresponding to the food segmentation mask. Based on the two-dimensional image coordinates of the centroid pixel, the depth pixel value of the corresponding pixel of the centroid pixel is read from the depth image; The three-dimensional image coordinates of the target access location in the image coordinate system are determined based on the two-dimensional image coordinates of the centroid pixel and the depth pixel value. Based on the camera intrinsic parameter matrix of the RGB-D camera and the pre-calibrated pose matrix of the RGB-D camera relative to the refrigerator, the three-dimensional image coordinates are converted into three-dimensional spatial coordinates in the refrigerator coordinate system. The three coordinate axes of the three-dimensional spatial coordinates correspond to the length direction, width direction and height direction of the refrigerator, respectively, and the height direction is parallel to the line of sight of the RGB-D camera. The coordinate components in the height direction are extracted from the three-dimensional spatial coordinates and used as the initial height coordinates.

4. The refrigerator according to claim 3, characterized in that, The controller is also configured to: Based on the coordinate components of the length direction and the width direction in the three-dimensional spatial coordinates, the target access position is determined to be the target partition position on the same storage surface of the refrigerator, wherein the storage surface is perpendicular to the height direction; The controller, based on the calibration height coordinates, determines the storage location information corresponding to the target access location, and is configured as follows: The storage location identifier of the target storage location that has a mapping relationship with both the calibration height coordinate and the target partition location is found from the pre-stored storage location mapping table. This identifier is used as the storage location information corresponding to the target access location. The storage location mapping table represents the mapping relationship between each storage location and the multiple reference height coordinates and multiple partition locations.

5. The refrigerator according to any one of claims 1-4, characterized in that, The refrigerator also includes: The door status sensor is configured to detect door opening and door closing events of the refrigerator door; The controller is also communicatively connected to the door status sensor and is configured to: Based on the door opening and closing events detected by the door status sensor, the time axis is divided into a steady-state segment and a variable segment, wherein the variable segment is the time period from the detection of the door opening event to the detection of the door closing event; During the change segment, the image acquisition device is controlled to acquire images of the refrigerator at a preset frequency to obtain multiple visual images; For each visual image, input the visual image into the semantic segmentation model to obtain the hand segmentation mask of the user's hand, the food segmentation mask of each food item in the visual image, and the food category corresponding to the food segmentation mask; Based on the relative positional relationship between the hand segmentation mask and each food segmentation mask, the food contact state corresponding to the visual image is determined, wherein the food contact state includes an uncontact state indicating that the user's hand is not in contact with any food, and a contact state indicating that the user's hand is in contact with food. Based on the difference in the hand segmentation mask of adjacent visual images in the plurality of visual images, the hand movement direction corresponding to the adjacent visual images is determined, wherein the hand movement direction includes an insertion direction indicating that the hand is extending into the refrigerator and an withdrawal direction indicating that the hand is withdrawing from the refrigerator. Based on the hand movement state corresponding to each adjacent visual image in the plurality of visual images, the food contact state corresponding to each visual image, and the food category of the contacted food, a visual change feature sequence of user hand-food interaction within the change segment is generated. The visual change feature sequence is a sequence of multiple event elements, each event element corresponding to a continuous unidirectional hand movement event. The value of the event element includes the hand movement direction and timestamp of the hand movement event, the food contact state corresponding to each visual image within the timestamp, and the food category of the contacted food. Based on the visual change feature sequence, the user's handling status of the food within the change segment is determined, wherein the handling status includes the following: adding food status, taking out food status, storing and taking out food status indicating adding one food and taking out another food status, consuming food status, and food status with no change. The controller acquires the pick-up and drop-down image and is configured as follows: When the processing state is the state of adding ingredients, taking out ingredients, storing ingredients, or consuming ingredients, one or more of the multiple visual images are selected as the picking and placing images according to the processing state. The controller updates the refrigerator's food inventory data based on the category information and storage location information of the target food ingredient, and is configured as follows: The refrigerator's food inventory data is updated based on the category information of the target food ingredient, the storage location information, and the user's handling status of the target food ingredient.

6. The refrigerator according to claim 5, characterized in that, The controller, based on the visual change feature sequence, determines the user's handling status of the food within the change segment, and is configured as follows: Perform at least one of the following decision operations: In the visual change feature sequence, the number of event elements is 2, and the value of the first event element corresponds to the insertion direction and contact state, and the value of the second event element corresponds to the withdrawal direction and non-contact state, the handling state is determined to be the food addition state. In the visual change feature sequence, the number of event elements is 2, and the value of the first event element corresponds to the insertion direction and the non-contact state, while the value of the second event element corresponds to the withdrawal direction and the contact state. In this case, the handling state is determined to be the food removal state. If the number of event elements in the visual change feature sequence is 4, and the value of the first event element corresponds to the insertion direction and non-contact state, the value of the second event element corresponds to the withdrawal direction and contact state, the value of the third event element corresponds to the insertion direction and contact state, and the value of the fourth event element corresponds to the withdrawal direction and non-contact state, and the food categories of the food in contact with the second and third event elements are the same, then the disposal state is determined to be the food consumption state. In the visual change feature sequence, the number of event elements is 2, and the value of the first event element corresponds to the insertion direction and contact state, the value of the second event element corresponds to the withdrawal direction and contact state, and the food categories of the contacted food corresponding to the first event element and the second event element are different, the handling state is determined to be the food storage and retrieval state. If the values ​​of all event elements in the visual change feature sequence correspond to the untouched state, the handling state is determined to be the state where the food has not changed. and / or The controller, based on the processing state, selects one or more from the plurality of visual images as the pick-up and put-down images, and is configured as follows: When the processing state is the state of adding ingredients, taking out ingredients, storing ingredients, or consuming ingredients, the visual image with the same timestamp as the latest timestamp of the first target event element and / or the visual image with the same timestamp as the earliest timestamp of the second target event element is used as the pick-up and put-down image, wherein the first target event element and the second target event element are located in the visual change feature sequence, the value of the first target event element corresponds to the insertion direction and contact state, and the value of the second target event element corresponds to the withdrawal direction and contact state.

7. The refrigerator according to claim 5, characterized in that, The refrigerator also includes a weight sensor array and / or a microphone array, wherein the weight sensor array is disposed below the load-bearing surface of each storage compartment and is configured to detect the total weight of the food items carried in the corresponding storage compartment; the microphone array is configured to collect the user's voice commands. The controller is also communicatively connected to the weight sensor array and / or the microphone array, and is configured to: Based on the difference in the total weight of each storage location detected by the weight sensing array at the beginning and end of the change segment, the weight change characteristics of each storage location are determined. and / or Speech recognition and semantic encoding are performed on the speech commands within the changed segment to obtain the user's speech semantic features; The controller, based on the visual change feature sequence, determines the user's handling status of the food within the change segment, and is configured as follows: The visual change feature sequence, the weight change feature and / or the speech semantic feature are input into the disposal state transition probability model based on multimodal fusion, and the probability distribution of the user's disposal state of food within the change segment is each of the following: adding food, taking out food, consuming food, and no change in food. The disposal state corresponding to the maximum probability value in the probability distribution is determined as the user's disposal state of the food ingredients.

8. The refrigerator according to claim 5, characterized in that, The controller is also configured to: Based on the detection signal from the door status sensor, a finite state machine for the door status of the refrigerator is established, wherein the state set of the finite state machine includes closed state, open state, open state, and closed state; When the state of the finite state machine transitions from the open state to the open state, a context freeze request is sent to the cloud server to pause the reasoning process of the large language model for the current dialogue and cache the current dialogue history. When the state of the finite state machine transitions from the closed state to the closed state, the updated food inventory data, along with the visual change feature sequence, is uploaded to the cloud server. This allows the cloud server to use the food inventory data and the visual change feature sequence as a context summary, inject them into the prompt words of the large language model, and resume the reasoning process of the large language model.

9. The refrigerator according to any one of claims 1-4, characterized in that, The controller is also configured to: Based on the updated food inventory information, determine the feedback information related to the user's retrieval and placement of the target food. Output the feedback information; The feedback information includes at least one of the following: Event broadcast information, wherein the event broadcast information includes at least one of the following: the time of taking or placing the target food ingredient, the type of food ingredient, the quantity taken or placed, and the target storage location; Inventory status information, wherein the inventory status information includes the remaining quantity of the target food ingredient in the refrigerator; Intelligent recommendation information, wherein the intelligent recommendation information includes at least one of recipe recommendation information, purchase suggestion information, or preservation suggestion information related to the target ingredient; Alarm notification information, wherein the alarm notification information includes an expiration date notification indicating that the target ingredient is nearing its expiration date and / or a storage notification indicating that the target ingredient is improperly stored.

10. A refrigerator control method, characterized in that, The refrigerator includes an image acquisition device disposed on the top front side inside the refrigerator, configured to capture images of a user taking or placing food items from above. The images include the user's hand and the imaging area of ​​the target food item being taken or placed. The method includes: Acquire the pick-and-place image; Based on the depth information of the retrieval and placement image, the initial height coordinates of the target food item's target storage and retrieval location in the refrigerator are determined, wherein the height coordinates are coordinates along the height direction of the refrigerator; Based on the initial height coordinates and the set of reference height coordinates for each storage location inside the refrigerator, the calibration height coordinates of the target access location are determined. The set of reference height coordinates includes multiple reference height coordinates corresponding to each storage location. Based on the calibration height coordinates, determine the storage location information corresponding to the target access location; Based on the category information of the target food ingredient and the storage location information, update the food inventory data of the refrigerator.