Method and apparatus for detecting food freshness, controller, and food storage device
By integrating a camera and gas sensor in a refrigerator and using a cross-attention network to fuse image and gas data, the system addresses the limitations of conventional methods, providing a more accurate and reliable assessment of food freshness.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- BSH HAUSGERATE GMBH
- Filing Date
- 2025-10-23
- Publication Date
- 2026-05-15
AI Technical Summary
Conventional methods for detecting food freshness in refrigerators rely heavily on subjective sensory evaluation, which is prone to human error, and existing camera and gas sensor systems suffer from low accuracy when food is packaged, leading to reduced detection reliability and increased food waste.
A system that integrates a camera and gas sensor in a refrigerator to obtain images and gas data, using a cross-attention network to fuse this information and determine food freshness, thereby improving accuracy by combining visual and chemical indicators.
The system provides a more comprehensive and accurate assessment of food freshness by reducing limitations and errors from single data sources, enhancing the reliability of freshness determination.
Smart Images

Figure EP2025080679_15052026_PF_FP_ABST
Abstract
Description
[0001] METHOD AND APPARATUS FOR DETECTING FOOD FRESHNESS, CONTROLLER, AND FOOD STORAGE DEVICE
[0002] TECHNICAL FIELD
[0003] The present disclosure relates to the field of computers, and more specifically, to a method and an apparatus for detecting food freshness, a controller, and a food storage device.
[0004] BACKGROUND
[0005] Food freshness refers to a status in which food maintains optimal quality and nutritional value of the food. Fresh food usually has a good appearance, taste, smell, and nutritional components. As time goes by, the food gradually becomes less fresh due to various factors (for example, microbial activity, enzymatic action, and oxidation), and eventually becomes spoiled or rotten.
[0006] When a food in a refrigerator becomes spoiled, an appearance of the food may change, such as discoloration, a slimy surface, mold growth, or a loss of firmness. In addition, there is an increase in gases (such as a spoilage gas) emitted by the food, which may cause a smell in the refrigerator to become pungent. These changes not only affect edibility of the food, but also may affect quality of other foods in the refrigerator.
[0007] SUMMARY
[0008] According to a first aspect of embodiments of the present disclosure, a method for detecting food freshness is provided. The method includes obtaining an image of a food placed in a storage space that is acquired by a camera placed in the storage space. The method further includes obtaining gas data in the storage space that is acquired by a gas sensor placed in the storage space. The method further includes generating fused data based on the image and the gas data of the storage space by using a cross-attention network. In addition, the method further includes determining freshness of the food based on the fused data.
[0009] According to a second aspect of the embodiments of the present disclosure, an apparatus is provided. The apparatus includes a food image obtaining module, configured to obtain an image of a food placed in a storage space that is acquired by a camera placed in the storage space. The apparatus further includes a gas data obtaining module, configured to obtain gas data in the storage space that is acquired by a gas sensor placed in the storage space. The apparatus further includes a fused data generation module, configured to generate fused data based on the image and the gas data of the storage space by using a cross-attention network. In addition, the apparatus further includes a freshness determining module, configured to determine freshness of the food based on the fused data.
[0010] According to a third aspect of the embodiments of the present disclosure, a controller is provided. The controller includes one or more processors; and a storage apparatus, configured to store one or more programs, where the one or more programs, when executed by the one or more processors, cause the one or more processors to implement a method for detecting food freshness. The method includes obtaining an image of a food placed in a storage space that is acquired by a camera placed in the storage space. The method further includes obtaining gas data in the storage space that is acquired by a gas sensor placed in the storage space. The method further includes generating fused data based on the image and the gas data of the storage space by using a cross-attention network. In addition, the method further includes determining freshness of the food based on the fused data.
[0011] According to a fourth aspect of the embodiments of the present disclosure, a food storage device is provided. The food storage device includes a storage space. The food storage device further includes a camera, disposed at a top of the storage space, where the camera is configured to photograph a food in the storage space. The storage space further includes a gas sensor, disposed inside the storage space, where the gas sensor is configured to acquire gas data in the storage space. In addition, the food storage device further includes the controller provided in the third aspect of the present disclosure.
[0012] According to a fifth aspect of the embodiments of the present disclosure, a computer program product is provided. The computer program product is tangibly stored in a non- transitory computer-readable medium and includes machine-executable instructions. The machine-executable instructions, when executed, cause a machine to implement the method for detecting food freshness provided in the first aspect of the present disclosure.
[0013] According to a sixth aspect of the embodiments of the present disclosure, a method for detecting food freshness is provided. The method includes obtaining an image of a food placed in a storage space that is acquired by a camera placed in the storage space. The method further includes obtaining gas data in the storage space that is acquired by a gas sensor placed in the storage space. The method further includes generating a first freshness result based on the image of the food. The method further includes generating a second freshness result based on the gas data, where the first freshness result and the second freshness result include a plurality of freshness levels and a plurality of confidences corresponding to the plurality of freshness levels. The method further includes generating a fused freshness result based on the first freshness result and the second freshness result by using a random forest model. In addition, the method further includes determining freshness of the food based on the fused freshness result.
[0014] It should be understood that the content described in the summary part is not intended to limit key or important features of the embodiments of the present disclosure, and is not intended to limit the scope of the present disclosure. Other features of the present disclosure will be easily understood by using the following descriptions.
[0015] BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The foregoing and other features, advantages, and aspects of embodiments of the present disclosure will become more obvious with reference to the accompanying drawings and the following detailed descriptions. In the accompanying drawings, the same or similar reference numerals represent the same or similar elements, where
[0017] FIG. 1 is a schematic diagram of an exemplary environment in which a plurality of embodiments of the present disclosure can be implemented;
[0018] FIG. 2 is a flowchart of a method for detecting food freshness according to some embodiments of the present disclosure;
[0019] FIG. 3 is a flowchart of an exemplary process for detecting food freshness according to some embodiments of the present disclosure;
[0020] FIG. 4 is a schematic diagram of an exemplary system architecture for detecting food freshness by fusing image data and gas data according to some embodiments of the present disclosure;
[0021] FIG. 5 is a schematic diagram of a structure of an exemplary cross-attention network according to some embodiments of the present disclosure;
[0022] FIG. 6 is a schematic diagram of a structure of another exemplary cross-attention network according to some embodiments of the present disclosure;
[0023] FIG. 7 is a schematic diagram of an exemplary system architecture for detecting food freshness by fusing a plurality of types of time series data according to some embodiments of the present disclosure;
[0024] FIG. 8 is a flowchart of a method for detecting food freshness based on a random forest model according to some embodiments of the present disclosure;
[0025] FIG. 9 is a schematic diagram of an exemplary system architecture for fusing an image freshness result and a gas freshness result by using a random forest model according to some embodiments of the present disclosure;
[0026] FIG. 10 is a block diagram of an apparatus for detecting food freshness according to some embodiments of the present disclosure; and
[0027] FIG. 11 is a block diagram of a controller in which a plurality of embodiments of the present disclosure can be implemented.
[0028] DETAILED DESCRIPTION
[0029] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure may be implemented in various forms and should not be construed as being limited to the embodiments described herein, but these embodiments are provided for a more thorough and complete understanding of the present disclosure. It should be understood that the accompanying drawings and the embodiments of the present disclosure are merely used as examples, but are not intended to limit the protection scope of the present disclosure. The following embodiments described with reference to the accompanying drawings are merely used as examples.
[0030] In the context of the current food industry, maintaining freshness of perishable foods, especially meat and fruit, during storage has become a major challenge. This is because these foods are extremely vulnerable to microbes and environmental factors. If not consumed timely, the perishable foods stored in a refrigerator are at a risk of spoilage, resulting in reduced quality and a potential food safety hazard. In addition, a user often forgets or ignores storage time of the foods, and lacks an effective management method, which causes food waste, and increases family expenditure and environmental load.
[0031] Monitoring the freshness of the perishable foods is a major concern for both suppliers and users in the current food industry. The user expects to obtain fresh and safe food. However, a conventional food freshness evaluation method excessively relies on subjective sensory evaluation such as sight, smell, and touch. This easily causes a human error, and may affect quality and safety of the food.
[0032] In some related technologies, the refrigerator may use a camera for food recognition and management. However, a recognition rate of the camera is low, and the camera tends to be obstructed. In some related technologies, a gas sensor placed in the refrigerator is configured to detect freshness of a food in the refrigerator. When gas data acquired by the gas sensor exceeds a safety threshold, it represents that the food in the refrigerator has spoiled. In this method, the freshness of the food can be theoretically determined by monitoring gas composition released by the food. However, when the spoiled food is placed in a packaging bag or a packaging box, the gas data acquired by the gas sensor may not reach a set threshold, resulting in reduced accuracy of a detection result.
[0033] For this purpose, in this embodiment of the present disclosure, a solution for detecting food freshness is provided. In this solution, a food storage device (for example, a refrigerator) may include a controller and at least one storage space. A camera and a gas sensor are disposed in the storage space. The controller may obtain an image of a food in the storage space that is acquired by the camera, and gas data in the storage space that is acquired by the gas sensor. Then, the controller may fuse image data and the gas data by using a cross-attention network, to generate fused data that indicates both information in the image and information in the gas data. Then, the controller may determine freshness of the food based on the generated fused data.
[0034] In this manner, the image data can provide intuitive information about an appearance of the food, and the gas data can reflect a chemical change of the food. By fusing the two types of data, more comprehensive and accurate food freshness evaluation can be implemented. In addition, determining the freshness of the food based on the fused data can reduce limitations and errors caused by a single data source, thereby improving accuracy of the determined freshness.
[0035] In various embodiments of this application, the gas sensor may include a gas sensor configured to detect existence, a concentration, or a composition change of a specific gas. For example, a sensor sensitive to volatile organic compounds (VOC) and other gas compositions of different types is provided. Such a gas sensor may be, for example, a semiconductor gas sensor, a metal oxide gas sensor, an electrochemical gas sensor, an infrared gas sensor, or a photoionization detector. In addition, the gas sensor may further include a smell sensor configured to recognize and distinguish between complex smells (for example, configured to recognize the VOC). Such a gas sensor may be, for example, an electronic nose, a mass spectrometry smell sensor, or a gas chromatography-mass spectrometry sensor. The "gas data" may be a gas concentration, gas composition, and / or smell data.
[0036] FIG. 1 is a schematic diagram of an exemplary environment 100 in which a plurality of embodiments of the present disclosure can be implemented. As shown in FIG. 1, the environment 100 includes a controller 102 and a storage space 104, and the controller 102 and the storage space 104 may be included in a same food storage device. For example, such a food storage device may be a refrigerator, a freezer, a fresh-keeping cabinet, a cold storage facility, a display cabinet, a cold chain transportation equipment, or a storage room. The controller 102 may be any device having a computing capability or a processing capability. It should be understood that the controller in the food storage device is usually an embedded device, which has only a limited processing capability, and due to limitations of space and costs, storage capacity of the embedded device is usually small. The food storage device may include a plurality of storage spaces 104. The storage space 104 may be, for example, a shelf or a drawer in a fresh food compartment of the refrigerator, a shelf or a drawer in a freezer compartment of the refrigerator, the freezer, the storage room, or the like. In some implementations, the storage space 104 may be the food storage device.
[0037] In the environment 100, a camera 106 is disposed at a top of the storage space 104, and is configured to photograph a food in the storage space 104. For example, the camera 106 may acquire an image 112 inside the storage space 104, and the image 112 may include a food 110- 1, a food 110-2, and a food 110-3. In addition, a gas sensor 108 is further disposed inside the storage space 104, and is configured to acquire gas data 114 in the storage space 104. A sensor array configured to acquire a gas sample may be built in the gas sensor 108. A plurality of sensors in the sensor array may be sensitive to a plurality of types of volatile organic compounds. After the gas sample enters these sensors, each sensor may respond to a specific component in the gas, and a response degree is related to a concentration and a type of the gas component. The gas data 114 may include a plurality of components and a corresponding plurality of concentrations. For example, the gas data 114 may include a concentration of ammonia, a concentration of carbon dioxide, a concentration of hydrogen sulfide, a concentration of alcohols (for example, ethanol), a concentration of ketones (for example, acetone), a concentration of aldehydes (for example, acetaldehyde), and the like. In the environment 100, when freshness of the food in the storage space 104 is detected, the controller 102 may obtain the image 112 of the storage space 104 that is acquired by the camera 106. In some embodiments, the controller 102 may send an instruction to the camera 106, to obtain the image of the food in the storage space 104 in real time. In some embodiments, the camera 106 may periodically acquire the image of the food in the storage space 104, and store the acquired image in a corresponding memory for the controller 102 to read. For example, the camera 106 may acquire images of the storage space 104 at a predetermined time interval (for example, 30 minutes, 1 hour, or 2 hours), and store the images in the memory. The controller 102 may detect the freshness of the food based on the images. For example, the controller 102 may read a latest image from the memory for freshness detection.
[0038] When detecting the freshness of the food, the controller 102 may further obtain the gas data 114 in the storage space 104 that is acquired by the gas sensor 108. In some embodiments, the controller 102 may send an instruction to the camera 106, to obtain the gas data in the storage space 104 in real time. In some embodiments, the gas sensor 108 may periodically acquire the gas data in the storage space 104, and store the acquired gas data in a corresponding memory for the controller 102 to read. Because a data volume of the gas data is small, the gas sensor 108 may acquire the gas data at a time interval (for example, 5 minutes, 10 minutes, or 15 minutes) smaller than an acquisition time interval of the camera 106. The controller 102 may detect the freshness of the food based on the gas data.
[0039] As shown in FIG. 1, after obtaining the image 112 and the gas data 114, the controller 102 may fuse image data and the gas data by using a cross-attention network 116, to generate fused data 118. The cross-attention network 116 is a neural network based on a cross-attention mechanism. The cross-attention network 116 may calculate a similarity between the image 112 and the gas data 114, and map information in the image 112 to the gas data 114 or map information in the gas data 114 to the image 112 by using an attention weight. The process may be implemented by constructing a relationship between a query, a key, and a value. In this way, the cross-attention network 116 can establish an association between the image data and the gas data. The generated fused data 118 may be, for example, in an embedded form (that is, a high-dimensional vector), and includes feature information of both the image 112 and the gas data 114. The fused data 118 may be used for a subsequent classification task, to generate the freshness of the food.
[0040] In the environment 100, the controller 102 may generate freshness 120 based on the fused data 118. In some embodiments, the controller 102 may input the fused data 118 into a trained classifier, to output a classification result, that is, the freshness 120. The freshness 120 may be, for example, one of a plurality of freshness levels. For example, the freshness of the food may be divided into a plurality of levels, for example, six levels. A level 1 represents the freshest, and a level 6 represents the least fresh. The freshness 120 may be one of the six levels.
[0041] In this manner, information in the image data and information in the gas data may be fused, and the freshness of the food may be determined based on fused data, so that limitations and errors caused by a single data source can be reduced, and accuracy of the determined freshness can be improved.
[0042] FIG. 2 is a flowchart of a method 200 for detecting food freshness according to some embodiments of the present disclosure. The method 200 may be performed by a controller of a food storage device, for example, the controller 102 in FIG. 1. As shown in FIG. 2, at block 202, the controller may obtain an image of a food placed in a storage space that is acquired by a camera placed in the storage space. For example, in the environment 100 shown in FIG. 1, the controller 102 may obtain the image 112 of the storage space 104 that is acquired by the camera 106, and the image 112 may include the food 110-1, the food 110-2, and the food 110- 3. In some embodiments, the controller 102 may send the instruction to the camera 106, to obtain the image of the food in the storage space 104 in real time.
[0043] In some embodiments, the controller 102 may read, from the memory, the image previously acquired by the camera 106.
[0044] At block 204, the controller may obtain gas data in the storage space that is acquired by a gas sensor placed in the storage space. For example, in the environment 100 shown in FIG. 1, the controller 102 may obtain the gas data 114 of the storage space 104 that is acquired by the gas sensor 108. The gas data 114 may include the plurality of gas components and the corresponding plurality of concentrations. In some embodiments, the controller 102 may send the instruction to the camera 106, to obtain the gas data in the storage space 104 in real time. In some embodiments, the controller 102 may read, from the memory, the gas data previously acquired by the gas sensor 108.
[0045] At block 206, the controller may generate fused data based on the image and the gas data by using a cross-attention network. For example, in the environment 100 shown in FIG. 1, the controller 102 may fuse the image data and the gas data by using the cross-attention network 116, to generate the fused data 118. The cross-attention network 116 is the neural network based on the cross-attention mechanism. The cross-attention network 116 can establish an association between the image data and the gas data, so that the fused data 118 can include the feature information of both the image 112 and the gas data 114. For example, the fused data 118 may be in the embedded form.
[0046] At block 208, the controller may determine freshness of the food based on the fused data. For example, in the environment 100 shown in FIG. 1, the controller 102 may generate the freshness 120 based on the fused data 118. For example, the controller 102 may input the fused data 118 to the trained classifier, to output the classification result, that is, the freshness 120. The freshness 120 may be, for example, one of the plurality of freshness levels.
[0047] In this manner, the information in the image data and the information in the gas data may be fused, and the freshness of the food may be determined based on the fused data, so that limitations and errors caused by the single data source can be reduced, and the accuracy of the determined freshness can be improved.
[0048] In some embodiments, an image of the storage space may be obtained based on a first time interval, the gas data includes a plurality of groups of gas data obtained based on a second time interval, and the first time interval is longer than the second time interval. In some embodiments, the controller may obtain a time stamp of the image of the storage space, and determine a target time period based on the time stamp and the first time interval. Then, the controller may determine a plurality of groups of gas data within the target time period from the gas data, and generate fused data based on the image of the storage space and the plurality of groups of gas data by using the cross-attention network. In some embodiments, the controller may determine a position of the food in the image of the storage space by using an object detection algorithm, and obtain the image of the food from the image of the storage space based on the position of the food. Then, the controller may generate the fused data based on the image of the food and the gas data by using the cross-attention network. In some embodiments, the controller may divide the image of the storage space into a plurality of sub-images. Then, the controller may determine, by using the object detection algorithm, that a first sub-image in the plurality of subimages includes a food, determine a position of the food in the first sub-image, and obtain an image of the food from the first sub-image based on the position of the food in the first subimage.
[0049] FIG. 3 is a flowchart of an exemplary process 300 for detecting food freshness according to some embodiments of the present disclosure. The process 300 may be performed by a controller in a food storage device. As shown in FIG. 3, at block 302, the controller may obtain an image of a storage space that is acquired by a camera, and gas data in the storage space that is acquired by a gas sensor. For example, the image of the storage space can be acquired by the camera in the storage space every one hour, where the image of the storage space may include all foods in the storage space. In addition, the gas sensor in the storage space may acquire the gas data in the storage space every 10 minutes. The controller may obtain a latest acquired image and a time stamp of the image. Because an acquisition time interval of image data is usually longer than an acquisition time interval of the gas data, the controller can determine a target time period based on the time stamp of the image and the acquisition time interval (for example, one hour) of the image data, to determine freshness of the food based on image data and a plurality of groups of gas data within the target time period. A time length of the target time period may be the acquisition time interval (for example, one hour) of the image data, an end moment of the target time period may be the time stamp of the image, and a start moment may be a moment obtained by subtracting the time length of the target time period from the end moment. For example, the controller may obtain, based on the time stamp of the latest acquired image, a plurality of groups of gas data (for example, six groups of gas data) within one hour before the time stamp. The controller may detect the freshness of the food in the storage space based on the image and the plurality of groups of gas data. In some other embodiments, the time length of the target time period may alternatively include an acquisition time interval of a plurality of pieces of image data. In this way, the freshness of the food may be determined based on a plurality of images and a corresponding plurality of groups of gas data. In some other embodiments, the start moment of the target time period may be the time stamp of the image, and the end moment may be a moment obtained by adding the start moment to the time length of the target time period. In this manner, the freshness of the food is detected by using the plurality of groups of gas data, so that a gas change of the food in the storage space can be known more comprehensively by using data at different time points, thereby improving accuracy of the determined freshness.
[0050] At block 304, the controller may mark a food in the image, to obtain a position and a category of the food. For example, the controller may mark a plurality of foods in the image by using an object detection algorithm, to obtain respective positions and categories of the plurality of foods. The object detection algorithm is a computer vision technology, and aims at recognizing and positioning a specific object in an image or a video. The object detection algorithm can determine a category of the object in the image, and can provide a precise position of each object. For example, the object detection algorithm may be YOLO, Faster R- CNN, an SSD algorithm, or the like. In the object detection algorithm, the position of the food may be represented, for example, by using an upper-left coordinate and a lower-right coordinate of a boundary box that defines a food boundary. The category of the food may be, for example, a vegetable, a fruit, a meat, or seafood. Because a processing capability of the controller is limited, the controller may first divide the image of the storage space into a plurality of subimages (for example, 2 x 2, 3 x 3, or 4 x 4 sub-images), and then sequentially perform object detection on the sub-images, to recognize a food included in each sub-image and mark a position of the food in the sub-image. In this manner, a processing resource of the controller can be saved, which is crucial for an embedded system.
[0051] At block 306, the controller may obtain an image of the food that is included in the image of the storage space based on the position of the food. For example, in the image (or a subimage) of the storage space, a position of an apple is represented, for example, by upper-left coordinates (xl, yl) and lower-right coordinates (x2, y2) of the boundary box. The controller may obtain a part of the image of the storage space based on the two sets of coordinates, and the part only includes or mainly includes the apple. In a subsequent process, the controller may detect freshness of the apple based on the image. In this manner, noise and interference information in the image can be reduced, thereby improving accuracy of the detected freshness.
[0052] At block 308, the controller may obtain a part of gas data corresponding to the category of the food in the gas data based on the category. For example, in the plurality of groups of gas data corresponding to the acquired image, each group of gas data may include concentrations of a plurality of gas components (for example, 60 or more). A part of the gas components may be related to freshness of the fruit, and a part of the gas components may be related to freshness of the meat. The two parts may include overlapping gas components. For example, the meat may release gases such as ammonia, hydrogen sulfide, and carbon dioxide in a spoilage process, and the apple may release gases such as ethylene, carbon dioxide, and aldehydes in a spoilage process. The controller may select several gas components associated with the apple from the plurality of gas components, and determine the freshness of the apple in the image based on the concentrations of the gas components. In this manner, corresponding gas data can be selected for the category of the food in the image to detect the freshness of the food, so that interference caused by changes of other gas components can be reduced, and the accuracy of the determined freshness can be improved.
[0053] At block 310, the controller may fuse the image of the food and the part of gas data by using a cross-attention network, to generate the freshness of the food. For example, the controller may fuse the image data and the gas data for an image only including the apple and a part of gas data associated with the apple in the plurality of groups of gas data by using the cross-attention network, to generate fused data including image feature information and gas feature information. Then, the controller may generate the freshness of the apple based on the fused data.
[0054] At block 312, the controller may determine whether the freshness of the food satisfies a predetermined freshness level threshold. The predetermined freshness level threshold may be manually calibrated, may be determined by the controller in real time, or may be input by a user. For example, if freshness levels include a total of six levels (for example, a level 1 to a level 6), the level 1 represents the freshest, and the level 6 represents the least fresh. The user may set the freshness level threshold to a level 3. If the detected freshness of the apple is lower than the level 3 (for example, a level 2), the controller may determine that the apple is still fresh at a current moment, so that the process 300 may return to 302 to wait for next detection. If the detected freshness of the apple reaches the level 3 or exceeds the level 3, the controller may determine that the apple is not fresh, and the process 300 may proceed to block 314.
[0055] At block 314, the controller may send a notification to prompt the user that the food has spoiled. For example, the controller may display the freshness of the apple by using a user interaction assembly. The user interaction assembly may include an indicator or a display screen on the food storage device, or the controller may send a notification indicating the freshness of the food to a user equipment. In this manner, the user may deal with the spoiled food in time based on a status of the user interaction assembly or the received notification, so that user experience can be improved.
[0056] In some embodiments, during generation of the fused data, the controller may generate an image embedding based on the image of the food, and generate a gas embedding based on the gas data. Then, the controller may fuse the image embedding and the gas embedding by using the cross-attention network, to generate a fused embedding as the fused data. In some embodiments, the controller may determine the category of the food based on the image of the food, and select a part of gas data corresponding to the category of the food from the gas data. Then, the controller may generate, based on the part of gas data, the gas embedding by using a feature extractor corresponding to the category of the food. In some embodiments, the controller may determine, based on the fused embedding by using a classifier network, a freshness level corresponding to the food from a plurality of freshness levels as the freshness of the food.
[0057] FIG. 4 is a schematic diagram of an exemplary system architecture 400 for detecting food freshness by fusing image data and gas data according to some embodiments of the present disclosure. As shown in FIG. 4, the architecture 400 includes a feature extractor 406 used for an image, a feature extractor 408 used for the gas data, a cross-attention network 414, and a classifier 418. In the architecture 400, an image 402 may be an image of a food (for example, an image of an apple) obtained from an image of a storage space.
[0058] A gas data 404 may include a plurality of groups of gas data corresponding to the image 402, and the gas data 404 may be a part of gas data determined from the plurality of groups of gas data based on a category of the food (for example, including a part of gas composition selected for a fruit category or further, an apple category).
[0059] In the architecture 400, the feature extractor 406 may be a visual feature extractor, configured to extract meaningful features from the image. The features are usually information such as a key point, an edge, a texture, and a shape that can represent image content, to be used by a subsequent task (for example, a classification, detection, or segmentation task). For example, the feature extractor 406 may be a convolutional neural network that extracts a feature in the image by using a multilayer convolution operation. As shown in FIG. 4, the feature extractor 406 may generate an image embedding 410 based on the image 402, and the image embedding 410 may include feature information in the image 402.
[0060] In the architecture 400, the feature extractor 408 is configured to extract feature information from the gas data. For example, the feature extractor 408 may be a recurrent neural network, which may process time series data to extract a dynamic feature from the time series data. As shown in FIG. 4, the feature extractor 408 may generate a gas embedding 412 based on the gas data 404.
[0061] In the architecture 400, the cross-attention network 414 (for example, the cross-attention network 116 in FIG. 1) is a neural network based on a cross-attention mechanism. The crossattention network 414 may fuse the image embedding and the gas embedding, to generate a fused embedding 416. The fused embedding 416 may include an image feature in the image embedding 410 and a gas data feature in the gas embedding 412. In comparison with another fusion mechanism, the cross-attention network can dynamically adjust weights between different modalities based on an input context, to ensure that important information can be highlighted, thereby improving fusion flexibility and accuracy. In addition, through cross attention, a model can capture a complex relationship between the image and the gas data, to help understand interactions between different modalities, thereby improving overall understanding of a scene. In addition, when processing noise or missing data, the crossattention network can perform adaptive adjustment by using an attention mechanism, so that a system is more robust to data.
[0062] After the fused embedding 416 is generated, the architecture 400 may input the fused embedding 416 into the classifier 418, to generate freshness 420 of the food (for example, the apple) in the image 402. For example, the classifier 418 may be a deep neural network, which includes a plurality of fully connected layers, and may perform feature transformation by using a non-linear activation function (for example, ReLU or Sigmoid). An output layer of the classifier 418 may transform an output of the network into a probability distribution by using a Softmax activation function, and determine a freshness level (for example, fresh, secondarily fresh, not fresh, a level 1, or a level 2) of the food based on an output probability value. In this manner, the overall understanding of the scene can be improved, flexibility and accuracy of data fusion can be enhanced, and robustness of the system can be improved when the data includes the noise or the data is missing.
[0063] In some embodiments, the cross-attention network is a first cross-attention network, the fused embedding is a first fused embedding, and the first cross-attention network may include a second image-to-gas cross-attention network and a third gas-to-image cross-attention network. The controller may generate a second fused embedding based on the image embedding and the gas embedding by using the second cross-attention network, where the image embedding is used as a query for the second cross-attention network, and the gas embedding is used as a key and a value of the second cross-attention network. Then, the controller may generate a third fused embedding based on the image embedding and the gas embedding by using the third cross-attention network, where the gas embedding is used as a query for the third cross-attention network, and the image embedding is used as a key and a value of the third cross-attention network. Then, the controller may generate the first fused embedding based on the second fused embedding and the third fused embedding.
[0064] FIG. 5 is a schematic diagram of a structure of an exemplary cross-attention network 500 according to some embodiments of the present disclosure. The cross-attention network 500 may be, for example, the cross-attention network 414 in FIG. 4. As shown in FIG. 5, the crossattention network 500 includes a gas-to-image cross-attention network 506 and an image-togas cross-attention network 508. In FIG. 5, an image embedding 502 may be, for example, the image embedding 410 in FIG. 4, which includes feature information in an image of a food. A gas embedding 504 may be, for example, the gas embedding 412 in FIG. 4, which includes feature information in gas data.
[0065] The gas-to-image cross-attention network 506 may guide a focus of an image feature by using a feature of the gas data, to dynamically adjust a part of the image that is related to the gas data. For example, in the gas-to-image cross-attention network 506, the gas embedding 504 may be used as a query, and the image embedding 502 may be used as a key and a value. The gas-to-image cross-attention network 506 may generate an attention weight by calculating a similarity between the query and the key (for example, by using a dot product operation), and weight a value by using the generated attention weight, to enhance an image feature related to the gas data. The gas-to-image cross-attention network 506 may fuse a weighted image feature and a gas feature, to generate a fused embedding 510.
[0066] The image-to-gas cross-attention network 508 may guide a focus of the gas feature by using the image feature, to dynamically adjust a part of the gas data that is related to image data. For example, in the image-to-gas cross-attention network 508, the image embedding 502 may be used as a query, and the gas embedding 504 may be used as a key and a value. The image-to-gas cross-attention network 508 may generate an attention weight by calculating a similarity between the query and the key, and weight a value by using the generated attention weight, to enhance a gas feature related to the image data. The image-to-gas cross-attention network 508 may fuse a weighted gas feature and the image feature, to generate a fused embedding 512.
[0067] As shown in FIG. 5, after the fused embedding 510 and the fused embedding 512 are generated, a fusion module 514 may generate a fused embedding 516 based on the fused embedding 510 and the fused embedding 512. In some embodiments, the fusion module 514 can directly concatenate the fused embedding 510 and the fused embedding 512 together, to generate the fused embedding 516. In this manner, a processing resource can be saved. In some embodiments, the fusion module 514 may be a network (for example, a deep neural network) used for feature fusion, which can learn a complex relationship between two fused embeddings. For example, the fusion module 514 can concatenate the fused embedding 510 and the fused embedding 512 together, and then generate the fused embedding 516 based on a concatenated embedding by using the deep neural network. In this manner, a relationship between the gas feature and the image feature can be fully revealed, and an optimal fusion policy can be automatically learned, thereby improving accuracy of detected food freshness.
[0068] In some embodiments, the controller may generate a first fused embedding based on a second fused embedding, a third fused embedding, the image embedding, and the gas embedding. FIG. 6 is a schematic diagram of a structure of another exemplary cross-attention network 600 according to some embodiments of the present disclosure. The cross-attention network 600 may be, for example, the cross-attention network 414 in FIG. 4. An image embedding 602, a gas embedding 604, a gas-to-image cross-attention network 606, an image- to-gas cross-attention network 608, a fused embedding 610, and a fused embedding 612 may be, for example, the image embedding 502, the gas embedding 504, the gas-to-image crossattention network 506, the image-to-gas cross-attention network 508, the fused embedding 510, and the fused embedding 512 in FIG. 5.
[0069] As shown in FIG. 6, in the cross-attention network 600, in addition to the fused embedding 610 and the fused embedding 612, a fusion module 614 may further generate a fused embedding 616 by fusing the image embedding 602 and the gas embedding 604. In some embodiments, the fusion module 614 can directly concatenate the fused embedding 610, the fused embedding 612, the image embedding 602, and the gas embedding 604 together, to generate the fused embedding 616. In this manner, a processing resource can be saved. In some embodiments, the fusion module 614 can concatenate the fused embedding 610, the fused embedding 612, the image embedding 602, and the gas embedding 604 together, and then generate the fused embedding 616 based on a concatenated embedding by using a deep neural network. In this manner, accuracy of detected food freshness can be improved.
[0070] Compared to only combining the fused embedding 610 and the fused embedding 612 to generate a new fused embedding, combining the original image embedding 602 and the gas embedding 604 with the fused embedding 610 and the fused embedding 612 together can retain basic information of gas and image data in original features that are not completely captured in a cross-attention fusion process, thereby ensuring that important low-level features are not lost. In addition, the original feature is combined with a fused feature, information of different levels can be fused, so that a capability of the cross-attention network 600 for understanding a complex relationship can be improved, and the accuracy of the determined food freshness can be improved.
[0071] At a training stage, a training data set may include an image data set of a storage space, a corresponding gas data set from a gas sensor, and corresponding ground truth freshness. The ground truth freshness may be annotated in collected training data manually, through expert evaluation, or by using an existing standard. In other words, the corresponding ground truth freshness may be added to each group of image and gas data as a label, to form a complete training sample.
[0072] During training, a feature extractor that has been trained may be used to extract the image embedding from the image training data and extract the gas embedding from the gas training data. Then, the fused embedding may be generated by using the image-to-gas cross-attention network and the gas-to-image cross-attention network, and predicted freshness may be generated based on the generated fused embedding by using a classifier. After the predicted freshness is generated, a loss is determined based on the predicted freshness and the ground truth freshness in the training data set. For example, the loss may be determined based on the predicted freshness and the ground truth freshness by using a cross-entropy loss function. Then, the loss may be minimized by simultaneously adjusting the image-to-gas cross-attention network, the gas-to-image cross-attention network, and the classifier.
[0073] In some embodiments, a controller may obtain temperature data in the storage space that is acquired by a temperature sensor placed in the storage space. The controller may further obtain humidity data in the storage space that is acquired by a humidity sensor placed in the storage space. Then, the controller may generate fused time series data based on the gas data, the temperature data, and the humidity data, and generate fused data based on an image of the storage space and the fused time series data by using the cross-attention network.
[0074] FIG. 7 is a schematic diagram of an exemplary system architecture 700 for detecting food freshness by fusing a plurality of types of time series data according to some embodiments of the present disclosure. As shown in FIG. 7, the architecture 700 includes a feature extractor 710 used for an image, a feature extractor 712 used for the time series data, a cross-attention network 718, and a classifier 722. The feature extractor 710 used for an image may be, for example, the feature extractor 406 in FIG. 4, which may generate an image embedding 714 based on an image 702.
[0075] In the architecture 700, the feature extractor 712 is configured to extract feature information from the time series data. The time series data may include gas data 704, temperature data 706, and humidity data 708. The gas data 704 may be data acquired by a gas sensor at a specific time interval, which includes a plurality of groups of gas data arranged chronologically within a specific time period; the temperature data 706 may be data acquired by a temperature sensor disposed in a storage space at the specific time interval, which includes a plurality of groups of temperature data arranged chronologically within the specific time period; and the humidity data 708 may be data acquired by a humidity sensor disposed in the storage space at the specific time interval, which includes a plurality of groups of humidity data arranged chronologically within the specific time period. Because the gas data 704, the temperature data 706, and the humidity data 708 are time series data, the feature extractor 712 may extract a time series feature from the data to generate a time series data embedding 716. The time series data embedding 716 includes feature information of the gas data 704, the temperature data 706, and the humidity data 708.
[0076] In the architecture 700, the cross-attention network 718 may fuse the image embedding 714 and the time series data embedding 716, to generate a fused embedding 720. The fused embedding 720 may include feature information of the image 702, the gas data 704, the temperature data 706, and the humidity data 708. Then, the classifier 722 may generate freshness 724 of a food in the image 702 based on the fused embedding 720.
[0077] In this manner, because humidity and temperature are important environmental factors affecting food freshness, after the data is fused, a model can obtain more comprehensive context information, helping evaluate the food freshness more accurately. In addition, different foods may have different spoilage rates under different humidity and temperature conditions. By taking the factors into consideration, the model can better adapt to changes in an environment, thereby improving performance of the model in a real scenario.
[0078] According an embodiment of the present disclosure, a method for detecting food freshness based on a random forest model is further provided. FIG. 8 is a flowchart of a method 800 for detecting food freshness based on a random forest model according to some embodiments of the present disclosure. As shown in FIG. 8, at block 802, a controller may obtain an image of a food that is acquired by a camera placed in a storage space. For example, the controller may obtain an image of the storage space, divide the image of the storage space into a plurality of sub-images, and then separately apply an object detection algorithm to each of the plurality of sub-images to mark a position of the food in the sub-image. Then, the controller may obtain the image of the food based on the position of the food.
[0079] At block 804, the controller may obtain gas data in the storage space that is acquired by a gas sensor placed in the storage space. For example, the controller may obtain a plurality of groups of gas data within a time period based on a time stamp of image data, and determine a part of gas data associated with a category of the food from the plurality of groups of gas data based on the category of the food determined through the object detection algorithm.
[0080] At block 806, the controller may generate a first freshness result based on the image of the food. For example, the controller may input the image of the food into a trained first freshness detection model, to output the first freshness result. The first freshness result includes a plurality of freshness levels and a corresponding plurality of confidences. For example, the first freshness result may include a level 1, and a corresponding confidence is 40%; the first freshness result may include a level 2, and a corresponding confidence is 60%; and the first freshness result may include a level 3, and a corresponding confidence is 80%. The first freshness detection model may be any model that determines food freshness based on the image. For example, the first freshness detection model may be a convolutional neural network or another model (for example, ResNet, MobileNet, VGGNet, or EfficientNet) based on the convolutional neural network. At a training stage, a food image with a label may be collected, and the label may cover the plurality of freshness levels, for supervised learning. An initial weight of the model may be randomly initialized, or fine adjustment is performed on a pretrained weight, to accelerate convergence of the model. Then, the food image may be input into the model, to output a freshness result. The output freshness result may be compared with the label of the food image to calculate a loss, and then the weight of the model is updated through backward propagation.
[0081] At block 808, the controller may generate a second freshness result based on the gas data. For example, the controller may input the gas data into a trained second freshness detection model, to output the second freshness result. The second freshness result includes the same plurality of freshness levels and correspondingly different confidences. The second freshness detection model may be any model that determines the food freshness based on the gas data. For example, the second freshness detection model may be a linear regression model, a logistic regression model, a support vector machine, a deep neural network, or the like. At the training stage, gas sensor data may be collected, and a feature (for example, a gas concentration change) related to the freshness is extracted. Then, a freshness label may be used as a target variable, and the second freshness detection model is trained by minimizing a mean square error between a predicted value and a real value.
[0082] At block 810, the controller may generate a fused freshness result based on the first freshness result and the second freshness result by using the random forest model, to determine freshness of the food. For example, the controller may input the first freshness result and the second freshness result into the random forest model, and the random forest model may generate the fused freshness result by fusing the first freshness result and the second freshness result. The fused freshness result includes the same plurality of freshness levels and correspondingly different confidences.
[0083] In some embodiments, the method 800 may further include: At block 812, the controller may determine the freshness of the food based on the fused freshness result. For example, the controller may determine a freshness level having the highest confidence in the fused freshness result as the freshness of the food.
[0084] In this manner, freshness detection results of different data sources can be fused, thereby improving accuracy of the determined freshness. In addition, in this manner, two independent freshness detection models can be used for different data sources, thereby simplifying an overall architecture design, and reducing model complexity. In addition, each freshness detection model may focus on feature extraction and learning of specific data of the freshness detection model, reducing design and debugging complexity. In addition, the random forest model requires fewer processing resources and space resources, and is more suitable for being applied to an embedded device.
[0085] FIG. 9 is a schematic diagram of an exemplary system architecture 900 for fusing an image freshness result and a gas freshness result by using a random forest model according to some embodiments of the present disclosure. As shown in FIG. 9, an image 902 is an image of a food in a storage space (for example, an image of an apple in the storage space), and gas data 904 is a part of gas data associated with the food in the image 902, which may include a plurality of groups of gas data arranged in a time series.
[0086] In an example of the architecture 900, freshness of the food is divided into three levels, that is, freshness levels 910-1, 910-2, and 910-3. A first freshness detection model 906 may generate a first freshness result based on the image 902, including a confidence 912 (for example, 40%) of the freshness level 910-1, a confidence 912-2 (for example, 60%) of the freshness level 910-2, and a confidence 912-3 (for example, 80%) of the freshness level 910- 3. In addition, a second freshness detection model 908 may generate a second freshness result based on the gas data 904, including a confidence 914-1 (for example, 20%) of the freshness level 910-1, a confidence 914-2 (for example, 30%) of the freshness level 910-2, and a confidence 914-3 (for example, 60%) of the freshness level 910-3.
[0087] In the architecture 900, two groups of freshness levels generated by the first freshness detection model and the second freshness detection model and corresponding confidences may be input into a random forest model 916. The random forest model 916 may fuse two groups of freshness results, to generate a fused freshness result, and generate freshness 918 based on the fused freshness result.
[0088] In some embodiments, in addition to the two groups of freshness results from image data and the gas data, an input of the random forest model 916 may further include a food category 920 (for example, a fruit, a vegetable, or a meat), a food position 922 (for example, a two- dimensional coordinate of the food in the image 902), and a food size 924 (for example, a size of the food in the image 902). Then, the random forest model 916 may generate the freshness 918 of the food based on the freshness result from the first freshness detection model 906, the freshness result from the second freshness detection model 908, the food category 920, the food position 922, and the food size 924.
[0089] In this manner, foods of different categories may have different freshness features, and adding the food category can help the model to better capture the features, thereby improving prediction accuracy. A size of the food may affect a spoilage rate of the food, and a larger food may take a longer time to become spoiled completely. Providing size information helps the model to perform more accurate evaluation. In addition, a plurality of input features can make the model more robust, can maintain good prediction performance under different conditions, and can reduce misjudgment caused by a single feature.
[0090] FIG. 10 is a block diagram of an apparatus 1000 for detecting food freshness according to some embodiments of the present disclosure. As shown in FIG. 10, the apparatus 1000 includes a food image obtaining module 1002, configured to obtain an image of a food placed in a storage space that is acquired by a camera placed in the storage space. The apparatus 1000 further includes a gas data obtaining module 1004, configured to obtain gas data in the storage space that is acquired by a gas sensor placed in the storage space. The apparatus 1000 further includes a fused data generation module 1006, configured to generate fused data based on the image and the gas data by using a cross-attention network. In addition, the apparatus 1000 further includes a freshness determining module 1008, configured to determine freshness of the food based on the fused data.
[0091] In some embodiments, the image is obtained based on a first time interval, the gas data includes a plurality of groups of gas data obtained based on a second time interval, and the first time interval is longer than the second time interval.
[0092] In some embodiments, the fused data generation module 1006 includes: a time stamp obtaining module, configured to obtain a time stamp of the image; a target time period determining module, configured to determine a target time period based on the time stamp of the image and the first time interval; a target time period using module, configured to determine a plurality of groups of gas data within the target time period from the gas data; and a gas data using module, configured to generate the fused data based on the image and the plurality of groups of gas data by using the cross-attention network.
[0093] In some embodiments, the food image obtaining module 1002 includes: a spatial image obtaining module, configured to obtain an image of the storage space that is acquired by the camera placed in the storage space, where the image of the storage space includes the food; and the fused data generation module 1006 includes: a food position determining module, configured to determine a position of the food in the image of the storage space by using an object detection algorithm; a food position using module, configured to obtain the image of the food from the image of the storage space based on the position of the food; and a food image using module, configured to generate the fused data based on the image of the food and the gas data by using the cross-attention network.
[0094] In some embodiments, the food position determining module includes: a sub-image division module, configured to divide the image of the storage space into a plurality of subimages; and a sub-image using module, configured to determine, by using the object detection algorithm, that a first sub-image in the plurality of sub-images includes the food and determine a position of the food in the first sub-image; and the food position using module includes: a sub-image position using module, configured to obtain the image of the food from the first subimage based on the position of the food in the first sub-image.
[0095] In some embodiments, the fused data is a fused embedding, and the food image using module includes: an image embedding generation module, configured to generate an image embedding based on the image of the food; a gas embedding generation module, configured to generate a gas embedding based on the gas data; and a fused embedding generation module, configured to fuse the image embedding and the gas embedding by using the cross-attention network, to generate the fused embedding.
[0096] In some embodiments, the gas embedding generation module includes: a food category determining module, configured to determine a category of the food based on the image of the food; a partial gas data selection module, configured to select a part of gas data corresponding to the category of the food from the gas data; and a partial gas data using module, configured to generate, based on the part of gas data, the gas embedding by using a feature extractor corresponding to the category of the food.
[0097] In some embodiments, the cross-attention network is a first cross-attention network, the fused embedding is a first fused embedding, the first cross-attention network includes a second image-to-gas cross-attention network and a third gas-to-image cross-attention network, and the fused embedding generation module includes: a second fused embedding generation module, configured to generate, based on the image embedding and the gas embedding, a second fused embedding by using the second cross-attention network, where the image embedding is used as a query for the second cross-attention network, and the gas embedding is used as a key and a value of the second cross-attention network; a third fused embedding generation module, configured to generate, based on the image embedding and the gas embedding, a third fused embedding by using the third cross-attention network, where the gas embedding is used as a query for the third cross-attention network, and the image embedding is used as a key and a value of the third cross-attention network; and a first fused embedding generation module, configured to generate the first fused embedding based on the second fused embedding and the third fused embedding.
[0098] In some embodiments, the first fused embedding generation module includes: a fourth fused embedding generation module, configured to generate the first fused embedding based on the second fused embedding, the third fused embedding, the image embedding, and the gas embedding.
[0099] In some embodiments, the fused data generation module 1006 includes: a temperature data obtaining module, configured to obtain temperature data in the storage space that is acquired by a temperature sensor placed in the storage space; a humidity data obtaining module, configured to obtain humidity data in the storage space that is acquired by a humidity sensor placed in the storage space; a fused time series data generation module, configured to generate fused time series data based on the gas data, the temperature data, and the humidity data; and a fused time series data using module, configured to generate, based on the image of the storage space and the fused time series data, fused data by using the cross-attention network.
[0100] In some embodiments, the freshness of the food is one of a plurality of freshness levels, and the freshness determining module 1008 includes: a classifier network using module, configured to determine, based on the fused data by using a classifier network, a freshness level corresponding to the food from the plurality of freshness levels as the freshness of the food.
[0101] It can be understood that, by using the apparatus 1000 of the present disclosure, at least one of a large quantity of advantages that can be implemented by using the foregoing described methods or processes may be implemented. For example, in this manner, information in the image data and information in the gas data may be fused, and the freshness of the food may be determined based on fused data, so that limitations and errors caused by a single data source can be reduced, and accuracy of the determined freshness can be improved.
[0102] FIG. 11 is a block diagram of a controller 1100 in which a plurality of embodiments of the present disclosure can be implemented.
[0103] The controller 1100 may be, for example, the controller 102 shown in FIG. 1. As shown in the figure, the controller 1100 includes a processor 1101, which can perform various appropriate actions and processing according to computer program instructions that are stored in a read-only memory (ROM) 1102 and that are loaded into a random access memory (RAM)
[0104] 1103. The RAM 1103 can further store various programs and data required for operating the controller 1100. The processor 1101, the ROM 1102, and the RAM 1103 are connected to each other through a bus 1104. An input / output (I / O) interface 1105 is also connected to the bus
[0105] 1104.
[0106] The processor 1101 may be various general-purpose and / or special-purpose processing components having processing and computing capabilities. Some examples of the processor 1101 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (Al) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, and the like. The processor 1101 performs the methods and processing described above, for example, the method 200. For example, in some embodiments, the method 200 may be implemented as a computer software program that is tangibly included in a machine-readable medium. In some embodiments, a part or all of the computer program may be loaded into and / or installed on the controller 1100 though the ROM 1102. When the computer program is loaded into the RAM 1103 and executed by the processor 1101, one or more steps of the method 200 described above may be performed. Alternatively, in other embodiments, the processor 1101 may be configured to perform the method 200 in any other suitable manner (for example, with the help of firmware).
[0107] The functions described above in this specification may be at least partially performed by one or more hardware logic components. For example, but not limited to, exemplary types of hardware logic components that may be used include: a field programmable gate array (FPGA), an application-specific integrated circuit (ASIC), an application-specific standard product (ASSP), a system-on-chip (SOC), a complex programmable logic device (CPLD), and the like.
[0108] Program code used for implementing the method of the present disclosure may be written in any combination of one or more programming languages. The program code may be provided for a processor or a controller of a general-purpose computer, a special-purpose computer, or another programmable data processing apparatus, so that the program code, when executed by the processor or the controller, enables implementation of the functions / operations specified in the flowcharts and / or the block diagrams. The program code may be completely executed on a machine, partially executed on the machine, partially executed on the machine as an independent software package and partially executed on a remote machine, or completely executed on a remote machine or server.
[0109] In the context of the present disclosure, the machine-readable medium may be a tangible medium that may include or store a program used by an instruction execution system, apparatus, or device or used in combination with an instruction execution system, apparatus, or device. The machine-readable medium may be a machine-readable signal medium or a machine- readable storage medium. The machine-readable medium may include, but is not limited to, an electric, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the above. A more specific example of the machine- readable storage medium may include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In addition, although the operations are described in a specific order, it should be understood that the operations are required to be performed in the shown specific order or in a sequential order, or all the shown operations are required to be performed to obtain an expected result. In an environment, multitasking and parallel processing may be advantageous. Similarly, although several specific implementation details are included in the foregoing discussion, the details should not be construed as a limitation on the scope of the present disclosure. Some features described in the context of the separate embodiments may further be implemented in a single implementation in combination. On the contrary, various features described in the context of the single implementation may alternatively be implemented in a plurality of implementations separately or in any suitable sub-combination.
[0110] Although the subject has been described by using languages specific to structural features and / or methods and logical actions, it should be understood that the subject defined in the appended claims is not necessarily limited to the specific features or actions described above. On the contrary, the specific features and actions described above are merely example forms for implementing the claims.
Claims
CLAIMSWhat is claimed is:
1. A method (200) for detecting food freshness, comprising: obtaining (202) an image of a food placed in a storage space that is acquired by a camera placed in the storage space; obtaining (204) gas data in the storage space that is acquired by a gas sensor placed in the storage space; generating (206) fused data based on the image and the gas data by using a cross-attention network; and determining (208) freshness of the food based on the fused data.
2. The method (200) according to claim 1, characterized in that the image is obtained based on a first time interval, the gas data comprises a plurality of groups of gas data obtained based on a second time interval, and the first time interval is longer than the second time interval.
3. The method (200) according to claim 2, characterized in that the generating (206) fused data based on the image and the gas data of the storage space by using a cross-attention network comprises: obtaining a time stamp of the image; determining a target time period based on the time stamp of the image and the first time interval; determining a plurality of groups of gas data within the target time period from the gas data; and generating the fused data based on the image and the plurality of groups of gas data by using the cross-attention network.
4. The method (200) according to claim 1, characterized in that the obtaining (202) an image of a food placed in a storage space that is acquired by a camera placed in the storage space comprises: obtaining an image of the storage space that is acquired by the camera placed in the storage space, wherein the image of the storage space comprises the food; and the generating (206) fused data based on the image and the gas data by using a crossattention network comprises: determining a position of the food in the image of the storage space by using an objectdetection algorithm; obtaining the image of the food from the image of the storage space based on the position of the food; and generating the fused data based on the image of the food and the gas data by using the cross-attention network.
5. The method (200) according to claim 4, characterized in that the determining a position of the food in the image of the storage space by using an object detection algorithm comprises: dividing the image of the storage space into a plurality of sub-images; and determining, by using the object detection algorithm, that a first sub-image in the plurality of sub-images comprises the food, and determining a position of the food in the first sub-image; and the obtaining the image of the food from the image of the storage space based on the position of the food comprises: obtaining the image of the food from the first sub-image based on the position of the food in the first sub-image.
6. The method (200) according to claim 4, characterized in that the fused data is a fused embedding, and the generating the fused data based on the image of the food and the gas data by using the cross-attention network comprises: generating an image embedding based on the image of the food; generating a gas embedding based on the gas data; and fusing the image embedding and the gas embedding by using the cross-attention network, to generate the fused embedding.
7. The method (200) according to claim 6, characterized in that the generating a gas embedding based on the gas data comprises: determining a category of the food based on the image of the food; selecting a part of gas data corresponding to the category of the food from the gas data; and generating, based on the part of gas data, the gas embedding by using a feature extractor corresponding to the category of the food.
8. The method (200) according to claim 6, characterized in that the cross-attention network is a first cross-attention network, the fused embedding is a first fused embedding, the first crossattention network comprises a second image-to-gas cross-attention network and a third gas-to-image cross-attention network, and the fusing the image embedding and the gas embedding by using the cross-attention network, to generate the fused embedding comprises: generating a second fused embedding based on the image embedding and the gas embedding by using the second cross-attention network, wherein the image embedding is used as a query for the second cross-attention network, and the gas embedding is used as a key and a value of the second cross-attention network; generating a third fused embedding based on the image embedding and the gas embedding by using the third cross-attention network, wherein the gas embedding is used as a query for the third cross-attention network, and the image embedding is used as a key and a value of the third cross-attention network; and generating the first fused embedding based on the second fused embedding and the third fused embedding.
9. The method (200) according to claim 8, characterized in that the generating the first fused embedding based on the second fused embedding and the third fused embedding comprises: generating the first fused embedding based on the second fused embedding, the third fused embedding, the image embedding, and the gas embedding.
10. The method (200) according to claim 1, characterized in that the generating (206) fused data based on the image and the gas data by using a cross-attention network comprises: obtaining temperature data in the storage space that is acquired by a temperature sensor placed in the storage space; obtaining humidity data in the storage space that is acquired by a humidity sensor placed in the storage space; generating fused time series data based on the gas data, the temperature data, and the humidity data; and generating the fused data based on the image of the storage space and the fused time series data by using the cross-attention network.
11. The method (200) according to claim 1, characterized in that the freshness of the food is one of a plurality of freshness levels, and the determining (208) freshness of the food based on the fused data comprises: determining, based on the fused data by using a classifier network, a freshness level corresponding to the food from the plurality of freshness levels as the freshness of the food.
12. A method (800) for detecting food freshness, comprising: obtaining (802) an image of a food placed in a storage space that is acquired by a camera placed in the storage space; obtaining (804) gas data in the storage space that is acquired by a gas sensor placed in the storage space; generating (806) a first freshness result based on the image of the food; generating (808) a second freshness result based on the gas data, wherein the first freshness result and the second freshness result comprise a plurality of freshness levels and a plurality of confidences corresponding to the plurality of freshness levels; and generating (810) a fused freshness result based on the first freshness result and the second freshness result by using a random forest model, to determine freshness of the food.
13. An apparatus (1000) for detecting food freshness, comprising: a food image obtaining module (1002), configured to obtain an image of a food placed in a storage space that is acquired by a camera placed in the storage space; a gas data obtaining module (1004), configured to obtain gas data in the storage space that is acquired by a gas sensor placed in the storage space; a fused data generation module (1006), configured to generate fused data based on the image and the gas data by using a cross-attention network; and a freshness determining module (1008), configured to determine freshness of the food based on the fused data.
14. A controller (1100), comprising: at least one processor (1101); and a memory (1102), coupled to the at least one processor and having instructions stored therein, wherein the instructions, when executed by the at least one processor (1101), cause the controller (1100) to perform the method according to any one of claims 1 to 12.
15. A food storage device, comprising: a storage space (104); a camera (106), disposed at a top of the storage space (104), wherein the camera is configured to photograph a food in the storage space; a gas sensor (108), disposed inside the storage space (104), wherein the gas sensor (108) is configured to acquire gas data in the storage space (104); and the controller (1100) according to claim 14.