Freezer goods identification system based on artificial intelligence

By using high-definition cameras and infrared cameras to collect image data in collaboration, and combining deep occlusion recognition and multimodal image fusion recognition mechanisms, the problem of low accuracy in identifying goods in complex environments like refrigerated display cases has been solved, achieving high-quality imaging and high-accuracy goods identification.

CN120997816AInactive Publication Date: 2025-11-21WANBAO ELECTRICAL APPLIANCE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511395922.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-28
Publication Date
2025-11-21
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional refrigerated display case product identification systems have low accuracy in complex environments, making it difficult to meet commercial requirements. Existing AI identification systems also have insufficient accuracy when dealing with complex internal structures of refrigerated display cases, mirror reflections, frost obstruction, and other factors.

Method used

High-definition cameras and infrared cameras are used to collect image data in collaboration. Combined with a deep occlusion recognition algorithm and dynamic shooting angle adjustment, a multimodal image fusion recognition mechanism is used to extract features through ResNet50+FPN and U-Net++ network structures. The recognition is performed by combining YOLOv8 object detection and ViT fusion structure, and the image enhancement process is triggered when the occlusion is severe.

Benefits of technology

Achieving high-quality imaging in complex environments such as strong occlusion, light reflection, and condensation interference significantly improves the accuracy of goods identification and system robustness, solving the problem of low accuracy in traditional identification systems under image quality fluctuations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120997816A_ABST
    Figure CN120997816A_ABST
Patent Text Reader

Abstract

The invention, which relates to the technical field of artificial intelligence and image identification, discloses an artificial intelligence-based refrigerator goods identification system comprising a refrigerator acquisition module, a goods image identification module, an identification processing module and an artificial intelligence module. A high-definition camera and an infrared thermal imaging device are combined to collect image data, and in combination with depth shielding recognition, image stability judgment and a dynamic shooting angle adjustment mechanism, it is ensured that high-quality images are obtained under the complex conditions of light reflection, shielding and frosting; the image recognition module fuses visible light and infrared image features, and extracts multi-dimensional information for goods positioning and classification; when the condition of low recognition confidence or serious image occlusion exists, an enhanced recognition mechanism is triggered, and a historical image is called for alignment and compensation recognition, so that the recognition accuracy and the system stability in a real retail environment are improved; the method has the advantages of being high in response intelligence degree, high in recognition efficiency and complex in adaptive scene, and is suitable for various refrigerator application scenes such as new retail and self-service vending.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of artificial intelligence and image recognition technology, specifically to an artificial intelligence-based refrigerated display case goods recognition system. Background Technology

[0002] With the development of new retail and unmanned retail, refrigerated display cases, as important terminal equipment for displaying and storing food and beverage products, are gradually being widely deployed in convenience stores, supermarkets, office buildings, transportation hubs, and other scenarios. The management and operation of traditional refrigerated display cases largely rely on manual inventory, regular maintenance, and self-service purchases by users, which suffers from low efficiency, low accuracy, and high operating costs. Especially in refrigerated display case scenarios with a wide variety of products, complex stacking, easy obstruction, and significant changes in ambient light, how to accurately and efficiently identify the product information inside the refrigerated display case has become a key challenge for the development of smart retail.

[0003] Currently, some refrigerated display cases have attempted to integrate RFID, barcode scanning, or gravity sensing for product identification. However, these methods generally suffer from limitations such as high cost, limited recognition accuracy, and complex maintenance, making them unsuitable for applications requiring rapid product turnover and seamless user interaction. In recent years, the development of artificial intelligence technology, especially computer vision and deep learning algorithms, has provided new solutions for image recognition of refrigerated display case products. However, most existing AI recognition systems rely on standardized shooting and ideal environments with minimal obstruction. When faced with the complex structure inside the refrigerated display case, mirror reflections, frost obstruction, and tilted / overlapping products, the recognition accuracy still falls short of commercial requirements. Summary of the Invention

[0004] To address the aforementioned issues, an artificial intelligence-based refrigerated display case goods recognition system is proposed.

[0005] The objective of this invention can be achieved through the following technical solution: a refrigerated display case goods recognition system based on artificial intelligence, comprising: a refrigerated display case acquisition module, an artificial intelligence module, and a goods image recognition module; The freezer's data acquisition module uses an artificial intelligence module to intelligently acquire image data from the freezer via a high-definition camera and an infrared camera mounted on the freezer. The product image recognition module performs deep identification of the product category from the image data, specifically: Perform joint encoding of multimodal images to obtain aligned RGB images and thermal imaging images; RGB images are input into a ResNet50+FPN structure to extract multi-scale semantic feature maps; thermal images are input into a U-Net++ network structure to output thermal feature maps; the output multi-scale semantic feature maps and thermal feature maps are fed into an encoder, and after fusion, a cross-modal feature representation fusion feature map is obtained, which serves as the basic feature map for product recognition. According to the fusion feature map obtained after fusion, the YOLOv8 target detection network is transmitted, and a candidate commodity region set is output; for each candidate commodity region set, the corresponding region feature tensor is cropped from the fusion feature map as a candidate region, and each candidate region is sent to a multi-channel fusion classifier for commodity category identification to output the final commodity category.

[0006] As a preferred embodiment of the present application, it further comprises an identification processing module; the identification processing module is used to execute an enhanced mechanism, which comprises: S1: based on the three-dimensional depth map, the image region of the same goods location in the historical image frame is searched; from the refrigerator acquisition module, according to the time of the current frame image and the goods location number, the image frame set of the same goods location in the past 5 minutes is queried from the historical image database; the previous frame with close time and good image clarity in the image frame set is selected as the candidate reference image; S2: the ECC image registration algorithm based on gray difference is performed on the current candidate region and the corresponding region in the reference image to align, and the transformation matrix is obtained; the affine transformation is performed on the corresponding region using the transformation matrix, so that it is strictly aligned with the current image candidate region, and the image block is output; S3: the occlusion difference degree analysis is performed to determine whether to use the reference image to complete the current identification process: The structural difference degree of the current image block and the registration image block is calculated, and an empirical threshold is set; if the structural difference degree is less than or equal to the empirical threshold, the existing occlusion or blur degree is lighter, and the reference image can be used as the auxiliary identification basis of the current region; S4: the reference image is used as the proxy input, and the category of the depth identification commodity is re-executed; S5: all identification tasks completed by the reference image completion are marked with a historical auxiliary identification label.

[0007] As a preferred embodiment of the present application, the specific process of intelligent acquisition is: When the refrigerator is in the identification state, the artificial intelligence module obtains the identification time point, issues an acquisition control instruction, and starts the synchronous acquisition of the high-definition camera and the infrared camera; the high-definition camera is responsible for acquiring the visible light image of the goods in the refrigerator; the infrared camera acquires the thermal imaging image to obtain the thermal feature distribution and occlusion information of the goods; if there is a depth occlusion, a specific time window and a shooting angle are set; after the image data is preliminarily processed by the image preprocessing unit, it is sent to the goods image recognition module for analysis.

[0008] As a preferred embodiment of the present application, the specific process of the artificial intelligence module obtaining the identification time point is: The identification state is continuously monitored by the artificial intelligence module to detect the running state and interaction events of the refrigerator, and when any trigger condition is detected, it is determined that the refrigerator enters the identification state: The triggering condition includes: the door body switch sensor detects that the refrigerator door body is switched from a closed state to an open state; the weight sensor detects that the weight of a certain product site inside the refrigerator changes, indicating that a product is taken out or put in; the infrared human sensor or the camera equipment detects that the user stays in front of the refrigerator for more than a preset time, triggering a timer or a remote control command; after any condition is met, the artificial intelligence module starts the recognition preparation process, controls the refrigerator acquisition module to enter the image acquisition preheating stage, and marks the current time as the time point to be recognized.

[0009] As a preferred embodiment of the present application, the specific process of depth occlusion is as follows: The depth occlusion is determined by multi-dimensional feature extraction and occlusion judgment algorithm on the visibility of the current product area. The specific determination includes: method one, method two and method three. Method one: the visual features of each occlusion area are extracted by a convolutional neural network, and compared with a preset complete product template. If there is a missing, it is preliminarily marked as an occlusion candidate area. Otherwise, based on the occlusion candidate area, a three-dimensional depth map of the inside of the refrigerator is constructed using binocular imaging, and the front and back occlusion relationship between products is identified by three-dimensional point cloud reconstruction and occlusion edge extraction algorithm. If there is a gradient mutation behind a certain product depth area or thermal imaging display, and there is no corresponding structure in the visible light image, it is judged as depth occlusion. Otherwise, method two is performed. Method two: the visible light image and the infrared image obtained under the same angle are spatially aligned, and mapped into a unified three-dimensional depth map. The heat source distribution area in the infrared image is extracted and matched with the identified product boundary area in the visible light image at the pixel level. If there is a clear heat feature in the infrared image, but there is a lack of edge structure, packaging color or no recognition result in the corresponding area in the visible light image, it is judged that the product is occluded by the front object or the refrigerator structure, resulting in a lack of visual information. Otherwise, if there is no clear heat feature in the infrared image, method three is performed. Method three: the image state of the refrigerator at the time of the last identification is completed is saved as a comparison reference image. After the current frame is collected, the historical frame image is analyzed by time difference, and the new or changed area is detected. If a new heat source feature is added in the thermal imaging, but the newly appeared product boundary in the visible light image is incomplete and the shape is blurred, it is further inferred that the area is in a depth occlusion state. When the judgment condition of any of the above methods is met, it is determined that there is depth occlusion in the current frame.

[0010] As a preferred embodiment of the present application, the specific process of setting a specific obtained time window and shooting angle is as follows: The position information of the depth occlusion corresponding area, the occlusion flag bit and the timestamp of the current image frame are obtained. The artificial intelligence module starts an image stability monitoring mechanism according to the timestamp, collects the next 3-5 frames of low frame rate image sequences, and analyzes the pixel change rate and image definition of the continuous frames; if the image pixel change rate is greater than 5%, it indicates that the refrigerator environment has not stabilized, and the time window is continued to be extended; otherwise, an image stability signal is output; When the stability signal is triggered, the current image is automatically executed to recognize the light feature to obtain a detection highlight area and a mirror reflection; then low-contrast feature analysis and judgment are performed to obtain frost or fogging phenomenon; if the image has a highlight area, the collection is suspended, and the time window is extended; when the judgment is passed, an image collectable signal is generated; if the image has a low-contrast area, a demisting mechanism is activated, and when the judgment is passed, an image collectable signal is generated; After receiving the collectable signal, a refrigerator three-dimensional structure model and a camera internal parameter matrix are called to map the occlusion area coordinates to the actual refrigerator space coordinates; and according to the region position and the visual angle accessibility, an optimal visual angle sequence is calculated; if the current angle is A, the angle adjustment task is handed over to the camera module controller to complete, and the current collection angle θ is recorded; after each visual angle adjustment, a new image is collected and analyzed to obtain the image definition, the occlusion area structure integrity, and the heat source form consistency; if any angle image meets the three conditions, the corresponding angle is determined as the optimal angle, and a high-quality image frame flag is sent out to trigger formal image collection; if all angles do not meet the three conditions, an enhanced mechanism is triggered.

[0011] As a preferred embodiment of the present application, the specific process of the multi-channel fusion classifier for identifying the product category and outputting the final product category is as follows: Step one: input each candidate region into a lightweight ViT unit to output a high semantic representation vector e i through the lightweight ViT unit; and simultaneously input a convolution channel to supplement the encoding of local texture details to obtain an auxiliary vector v i ; the final feature vector is z i =e i +v i +Stats(F i ), wherein Stats represents the region heat distribution and color histogram statistics; Step two: send the final feature vector z i into a fully connected classifier to obtain, through a softmax function, P , wherein W1 and b1 are the weights and bias of the first fully connected layer, PeLU(·) is an activation function, W2 and b2 are the weights and bias of the second fully connected layer, and softmax(·) is an output converted into a probability distribution P i ∈R C , each dimension represents the probability of belonging to a certain product; and the final output is Class.i =arg max(P i ); Category confidence level: Conf i =max(P i If the category confidence is less than 0.6 or the category difference is less than 0.1, the enhancement mechanism is triggered; otherwise, the final product category is output.

[0012] Compared with the prior art, the beneficial effects of the present invention are: 1. This invention utilizes a high-definition camera and an infrared thermal imager to collaboratively acquire images of goods, and combines a deep occlusion recognition algorithm with a dynamic shooting angle adjustment mechanism to achieve high-quality imaging of goods inside refrigerated cabinets under complex environments such as strong occlusion, light reflection, and condensation interference. This effectively improves the stability and clarity of the image input stage and solves the problem of low accuracy in traditional recognition systems when image quality fluctuates.

[0013] 2. This invention further introduces a multimodal image fusion recognition mechanism, which deeply fuses visible light and thermal imaging features and realizes multidimensional feature modeling through encoder and ViT fusion structure; when the recognition confidence is insufficient or the occlusion is severe, the image enhancement process is automatically triggered, and historical image frames are used for reference recognition compensation, which significantly improves the accuracy of goods recognition and system robustness in complex situations, which is something that traditional single image analysis methods cannot achieve. Attached Figure Description

[0014] To facilitate understanding by those skilled in the art, the present invention will be further described below with reference to the accompanying drawings.

[0015] Figure 1 This is a schematic diagram of the principle of the present invention. Detailed Implementation

[0016] The technical solution of the present invention will be clearly and completely described below with reference to the embodiments. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.

[0017] It should be understood that the terms “comprising” and “including” used in this disclosure and claims indicate the presence of the described features, integrals, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.

[0018] It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the disclosure. As used in this disclosure and the claims, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms "and / or," as used herein, refers to and encompasses any and all possible combinations of one or more of the associated listed items. It is also to be understood that the use of "about" to modify a particular recited parameter or claim element, in and of itself, does not constitute a disclaimer or admission that the modified recited parameter or claim element is not exact.

[0019] Referring to Figure 1 As shown in the figure, an artificial intelligence-based refrigerator goods identification system comprises a refrigerator acquisition module, a goods image identification module, an identification processing module, and an artificial intelligence module. The refrigerator acquisition module is controlled by the artificial intelligence module to perform intelligent acquisition. The intelligent acquisition process collects image data of the refrigerator through a high-definition camera and an infrared camera mounted on the refrigerator. In one specific example, when the refrigerator is in a to-be-identified state, the artificial intelligence module obtains a to-be-identified time point, issues an acquisition control instruction, and starts the high-definition camera and the infrared camera to perform synchronous acquisition. The high-definition camera is responsible for acquiring a visible light image of goods inside the refrigerator, which is used to extract appearance information such as color, shape, and packaging. The infrared camera acquires a thermal imaging image to obtain thermal feature distribution and shielding information of the goods. If there is deep shielding, a specific obtained time window and a shooting angle are set to avoid the influence of light reflection, frost interference, and the like, so as to ensure that the image is clear and complete. The obtained image data is sent to the goods image identification module for subsequent analysis after being preliminarily processed by an image preprocessing unit. It should be noted that the preliminary processing includes denoising, correction, image enhancement, and the like.

[0020] The to-be-identified state is that the artificial intelligence module continuously monitors the running state and interaction events of the refrigerator. When any trigger condition is detected, it is determined that the refrigerator enters the to-be-identified state. The trigger conditions include that a door body switch sensor detects that the refrigerator door body changes from a closed state to an open state; a weight sensor detects that the weight of a certain goods site inside the refrigerator changes, indicating that goods are taken out or put in; an infrared human sensor or a camera device detects that a user stays in front of the refrigerator for more than a preset time length; the system is triggered by timing or a remote control command is issued; After any condition is established, the artificial intelligence module starts an identification preparation process, controls the refrigerator acquisition module to enter an image acquisition preheating stage, and marks the current time as the to-be-identified time point. Further, deep shielding is determined by a multi-dimensional feature extraction and shielding judgment algorithm on the visibility of the current goods region, specifically including method one, method two, and method three. Method one: the visual features of each occlusion region are extracted by convolutional neural network, and compared with the preset complete product template; if there is a missing, it is preliminarily marked as an occlusion candidate region; otherwise, based on the occlusion candidate region, the three-dimensional depth map of the refrigerator interior is constructed by binocular imaging, and the front and rear occlusion relationship between the goods is identified by three-dimensional point cloud reconstruction and occlusion edge extraction algorithm; if there is a gradient mutation behind the depth region of a certain product or thermal imaging display, and there is no corresponding structure in the visible light image, it is judged as depth occlusion; otherwise, method two is performed; It should be noted that the visual features of the occlusion region include edge contour, color texture, packaging identification, etc.

[0021] Method two: the visible light image and the infrared image obtained under the same angle are spatially aligned and processed, and the two are mapped into a unified three-dimensional depth map; on this basis, the heat source distribution region in the infrared image is extracted and pixel-level matching is performed with the identified product boundary region in the visible light image; if there is an explicit heat energy feature in the infrared image, such as a product just put in by hand, a beverage at room temperature, etc., but the corresponding region in the visible light image lacks edge structure, packaging color or the recognition result is empty, it is judged that the product is occluded by the front object or the refrigerator structure, resulting in the lack of visual information; otherwise, if there is no explicit heat energy feature in the infrared image, method three is performed; Method three: the image state of the refrigeration refrigerator at the time of the last identification is completed is saved as a comparison reference image; after the current frame is collected, time sequence difference analysis is performed with the historical frame image to detect the added or changed region; if a new heat source feature is added in the thermal imaging in a certain region, but the newly appeared product boundary in the visible light image is incomplete and the shape is blurred, it is further inferred that the region is in a depth occlusion state; When the judgment condition of any of the above methods is met, it is determined that there is a depth occlusion in the current frame; Further, a specific obtained time window and shooting angle are set, specifically: The position information (x, y) of the depth occlusion corresponding region, the occlusion flag and the timestamp T0 of the current image frame are obtained; The artificial intelligence module starts the image stability monitoring mechanism according to the timestamp T0, collects the next 3-5 frames of low frame rate image sequence, and analyzes the pixel change rate ΔP and the image definition of the continuous frames; if the image pixel change rate ΔP is greater than 5%, it indicates that the refrigerator environment has not stabilized, for example, the user is still operating, the goods have not yet stabilized, then the time window is continued to be extended; otherwise, an image stability signal is output; When the stable signal is triggered, the current image is automatically subjected to light feature recognition to obtain a detection highlight area and a mirror reflection; low-contrast feature analysis is then performed to determine whether there is frost or fogging; if the image contains a highlight area, image acquisition is suspended and the time window is extended; when the determination is passed, an image acquisition signal is generated; if the image contains a low-contrast area, a demisting mechanism is activated (the demisting device is started), and when the determination is passed, an image acquisition signal is generated; Upon receiving the acquisition signal, the refrigerator three-dimensional structure model and the camera internal parameter matrix are called to map the occlusion region coordinates (x, y) to the actual refrigerator space coordinates; and based on the region position and the reachability of the viewing angle, the optimal viewing angle sequence is calculated, for example, angle A→ angle B→ angle C; if the current angle is A, the angle adjustment task is handed over to the camera module controller, for example, the pan-tilt rotates ±10°, and the current acquisition angle θ is recorded; after each viewing angle adjustment, a new image is acquired and analyzed to obtain the image clarity, the occlusion region structural integrity (contour closure, LOGO restoration, etc.), and the heat source form consistency; if any angle image meets the three conditions, the corresponding angle is determined as the optimal angle, and a high-quality image frame flag is issued to trigger formal image acquisition; if all angles do not meet the three conditions, an enhanced mechanism is triggered. It should be noted that the three conditions are: image clarity, occlusion region structural integrity, and heat source form consistency, all of which are greater than the respective preset thresholds.

[0022] The goods image recognition module performs deep recognition of the category of the goods based on the image data, specifically: Joint encoding of multi-modal images is performed to obtain an aligned visible light image I vis (i.e., an RGB image I vis ) and an infrared thermal image I ir (i.e., a thermal imaging image I ir ), which respectively contain the appearance and thermal feature information of the goods; The RGB image I vis is input into a ResNet50+FPN structure to extract multi-scale semantic features F vis ∈R 256×H×W ; the thermal imaging image I ir is input into a U-Net++ network structure to output a thermal feature map F ir ∈R 128×H×W ; The output multi-scale semantic features F vis and the thermal feature map F ir are sent to an encoder to obtain a cross-modal feature representation fusion feature map F cm ∈R 512×H×W after fusion; The fusion feature map F cm is transmitted to a YOLOv8 target detection network to output a candidate goods region set: wherein each candidate box contains a location and an initial confidence value s i ; the initial confidence value s i is a quantitative value output by a candidate region detection network (e.g., YOLOv8, CenterNet, or Anchor-Free type detector) for measuring the "target sensitivity" of whether a candidate region contains a target product object; in the detection network, for each candidate region, its confidence value s i comes from the comprehensive results of three parts: s i = σ (S (x i , y i )) · IOU pred_anchor · A ctscore , wherein: σ (S (x i , y i )) : the "target existence value" after sigmoid activation; IOU pred_anchor : the IOU of the current candidate box and the best anchor box (used to assist in judging the matching quality); A ctscore : the activation degree of the activation channel (such as heat texture, contour response, etc.), which supplements the spatial response information).

[0023] The encoder models the global dependency within the same modality through a self-attention mechanism, and achieves information alignment and feature fusion between modalities through a cross-attention mechanism. In the fusion process, Query, Key, and Value come from different modal features, so that the model can selectively focus on the corresponding regions or semantics across modalities. In this application, it is specially used for jointly modeling and deeply fusing the input multi-scale semantic features Fvis and the infrared thermal feature map Fir, achieving semantic alignment and complementary enhancement of the visual and thermal modalities. After processing by the encoder, a unified cross-modal feature representation is output.

[0024] For each candidate product region set B, the corresponding region feature tensor is cropped from F cm , as the candidate region F i Each candidate region F i is sent to a multi-channel fusion classifier (Multi-Modal Classifier) for product category recognition, including: Step one: input each F i into a lightweight ViT unit (Vision Transformer), for example, the patch size is 16x16, the number of layers is 4, and the dimension is 256; output a high semantic representation vector e i ∈R 256; and meanwhile input a convolution path (such as MobileNet) to supplement the encoding of local texture details, to obtain an auxiliary vector v i ; the final feature is: z i = e i + v i + Stats(F i ), wherein Stats represents a statistical quantity such as a regional heat distribution or a color histogram; Step two: send z i to a fully connected classifier (2-layer MLP, the output dimension is equal to the number of product categories C), and obtain through softmax: , wherein W1, b1: the weight and bias of the first fully connected layer (used for feature transformation), PeLU(·) is an activation function, W2, b2: the weight and bias of the second fully connected layer, which maps the dimension to the number of product categories C, and softmax(·) is an output transformation into a probability distribution P i ∈ R C , each dimension represents the probability of belonging to a certain product category; the final output is: Class i = arg max(P i ); the category confidence is: Conf i = max(P i ); If the category confidence Conf i < 0.6 or the category Class i differs by less than 0.1 (i.e. the classification result is not obvious), the enhancement mechanism is triggered; otherwise, the final product category is output. The recognition processing module is used to execute the enhancement mechanism, including: S1: based on the three-dimensional depth map, search for the image area of the same location in the historical image frame; from the image acquisition management module, according to the timestamp t0 of the current frame image and the location number p i , query the image frame set of the same location in the past 5 minutes from the historical image database , wherein j = 1, 2, …, k; select the previous frame I ref = as the candidate reference image with close time and good image clarity in the image frame set . It should be noted that the image clarity can be judged by local variance, edge response and other indicators (such as Laplace clarity score method), and blurred image frames are excluded.

[0025] S2: compare the current candidate region F i ⊂ I i (current frame I i ) with the corresponding region R ref ⊂ Iref , execute the ECC image registration algorithm based on gray difference to align and get the transformation matrix T align , use the transformation matrix to perform affine transformation on R ref , so that it is strictly aligned with the current image candidate region, and output the image block .

[0026] S3: Perform occlusion difference analysis to determine whether the reference image can be used to complete the current recognition process: Calculate the structural difference Δs of the current image block and the registered image block by: ; set an empirical threshold θ occlusion (the threshold is statically set, and can also be fine-tuned according to the variability of different product categories), if Δs≤θ occlusion , the existing occlusion or blur is relatively light, and the reference image can be used as auxiliary identification basis for the current region.

[0027] S4: Use the reference image as a proxy input to re-execute the depth recognition of the product category; S5: All recognition tasks completed by the reference image are labeled as "historical auxiliary identification".

[0028] The preferred embodiments of the application disclosed above are only used to help explain the application. The preferred embodiments do not describe all the details and do not limit the application to the specific embodiments. Obviously, according to the content of the specification, many modifications and changes can be made. The specification selects and describes these embodiments in order to better explain the principles and practical applications of the application, so that those skilled in the art can well understand and utilize the application. The application is limited by the claims and their entire scope and equivalents.

Claims

1. An artificial intelligence-based refrigerated display case goods recognition system, comprising: The freezer data acquisition module, the artificial intelligence module, and the product image recognition module are characterized by: The freezer's data acquisition module uses an artificial intelligence module to intelligently acquire image data from the freezer via a high-definition camera and an infrared camera mounted on the freezer. The product image recognition module performs deep identification of the product category from the image data, specifically: Perform joint encoding of multimodal images to obtain aligned RGB images and thermal imaging images; RGB images are input into a ResNet50+FPN structure to extract multi-scale semantic feature maps; thermal images are input into a U-Net++ network structure to output thermal feature maps; the output multi-scale semantic feature maps and thermal feature maps are fed into an encoder, and after fusion, a cross-modal feature representation fusion feature map is obtained, which serves as the basic feature map for product recognition. The fused feature map is fed into the YOLOv8 object detection network to output a set of candidate product regions. For each set of candidate product regions, the corresponding region feature tensor is cropped from the fused feature map as a candidate region. Each candidate region is then fed into a multi-channel fusion classifier to identify the product category and output the final product category.

2. The refrigerated display case goods identification system based on artificial intelligence according to claim 1, characterized in that, The specific process of intelligent data collection is as follows: When the freezer is in the state of waiting to be identified, the artificial intelligence module obtains the time point to be identified, issues a collection control command, and starts the high-definition camera and infrared camera to collect data synchronously; the high-definition camera is responsible for acquiring visible light images of the goods inside the freezer; the infrared camera acquires thermal imaging images to obtain thermal feature distribution and occlusion information of the goods; If there is deep occlusion, a specific time window and shooting angle are set; then, the image data is pre-processed by the image preprocessing unit before being sent to the product image recognition module for analysis.

3. The refrigerated display case goods recognition system based on artificial intelligence according to claim 2, characterized in that, The specific process by which the artificial intelligence module obtains the time point to be identified is as follows: The waiting-to-be-identified state is determined by the continuous monitoring of the freezer's operating status and interaction events by the artificial intelligence module. When any trigger condition is detected, the freezer is determined to have entered the waiting-to-be-identified state. Triggering conditions include: the door switch sensor detects that the freezer door has changed from closed to open; the weight sensor detects a change in the weight of a certain item inside the freezer, indicating that an item has been taken out or put in; the infrared human body sensor or camera detects that a user has stayed in front of the freezer for more than a preset time, triggering a timed event or issuing a remote control command; once any condition is met, the artificial intelligence module starts the recognition preparation process, controls the freezer acquisition module to enter the image acquisition preheating stage, and marks the current time as the time point to be recognized.

4. The refrigerated display case goods recognition system based on artificial intelligence according to claim 3, characterized in that, The specific process of depth occlusion is as follows: Deep occlusion uses multi-dimensional feature extraction and occlusion judgment algorithms to determine the visibility of the current product area. The specific judgment methods include: Method 1, Method 2 and Method 3. Method 1: Extract visual features of each occluded region using a convolutional neural network and compare them with a pre-set complete product template. If any features are missing, they are initially marked as occlusion candidate regions. Otherwise, based on the occlusion candidate regions, construct a 3D depth map of the freezer's interior using binocular imaging. Identify the front-to-back occlusion relationship between products using 3D point cloud reconstruction and occlusion edge extraction algorithms. If a gradient abrupt change or thermal imaging shows behind a product's depth region, and there is no corresponding structure in the visible light image, it is determined to be a depth occlusion. Otherwise, proceed to Method 2. Method 2: Spatial alignment processing is performed on the visible light image and infrared image acquired from the same viewpoint, mapping them to a unified 3D depth map; and the heat source distribution area in the infrared image is extracted and matched pixel-level with the identified product boundary area in the visible light image; if clear thermal energy features appear in the infrared image, but the corresponding area in the visible light image lacks edge structure, packaging color, or the recognition result is empty, it is determined that the product is obstructed by the object in front or the freezer structure, resulting in a lack of visual information; otherwise, if no clear thermal energy features appear in the infrared image, Method 3 is used; Method 3: Save the image state of the freezer when it was last identified as a comparison reference; after the current frame is acquired, perform temporal difference analysis with the historical frame images to detect new or changed areas; if a certain area has new heat source features in the thermal imaging, but the boundaries of the newly appearing goods in the visible light image are incomplete and the shape is blurred, then it is further inferred that the area is in a state of deep occlusion. If any of the above methods are satisfied, then the current frame is determined to have depth occlusion.

5. The refrigerated display case goods recognition system based on artificial intelligence according to claim 4, characterized in that, The specific process of setting a specific time window and shooting angle is as follows: Obtain the location information of the region corresponding to the deep occlusion, the occlusion flag, and the timestamp of the current image frame; The artificial intelligence module activates the image stability monitoring mechanism based on the timestamp, and collects the next 3 to 5 frames of low frame rate image sequence. It performs pixel change rate analysis and image sharpness analysis on the consecutive frames. If the image pixel change rate is >5%, it means that the freezer environment has not yet stabilized, so the time window is extended. Otherwise, an image stabilization signal is output. Once a stable signal is triggered, the system automatically performs illumination feature recognition on the current image to detect bright areas and specular reflections; then it performs low-contrast feature analysis to determine if frost or fogging occurs; if bright areas exist in the image, the acquisition is paused and the time window is extended; once the judgment is passed, an image acquisition signal is generated; if low-contrast areas exist in the image, the defogging mechanism is activated, and once the judgment is passed, an image acquisition signal is generated. After receiving a collectable signal, the system calls the 3D structural model of the freezer and the intrinsic parameter matrix of the camera to map the coordinates of the obstructed area to the actual spatial coordinates of the freezer; and calculates the optimal viewing angle sequence based on the location of the area and the accessibility of the viewing angle; if the current angle is A, the angle adjustment task is handed over to the camera module controller and the current acquisition angle θ is recorded. After each viewpoint adjustment, a new image is acquired and analyzed to obtain image sharpness, structural integrity of the occluded area, and consistency of heat source morphology; If any angle of the image meets all three conditions, the corresponding angle is determined to be the optimal angle, and a high-quality image frame flag is issued, triggering formal image acquisition; if none of the angles meet all three conditions, the enhancement mechanism is triggered.

6. The refrigerated display case goods identification system based on artificial intelligence according to claim 1, characterized in that, The specific process of a multi-channel fusion classifier identifying and outputting the final product category is as follows: Step 1: Input each candidate region into a lightweight ViT unit, and output a high semantic representation vector e through the lightweight ViT unit. i Simultaneously, a convolutional path is input to supplement the encoding of local texture details, resulting in an auxiliary vector v. i The final eigenvector is then z. i =e i +v i +Stats(F i ), where Stats represents the statistics of the regional heat distribution and color histogram; Step 2: Convert the final feature vector z i The data is fed into a fully connected classifier and processed by the softmax function to obtain: Where W1, b1: weights and biases of the first fully connected layer, PeLU(·) is the activation function, W2, b2: weights and biases of the second fully connected layer, and softmax(·) is the output transformed into a probability distribution P. i ∈R C Each dimension represents the probability of belonging to a certain product category; the final output is: Class i =arg max(P i ); Category confidence level: Conf i =max(P i If the category confidence is less than 0.6 or the category difference is less than 0.1, the enhancement mechanism is triggered; otherwise, the final product category is output.