Interaction method and device of refrigeration equipment, refrigeration equipment and interaction system of refrigeration equipment

By using multimodal fusion of video big data models and voice interaction data in smart refrigerators, the problems of insufficient recognition accuracy and interaction precision in smart refrigerators have been solved, achieving more efficient food recognition and user intent inference, and improving the level of intelligence.

CN121804159APending Publication Date: 2026-04-07QINDAO HAIER REFRIGERATOR CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610003636.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-04
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

The existing video interaction methods of smart refrigerators are insufficient in terms of recognition accuracy and interaction precision, making it difficult to meet users' demands for improved intelligence.

Method used

By acquiring video interaction data between users and refrigeration equipment, the target food is tracked and detected using a large video model. The image feature information is extracted by combining the relative positional relationship between the food and the storage space, and multimodal information fusion is performed by combining voice interaction data to output the target interaction information.

Benefits of technology

It improves the accuracy of food identification and the precision of interaction, enabling more accurate inference of user intent and providing efficient intelligent interactive responses.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121804159A_ABST
    Figure CN121804159A_ABST
Patent Text Reader

Abstract

The invention discloses an interaction method and device of refrigeration equipment, the refrigeration equipment and an interaction system of the refrigeration equipment, and belongs to the technical field of refrigeration equipment. The interaction method of the refrigeration equipment comprises the steps of obtaining interaction data of a user and the refrigeration equipment, wherein the interaction data at least comprises video interaction data; inputting the video interaction data into a video large model, tracking and detecting a target food material in the video interaction data through the video large model, and outputting image feature information of the target food material based on a relative position relationship between the target food material and a storage space of the refrigeration equipment; and outputting target interaction information based on the image feature information. The interaction recognition accuracy and the interaction precision can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of refrigeration equipment technology, and particularly relates to an interaction method, device, refrigeration equipment and its interaction system for refrigeration equipment. Background Technology

[0002] With the continuous development of artificial intelligence technology, intelligent interaction technology has been widely applied in smart refrigerators, and users' expectations for the level of intelligence in smart refrigerators continue to rise. Improving the interaction efficiency between users and smart refrigerators has become key to promoting their intelligent upgrade. However, current interaction methods based on video and other formats need improvement in terms of recognition accuracy and interaction precision. Summary of the Invention

[0003] This invention aims to at least solve one of the technical problems existing in the prior art. To this end, this invention proposes an interaction method, apparatus, refrigeration equipment, and interaction system for refrigeration equipment, which can improve the recognition accuracy and interaction precision of the interaction.

[0004] In a first aspect, this application provides an interaction method for a refrigeration device, the method comprising: Acquire user interaction data with the cooling device, the interaction data including at least video interaction data; The video interaction data is input into the video big model, and the target food in the video interaction data is tracked and detected by the video big model. Based on the relative positional relationship between the target food and the storage space of the refrigeration equipment, the image feature information of the target food is output. Based on the image feature information, target interaction information is output.

[0005] According to the interaction method of the refrigeration equipment in this application, by acquiring video interaction data between the user and the refrigeration equipment and inputting it into a large video model, the target food is accurately tracked and detected using the large video model. By combining the relative positional relationship between the target food and the storage space of the refrigeration equipment to extract image feature information, clear spatial context constraints can be provided for the image feature information, thereby improving the accuracy of food identification and more accurately inferring the user's intention. Finally, the target interaction information is output, which can improve the recognition accuracy and interaction precision of the interaction.

[0006] According to one embodiment of this application, the video big model includes a target tracking model and a target detection model. The step of inputting the video interaction data into the video big model, tracking and detecting target food items in the video interaction data through the video big model, and outputting image feature information of the target food items based on the relative positional relationship between the target food items and the storage space of the refrigeration equipment includes: The video interaction data is input into the target tracking model. When the target food is detected to be outside the storage space, the target tracking model continuously tracks the target food, or when the target food is detected to be inside the storage space, it outputs the location information of the target food. When the target tracking model detects that the target food is located in the storage space, the video interaction data is input to the target detection model, which detects the type and quantity of the target food and outputs the category count information of the target food. The image feature information is determined based on the location information and the category count information.

[0007] According to one embodiment of this application, the step of outputting target interaction information based on the image feature information includes: The image feature information is input into the target classification model, and the target food is classified by the target classification model to obtain target classification information. A deep classification network is set before the neck network of the target classification model. Based on the target classification information, the target interaction information is output.

[0008] According to one embodiment of this application, the interaction data further includes voice interaction data, and before outputting the target interaction information based on the image feature information, the method further includes: The voice interaction data is converted into text format to obtain text interaction data; The text interaction data is input into a language model, and the semantic features of the text interaction data are analyzed by the language model to output semantic information. The step of outputting target interaction information based on the image feature information includes: Based on the image feature information and the semantic information, the target interaction information is output.

[0009] According to one embodiment of this application, the step of outputting the target interaction information based on the image feature information and the semantic information includes: The image feature information is input into the target verification model, and the target verification model outputs the false negative and false positive information of the target food ingredient. The missed detection and false detection information and the semantic information are input into the target correction model. The missed detection and false detection of the target food ingredient are corrected by the target correction model to obtain the corrected image feature information. Based on the corrected image feature information, the target interaction information is output.

[0010] According to one embodiment of this application, the video big data model includes a first network and a second network, wherein the sampling frame rate of the first network in the time dimension is lower than that of the second network, and the step of inputting the video interaction data into the video big data model and tracking and detecting the target food ingredients in the video interaction data through the video big data model includes: The first network detects the target food ingredients in a single frame of the video interaction data, and the second network tracks and detects the target food ingredients in a continuous video of the video interaction data.

[0011] According to one embodiment of this application, the training process of the video large model uses the external cache of the storage device of the cooling device, and the prediction process of the video large model uses the internal cache of the storage device.

[0012] Secondly, this application provides an interaction device for a refrigeration device, the device comprising: An acquisition module is used to acquire interaction data between the user and the cooling device, the interaction data including at least video interaction data; The first processing module is used to input the video interaction data into the video big model, track and detect the target food in the video interaction data through the video big model, and output the image feature information of the target food based on the relative positional relationship between the target food and the storage space of the refrigeration equipment. The second processing module is used to output target interaction information based on the image feature information.

[0013] According to the interactive device of the refrigeration equipment of this application, by acquiring video interaction data between the user and the refrigeration equipment and inputting it into a large video model, the target food is accurately tracked and detected using the large video model. By combining the relative positional relationship between the target food and the storage space of the refrigeration equipment to extract image feature information, clear spatial context constraints can be provided for the image feature information, thereby improving the accuracy of food identification and more accurately inferring the user's intention. Finally, the target interaction information is output, which can improve the recognition accuracy and interaction precision of the interaction.

[0014] Thirdly, this application provides a refrigeration device, comprising: The interactive device of the refrigeration equipment as described in the second aspect above; A data acquisition device is connected to the interaction device. The data acquisition device is used to collect interaction data between the user and the cooling equipment. The interaction data includes at least video interaction data.

[0015] According to the refrigeration device of this application, by acquiring video interaction data between the user and the refrigeration device and inputting it into a large video model, the target food is accurately tracked and detected using the large video model. By combining the relative positional relationship between the target food and the storage space of the refrigeration device to extract image feature information, clear spatial context constraints can be provided for the image feature information, thereby improving the accuracy of food identification and more accurately inferring the user's intention. Finally, the target interaction information is output, which can improve the recognition accuracy and precision of the interaction.

[0016] Fourthly, this application provides an interactive system for a refrigeration device, comprising: The refrigeration equipment as described in the third aspect above; The server is connected to the refrigeration equipment and is used to output target interaction information.

[0017] According to the interactive system of the refrigeration equipment of this application, by acquiring video interaction data between the user and the refrigeration equipment and inputting it into a large video model, the system accurately tracks and detects the target food ingredients using the large video model. By combining the relative positional relationship between the target food ingredients and the storage space of the refrigeration equipment to extract image feature information, it can provide clear spatial context constraints for the image feature information, thereby improving the accuracy of food ingredient recognition and more accurately inferring the user's intention. Finally, it outputs the target interaction information, which can improve the recognition accuracy and interaction precision of the interaction.

[0018] Fifthly, this application provides an electronic device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the interaction method of the cooling device as described in the first aspect above.

[0019] In a sixth aspect, this application provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the interaction method of the refrigeration device as described in the first aspect above.

[0020] In a seventh aspect, this application provides a computer program product, including a computer program that, when executed by a processor, implements the interaction method of the refrigeration device as described in the first aspect above.

[0021] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description

[0022] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the description of the embodiments taken in conjunction with the following drawings, in which: Figure 1This is a flowchart illustrating the interaction method of the refrigeration equipment provided in the embodiments of this application; Figure 2 This is a schematic diagram of the structure of the interactive system of the refrigeration equipment provided in the embodiments of this application; Figure 3 This is a schematic diagram of the structure of the interactive device of the refrigeration equipment provided in the embodiments of this application; Figure 4 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation

[0023] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.

[0024] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.

[0025] The following description, in conjunction with the accompanying drawings, details the interaction method, interaction device, electronic device, and readable storage medium of the refrigeration equipment provided in this application, through specific embodiments and application scenarios.

[0026] The interaction method of the refrigeration equipment can be applied to the terminal, and can be executed by the hardware or software in the terminal.

[0027] The interaction method of the refrigeration device provided in this application embodiment can be executed by an electronic device or a functional module or functional entity in an electronic device that can implement the interaction method of the refrigeration device. The electronic devices mentioned in this application embodiment include, but are not limited to, mobile phones, tablets, computers, cameras and wearable devices. The interaction method of the refrigeration device provided in this application embodiment will be described below using an electronic device as the execution subject.

[0028] like Figure 1 As shown, the interaction method of the refrigeration device includes steps 110, 120 and 130.

[0029] Step 110: Obtain the interaction data between the user and the cooling equipment. The interaction data shall include at least video interaction data.

[0030] Interactive data refers to various behavioral information data generated by users during their interaction with the cooling equipment. Video interactive data refers to interactive data in the form of video. Interactive data may also include voice, images, touch, and clicks, which are used to reflect the user's intentions, behaviors, or states.

[0031] In this step, video acquisition devices can be set up in the cooling equipment to collect video interaction data.

[0032] Step 120: Input the video interaction data into the video big model, track and detect the target food in the video interaction data through the video big model, and output the image feature information of the target food based on the relative positional relationship between the target food and the storage space of the refrigeration equipment.

[0033] Among them, the video big model is a deep learning model trained on large-scale data, which can understand the temporal and spatial information in the video and is used to identify, track and analyze dynamic visual content.

[0034] The target ingredient is a specific food item that the user is currently paying attention to or interacting with, such as milk or apples, and is the core object detected by the system.

[0035] Relative positional relationship is used to describe the spatial correspondence between the target food and the equipment structure, for example, located outside the storage space, located on the left side of the second floor of the cold storage room.

[0036] Image feature information is information extracted from video interaction data that characterizes the visual attributes of the target food, such as shape, color, texture, and semantic embedding vectors, and is used for subsequent recognition and reasoning.

[0037] In this embodiment, a large video model is used to analyze the video interaction data frame by frame, identify and continuously track the target food that the user is interested in or operates on, and combine the structural information of the storage space inside the refrigeration equipment to determine the specific location of the target food and obtain the relative positional relationship.

[0038] Different detection and tracking operations can be performed when the target food is located inside or outside the storage space. While identifying the target food, the video big model can combine the structural priors of the storage space inside the refrigeration equipment, such as area division, to map the spatial coordinates of the target food in the video frame to the specific location of the storage space, and extract image feature information such as the appearance, category, quantity and status of the target food in this context.

[0039] Step 130: Output target interaction information based on image feature information.

[0040] Among them, the target interaction information is semantic feedback used to respond to user intent, such as expiration reminders, recipe recommendations, or restocking prompts.

[0041] In this embodiment, based on image feature information, and by integrating context such as user history, current time, and recipe preferences, the system accurately understands the user's interaction intent with the target ingredient, thereby generating targeted target interaction information.

[0042] This may trigger intelligent responses, such as voice prompts like, "The milk you purchased last time will expire in 2 days"; app push notifications like, "Tomatoes are running low in stock, would you like to add them to your shopping list?"; and automatic recording of food consumption logs.

[0043] According to the interaction method of the refrigeration equipment provided in the embodiments of this application, by acquiring video interaction data between the user and the refrigeration equipment and inputting it into a large video model, the target food is accurately tracked and detected using the large video model. By combining the relative positional relationship between the target food and the storage space of the refrigeration equipment to extract image feature information, clear spatial context constraints can be provided for the image feature information, thereby improving the accuracy of food identification and more accurately inferring the user's intention. Finally, the target interaction information is output, which can improve the recognition accuracy and interaction precision of the interaction.

[0044] In some embodiments, the video big model includes a target tracking model and a target detection model. Video interaction data is input into the video big model, which tracks and detects target food items in the video interaction data. Based on the relative positional relationship between the target food items and the storage space of the refrigeration equipment, the model outputs image feature information of the target food items, including: The video interaction data is input into the target tracking model. The target tracking model continuously tracks the target food when it is detected that the target food is outside the storage space, or outputs the location information of the target food when it is detected that the target food is inside the storage space. When the target tracking model detects that the target food is located in the storage space, the video interaction data is input into the target detection model. The target detection model detects the type and quantity of the target food and outputs the category count information of the target food. Image feature information is determined based on location information and category counting information.

[0045] The location information refers to the specific coordinates or area identification information of the target food in the storage space, such as the left side of the second floor of the refrigerator.

[0046] Category counting information refers to the type of target ingredient and its corresponding quantity, such as 2 boxes of milk and 3 apples.

[0047] In this embodiment, the video big data model includes a target tracking model and a target detection model. First, the video interaction data is input into the target tracking model. If the target food is detected to be outside the storage space of the refrigeration equipment, its movement trajectory is continuously tracked. If it is determined to be inside the storage space, its specific location information is output. Subsequently, when the target food is inside the storage space, the corresponding video frame is input into the target detection model to identify the type of target food and count its quantity, generating category count information. Finally, the location information and category count information are fused to construct image feature information containing spatial semantics and content attributes.

[0048] In some embodiments, target interaction information is output based on image feature information, including: Image feature information is input into the target classification model, and the target food is classified by the target classification model to obtain target classification information. A deep classification network is set before the neck network of the target classification model. Based on the target classification information, output the target interaction information.

[0049] Among them, the target classification information is the semantic label output after the target food is identified by the target classification model, including the specific category, attributes or status of the food, such as fresh milk, expired eggs or organic apples.

[0050] In this embodiment, the image feature information is input into the target classification model. The target classification model introduces a deep classification network in front of the neck network to enhance the fine-grained food identification capability, thereby accurately classifying the target food and generating target classification information.

[0051] Based on target classification information, such as food category, status, or shelf-life related attributes, corresponding semantic target interaction information is generated, such as "Eggs are about to expire, we recommend using them first" or "Yogurt stock is low, do you want to add it to your shopping list?", to achieve high-precision, context-aware intelligent interactive response.

[0052] In some embodiments, the interaction data further includes voice interaction data, and the method further includes the following steps before outputting the target interaction information based on image feature information: The voice interaction data is converted into text format to obtain text interaction data; The text interaction data is input into the language model, and the semantic features of the text interaction data are analyzed by the language model to output semantic information. Based on image feature information, the target interaction information is output, including: Based on image feature information and semantic information, the target interaction information is output.

[0053] Semantic information is a structured or abstract linguistic representation that is parsed from text interaction data and reflects the user's intent, needs, or contextual meaning, such as a user wanting to find milk or asking if ingredients are expired.

[0054] In this embodiment, the interaction data also includes voice interaction data. The system first converts the voice interaction data into text to obtain text interaction data, and then inputs the text interaction data into the language big model to parse semantic features and output semantic information.

[0055] Image feature information can reflect the visual and spatial attributes of the target food, while semantic information can reflect the user's intention or needs expressed in language. By fusing image feature information and semantic information, and using multimodal alignment or joint reasoning mechanisms, target interaction information that accurately matches the user's current interaction scenario can be generated.

[0056] In some embodiments, based on image feature information and semantic information, target interaction information is output, including: Image feature information is input into the target verification model, and the target verification model outputs the false negative and false positive information of the target food. The missed detection and false detection information and semantic information are input into the target correction model. The target correction model corrects the missed detection and false detection of the target food and obtains the corrected image feature information. Based on the corrected image feature information, the target interaction information is output.

[0057] Among them, the information on missed detections and false detections refers to the specific content that was missed (missed detection) or incorrectly identified (false detection) during the food detection process, as well as its confidence level, which is used to indicate the deviation between the identification results and the actual scenario.

[0058] In this embodiment, image feature information is input into the target verification model. The target verification model can identify missed or false detections by comparing the image feature information with preset food distribution patterns, storage space logic, or historical records, and output the corresponding missed or false detection information.

[0059] The target correction model combines missed and false detection information with semantic information, and uses contextual reasoning and multimodal alignment mechanisms to supplement, delete or adjust the original image features, thereby generating more accurate and consistent corrected image feature information, and outputting high-precision target interaction information accordingly.

[0060] In some embodiments, the video big model includes a first network and a second network, wherein the first network has a lower sampling frame rate in the time dimension than the second network. Video interaction data is input into the video big model, and the video big model tracks and detects target ingredients in the video interaction data, including: The first network detects target ingredients in single-frame images of the video interaction data, and the second network tracks and detects target ingredients in continuous videos of the video interaction data.

[0061] In this embodiment, the video big model includes a first network and a second network with different frame rates: the first network performs high-precision target food detection on single-frame images of video interaction data with a low temporal sampling rate, while the second network performs temporal tracking and dynamic detection of target food in continuous video sequences with a high frame rate. The two work together to achieve food perception that balances efficiency and accuracy.

[0062] In some embodiments, the training process of the video large model uses an external cache of the storage device of the cooling device, and the prediction process of the video large model uses an internal cache of the storage device.

[0063] In this embodiment, during the training of the large video model, the external cache of the cooling device storage device, such as cloud or local extended storage, is used to process large-scale training data; while in the prediction stage, the internal cache of the storage device, such as the device-side high-speed cache, is called to achieve low latency and high efficiency in real-time food detection and interactive response.

[0064] The following describes a specific embodiment of an interaction method for a refrigeration device.

[0065] like Figure 2 As shown, Step 1, executed through the data acquisition module 101, primarily involves acquiring voice, audio-visual, and image data for prediction and training. This includes data acquisition related to refrigerator-user interaction, internal and external environmental information, refrigerator space video, human-computer interaction, multi-turn dialogue, intelligent agents, multi-agent systems, personal intelligent assistants, data transmission, data generation tools, and data uploading. The focus is on voice, image, and audio-visual data, forming effective voice data or voice data with and without wake-up words. Additional data requires labeling or tagging, and image / visual and video data need to be categorized and processed. The acquisition includes real-time and offline voice, image, video, and audio-visual data from various input methods, including but not limited to smart devices or home appliances, such as 5G / 6G, WiFi, APP, microphones, mobile phones, Bluetooth, large screens, and other channels. If the hardware acquisition device is an independent peripheral, the data acquisition is completed by the external device.

[0066] Step 2: The data processing and acquisition module 102 is used to perform the processing of voice, image, and video data, such as the video refrigerator. The main tasks are voice, image, audio and video processing, and refrigerator space processing.

[0067] Step 2.1: The speech recognition module 102-1 is used. Due to the high error rate of speech in the vertical field, based on the speech data, firstly, the speech front-end is processed, mainly to process empty speech and noisy speech, and extract speech features. For empty speech, speech tools are used to extract speech in empty periods; for noisy speech, a binary classification algorithm is used to separate noise and save pure speech features respectively; next, a keyword prompting algorithm with keywords such as refrigerator, muscle, magnetic control, gear position, etc. is implemented to improve the recognition and transcription rate; finally, the speech features are decoded and the text content is output.

[0068] Step 2.2: The video data processing module 102-2 is used to process the audio and video data. For video data, the process involves initialization and loading, streaming binary video decoder, separation and extraction of frame information, and saving of frame data. For audio and video data, the process involves audio and video separation tools. Then, the audio and video data are saved and then enter their respective processing flows.

[0069] Step 2.3: Execute the frame data acquisition module 102-3 to complete the video frame data storage. Distributed storage methods are used, such as Ceph distributed storage, combined with internal and external caching methods. For example, when the model is training and the data volume is large, external distributed storage is used, which can store a large amount of data and has high read and write efficiency. When the model is predicting, for consecutive frames such as 30fps, binary Redis and video decoding tools are used.

[0070] Step 3, Multimodal large model or large model fusion module 103, which completes the training and prediction tasks of large models with text data, images, and videos as data forms.

[0071] Step 3.1: Using tracking model 103-1, complete the training and prediction task of the tracking model for the hand and drawer holding the target object. First, a target tracking model for the hand-held food object is constructed with the hand and drawer as the regions. This target tracking model can be a deep convolutional neural network, a transformer architecture, a diffusion network model, or a deep network model that combines these. The drawer frame is used as the dividing line. When the object is inside the drawer frame, it needs to be identified and classified as "what it is" and "where it is". When the object is outside the drawer frame, the model tracks it through video, that is, it obtains the tracking target. The target objects are divided into hand-held food or ingredients that are outside the dividing line and those that are inside the dividing line. The hand posture can be flat / oblique, and the target object or object can be at various angles. Next, the position and target bounding box of the hand-held food object are obtained by tracking. Finally, the target feature information of the hand-held object is continuously tracked.

[0072] Step 3.2: The food object detection model 103-2 is used to train and predict target models for objects held in hands and drawers, drawer objects, shelves, storage boxes, and unheld objects within the refrigerator space. The model primarily obtains feature information (what it is, how much of it) or location information of the target object and outputs it to the classification model 103-3 as input data. It also compensates for or corrects unclear or occluded images or frames, helping to address occlusion issues and long-tailed image problems such as unclear or partially clear images. For example, it might show 5 eggs, 2 tomatoes, etc.

[0073] Step 3.3: Execute the image or target classification model module to complete quantitative tasks such as identifying the specific type and quantity of common ingredients, easily confused food targets, small food targets (i.e., small size), and multi-size features. Here, the model is represented by a deep convolutional network model. For example, convolutional layers form the backbone network, followed by a neck network, and finally the detection head. To better identify easily confused targets, a deep classification model can be added before the neck network. For example, if saffron is a broad category, it can be further subdivided into smaller categories, i.e., fine-grained network model features. This improves the classification of easily confused targets. Similarly, for unclear or noisy images, data augmentation, fusion of shallow and deep features of the model, and noise reduction methods are used to improve feature acquisition capabilities.

[0074] Step 3.4: Execute the large language model 103-4 to complete tasks such as text semantic matching, inference, and generation. Especially when image object tracking, recognition, or detection is inaccurate, the target ID of the interactive content generated by the language model can be associated with the video object, such as model alignment, thus compensating for the error correction mechanism in video tracking or recognition and improving the efficiency and experience of user interaction. This large language model supports three types: decoder, encoder, and encoder-decoder end-to-end large model network structure. It is recommended to optimize non-inference algorithms: that is, to complete the optimization and improvement of the main core module algorithms of the large model, forming a hybrid modular large model system. Modular algorithm methods can be adopted, such as targeted improvements to the attention mechanism, such as multi-head attention (MHA). First, according to the target task, the decoder is processed using a feed-forward network (FFN) and normalized for residual direct connection. It also includes a hybrid expert model, where one shared expert and 255 others are used to handle different application scenarios. This approach has the advantage of division of labor and coordination among experts, improving interaction efficiency. Similarly, this includes model compression and multimodal temporal prediction (MTP) multi-token prediction and multimodal position rotation encoding, as well as improvements to context learning algorithms and reinforcement learning, with reinforcement learning employing training strategies such as DPPO. A proposed inference optimization module, which completes the inference algorithm task for a hybrid modular large model, should first employ procedural supervised learning (PSI); then perform reinforcement learning using sampling training strategies (PPO / DPO), including optimizing training preferences by combining synthetic data obtained from knowledge retrieval. This yields a complex task inference model. The underlying inference model can utilize DeepSeek-R1 or LLama, Google open-source models, and training methods, along with comparative results, to select the optimal model base.

[0075] Step 4: Execute the large image module 104 to complete the target object classification training and prediction task. First, obtain multi-scale convolutional neural networks (MCNNs), such as those beneficial for low-resolution, multi-target image classification, or easily confused feature information, and extract long-tail image characteristic information; second, perform multi-channel convolutions on the neural convolutional network; third, obtain similar features and low-resolution features respectively, and further obtain video feature information; then, perform 1×1 convolutions for fusion; finally, the loss function stabilizes or generalizes the model and outputs the classification results.

[0076] Step 5: Execute the visual module 105 to complete the target object tracking, detection, localization training, and prediction tasks. This network model uses a backbone network such as the YOLOx series and is divided into a slow network to acquire low-resolution images with difficult-to-distinguish features, and a fast network to acquire high-resolution images and videos with better-distinguishable features or high weights. The fast and slow network structure is a deep neural network model of deep convolutional network -> multi-channel -> fused features. Next, the neck area undergoes interactive trajectory overlap processing. The purpose is to process images and videos with interactive trajectories. This can be achieved using Intersection over Union (IOU) or Generalized Intersection over Union (GIOU) loss methods or methods, or by considering the positional deviation information from the center point or key points. For example, if the positional deviation information is less than a certain value, it represents the optimal feature region position information, thus obtaining the best and most effective image and video. Finally, head target detection processing, i.e., classification and localization, is performed.

[0077] Step 6: The target object tracking, result verification and correction tasks are completed by executing the target tracker, verifier and corrector 106.

[0078] Step 6.1: Execute the target object tracker 106-1 to complete the tracking and matching tasks of the target bounding box. The main tasks include the unique identifier (ID) of the target tracking, similar trajectories, similar features, reconstruction, difference between the predicted bounding box and the actual bounding box, and best matching. For example, first obtain the state and position information of the predicted target in the next frame, and then perform the best matching between the target bounding box or state of the current frame and the actual / current target tracking bounding box.

[0079] Step 6.2: Execute through the target object corrector 106-2 to complete the verification or validation and correction of the object target box matching results. This mainly solves problems such as the next frame not being detected, the next frame's state and position information being incorrect, or frame information not being detected, or even frame loss. In such cases, a step or process is initiated, such as the video object message driver 107.

[0080] Step 7: Executed via the video object message driver or video event manager 107, this involves driving the video object prediction results after validity verification, judging and analyzing interactive behavior, and generating tracking events. This generally includes tasks such as judging the dynamics of the tracked object, storing, analyzing, and generating results. Specific steps include: ① If the predicted position and state of the next frame are valid, then record information or driving information for the drawer food object is generated, and the threshold for the predicted position and state information of the target object in the next frame and the re-identification mechanism are analyzed and judged. Figure 2 If the value is not within this range, continue to obtain the prediction for the next frame for re-identification or loop tracking identification. Figure 2102-3; If the frame has a valid value, proceed to step 107-3 in Figure 2; ② Judge the interactive behaviors such as accessing and storing the target object; ③ Build an event manager or driver for the predicted video object, that is, complete the complete video prediction task, and build an event manager with behaviors such as accessing and storing objects in the drawer, holding objects in the drawer, and shelves, that is, the prediction result is similar to the message management queue mechanism.

[0081] Step 8: Execute the video results presentation to complete the presentation of results such as food ingredient recognition, food ingredient classification, fruit and vegetable products, and packaged foods. This includes voice broadcasting, text and image presentation tasks. The voice and text output processing is as follows.

[0082] Step 8.1: For video output data such as short videos, an attention-based convolutional block (Att conV block) is formed by a transformer deep network model consisting of a full attention module (Convolution, conV), a rectified linear unit (ReLU), and layer normalization (LayerNorm) to complete the decoding task. The specific implementation process is as follows: first, the full attention mechanism is used to form a unified feature vector or expression for the features themselves and between features, and then convolution operation, ReLU calculation, and LayerNorm are performed. In this way, semantic features with high-level and long-dependency distance relationships are obtained to realize the decoding task and generate the corresponding video data.

[0083] Step 8.2 Output data in text, image, or voice formats, such as the name of the food ingredient, location information, food ingredient freshness information, and food ingredient expiration information. It can also be presented in a combination of text and images.

[0084] Refrigerator: A multi-layered refrigerator supporting multimodal video, including video, audio, and 3D video, as well as intelligent refrigeration refrigerators with constant temperature functions such as freezing and refrigeration. It features human-computer interaction interfaces such as gestures, voice, text, audio and video, image acquisition, microphone arrays, and radar signal acquisition (e.g., images). Tools include Optical Character Recognition (OCR) devices, documents, Virtual Reality (VR), and Augmented Reality (AR). Presentation formats include APP / PC / web / mini-programs, and large screens. Supported protocols include TCP / IP, Wi-Fi, and Bluetooth. Supported networks include 5G, Wi-Fi, wired network, and 6G. Application areas include home, hotel, industrial, indoor and outdoor, military, and aerospace. Embedding methods include flat embedding, automotive embedding, and integration.

[0085] Video refrigerator scenarios include integrated or specialized solutions for food identification and classification, food video storage, preservation, recipes, and dietary health, food video management, dietary health, the entire process of food retrieval / storage, purchase lists and intelligent replenishment, and video interactive services and modes such as diet planning.

[0086] The interactive method for refrigeration equipment provided in this application addresses the issues of low accuracy and slow response speed in target object recognition, tracking, detection, positioning, and classification in video refrigerators, thereby improving user experience and intelligence. Specifically, it addresses the problem of low accuracy and precision in video recognition of food ingredients, packaged food, etc., in application scenarios such as video recognition, video detection and positioning, and video image classification. The video refrigerator is composed of core components including video data acquisition and processing, video data processing, a multimodal large-scale model vertical video classification model, a large-scale visual model, a target object tracker and verifier, and a target object message event manager. Figure 2 As shown. To address the low recognition rate of video object tracking, such as handheld food, hand and drawer area object recognition, and object recognition in refrigerator drawers or shelves, a video object recognition, tracking, and localization method is proposed. Figure 2 As shown in Figures 103-1 and 103-2, this system meets the needs of refrigerator users to randomly and arbitrarily retrieve and store food items; what is placed is what is seen. Addressing the issue of low classification accuracy for easily confused food items, fruits, vegetables, and low-resolution images in videos, a video refrigerator image / target classification model is proposed, such as... Figure 2As shown in Figure 103-3, this method allows users to store ingredients freely, with what they place being what they see. Addressing the issue of low interaction efficiency due to the complementarity and integration of audio and video, a method integrating validity verification and message event management that deeply fuses a large text language model and an image / object model is proposed. Figure 2 As shown in Figures 103, 106, and 107.

[0087] Related technologies suffer from a series of drawbacks, including unclear video recordings of refrigerator storage / retrieval operations, blurry videos, obstructed videos, inaccurate tracking and identification of food items with trademarks / barcodes, and inaccurate target positioning. For example, plums are easily confused with cherry tomatoes, pears with citrus fruits, and star fruit is particularly difficult to distinguish from other fruits due to its texture and shape.

[0088] Target detection and localization are prone to overlap or occlusion with the trajectory, and the target object and the tracking trajectory are easily confused or the behavior trajectory is confused. For example, videos of fruit held with one or two hands in a flat, tilted, or large angle situation, as well as food storage / retrieval in low-resolution video images.

[0089] When refrigerator storage and retrieval actions involve audio and video, or when video and audio are separated, the recognition rate is low and the interactive experience is poor in terms of interactivity and integration. For example, if there is a problem in recognizing food items such as apples in the video, voice interaction can be used to correct or rectify the problem of low recognition rate of food items such as apples. If the video recognition is insufficient, voice interaction can make up for the problem of poor experience.

[0090] In the refrigerator video, food can be stored with one hand and retrieved with the other, and vice versa. How can we quickly identify the food and display it on the refrigerator app or a large screen?

[0091] Other scenarios: such as how to quickly visualize food ingredients and results, present the specific status of refrigerator space, and enable interaction and status information between food ingredients and their location, etc., using voice recognition, cameras, apps, large screens, etc.

[0092] While the core technologies of large-scale artificial intelligence models are rapidly developing and being applied, video interaction and effects still face numerous challenges and difficulties. For example, in smart home appliances such as refrigerators and the refrigeration industry, poor interactive experiences and low adoption rates pose significant challenges, especially given the substantial discrepancy between user expectations and actual interaction. Failure to address these issues can lead to unsatisfactory user experiences and limitations on product competitiveness. Currently, no solutions address the problems described above regarding video refrigerators, systems, storage media, and devices used in smart video refrigerator scenarios, downstream tasks, and video and audio-visual interaction functions. Firstly, single-interaction technologies or frameworks, such as voice or video, or single-modal interaction, single-modal learning, and pattern recognition methods, are insufficient to meet the requirements for interaction accuracy and request response time. This results in inconvenience and inaccurate recognition for users in various scenarios, impacting product brand influence. For example, real-time identification and management of ingredients such as vegetables, food, meat, and fruits and vegetables not only requires accurate identification, but also the integration of multiple learning modes, such as video image recognition, target detection and localization, and tracking, to meet the needs of storing and retrieving ingredients anytime, anywhere, and as needed. Secondly, the accuracy and speed of single-target and multi-target tracking, identification, and localization in video refrigerators are poor: either they cannot track a specific frame of data of a particular target, or they can track a certain frame, or they are accurate for single targets but poor for multi-target targets. In short, the performance of "precision, location accuracy, and recognition speed" of target tracking, identification, localization, and video image classification is unsatisfactory. Therefore, how to form or solve the optimal performance and effect of each module and the overall model is a challenge. Finally, although deep network model structures represented by Convolutional Neural Network (CNN), transformer, MLP, and diffusion network model SM have achieved good results, these models have limitations, such as large computational load or unsatisfactory accuracy in adapting to downstream tasks. With the rapid development of large-scale models, the challenge lies in how to integrate and construct the best combination of effective model modules, including how to efficiently integrate object detection models, text and vision models to obtain the best model performance.

[0093] The interactive method for refrigeration equipment provided in this application makes video refrigerator interaction as simple and natural as human-to-human communication.

[0094] Based on this, the interaction method for refrigeration equipment provided in this application embodiment innovates in various aspects, including multiple interaction methods, large model post-training fine-tuning, reinforcement learning, supervised / semi-supervised / unsupervised / self-learning, etc. It not only utilizes sensors, acquisition devices, and acquisition equipment to collect multi-source heterogeneous data such as real-time and offline voice, video, images, and text from users, but also adapts to downstream multi-task business application systems, intelligent device interaction, and a one-stop service model for content presentation. Therefore, the interaction method for refrigeration equipment provided in this application embodiment is of great significance and creates core value for improving the applicability and interaction efficiency of refrigerator video interaction, effects, intelligent device entry, and interaction modes, as well as promoting demand functions, scientific research and implementation, and socio-economic benefits.

[0095] Addressing the key issue of improving the accuracy and effectiveness of video interaction and generation in refrigerators through multimodal large models, specialized and vertical large models such as visual large models, textual large language models, and image large models, this application's embodiment of the interaction method for refrigeration equipment proposes a new perspective on adapting artificial intelligence large models to downstream or target tasks. It fully utilizes various interaction types such as voice and audio-visual to construct and train large models, and improves and optimizes them through multi-technology fusion. This method is a pioneering and novel service model for solving application scenarios using large model application systems, and is expected to improve the efficiency of big data technology development, such as the fusion of data elements and large models, multi-task joint applications, and vertical field applications.

[0096] The interactive method for refrigeration equipment provided in this application addresses the issues of low accuracy and slow response speed in target object recognition, tracking, detection, positioning, and classification in video refrigerators, thereby improving user experience and intelligence. Specifically, it addresses the problem of low accuracy and precision in video recognition of food ingredients, packaged food, etc., in application scenarios such as video recognition, video detection and positioning, and video image classification of video data collected or acquired by the refrigerator.

[0097] ① To address the low recognition rate of video object tracking, such as handheld food objects, hand and drawer area object recognition, and object recognition in refrigerator drawers or shelves, a video object recognition, tracking, and localization method is proposed. Figure 2 As shown in Figures 103-1 and 103-2, this model meets the needs of refrigerator users to randomly and arbitrarily retrieve and store food items, ensuring that what is placed is what is visible; ② To address the problem of low classification accuracy for easily confused food items, fruits and vegetables, and low-resolution images in videos, a video refrigerator image / target classification model is proposed, such as... Figure 2As shown in Figure 103-3, it allows users to store ingredients freely, with what they place being what they see; ③ To address the problem of low interaction efficiency in audio-visual complementarity and fusion, a method integrating tracking, correction, and re-identification mechanism and message event manager that deeply integrates a large text language model and an image / target model is proposed, as shown in Figure 103-3. Figure 2 As shown in Figures 103, 106, and 107.

[0098] The overall technical solution is applied in a practical and creative way, with each module being interdependent and closely related, inseparable and interconnected. Each module and process is sequential, including the model training and prediction stages. The focus is on acquiring frame data, data processing, the hand and drawer holding the target object, the target detection model, the classification model, the target tracker and corrector, the re-identification mechanism, and the target object message manager, forming a complete system of modules. The learning methods include unsupervised, semi-supervised, supervised, and self-supervised multi-source data, multimodal large-scale model construction, large-scale model fine-tuning, visual large-scale models, image recognition large-scale models, and dedicated models for target detection, tracking, and classification, as well as human-computer interaction and human feedback reinforcement learning. In other words, it presents a complete and systematic solution for this problem, outlining specific steps and applications for a targeted video refrigerator method and system.

[0099] The interactive method for refrigeration equipment provided in this application includes visual / target / image / video prompting engineering in various modules, N-shot, hybrid expert models, self-attention mechanisms including mutual attention and self-attention and cross-attention networks, and other fusion deep learning networks such as deep convolutional network models such as YOLOx and Mamba, transformers, diffusion network models and their derived deep neural network models, system applications and devices, storage media, including methods for constructing and training large models of text and visual language. Among them, the deep neural network model with end-to-end encoder-decoder structure or LLM construction with multi-head attention only-decoder / encoder transformer, cross-attention mechanism / group attention, and self-attention blocks of transformer is used to construct a large deep neural network language model as an application development scheme for generative content such as controllers / services.

[0100] The interactive method for refrigeration equipment provided in this application addresses the construction and training of a large-scale integrated model for refrigerator video interaction and processing, acquisition methods, applications, and downstream tasks. It utilizes CNN / Mamba / transformer models, compression and semantic layers, feedforward neural networks (FFN), and normalization calculations to construct an efficient large-scale model, reward model, and training strategy. This method not only employs optimized and improved end-to-end visual / text / image large-scale models with CNN / Mamba / transformer models, but also incorporates deep integrated neural network models such as distillation, bidirectional long short-term memory (BiLSTM), diffusion models, pre-trained LLM models, attention mechanisms, and text / speech / audio / video encoding / decoding, based on rich theoretical foundations, ultimately forming a Gaussian mixture deep neural network.

[0101] The interaction method for the refrigeration equipment provided in this application involves the construction of large CNN / Mamba / transformer models, which can employ semantically compressed visual / text large model LLMs with decoder-ONLY, encoder-ONLY, and encoder-decoder configurations. These include, but are not limited to, deep network models such as LLMs with graph neural networks, attention mechanisms, CNN / Mamba / transformer models and their variants or improvements, distillation network LLMs, latent / diffusion model LLMs, and U-net networks, as well as fused deep network models such as those based on CNN / Mamba / transformer deep convolutional networks and their derivative variants, such as BiLSTM+Multi-Stream Multi-Scale Convolutional Neural Network (MMCNN)-RNN, BiLSTM+MMCNN+Attention, MMCNN+BiLSTM+Attention fusion model LLMs, and gated recurrent units (Gated Recurrent Units). The models include GRU+CNN, GRU+CNN+Attention models, deep reinforcement learning and reward models, RAG's LLM series of neural networks or deep neural network models, as well as Gaussian mixture deep neural network models.

[0102] The interaction method for a refrigeration device provided in this application can be executed by an interaction device of the refrigeration device. This application uses the example of an interaction device of the refrigeration device executing the interaction method to illustrate the interaction device of the refrigeration device provided in this application.

[0103] This application also provides an interactive device for a refrigeration equipment.

[0104] like Figure 3 As shown, the interaction device of the refrigeration equipment includes: The acquisition module 310 is used to acquire interaction data between the user and the cooling equipment, and the interaction data includes at least video interaction data. The first processing module 320 is used to input video interaction data into the video big model, track and detect the target food in the video interaction data through the video big model, and output the image feature information of the target food based on the relative positional relationship between the target food and the storage space of the refrigeration equipment. The second processing module 330 is used to output target interaction information based on image feature information.

[0105] According to the interactive device for refrigeration equipment provided in the embodiments of this application, by acquiring video interaction data between the user and the refrigeration equipment and inputting it into a large video model, the large video model is used to accurately track and detect the target food. By combining the relative positional relationship between the target food and the storage space of the refrigeration equipment to extract image feature information, clear spatial context constraints can be provided for the image feature information, thereby improving the accuracy of food identification and more accurately inferring the user's intention. Finally, the target interaction information is output, which can improve the recognition accuracy and interaction precision of the interaction.

[0106] In some embodiments, the video big model includes a target tracking model and a target detection model. The first processing module 320 is used to input video interaction data into the target tracking model, and to continuously track the target food when the target food is detected to be outside the storage space, or to output the location information of the target food when the target food is detected to be inside the storage space. When the target tracking model detects that the target food is located in the storage space, the video interaction data is input into the target detection model. The target detection model detects the type and quantity of the target food and outputs the category count information of the target food. Image feature information is determined based on location information and category counting information.

[0107] In some embodiments, the second processing module 330 is used to input image feature information into a target classification model, classify the target food ingredients through the target classification model, and obtain target classification information. A deep classification network is set before the neck network of the target classification model. Based on the target classification information, output the target interaction information.

[0108] In some embodiments, the interactive data further includes voice interactive data, and the first processing module 320 is further configured to convert the voice interactive data into text form to obtain text interactive data; The text interaction data is input into the language model, and the semantic features of the text interaction data are analyzed by the language model to output semantic information. Based on image feature information, the target interaction information is output, including: Based on image feature information and semantic information, the target interaction information is output.

[0109] In some embodiments, the first processing module 320 is used to input image feature information into the target verification model and output the false detection information of the target food through the target verification model; The missed detection and false detection information and semantic information are input into the target correction model. The target correction model corrects the missed detection and false detection of the target food and obtains the corrected image feature information. Based on the corrected image feature information, the target interaction information is output.

[0110] In some embodiments, the video big model includes a first network and a second network. The sampling frame rate of the first network in the time dimension is lower than that of the second network. The first processing module 320 is used to detect target food in a single frame image of the video interaction data through the first network, and to track and detect target food in a continuous video of the video interaction data through the second network.

[0111] In some embodiments, the training process of the video large model uses an external cache of the storage device of the cooling device, and the prediction process of the video large model uses an internal cache of the storage device.

[0112] The interactive device of the cooling equipment in this application embodiment can be an electronic device or a component of an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other devices besides a terminal. For example, the electronic device can be a mobile phone, tablet computer, laptop computer, handheld computer, in-vehicle electronic device, mobile internet device (MID), augmented reality (AR) / virtual reality (VR) device, robot, wearable device, ultra-mobile personal computer (UMPC), netbook or personal digital assistant (PDA), etc. It can also be a server, network attached storage (NAS), personal computer (PC), television (TV), ATM or self-service machine, etc. The embodiments of this application do not specifically limit it.

[0113] The interactive device of the refrigeration equipment in this application embodiment can be a device with an operating system. The operating system can be Microsoft (Windows) operating system, Android operating system, iOS operating system, or other possible operating systems. This application embodiment does not specifically limit the specific operating system.

[0114] The interactive device for the refrigeration equipment provided in this application embodiment can achieve... Figure 1 The various processes implemented in the method implementation examples will not be described again here to avoid repetition.

[0115] This application also provides a refrigeration device.

[0116] The refrigeration equipment includes the interactive device and data acquisition device of the refrigeration equipment as described above.

[0117] The data acquisition device is connected to the interaction device. The data acquisition device is used to collect interaction data between the user and the cooling equipment. The interaction data includes at least video interaction data.

[0118] According to the refrigeration device provided in the embodiments of this application, by acquiring video interaction data between the user and the refrigeration device and inputting it into a large video model, the target food is accurately tracked and detected using the large video model. By combining the relative positional relationship between the target food and the storage space of the refrigeration device to extract image feature information, clear spatial context constraints can be provided for the image feature information, thereby improving the accuracy of food identification and more accurately inferring the user's intention. Finally, the target interaction information is output, which can improve the recognition accuracy and precision of the interaction.

[0119] This application also provides an interactive system for a refrigeration device.

[0120] The interactive system of the refrigeration equipment includes: Such as the refrigeration equipment and server mentioned above.

[0121] The server connects to the cooling equipment and is used to output target interaction information.

[0122] According to the interactive system of the refrigeration equipment provided in the embodiments of this application, by acquiring video interaction data between the user and the refrigeration equipment and inputting it into a large video model, the system accurately tracks and detects the target food ingredients using the large video model. By combining the relative positional relationship between the target food ingredients and the storage space of the refrigeration equipment to extract image feature information, it can provide clear spatial context constraints for the image feature information, thereby improving the accuracy of food ingredient recognition and more accurately inferring the user's intention. Finally, it outputs the target interaction information, which can improve the recognition accuracy and interaction precision of the interaction.

[0123] In some embodiments, such as Figure 4 As shown, this application embodiment also provides an electronic device 400, including a processor 401, a memory 402, and a computer program stored in the memory 402 and executable on the processor 401. When the program is executed by the processor 401, it implements the various processes of the above-described interaction method embodiment of the cooling device and achieves the same technical effect. To avoid repetition, it will not be described again here.

[0124] It should be noted that the electronic devices in the embodiments of this application include the mobile electronic devices and non-mobile electronic devices described above.

[0125] This application also provides a non-transitory computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it implements the various processes of the above-described interaction method embodiment of the refrigeration device and achieves the same technical effect. To avoid repetition, it will not be described again here.

[0126] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.

[0127] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the interaction method of the above-mentioned refrigeration device.

[0128] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.

[0129] This application embodiment also provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement the various processes of the above-described interaction method embodiment of the refrigeration device, and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0130] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.

[0131] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.

[0132] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the related technology, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0133] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

[0134] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "illustrative embodiment," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0135] Although embodiments of this application have been shown and described, those skilled in the art will understand that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of this application, the scope of which is defined by the claims and their equivalents.

Claims

1. An interaction method for a refrigeration device, characterized in that, include: Acquire user interaction data with the cooling device, the interaction data including at least video interaction data; The video interaction data is input into the video big model, and the target food in the video interaction data is tracked and detected by the video big model. Based on the relative positional relationship between the target food and the storage space of the refrigeration equipment, the image feature information of the target food is output. Based on the image feature information, target interaction information is output.

2. The interaction method of the refrigeration equipment according to claim 1, characterized in that, The video big data model includes a target tracking model and a target detection model. The video interaction data is input into the video big data model, which tracks and detects the target food items in the video interaction data. Based on the relative positional relationship between the target food items and the storage space of the refrigeration equipment, the model outputs the image feature information of the target food items, including: The video interaction data is input into the target tracking model. When the target food is detected to be outside the storage space, the target tracking model continuously tracks the target food, or when the target food is detected to be inside the storage space, it outputs the location information of the target food. When the target tracking model detects that the target food is located in the storage space, the video interaction data is input to the target detection model, which detects the type and quantity of the target food and outputs the category count information of the target food. The image feature information is determined based on the location information and the category count information.

3. The interaction method of the refrigeration equipment according to claim 1, characterized in that, The step of outputting target interaction information based on the image feature information includes: The image feature information is input into the target classification model, and the target food is classified by the target classification model to obtain target classification information. A deep classification network is set before the neck network of the target classification model. Based on the target classification information, the target interaction information is output.

4. The interaction method of the refrigeration equipment according to claim 1, characterized in that, The interaction data also includes voice interaction data. Before outputting the target interaction information based on the image feature information, the method further includes: The voice interaction data is converted into text format to obtain text interaction data; The text interaction data is input into a language model, and the semantic features of the text interaction data are analyzed by the language model to output semantic information. The step of outputting target interaction information based on the image feature information includes: Based on the image feature information and the semantic information, the target interaction information is output.

5. The interaction method of the refrigeration equipment according to claim 4, characterized in that, The step of outputting the target interaction information based on the image feature information and the semantic information includes: The image feature information is input into the target verification model, and the target verification model outputs the false negative and false positive information of the target food ingredient. The missed detection and false detection information and the semantic information are input into the target correction model. The missed detection and false detection of the target food are corrected by the target correction model to obtain the corrected image feature information. Based on the corrected image feature information, the target interaction information is output.

6. The interaction method of the refrigeration equipment according to any one of claims 1-5, characterized in that, The video big data model includes a first network and a second network. The first network has a lower sampling frame rate in the time dimension than the second network. The step of inputting the video interaction data into the video big data model and tracking and detecting target ingredients in the video interaction data using the video big data model includes: The first network detects the target food ingredients in a single frame of the video interaction data, and the second network tracks and detects the target food ingredients in a continuous video of the video interaction data.

7. The interaction method of the refrigeration equipment according to any one of claims 1-5, characterized in that, The training process of the large video model uses the external cache of the storage device of the cooling device, and the prediction process of the large video model uses the internal cache of the storage device.

8. An interactive device for a refrigeration equipment, characterized in that, include: An acquisition module is used to acquire interaction data between the user and the cooling device, the interaction data including at least video interaction data; The first processing module is used to input the video interaction data into the video big model, track and detect the target food in the video interaction data through the video big model, and output the image feature information of the target food based on the relative positional relationship between the target food and the storage space of the refrigeration equipment. The second processing module is used to output target interaction information based on the image feature information.

9. A refrigeration device, characterized in that, include: The interactive device for the refrigeration equipment as described in claim 8; A data acquisition device is connected to the interaction device. The data acquisition device is used to collect interaction data between the user and the cooling equipment. The interaction data includes at least video interaction data.

10. An interactive system for a refrigeration device, characterized in that, include: The refrigeration equipment as described in claim 9; The server is connected to the refrigeration equipment and is used to output target interaction information.