Interaction method, device and system of refrigeration equipment and refrigeration equipment

By acquiring and analyzing user interaction data and spatial structure information with the smart refrigerator, and combining it with a multimodal large model, the problem of low accuracy in voice interaction of the smart refrigerator was solved, achieving accurate recognition of user intent and improving the accuracy of interaction results.

CN121297344APending Publication Date: 2026-01-09QINDAO HAIER REFRIGERATOR CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511432039.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-30
Publication Date
2026-01-09

AI Technical Summary

Technical Problem

The voice interaction methods of existing smart refrigerators have low accuracy, leading to misunderstandings of user intentions and reducing the accuracy of interaction.

Method used

By acquiring user interaction data and spatial structure information with the refrigeration equipment, and combining voice interaction data and spatial structure information, the spatial interaction characteristics are analyzed. A multimodal large model is then used to accurately identify the specific item or area pointed to by the user, thus avoiding misunderstanding of the user's intentions.

Benefits of technology

It improves the accuracy of user interaction with the smart refrigerator, enabling precise identification of the specific item or area pointed to by the user, thus enhancing the accuracy and efficiency of the interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121297344A_ABST
    Figure CN121297344A_ABST
Patent Text Reader

Abstract

The invention discloses an interaction method, device and system of refrigeration equipment and the refrigeration equipment, and belongs to the technical field of refrigeration equipment. The interaction method of the refrigeration equipment comprises the steps of obtaining interaction data of a user and the refrigeration equipment and space structure information of the refrigeration equipment, wherein the interaction data at least comprises voice interaction data; based on the interaction data and the space structure information, obtaining space interaction features, the space interaction features being used for representing an association relationship between the interaction data and the space structure information; and determining and outputting a target interaction result of the user and the refrigeration equipment based on the space interaction characteristics. The specific article or area pointed by the user can be accurately recognized, and the situation that the interaction accuracy is reduced due to misunderstanding of the user intention is avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of refrigeration equipment technology, and particularly relates to an interaction method, device, system and refrigeration equipment for refrigeration equipment. Background Technology

[0002] With the continuous advancement of artificial intelligence technology, intelligent interaction technology has been widely applied in smart refrigerators. Users' requirements for the intelligence level of smart refrigerators are constantly increasing, and improving the interaction efficiency between users and smart refrigerators is the key to enhancing the intelligence level of smart refrigerators.

[0003] Currently, users typically interact with smart refrigerators through voice commands, but this method of interaction is not very accurate. Summary of the Invention

[0004] This application aims to address at least one of the technical problems existing in the prior art. To this end, this application proposes an interaction method, apparatus, system, and refrigeration equipment for a cooling device, which can accurately identify the specific item or area pointed to by the user, avoiding reduced interaction accuracy due to misunderstanding of the user's intention.

[0005] In a first aspect, this application provides an interaction method for a refrigeration device, the method comprising: Acquire user interaction data with the refrigeration device and spatial structure information of the refrigeration device, wherein the interaction data includes at least voice interaction data; Based on the interaction data and the spatial structure information, spatial interaction features are obtained, which are used to characterize the correlation between the interaction data and the spatial structure information. Based on the spatial interaction characteristics, the target interaction result between the user and the refrigeration equipment is determined and output.

[0006] According to the interaction method of the refrigeration equipment in this application, by associating voice interaction data with spatial structure information, spatial interaction features are obtained. The user's interaction intent in the spatial interaction features is analyzed, along with the area location corresponding to the interaction intent in the refrigeration equipment and the storage status of items in the area location. The target interaction result between the user and the refrigeration equipment is obtained. By combining voice interaction data with spatial structure information, the specific item or area pointed to by the user can be accurately identified, avoiding the reduction of interaction accuracy due to misunderstanding of the user's intent.

[0007] According to one embodiment of this application, obtaining spatial interaction features based on the interaction data and the spatial structure information includes: Obtain the target interaction location and target interaction task of the interaction data; The target interaction location is matched with the spatial structure information to obtain the target spatial label of the spatial structure information; The spatial label is bound to the target interaction task to obtain the spatial interaction feature.

[0008] According to one embodiment of this application, determining and outputting the target interaction result between the user and the cooling device based on the spatial interaction features includes: The spatial interaction features are input into a multimodal large model to obtain the target interaction result output by the multimodal large model.

[0009] According to one embodiment of this application, the multimodal large model includes a first system module and a second system module. The response speed of the first system module is greater than that of the second system module, and the reasoning ability of the second system module is stronger than that of the first system module. The step of inputting the spatial interaction features into the multimodal large model to obtain the target interaction result output by the multimodal large model includes: If the complexity of the spatial interaction features is less than a first complexity threshold, the target interaction result is inferred through the first system module. Alternatively, if the complexity of the spatial interaction features is greater than or equal to the first complexity threshold, the target interaction result can be inferred through the second system module.

[0010] According to one embodiment of this application, the multimodal large model is trained through the following steps: Construct a reward model for the multimodal large model, and score the output of the multimodal large model using the reward model to obtain a first score value; Based on the first score, the reinforcement learning process of the multimodal large model is initiated; Based on the first score and the reinforcement learning factors of the reinforcement learning process, the training factors are determined. The multimodal large model is trained based on the training factors.

[0011] According to one embodiment of this application, obtaining spatial interaction features based on the interaction data and the spatial structure information includes: The voice interaction data is then processed with voice prompts and enhancements. First speech feature information is extracted from the enhanced speech interaction data; Based on the first voice feature information and the context data of the voice interaction data, the voice interaction data is decoded; The spatial interaction features are obtained based on the decoded voice interaction data and the spatial structure information.

[0012] According to one embodiment of this application, the enhancement processing of the voice interaction data includes: The second speech feature information of the speech interaction data is obtained by multi-scale convolution in the temporal domain; Based on the second voice feature information, the voice interaction data is repaired. Based on the repaired voice interaction data, a third voice feature information is obtained by performing a stable adversarial process through the generator and discriminator and the loss function convergence. The voice interaction data is subjected to noise reduction processing based on the third voice feature information in order to enhance the voice interaction data.

[0013] According to one embodiment of this application, the interactive data further includes at least one of text data, image data, and video data.

[0014] Secondly, this application provides an interaction device for a refrigeration device, the device comprising: The acquisition module is used to acquire user interaction data with the refrigeration device and spatial structure information of the refrigeration device, wherein the interaction data includes at least voice interaction data. The first processing module is used to obtain spatial interaction features based on the interaction data and the spatial structure information, wherein the spatial interaction features are used to characterize the correlation between the interaction data and the spatial structure information; The second processing module is used to determine and output the target interaction result between the user and the refrigeration equipment based on the spatial interaction characteristics.

[0015] According to the interactive device of the refrigeration equipment of this application, by associating voice interaction data with spatial structure information, spatial interaction features are obtained. The user's interaction intent in the spatial interaction features, the area location corresponding to the interaction intent in the refrigeration equipment, and the storage status of items in the area location are analyzed to obtain the target interaction result between the user and the refrigeration equipment. By combining voice interaction data with spatial structure information, the specific item or area pointed to by the user can be accurately identified, avoiding the reduction of interaction accuracy due to misunderstanding of the user's intent.

[0016] Thirdly, this application provides a refrigeration device, comprising: The interactive device of the refrigeration equipment as described in the second aspect above; A data acquisition device is connected to the interaction device, and the data acquisition device is used to collect interaction data between the user and the refrigeration equipment and spatial structure information of the refrigeration equipment.

[0017] According to the refrigeration device of this application, by associating voice interaction data with spatial structure information, spatial interaction features are obtained. The user's interaction intent in the spatial interaction features is analyzed, along with the area location corresponding to the interaction intent in the refrigeration device and the storage status of items in the area location. The target interaction result between the user and the refrigeration device is obtained. By combining voice interaction data with spatial structure information, the specific item or area pointed to by the user can be accurately identified, avoiding the reduction of interaction accuracy due to misunderstanding of the user's intent.

[0018] Fourthly, this application provides an interactive system for a refrigeration device, comprising: The refrigeration equipment described in the third aspect above; The server is connected to the refrigeration equipment and is used to output the target interaction result obtained by the interaction device executing the interaction method.

[0019] According to the interactive system of the refrigeration equipment of this application, by associating voice interaction data with spatial structure information, spatial interaction features are obtained. The user's interaction intent in the spatial interaction features is analyzed, along with the area location corresponding to the interaction intent in the refrigeration equipment and the storage status of items in the area location. The target interaction result between the user and the refrigeration equipment is obtained. By combining voice interaction data with spatial structure information, the specific item or area pointed to by the user can be accurately identified, avoiding the reduction of interaction accuracy due to misunderstanding of the user's intent.

[0020] Fifthly, this application provides an electronic device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the interaction method of the cooling device as described in the first aspect above.

[0021] In a sixth aspect, this application provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the interaction method of the refrigeration device as described in the first aspect above.

[0022] In a seventh aspect, this application provides a computer program product, including a computer program that, when executed by a processor, implements the interaction method of the refrigeration device as described in the first aspect above.

[0023] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description

[0024] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the description of the embodiments taken in conjunction with the following drawings, in which: Figure 1This is a flowchart illustrating the interaction method of the refrigeration equipment provided in the embodiments of this application; Figure 2 This is a schematic diagram of the structure of the interactive system of the refrigeration equipment provided in the embodiments of this application; Figure 3 This is a schematic diagram of the structure of the interactive device of the refrigeration equipment provided in the embodiments of this application; Figure 4 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation

[0025] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.

[0026] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.

[0027] The following description, in conjunction with the accompanying drawings, details the interaction method, interaction device, refrigeration equipment, interaction system, electronic device, and readable storage medium of the refrigeration equipment provided in this application, through specific embodiments and application scenarios.

[0028] The interaction method of the refrigeration equipment can be applied to the terminal, and can be executed by the hardware or software in the terminal.

[0029] The terminal includes, but is not limited to, portable communication devices such as mobile phones or tablets with touch-sensitive surfaces (e.g., touchscreen displays and / or touchpads). It should also be understood that, in some embodiments, the terminal may not be a portable communication device, but rather a desktop computer with touch-sensitive surfaces (e.g., touchscreen displays and / or touchpads).

[0030] The following embodiments describe a terminal including a display and a touch-sensitive surface. However, it should be understood that the terminal may include one or more other physical user interface devices such as a physical keyboard, mouse, and joystick.

[0031] The interaction method of the refrigeration device provided in this application embodiment can be executed by an electronic device or a functional module or functional entity in an electronic device that can implement the interaction method of the refrigeration device. The electronic devices mentioned in this application embodiment include, but are not limited to, mobile phones, tablets, computers, cameras and wearable devices. The interaction method of the refrigeration device provided in this application embodiment will be described below using an electronic device as the execution subject.

[0032] like Figure 1 As shown, the interaction method of the refrigeration device includes steps 110, 120 and 130.

[0033] Step 110: Obtain the interaction data between the user and the refrigeration equipment and the spatial structure information of the refrigeration equipment.

[0034] Interactive data refers to the data input by the user to the refrigeration device, which can represent the user's intention to operate the refrigeration device. Interactive data includes at least voice interactive data, which can be represented as the data generated by the user communicating with the refrigeration device.

[0035] Spatial structure information is information used to characterize the spatial layout of refrigeration equipment. The spatial structure information can include the items stored in each spatial location of the refrigeration equipment. For example, the spatial structure information can include that the drawers of the refrigeration equipment contain fruit, and the first shelf of the refrigeration equipment contains vegetables.

[0036] In this step, a corresponding media acquisition device can be set up on the cooling device to obtain interactive data. For example, voice interaction data can be obtained through a microphone.

[0037] Spatial structural information of refrigeration equipment can be collected through methods such as laser scanning or ultrasonic ranging.

[0038] Step 120: Based on the interaction data and spatial structure information, spatial interaction features are obtained. Spatial interaction features are used to characterize the relationship between interaction data and spatial structure information.

[0039] Among them, spatial interaction features are features jointly extracted from interaction data and spatial structure information that can characterize the relationship between the two.

[0040] In this step, interactive data is combined with spatial structure information. For example, user interaction behavior is mapped to the specific location or area of ​​the refrigeration equipment space. Features that can reflect user behavior patterns and space usage are extracted from the fused data and used as spatial interaction features.

[0041] Step 130: Based on spatial interaction characteristics, determine and output the target interaction results between the user and the refrigeration equipment.

[0042] The target interaction result is the information returned by the refrigeration equipment to the user after the user interacts with the refrigeration equipment. The target interaction result includes the information about the refrigeration equipment that the user wants to know.

[0043] In this embodiment, the target interaction result can be obtained by analyzing the user's interaction intent in the spatial interaction features, the area location corresponding to the interaction intent in the refrigeration device, and the storage status of items in the area location.

[0044] According to the interaction method of the refrigeration equipment provided in the embodiments of this application, by associating voice interaction data and spatial structure information, spatial interaction features are obtained. The user's interaction intent in the spatial interaction features, the area location corresponding to the interaction intent in the refrigeration equipment, and the storage status of items in the area location are analyzed to obtain the target interaction result between the user and the refrigeration equipment. By combining voice interaction data with spatial structure information, the specific item or area pointed to by the user can be accurately identified, avoiding the reduction of interaction accuracy due to misunderstanding of the user's intent.

[0045] In some embodiments, spatial interaction features are obtained based on interaction data and spatial structure information, including: The target interaction location and target interaction task are used to obtain interaction data. Match the target interaction location with the spatial structure information to obtain the target spatial label of the spatial structure information; By binding the target spatial label with the target interaction task, spatial interaction features are obtained.

[0046] The target interaction location is used to represent the location where the user expects to interact with the refrigeration equipment. For example, if the user says to the refrigeration equipment, "View all the items in the freezer," then the target interaction location is the freezer.

[0047] The target interaction task is used to characterize the specific operation that the user expects to perform. For example, if the user says "view all the items in the freezer", then the target interaction task is to view all the items.

[0048] Target space tags are semantic identifiers for the location of areas within a refrigeration device. They describe the functional attributes or physical structure of that location, such as "upper shelf of the refrigerator compartment," "drawer of the freezer compartment," "right side of the door shelf," or "touchscreen area," and are used to associate the physical location of the refrigeration device with the user's intent.

[0049] In this embodiment, the target interaction location and target interaction task can be extracted and segmented from the interaction data using a corresponding model.

[0050] In this embodiment, the user-specified target interaction location is compared with the spatial structure information to find the corresponding actual physical location, and a target spatial label is assigned to it.

[0051] By combining the specific location of the user's operation corresponding to the target spatial label with the target interaction task, a task description unit with spatial semantics, namely spatial interaction feature, is constructed.

[0052] For example, when the system detects that the user's hand is close to the middle layer of the refrigerator, the target space is labeled as the middle layer of the refrigerator. Combined with the voice command, the target interaction task is to put the yogurt in. Thus, the middle layer of the refrigerator is bound to the task of storing food, forming a spatial interaction feature of storing yogurt in the middle layer of the refrigerator.

[0053] In some embodiments, based on spatial interaction characteristics, the target interaction result between the user and the cooling device is determined and output, including: Spatial interaction features are input into a multimodal large model to obtain the target interaction results output by the multimodal large model.

[0054] Among them, the multimodal large model is a deep learning model or a large-scale deep neural network model that can process and understand various types of media data, such as text, images, audio and video.

[0055] In this embodiment, spatial interaction features are input into a multimodal large model. The multimodal large model processes and analyzes the information in the spatial interaction features and finally outputs the target interaction result.

[0056] In some embodiments, the multimodal large model includes a first system module and a second system module. The response speed of the first system module is greater than that of the second system module, and the reasoning ability of the second system module is stronger than that of the first system module. Spatial interaction features are input into the multimodal large model to obtain the target interaction result output by the multimodal large model, including: When the complexity of spatial interaction features is less than the first complexity threshold, the target interaction result is inferred through the first system module; Alternatively, if the complexity of the spatial interaction features is greater than or equal to the first complexity threshold, the target interaction result can be inferred through the second system module.

[0057] The first complexity threshold is a preset complexity.

[0058] In this embodiment, the first system module has a fast response speed and is suitable for handling simple or routine interactive tasks. When the complexity of the spatial interaction features is less than the first complexity threshold, the first system module is activated. For relatively simple or common tasks, fast response is more important than in-depth reasoning and can provide an immediate user experience.

[0059] The second system module has strong reasoning ability and is suitable for handling complex or interactive tasks that require in-depth analysis. When the complexity of spatial interaction features is greater than or equal to the first complexity threshold, the second system module is activated, and its reasoning ability can ensure the accuracy and effectiveness of handling complex tasks.

[0060] In some embodiments, a multimodal large model is trained through the following steps: Construct a reward model for a multimodal large model, and score the output of the multimodal large model using the reward model to obtain the first score value; Based on the first score, initiate the reinforcement learning process of the multimodal large model; Based on the first score and the reinforcement learning factors in the reinforcement learning process, the training factors are determined. Training a large multimodal model based on training factors.

[0061] The first score reflects the quality and accuracy of the multimodal large model output.

[0062] In this embodiment, a reward model is constructed to evaluate the output of the multimodal large model and generate a first score. After obtaining the first score, the reinforcement learning process of the multimodal large model is started. During the reinforcement learning process, the multimodal large model adjusts its behavior according to the first score in order to obtain a higher first score in the future.

[0063] By combining the first score and the reinforcement learning factor in the reinforcement learning process, the training factor is determined. The training factor is a set of parameters used to guide the direction and intensity of training of multimodal large models, and depends on the spatial interaction trajectory and the interaction trajectory table.

[0064] The multimodal large model is trained using training factors, and the training process is iterated until the performance of the multimodal large model reaches the desired level.

[0065] In some embodiments, spatial interaction features are obtained based on interaction data and spatial structure information, including: Perform voice prompts and enhancement processing on voice interaction data; Extract the first speech feature information from the enhanced speech interaction data; Based on the first speech feature information and the context data of the speech interaction data, the speech interaction data is decoded; Spatial interaction features are obtained based on the decoded voice interaction data and spatial structure information.

[0066] In this embodiment, the quality of the voice interaction data is improved by enhancing the voice interaction data, thereby improving the clarity and recognizability of the voice interaction data in complex environments.

[0067] Key acoustic features and learning patterns for recognition are extracted from the enhanced voice interaction data to obtain the first voice feature information.

[0068] Based on the first speech feature information and contextual data such as historical dialogues, user preferences and time context, the speech interaction data is decoded, and the acoustic features of the speech interaction data are converted into semantically accurate text instructions.

[0069] Based on the decoded voice interaction data and spatial structure information, the spatial intent in the voice interaction data is aligned with the actual spatial labels of the cooling equipment, and combined with the operation actions in the voice interaction data to form spatial interaction features.

[0070] In some embodiments, enhancing the voice interaction data includes: Second speech feature information of speech interaction data is obtained by multi-scale convolution in the temporal domain; Based on the second speech feature information, the speech interaction data is repaired. Based on the repaired voice interaction data, the third voice feature information is obtained by performing a stable adversarial process through the generator and discriminator and the loss function convergence. Noise reduction is performed on the voice interaction data based on third-party speech feature information to enhance the voice interaction data.

[0071] In this embodiment, local dynamic patterns at different time scales are extracted from the speech interaction data through temporal multi-scale convolution to obtain second speech feature information that can characterize speech details and rhythmic structure.

[0072] Based on this second speech feature information, distorted or missing segments in the speech interaction data are identified and corrected, thereby restoring and improving the integrity of the speech interaction data.

[0073] The repaired speech interaction data is input into an adversarial network consisting of a generator and a discriminator. The generator further optimizes the speech features, while the discriminator evaluates the authenticity of the generated features. Through adversarial training, the generated third speech feature information is made closer to the natural and clear speech distribution.

[0074] By using third-party speech feature information to guide the noise reduction model, speech and background noise can be accurately separated, and speech interaction data can be enhanced, thereby significantly improving the robustness and accuracy of subsequent recognition.

[0075] In some embodiments, the interactive data also includes at least one of text data, image data, and video data.

[0076] In this embodiment, the interactive data not only includes voice data, but can also incorporate text data, such as user-inputted instructions, image data, such as user gestures captured by a camera, and video data, such as continuous user actions. Through multimodal information collaborative analysis, the understanding of user intent is enhanced.

[0077] The following describes a specific embodiment of an interaction method for a refrigeration device.

[0078] In this embodiment, the interaction method of the refrigeration device is based on, for example... Figure 2 The interactive system of the refrigeration equipment shown is executed.

[0079] Step 1: Acquire Language Interaction Data. This involves collecting language interaction data between the refrigeration device (such as a refrigerator) and the user. This includes internal and external environmental information, human-computer interaction, multi-turn dialogue, intelligent agents, multi-agent interactions, personal intelligent assistants, data transmission, data generation tools, data uploading, and various hardware and software tasks related to smart terminals. The focus is on voice, text, images, and audio / video data, including valid voice data (with and without wake-up words). Data also requires labeling or tagging, and image / visual data needs to be categorized and processed. In short, this involves the user's "speaking" and "storing / retrieving" actions during language interaction with the refrigerator. Image data is based on the refrigerator's spatial location information. This includes real-time and offline voice, text, and video data acquisition from various input methods, including but not limited to smart devices or home appliances, such as 5G / 6G, WiFi, apps, microphones, mobile phones, and Bluetooth. If the hardware collector is an independent peripheral, the data collection will be completed by the external device.

[0080] Step 2, the language interaction data processing module, is responsible for processing or converting the language interaction data between the refrigeration equipment, such as the refrigerator, and the user into text data. This mainly involves tasks such as acquiring voice, audio and video, text, and spatial location information.

[0081] Step 2.1, Speech Interaction Recognition: Due to the high error rate in current refrigerator speech recognition, a speech enhancement model based on adversarial networks and efficient training is adopted for the speech data. First, effective speech prompts are obtained. To further extract and refine pure speech, adversarial networks and trained speech enhancement methods are used. Next, speech features and a speech language model are extracted to obtain contextual speech feature information. Finally, the obtained speech features are used for context decoding. The training data for the context decoders is constructed from the output data of a large model, which enables the acquisition of a more effective context decoder. The transcription framework can adopt open-source speech recognition architectures such as Kaldi, Conformer, and Whisper, as well as novel end-to-end speech recognition architecture models.

[0082] The specific implementation steps of the speech enhancement method based on adversarial networks and efficient training are as follows: ① Obtain speech data features of channel information based on temporal multi-scale convolution, i.e., Convolutional Neural Network (CNN); ② Speech loss repair: repair severely lost signals, reverberant signals, etc., based on the obtained speech features to obtain smooth and stable speech feature data; ③ Train the generator in the adversarial network to obtain model stability and strong generalization model and its parameters; ④ Based on ③, generate the result by the discriminator from the pre-speech, thus obtaining more efficient speech data features; ⑤ Use the loss function to make the generator obtain the optimal model so that the discriminator can better obtain speech features, and then check for noise to further make the speech data features more effective.

[0083] Step 2.2: For text data, text processing methods including text parsing, text prompts during prediction, token algorithms, and text enhancement generation data obtained through synthetic data or model-based or data-driven methods are used to improve the model algorithm effect and obtain more effective text data. This is mainly because the text data contains noise, sparsity, etc., which are not conducive to text feature extraction or calculation.

[0084] Step 2.3: Audio and video text extraction, which involves extracting video data and implementing an image-text cross-attention mechanism. First, audio and video are separated; then, image encoding and image tokenization are performed to pair images with text. Step 2.1: Construct fused feature information to obtain image-text cross-attention features; then, cross-interaction attention features are fused to realize the relationship between images and corresponding text features in the video and input them into the large model / multimodal large model.

[0085] Step 2.4, Language Interaction Spatial Location and Angle Module, together with Step 2.3, forms an image-text-spatial multi-semantic information fusion, that is, to complete the task of generating results such as retrieving and storing spatial location data in the refrigerator and language interaction. The specific implementation is as follows: ① Obtain the spatial location information ID and token for storage and retrieval and spatial localization; then generate a spatial location description: construct spatial location fusion features from position and angle, structural information and task ID and analyze the spatial location description; finally, generate a language spatial decoder, that is, spatial location feature information.

[0086] Step 3: The large-scale language interaction model. This involves constructing the fast and slow systems of the large model, which is based on text data. The output data can also be used as training data for the speech-to-text transcription decoder in Step 2. If the large-scale language interaction model is slow, further model compression, pruning, and distillation methods can be applied to improve its time response efficiency.

[0087] Step 3.1: Complete the decoding task for the system, the first system, or the fast system. A transformer or similar model architecture can be used. For existing application scenarios, a self-developed model such as the BERT (Bidirectional Encoder Representations from Transformers) model can be adopted, which is suitable for scenarios such as food management, knowledge graphs for food health, and knowledge question answering. Alternatively, a privately deployed large model can be used, such as a model fine-tuned based on deepSeek-32B, or a large language model meta (llama) model distilled from deepSeek-761B. This not only reduces costs but also achieves the current requirement of fast response time. A retrieval-augmented generation (RAG) architecture can also be combined. For example, when faced with a query, results can be extracted through semantic similarity calculation, and the required results can be returned to the refrigerator.

[0088] Step 3.2: Complete the decoder task of the system, second system, or slow system. A transformer or similar model architecture can be used. For tasks requiring deep thinking, such as reinforcement learning, visual language models, and complex reasoning tasks involving deep thinking, learning without imitation, or deep exploration, reinforcement learning training strategies such as DeepSeek-V3, R1, or open-source Vision-Language Models (VLM) can be used for model fine-tuning and distillation into efficient models like LLama / Qwen. This can also be combined with a RAG architecture; for example, when faced with a query, results can be extracted through semantic similarity calculation and then returned to the repository. Additionally, timely knowledge updates and a rich knowledge base can be achieved.

[0089] By constructing the above system, application scenarios can be classified and processed according to factors such as the frequency or number of language interaction accesses, thereby improving the accuracy and efficiency of interaction.

[0090] Step 4: Language Interaction Large-Scale Model Training Strategy and Method: This involves implementing and optimizing the language interaction large-scale model training strategy and method. First, a reward model is built upon the language interaction large-scale model, and its training results are scored or rated. Feature scores are evaluated using historical data accumulated from past language interactions, including speech and action interaction data. If there is a high overlap with these historical language interaction data, an interaction trajectory table is generated. The interaction trajectory table is then used for the smallest unit semantic segmentation, and a set of interaction trajectory tables is constructed. If the interaction trajectory table construction meets the requirements for improving or enhancing reinforcement learning, reinforcement learning is then performed. During this process, training factors are used to select or adjust the training strategy. The training factors consist of two parts: the interaction trajectory score and the reinforcement learning factor. Expectation, mean, variance, covariance, or their matrices can be used to calculate the training factors. The interaction trajectory table and its constructed training set also need to be binary vectorized and normalized to improve computational efficiency.

[0091] Step 5, the model backend module, completes downstream tasks such as deploying and publishing pre-trained large-scale models and application services. It mainly handles large-scale model interface configuration, publishing and data protection strategies, model fine-tuning (Supervised Fine-Tuning, SFT) + reinforcement learning user interface (UI) / user experience (UE) interaction to facilitate large-scale model use. This includes features such as voice, image, text, and video content, model encapsulation, model plugins, model service activation, interface call instructions, interface result verification, and intelligent agents. It not only achieves a visual, user-friendly, and quantifiable multi-task joint application service interface but also visualizes post-model training and other algorithm operations.

[0092] Step 6, the text-to-speech module, completes the task of converting text into speech files. It adopts a natural language instruction-guided method and a context-customized synthesis model. Specifically, if there is a previously synthesized speaker with the same timbre, the learned speaker timbre or speech clone is used for synthesis, resulting in realistic sound effects. If there is contextual information, the synthesis is based on the guided content combined with natural language forms, resulting in rich and diverse synthesis effects, and a natural, fluent Mean Opinion Score (MOS) with high intelligibility.

[0093] Table 1

[0094] Step 7: Result presentation, including voice broadcast, text and image presentation tasks, with voice and text output processing as follows.

[0095] Step 7.1: For the video output data, an attention-based convolutional block (Att conV block) is formed by a transformer deep network model consisting of a full attention module, convolution (conV), rectified linear units (ReLU), and layer normalization (LayerNorm) to complete the decoding task. The specific implementation process is as follows: first, the full attention mechanism is used to form a unified feature vector or expression for the features themselves and between features, and then convolution operation, ReLU calculation, and LayerNorm are performed. In this way, semantic features with high-level and long-term dependent distance relationships are obtained to realize the decoding task and generate the corresponding video data.

[0096] Step 7.2: Output text data format data.

[0097] Language interaction scenarios include, but are not limited to, Chinese "speaking" voice interaction, "storing and retrieving" action interaction, "translation", "generation" such as text-to-image / text / voice / video / audio-video format food recognition, food storage and preservation, food management, dietary health, food retrieval and placement, the entire process of food preservation and health, shopping lists, diet planning, diet reminders and other interactive services and modes as shown in Table 1.

[0098] The interaction method for a refrigeration device provided in this application can be executed by an interaction device of the refrigeration device. This application uses the example of an interaction device of the refrigeration device executing the interaction method to illustrate the interaction device of the refrigeration device provided in this application.

[0099] This application also provides an interactive device for a refrigeration equipment.

[0100] like Figure 3 As shown, the interaction device of the refrigeration equipment includes: The acquisition module 310 is used to acquire user interaction data with the refrigeration equipment and spatial structure information of the refrigeration equipment. The interaction data includes at least voice interaction data. The first processing module 320 is used to obtain spatial interaction features based on interactive data and spatial structure information. The spatial interaction features are used to characterize the correlation between interactive data and spatial structure information. The second processing module 330 is used to determine and output the target interaction results between the user and the refrigeration equipment based on spatial interaction characteristics.

[0101] According to the interactive device of the refrigeration equipment provided in the embodiments of this application, by associating voice interaction data and spatial structure information, spatial interaction features are obtained. The user's interaction intention in the spatial interaction features, the area location corresponding to the interaction intention in the refrigeration equipment, and the storage status of items in the area location are analyzed to obtain the target interaction result between the user and the refrigeration equipment. By combining voice interaction data with spatial structure information, the specific item or area pointed to by the user can be accurately identified, avoiding the reduction of interaction accuracy due to misunderstanding of the user's intention.

[0102] In some embodiments, the first processing module 320 is used to acquire the target interaction location and target interaction task of the interaction data; Match the target interaction location with the spatial structure information to obtain the target spatial label of the spatial structure information; By binding the target spatial label with the target interaction task, spatial interaction features are obtained.

[0103] In some embodiments, the second processing module 330 is used to input spatial interaction features into a multimodal large model to obtain the target interaction result output by the multimodal large model.

[0104] In some embodiments, the multimodal large model includes a first system module and a second system module. The response speed of the first system module is greater than that of the second system module, and the reasoning ability of the second system module is stronger than that of the first system module. The second processing module 330 is used to reason about the target interaction result through the first system module when the complexity of the spatial interaction feature is less than a first complexity threshold. Alternatively, if the complexity of the spatial interaction features is greater than or equal to the first complexity threshold, the target interaction result can be inferred through the second system module.

[0105] In some embodiments, the second processing module 330 is further configured to construct a reward model for the multimodal large model, and score the output of the multimodal large model using the reward model to obtain a first score value; Based on the first score, initiate the reinforcement learning process of the multimodal large model; Based on the first score and the reinforcement learning factors in the reinforcement learning process, the training factors are determined. Training a large multimodal model based on training factors.

[0106] In some embodiments, the first processing module 320 is used to perform enhancement processing on the voice interaction data; Extract the first speech feature information from the enhanced speech interaction data; Based on the first speech feature information and the context data of the speech interaction data, the speech interaction data is decoded; Spatial interaction features are obtained based on the decoded voice interaction data and spatial structure information.

[0107] In some embodiments, the first processing module 320 is used to obtain the second speech feature information of the speech interaction data through temporal multi-scale convolution; Based on the second speech feature information, the speech interaction data is repaired. Based on the repaired voice interaction data, an adversarial process is performed by a generator and a discriminator to obtain third voice feature information; Noise reduction is performed on the voice interaction data based on third-party speech feature information to enhance the voice interaction data.

[0108] In some embodiments, the interactive data also includes at least one of text data, image data, and video data.

[0109] The interactive device of the cooling equipment in this application embodiment can be an electronic device or a component of an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other devices besides a terminal. For example, the electronic device can be a mobile phone, tablet computer, laptop computer, handheld computer, in-vehicle electronic device, mobile internet device (MID), augmented reality (AR) / virtual reality (VR) device, robot, wearable device, ultra-mobile personal computer (UMPC), netbook or personal digital assistant (PDA), etc. It can also be a server, network attached storage (NAS), personal computer (PC), television (TV), ATM or self-service machine, etc. The embodiments of this application do not specifically limit it.

[0110] The interactive device of the refrigeration equipment in this application embodiment can be a device with an operating system. This operating system can be Android, iOS, or other possible operating systems; this application embodiment does not specifically limit it.

[0111] The interactive device for the refrigeration equipment provided in this application embodiment can achieve... Figure 1 The various processes implemented in the method implementation examples will not be described again here to avoid repetition.

[0112] This application also provides a refrigeration device.

[0113] The refrigeration equipment includes the interactive device and data acquisition device of the aforementioned refrigeration equipment.

[0114] The data acquisition device is connected to the interaction device. The data acquisition device is used to collect the interaction data between the user and the refrigeration equipment and the spatial structure information of the refrigeration equipment.

[0115] According to the refrigeration device provided in the embodiments of this application, by associating voice interaction data and spatial structure information, spatial interaction features are obtained. The user's interaction intent in the spatial interaction features, the area location corresponding to the interaction intent in the refrigeration device, and the storage status of items in the area location are analyzed to obtain the target interaction result between the user and the refrigeration device. By combining voice interaction data with spatial structure information, the specific item or area pointed to by the user can be accurately identified, avoiding the reduction of interaction accuracy due to misunderstanding of the user's intent.

[0116] The refrigeration equipment provided in this application is not limited to refrigerators. It can be a smart refrigeration refrigerator with several layers, including constant temperature refrigerators for freezing and refrigeration, and can be equipped with human-computer interaction objects such as gestures, voice, text, audio and video, image acquisition, microphone arrays, radar signal acquisition such as images, tools such as optical character recognition (OCR) devices, documents, virtual reality (VR), augmented reality (AR), etc. The presentation forms include APP / PC / web / mini-program, large screen, etc.; supported protocols include TCP / IP, wifi, Bluetooth, etc.; supported networks include 5G, wifi, wired network, 6G, etc. The application scope or field is home, hotel, industrial field, indoor and outdoor occasions, military and aerospace, etc., and the embedding method is flat embedding, vehicle embedding, integration, etc.

[0117] This application also provides an interactive system for a refrigeration device.

[0118] The interactive system of refrigeration equipment includes the aforementioned refrigeration equipment and server.

[0119] The server is connected to the cooling equipment and is used to output the target interaction result obtained by the interactive device executing the interaction method.

[0120] According to the interactive system of the refrigeration equipment provided in the embodiments of this application, by associating voice interaction data with spatial structure information, spatial interaction features are obtained. The user's interaction intent in the spatial interaction features, the area location corresponding to the interaction intent in the refrigeration equipment, and the storage status of items in the area location are analyzed to obtain the target interaction result between the user and the refrigeration equipment. By combining voice interaction data with spatial structure information, the specific item or area pointed to by the user can be accurately identified, avoiding the reduction of interaction accuracy due to misunderstanding of the user's intent.

[0121] The interactive method for refrigeration equipment provided in this application is conceived, innovatively applied, and creatively described from the perspectives of large-scale models, speech processing, and spatial positioning technologies and their model training strategies. It describes these interdependent and interconnected, inseparable and closely related modules and processes as a coherent whole, including model pre-training and prediction stages. The application covers the system architecture design methods for each language interaction module, unsupervised, semi-supervised, and supervised multi-interaction modes, large-scale model construction and training methods, model response, intelligent agents / multi-agent systems, large-scale model fine-tuning, intelligent agent development processes (including their components), interactions, and human feedback reinforcement learning. In other words, it presents a complete and systematic solution to this problem, outlining specific steps and applications for a targeted language interaction method and system.

[0122] The interactive method for refrigeration equipment provided in this application focuses on core algorithms such as speech denoising, speech enhancement, large model training strategies, refrigerator spatial location positioning technology (from obtaining location to describing location information and its fusion feature relationship), large language interaction model and its hybrid modules, large inference model and training factors, interaction trajectory, and interaction trajectory table. It includes N-shots, prompt words, hybrid expert models, self-attention mechanisms including mutual attention and self-attention and cross-attention networks, optimization methods and system applications and devices, storage media, and methods for constructing and training large text, visual, and speech language models. It also includes application development schemes for generative content such as end-to-end encoder-decoder structure deep neural network models or LLM construction with multi-head attention only-decoder / encoder-transformer, cross-attention / group attention mechanisms, self-attention block transformers, or Kaldi or conformer, whisper, etc., to construct large-scale deep neural network models. Among them, this invention patent innovatively proposes training methods and strategies, model structures and optimizations for speech enhancement, speech denoising, large visual models, large text models, visual image and understanding and their tokenization modules, core computing, and speech and text processing technologies.

[0123] The interaction method for refrigeration equipment provided in this application is based on the construction and training of a large downstream task integrated model and the construction of an efficient large model with a transformer model, a feedforward neural network (FFN), and normalization calculation. It includes reinforcement learning, reward models, and training strategies. Here, not only can the optimization and improvement of the end-to-end LLM with the Transformer model, the Genetic Algorithm (GA) / graph / multi-head self / cross attention mechanism, diffusion network model, encoding and decoding, and other deep integrated fusion neural network models such as distillation, BiLSTM, diffusion model, pre-trained model LLM, attention mechanism, text / speech / audio / video encoding and decoding, etc. be used, which are based on rich theoretical foundations and form the complete system of this invention patent, which is the advantage of this invention. There is also a Gaussian mixture deep neural network, see Step 1-step7 for details.

[0124] The text transformer large model, training strategy method, and diffusion network model construction involved in the interaction method of the refrigeration equipment provided in this application can not only adopt semantic compression large language model LLM with decoder-ONLY, encoder-ONLY, and encoder-decoder, but also include, but are not limited to, deep network model LLM with graph neural network, attention mechanism, transformer model and its variants or improvements, distillation network LLM, latent / diffusion model LLM, U-net network and other deep network model LLM, fusion deep network model LLM such as transformer-based bidirectional long short-term memory network (BiLSTM) + multi-stream multi-scale convolutional neural network (MMCNN)-RNN, BiLSTM+MMCNN+Attention, MMCNN+BiLSTM+Attention fusion model LLM, gated recurrent unit (Gated) RecurrentUnit (GRU) + CNN, GRU + CNN + Attention model, deep reinforcement learning and reward model, RAG's LLM and other series of neural network or deep neural network models, as well as Gaussian mixture deep neural network models or similar deep network models.

[0125] The interaction method for refrigeration equipment provided in this application proposes a novel approach that integrates artificial intelligence large models, training strategies, spatial positioning, and other interdisciplinary technologies. It also fully utilizes and optimizes multiple natural language interaction methods such as voice, text, and audio-visual interaction. This method is a pioneering, novel, and creative overall and indivisible technical solution that is expected to improve the efficiency of multi-task joint application development of voice processing, spatial positioning, training data, large models, and diffusion network models, as well as its vertical fields.

[0126] In some embodiments, such as Figure 4 As shown, this application embodiment also provides an electronic device 400, including a processor 401, a memory 402, and a computer program stored in the memory 402 and executable on the processor 401. When the program is executed by the processor 401, it implements the various processes of the above-described interactive method embodiment of the cooling device and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0127] It should be noted that the electronic devices in the embodiments of this application include the mobile electronic devices and non-mobile electronic devices described above.

[0128] This application also provides a non-transitory computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it implements the various processes of the above-described interaction method embodiment of the refrigeration device and achieves the same technical effect. To avoid repetition, it will not be described again here.

[0129] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.

[0130] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the interaction method of the above-described refrigeration device.

[0131] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.

[0132] This application embodiment also provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement the various processes of the above-described interaction method embodiment of the refrigeration device, and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0133] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.

[0134] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.

[0135] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0136] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

[0137] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "illustrative embodiment," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0138] Although embodiments of this application have been shown and described, those skilled in the art will understand that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of this application, the scope of which is defined by the claims and their equivalents.

Claims

1. An interaction method for a refrigeration device, characterized in that, include: Acquire user interaction data with the refrigeration device and spatial structure information of the refrigeration device, wherein the interaction data includes at least voice interaction data; Based on the interaction data and the spatial structure information, spatial interaction features are obtained, which are used to characterize the correlation between the interaction data and the spatial structure information. Based on the spatial interaction characteristics, the target interaction result between the user and the refrigeration equipment is determined and output.

2. The interaction method of the refrigeration equipment according to claim 1, characterized in that, The process of obtaining spatial interaction features based on the interaction data and the spatial structure information includes: Obtain the target interaction location and target interaction task of the interaction data; The target interaction location is matched with the spatial structure information to obtain the target spatial label of the spatial structure information; The spatial label is bound to the target interaction task to obtain the spatial interaction feature.

3. The interaction method of the refrigeration equipment according to claim 1, characterized in that, The step of determining and outputting the target interaction result between the user and the cooling device based on the spatial interaction features includes: The spatial interaction features are input into a multimodal large model to obtain the target interaction result output by the multimodal large model.

4. The interaction method of the refrigeration equipment according to claim 3, characterized in that, The multimodal large model includes a first system module and a second system module. The response speed of the first system module is greater than that of the second system module, and the reasoning ability of the second system module is stronger than that of the first system module. The step of inputting the spatial interaction features into the multimodal large model to obtain the target interaction result output by the multimodal large model includes: If the complexity of the spatial interaction features is less than a first complexity threshold, the target interaction result is inferred through the first system module. Alternatively, if the complexity of the spatial interaction features is greater than or equal to the first complexity threshold, the target interaction result can be inferred through the second system module.

5. The interaction method of the refrigeration equipment according to claim 3, characterized in that, The multimodal large model is trained through the following steps: Construct a reward model for the multimodal large model, and score the output of the multimodal large model using the reward model to obtain a first score value; Based on the first score, the reinforcement learning process of the multimodal large model is initiated; Based on the first score and the reinforcement learning factors of the reinforcement learning process, the training factors are determined. The multimodal large model is trained based on the training factors.

6. The interaction method of the refrigeration equipment according to any one of claims 1-5, characterized in that, The process of obtaining spatial interaction features based on the interaction data and the spatial structure information includes: The voice interaction data is then processed with voice prompts and enhancements. First speech feature information is extracted from the enhanced speech interaction data; Based on the first voice feature information and the context data of the voice interaction data, the voice interaction data is decoded; The spatial interaction features are obtained based on the decoded voice interaction data and the spatial structure information.

7. The interaction method of the refrigeration equipment according to claim 6, characterized in that, The enhancement processing of the voice interaction data includes: The second speech feature information of the speech interaction data is obtained by multi-scale convolution in the temporal domain; Based on the second voice feature information, the voice interaction data is repaired. Based on the repaired voice interaction data, a third voice feature information is obtained by performing a stable adversarial process through the generator and discriminator and the loss function convergence. The voice interaction data is subjected to noise reduction processing based on the third voice feature information in order to enhance the voice interaction data.

8. The interaction method of the refrigeration equipment according to any one of claims 1-5, characterized in that, The interactive data also includes at least one of text data, image data, and video data.

9. An interactive device for a refrigeration equipment, characterized in that, include: The acquisition module is used to acquire user interaction data with the refrigeration device and spatial structure information of the refrigeration device, wherein the interaction data includes at least voice interaction data. The first processing module is used to obtain spatial interaction features based on the interaction data and the spatial structure information, wherein the spatial interaction features are used to characterize the correlation between the interaction data and the spatial structure information; The second processing module is used to determine and output the target interaction result between the user and the refrigeration equipment based on the spatial interaction characteristics.

10. A refrigeration device, characterized in that, include: The interactive device for the refrigeration equipment as described in claim 9; A data acquisition device is connected to the interaction device, and the data acquisition device is used to collect interaction data between the user and the refrigeration equipment and spatial structure information of the refrigeration equipment.

11. An interactive system for a refrigeration device, characterized in that, include: The refrigeration equipment as described in claim 10; The server is connected to the refrigeration equipment and is used to output the target interaction result obtained by the interaction device executing the interaction method.