Home appliances and their control methods

CN119511833BActive Publication Date: 2026-04-03WUHU MIDEA KITCHEN & BATH APPLIANCES MFG CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-31
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing smart home systems struggle to achieve complex cross-device scene linkages, resulting in insufficient scene-based control capabilities.

Method used

By acquiring user commands, using a large language model for scene recognition and task decomposition, and assigning scene agents to control corresponding device agents, cross-device scene linkage is achieved.

Benefits of technology

It enables complex cross-device scenario linkage, enhances scenario-based control capabilities, and improves the intelligent and personalized management of home appliances.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119511833B_ABST
    Figure CN119511833B_ABST
Patent Text Reader

Abstract

This application discloses a home appliance and its control method. The method includes: acquiring user commands; using a large language model to perform scene recognition and task decomposition on the user commands to obtain at least one scene corresponding to the user commands, and a task corresponding to each scene; wherein each scene corresponds to at least one home appliance, and each scene corresponds to a scene intelligent agent; using the scene intelligent agent to assign a corresponding device intelligent agent to each task, so that the device intelligent agent controls the corresponding home appliance; wherein each device intelligent agent is deployed in the corresponding home appliance. Through the above method, complex cross-device scene linkage can be achieved, improving scene-based control capabilities.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of household appliances, in particular to household appliances and their control methods. Background Art

[0002] With the rapid development of artificial intelligence and Internet of Things technologies, smart home systems are evolving towards greater intelligence, personalization, and adaptability. However, related smart home systems are difficult to achieve complex cross-device scenario linkages. Summary of the Invention

[0003] The household appliances and their control methods provided by this application can achieve complex cross-device scenario linkages and enhance the scenario control ability.

[0004] In a first aspect, this application provides a method for controlling household appliances. The method includes: obtaining a user instruction; using a large language model to perform scenario recognition and task decomposition on the user instruction to obtain at least one scenario corresponding to the user instruction and tasks corresponding to each scenario; wherein, each scenario corresponds to at least one household appliance, and each scenario corresponds to a scenario agent; using the scenario agent to allocate a corresponding device agent to each task so that the device agent controls the corresponding household appliance; wherein, each device agent is deployed in the corresponding household appliance.

[0005] Wherein, using the scenario agent to allocate a corresponding device agent to each task so that the device agent controls the corresponding household appliance includes: based on each task, performing scenario collaboration between at least two scenario agents to obtain an execution plan corresponding to each device agent; and allocating the execution plan to the corresponding device agent so that the device agent controls the corresponding household appliance.

[0006] Wherein, the at least two scenario agents include a first scenario agent and a second scenario agent. Based on each task, performing scenario collaboration between the at least two scenario agents includes: in response to a conflict existing between the execution plans formulated by the first scenario agent and the second scenario agent, coordinating the first scenario agent and the second scenario agent to reformulate the execution plan; and / or, in response to the first execution plan formulated by the first scenario agent and the second execution plan formulated by the second scenario agent corresponding to the same target household appliance, synthesizing the first execution plan and the second execution plan to obtain a third execution plan for the device agent corresponding to the target household appliance; and / or, in response to the first scenario agent formulating a fourth execution plan, using the second scenario agent to adjust the fourth execution plan.

[0007] Wherein, the scenarios include at least one of a living scenario, a work and study scenario, an entertainment scenario, a sleep scenario, a cooking scenario, a health scenario, a water usage scenario, and an environmental control scenario.

[0008] Specifically, the large language model is used to perform scene recognition and task decomposition on user commands to obtain at least one scene corresponding to the user command and the task corresponding to each scene. This includes: using the large language model to perform scene recognition and task decomposition on user commands to obtain work and study scene, living scene and health scene corresponding to the user command; assigning a first task to the work and study scene, a second task to the living scene and a third task to the health scene.

[0009] The scene intelligence agent includes a first scene intelligence agent, a second scene intelligence agent, and a third scene intelligence agent. Each task is assigned a corresponding device intelligence agent using the scene intelligence agents, including: using the first scene intelligence agent to formulate a first execution plan based on the first task, and determining the corresponding device intelligence agent based on the first execution plan; the first execution plan includes lighting control and / or audio control; using the second scene intelligence agent to formulate a second execution plan based on the second task, and determining the corresponding device intelligence agent based on the second execution plan; the second execution plan includes air conditioning control and / or curtain control; using the third scene intelligence agent to formulate a third execution plan based on the third task, and determining the corresponding device intelligence agent based on the third execution plan; the third execution plan includes exercise reminders and / or air purification control.

[0010] Specifically, the large language model is used to perform scene recognition and task decomposition on user commands to obtain at least one scene corresponding to the user command and a task corresponding to each scene. This includes: using the large language model to perform scene recognition and task decomposition on user commands to obtain cooking scene, living scene and health scene corresponding to the user command; assigning a fourth task to the cooking scene, a fifth task to the living scene and a sixth task to the health scene.

[0011] The scenario-based intelligent agents include a fourth scenario-based intelligent agent, a fifth scenario-based intelligent agent, and a sixth scenario-based intelligent agent. Each task is assigned a corresponding device intelligent agent using these scenario-based intelligent agents, including: using the fourth scenario-based intelligent agent to formulate a fourth execution plan based on the fourth task, and determining the corresponding device intelligent agent based on the fourth execution plan; the fourth execution plan includes food identification control, menu recommendation, and / or oven control; using the fifth scenario-based intelligent agent to formulate a fifth execution plan based on the fifth task, and determining the corresponding device intelligent agent based on the fifth execution plan; the fifth execution plan includes lighting control and / or air conditioning control; using the sixth scenario-based intelligent agent to formulate a sixth execution plan based on the sixth task, and determining the corresponding device intelligent agent based on the sixth execution plan; the sixth execution plan includes data analysis control and / or dietary recommendations.

[0012] Specifically, the large language model is used to perform scene recognition and task decomposition on user commands to obtain at least one scene corresponding to the user command and a task corresponding to each scene. This includes: using the large language model to perform scene recognition and task decomposition on user commands to obtain a water use scene, an environmental control scene, and a health scene corresponding to the user command; assigning a seventh task to the water use scene, an eighth task to the environmental control scene, and a ninth task to the health scene.

[0013] The scenario-based intelligent agents include a seventh, eighth, and ninth scenario-based intelligent agent. Each task is assigned a corresponding device intelligent agent using these scenario-based intelligent agents. This includes: using the seventh scenario-based intelligent agent to formulate a seventh execution plan based on the seventh task, and determining the corresponding device intelligent agent based on the seventh execution plan; the seventh execution plan includes water temperature regulation, water quality regulation, and / or water pressure regulation; using the eighth scenario-based intelligent agent to formulate an eighth execution plan based on the eighth task, and determining the corresponding device intelligent agent based on the eighth execution plan; the eighth execution plan includes lighting control and / or air conditioning control; using the ninth scenario-based intelligent agent to formulate a ninth execution plan based on the ninth task, and determining the corresponding device intelligent agent based on the ninth execution plan; the ninth execution plan includes data analysis control and / or dietary recommendations.

[0014] Specifically, after a scene ends, the scene agent assigns a scene end task to the device agent, enabling the device agent to control the corresponding home appliance to complete the scene end task.

[0015] Secondly, this application provides a control system for a home appliance, the control system comprising: a communication interface for communicating with the home appliance; and a processor coupled to the communication interface for implementing the method provided in the first aspect.

[0016] Thirdly, this application provides a method for controlling a household appliance, the method comprising: receiving a control command; wherein the control command is obtained by the method provided in the first aspect; and operating in accordance with the control command.

[0017] Fourthly, this application provides a home appliance that includes a communication interface and a processor, the processor being coupled to the communication interface for implementing the method provided in the first aspect.

[0018] The beneficial effects of this application are as follows: Unlike existing technologies, the home appliances and their control methods and systems provided in this application acquire user commands; utilize a large language model to perform scene recognition and task decomposition on the user commands, obtaining at least one scene corresponding to the user command, and a task corresponding to each scene; wherein each scene corresponds to at least one home appliance, and each scene corresponds to one scene intelligent agent; the scene intelligent agent is used to assign a corresponding device intelligent agent to each task, so that the device intelligent agent controls the corresponding home appliance; wherein each device intelligent agent is deployed in the corresponding home appliance, and by setting scene intelligent agents, scene management of a large number of home appliances can be achieved, enabling complex cross-device scene linkage and improving scene-based control capabilities. Attached Figure Description

[0019] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. Wherein:

[0020] Figure 1 This is a schematic diagram of the structure of an embodiment of the control system for a household appliance provided in this application;

[0021] Figure 2 This is a flowchart illustrating an embodiment of the multimodal data processing method provided in this application;

[0022] Figure 3 This is a flowchart illustrating another embodiment of the multimodal data processing method provided in this application;

[0023] Figure 4 This is a flowchart illustrating another embodiment of the multimodal data processing method provided in this application;

[0024] Figure 5 This is a flowchart illustrating another embodiment of the multimodal data processing method provided in this application;

[0025] Figure 6 This is a flowchart illustrating another embodiment of the multimodal data processing method provided in this application;

[0026] Figure 7 This application provides Figure 6 A flowchart illustrating an embodiment of step 62;

[0027] Figure 8 This is a flowchart illustrating another embodiment of the multimodal data processing method provided in this application;

[0028] Figure 9This is a flowchart illustrating an embodiment of the control method for home appliances provided in this application;

[0029] Figure 10 This is a schematic flowchart of another embodiment of the control method for home appliances provided in this application;

[0030] Figure 11 This is a schematic flowchart of another embodiment of the control method for home appliances provided in this application;

[0031] Figure 12 This is a schematic flowchart of another embodiment of the control method for home appliances provided in this application;

[0032] Figure 13 This is a schematic flowchart of an embodiment of the home appliance control method provided in this application;

[0033] Figure 14 This is a schematic flowchart of another embodiment of the home appliance control method provided in this application;

[0034] Figure 15 This is a schematic flowchart of another embodiment of the home appliance control method provided in this application;

[0035] Figure 16 This is a schematic flowchart of another embodiment of the home appliance control method provided in this application;

[0036] Figure 17 This is a schematic flowchart of another embodiment of the home appliance control method provided in this application;

[0037] Figure 18 This is a schematic diagram of another embodiment of the control system for the home appliance provided in this application;

[0038] Figure 19 This is a schematic diagram of the structure of an embodiment of the home appliance provided in this application. Detailed Implementation

[0039] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. It is understood that the specific embodiments described herein are only for explaining this application and not for limiting it. Furthermore, it should be noted that, for ease of description, only the parts related to this application are shown in the accompanying drawings, not all structures. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.

[0040] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0041] See Figure 1 , Figure 1 This is a schematic diagram of an embodiment of the control system for a home appliance provided in this application. The control system 1000 includes a communication interface 100 and a processor 200. The communication interface 100 is used to communicate with the home appliance. The processor 200 is coupled to the communication interface 100 and is used to implement corresponding methods based on relevant data from the home appliance. Specific implementation methods are described in any of the following embodiments.

[0042] See Figure 2 , Figure 2 This is a flowchart illustrating an embodiment of the multimodal data processing method provided in this application. The method includes:

[0043] Step 21: Acquire multimodal data.

[0044] The multimodal data includes at least one of text data, image data, and voice data, as well as sensor data; the sensor data is collected by the sensors corresponding to the home appliances.

[0045] In some embodiments, text data may be user-inputted commands and / or historical interaction records, etc. Image data may be environmental images captured by a camera on the home appliance, user gestures, etc. Voice data may be user voice commands and / or ambient sounds, etc. Sensor data may be data from various sensors such as temperature, humidity, light intensity, and energy consumption.

[0046] In some embodiments, multimodal data includes text data and sensor data.

[0047] In some embodiments, multimodal data includes image data and sensor data.

[0048] In some embodiments, multimodal data includes voice data and sensor data.

[0049] In some embodiments, multimodal data includes text data, voice data, and sensor data.

[0050] In some embodiments, multimodal data includes text data, image data, and sensor data.

[0051] In some embodiments, multimodal data includes voice data, image data, and sensor data.

[0052] In some embodiments, multimodal data includes text data, image data, voice data, and sensor data.

[0053] In some embodiments, sensor data can be temperature data, humidity data, or light intensity data. That is, corresponding data is collected using appropriate sensors. For example, a temperature sensor is used to collect temperature data, a humidity sensor is used to collect humidity data, and a light intensity sensor is used to collect light intensity data.

[0054] Step 22: Use at least two fusion methods to fuse the multimodal data to obtain multimodal fused data with a unified semantic space.

[0055] In some embodiments, the fusion method can be selected according to the specific type of multimodal data, and then the corresponding fusion method can be used to fuse the multimodal data to obtain multimodal fused data with a unified semantic space. That is, by fusing multimodal data, the multimodal data can be aligned in the semantic space, which enhances the correlation between multimodal data, is more conducive to the recognition of subsequent large language models, and improves the inference accuracy of large language models.

[0056] In this embodiment, at least two fusion methods are used to fuse multimodal data containing sensor data, resulting in multimodal fused data with a unified semantic space. This enables effective fusion of multimodal data in the home appliance field, breaking down data silos. Furthermore, precise alignment of different modal data within the unified semantic space effectively enhances the understanding capabilities of a large language model after inputting the multimodal fused data. Consequently, it significantly improves the generalization ability and robustness of the large language model in home appliance scenarios. This leads to a more natural and intelligent human-computer interaction experience during home appliance control, improving user satisfaction and enhancing the accuracy and efficiency of intelligent home appliance control, while reducing misoperation and energy waste.

[0057] See Figure 3 , Figure 3 This is a flowchart illustrating another embodiment of the multimodal data processing method provided in this application. The method includes:

[0058] Step 31: Acquire multimodal data; wherein, multimodal data includes at least text data and sensor data.

[0059] Sensor data is collected by the sensors corresponding to the home appliances.

[0060] Step 32: Preprocess the sensor data in the multimodal data to obtain the first vector representation; and preprocess the text data in the multimodal data to obtain the text vector.

[0061] In some embodiments, text data can be segmented, stop words removed, and a pre-trained word embedding model can be used to convert the text into vectors to obtain text vectors.

[0062] In some embodiments, text data can be segmented, stop words removed, and a pre-trained word embedding model can be used to convert the text into vectors to obtain text vectors.

[0063] In some embodiments, the sensor data can be normalized, time series features extracted, and dimensionality reduction techniques such as autoencoders or PCA (Principal Component Analysis) can be used to convert the sensor data into a low-dimensional vector representation to obtain a first vector representation.

[0064] Step 33: Using at least two fusion methods, perform data fusion on the text vector and the first vector representation to obtain multimodal fused data with a unified semantic space.

[0065] In this embodiment, at least two fusion methods are used to fuse multimodal data, including at least text data and sensor data, to obtain multimodal fused data with a unified semantic space. This enables effective fusion of multimodal data in the home appliance field, breaking down data silos. Furthermore, precise alignment of different modal data is achieved within the unified semantic space. After inputting the multimodal fused data into a large language model, the understanding ability of the large language model is effectively improved, thereby significantly enhancing its generalization ability and robustness in home appliance scenarios. Consequently, a more natural and intelligent human-computer interaction experience is achieved during home appliance control, improving user satisfaction and also enhancing the accuracy and efficiency of intelligent home appliance control, reducing misoperation and energy waste.

[0066] See Figure 4 , Figure 4 This is a flowchart illustrating another embodiment of the multimodal data processing method provided in this application. The method includes:

[0067] Step 41: Acquire multimodal data; where multimodal data includes image data and sensor data.

[0068] Sensor data is collected by the sensors corresponding to the home appliances.

[0069] Step 42: Preprocess the sensor data in the multimodal data to obtain the first vector representation; preprocess the image data in the multimodal data to obtain image features.

[0070] In some embodiments, image data can be subjected to operations such as image scaling, cropping, and data augmentation, and then pre-trained convolutional neural networks can be used to extract image features.

[0071] In some embodiments, the sensor data can be normalized, time series features extracted, and dimensionality reduction techniques such as autoencoders or PCA can be used to convert the sensor data into a low-dimensional vector representation to obtain a first vector representation.

[0072] Step 43: Using at least two fusion methods, perform data fusion on the image features and the first vector representation to obtain multimodal fused data with a unified semantic space.

[0073] In this embodiment, at least two fusion methods are used to fuse multimodal data, including image data and sensor data, to obtain multimodal fused data with a unified semantic space. This enables effective fusion of multimodal data in the home appliance field, breaking down data silos. Furthermore, precise alignment of different modal data is achieved within the unified semantic space. After inputting the multimodal fused data into a large language model, the understanding ability of the large language model is effectively improved, thereby significantly enhancing its generalization ability and robustness in home appliance scenarios. This leads to a more natural and intelligent human-computer interaction experience during home appliance control, improving user satisfaction, and also enhancing the accuracy and efficiency of intelligent home appliance control, reducing misoperation and energy waste.

[0074] See Figure 5 , Figure 5 This is a flowchart illustrating another embodiment of the multimodal data processing method provided in this application. The method includes:

[0075] Step 51: Acquire multimodal data; wherein, multimodal data includes at least voice data and sensor data.

[0076] Sensor data is collected by the sensors corresponding to the home appliances.

[0077] Step 52: Preprocess the sensor data in the multimodal data to obtain the first vector representation; preprocess the speech data in the multimodal data to obtain the second vector representation.

[0078] In some embodiments, the speech signal can be preprocessed to extract MFCC features, and a pre-trained speech recognition model can be used to convert the speech into text, which is then converted into a vector representation. This results in the second vector representation described above.

[0079] In some embodiments, the sensor data can be normalized, time series features extracted, and dimensionality reduction techniques such as autoencoders or PCA can be used to convert the sensor data into a low-dimensional vector representation to obtain a first vector representation.

[0080] Step 53: Using at least two fusion methods, fuse the second vector representation and the first vector representation to obtain multimodal fused data with a unified semantic space.

[0081] In this embodiment, at least two fusion methods are used to fuse multimodal data, including at least voice data and sensor data, to obtain multimodal fused data with a unified semantic space. This enables effective fusion of multimodal data in the home appliance field, breaking down data silos. Furthermore, precise alignment of different modal data is achieved within the unified semantic space. After inputting the multimodal fused data into a large language model, the understanding ability of the large language model is effectively improved, thereby significantly enhancing its generalization ability and robustness in home appliance scenarios. Consequently, a more natural and intelligent human-computer interaction experience is achieved during home appliance control, improving user satisfaction and also enhancing the accuracy and efficiency of intelligent home appliance control, reducing misoperation and energy waste.

[0082] See Figure 6 , Figure 6 This is a flowchart illustrating another embodiment of the multimodal data processing method provided in this application. The method includes:

[0083] Step 61: Acquire multimodal data; wherein, multimodal data includes at least text data, image data, voice data, and sensor data.

[0084] Step 62: Use at least two fusion methods to fuse the multimodal data to obtain multimodal fused data with a unified semantic space.

[0085] Among them, at least two fusion methods include: attention mechanism fusion method; weighted fusion method; and splicing fusion method.

[0086] Attention mechanisms are resource allocation schemes designed to allocate limited computational resources to more important tasks. They originate from research on human vision, which shows that humans selectively focus on certain information while ignoring others. There are two main types of attention mechanisms: focused attention and saliency-based attention.

[0087] How attention mechanisms work: In deep learning, attention mechanisms dynamically generate weights for different connections, allowing the model to selectively focus on different parts of the input sequence. For example, in the Transformer model, the degree of attention given to each input vector is dynamically determined by calculating the correlation between each input vector and other vectors. This mechanism can be applied to fields such as natural language processing and computer vision to help models better capture key information.

[0088] That is, attention mechanisms can be used to quickly capture key information from different modalities in multimodal data, and then fuse the key information.

[0089] Weighted fusion allows for the setting of weights for data from different modalities, and then the modal data is fused according to the corresponding weights. For example, data from different modalities can be processed to unify dimensions, and then the modal data can be fused according to the corresponding weights after dimension unification.

[0090] The concatenation and fusion method can combine data from different modalities. In some embodiments, the priorities of different modalities can be set, and concatenation is performed according to priority. For example, higher-priority data is concatenated first, and lower-priority data is concatenated later. That is, data is sorted from high to low priority and then concatenated. This allows for faster extraction of higher-priority data during subsequent data extraction, thereby improving the accuracy of high-priority data and facilitating more accurate feature acquisition and semantic information understanding of the modality by the subsequent large language model.

[0091] In some embodiments, different modal data can be initially fused using corresponding fusion methods, and then the initially fused data can be fused again to obtain multimodal fused data.

[0092] In some embodiments, see Figure 7 Step 62 can be the following process:

[0093] Step 621: Use the attention mechanism to fuse the multimodal data to obtain the first fused data.

[0094] In some embodiments, an attention mechanism is used to extract value vectors, key vectors, and query vectors from multimodal data. Then, the correlation between the key vectors and query vectors is calculated using the key vectors and query vectors. An attention score is then obtained using a softmax function, and a weighted sum is performed based on the attention scores to obtain a value vector with the attention score. This means that the first fused data can reflect the degree of matching between different modalities, thereby achieving data alignment.

[0095] In some embodiments, the attention mechanism may include a self-attention mechanism, a cross-attention mechanism, and / or a multi-head attention mechanism.

[0096] Step 622: Use a weighted fusion method to fuse the multimodal data to obtain the second fused data.

[0097] Step 623: Use a splicing and fusion method to fuse the multimodal data to obtain the third fused data.

[0098] In some embodiments, the weighted fusion method and the splicing fusion method may employ residual connection and layer normalization techniques.

[0099] Residual connections refer to the practice in deep neural networks of directly passing the input from one layer to the next, essentially adding cross-layer connections. This method can alleviate the vanishing and exploding gradient problems, allowing for greater effective depth in deep neural networks and thus improving network performance.

[0100] Layer normalization is a normalization technique that, compared to batch normalization, is more suitable for RNNs and long-term sequences. Its basic idea is to normalize all neurons in each sample of the neural network, that is, to normalize the mean and variance of neurons with respect to that sample, thereby reducing the dependencies between neurons, avoiding model instability caused by changes in batch size, effectively eliminating correlations between layers, and improving the convergence and performance of the network. Compared to batch normalization, layer normalization is more stable and more robust to noise and small batch sizes.

[0101] Step 624: Merge the first fused data, the second fused data, and the third fused data to obtain multimodal fused data.

[0102] In some embodiments, a fusion representation generation technique can be used to fuse first fused data, second fused data, and second fused data to obtain multimodal fused data.

[0103] In this embodiment, at least two fusion methods are used to fuse multimodal data, including text data, image data, voice data, and sensor data, to obtain multimodal fused data with a unified semantic space. This enables effective fusion of multimodal data in the home appliance field, breaking down data silos. Furthermore, precise alignment of different modal data is achieved within the unified semantic space. After inputting the multimodal fused data into a large language model, the understanding ability of the large language model is effectively improved, thereby significantly enhancing its generalization ability and robustness in home appliance scenarios. This leads to a more natural and intelligent human-computer interaction experience during home appliance control, improving user satisfaction, and also enhancing the accuracy and efficiency of intelligent home appliance control, reducing misoperation and energy waste.

[0104] See Figure 8 , Figure 8 This is a flowchart illustrating another embodiment of the multimodal data processing method provided in this application. The method includes:

[0105] Step 81: Obtain multimodal data.

[0106] The multimodal data includes at least one of text data, image data, and voice data, as well as sensor data; the sensor data is collected by the sensors corresponding to the home appliances.

[0107] Step 82: Use at least two fusion methods to fuse the multimodal data to obtain multimodal fused data with a unified semantic space.

[0108] In some embodiments, steps 81 to 82 have the same or similar technical solutions as any embodiment of this application, and will not be described in detail here.

[0109] Step 83: Input the multimodal fusion data into the large language model to obtain the contrastive loss and multi-task learning loss.

[0110] In some embodiments, the large language model can use a pre-trained GPT model as its core and fine-tune it for the home appliance domain. Key features include:

[0111] 1. It uses the Transformer architecture and has powerful sequence modeling capabilities.

[0112] 2. Conduct pre-training on a large-scale corpus of language related to the home appliance industry to master domain knowledge.

[0113] 3. Supports multimodal input and can process fused feature representations.

[0114] 4. It adopts an autoregressive generation method to achieve flexible text generation and dialogue capabilities.

[0115] In some embodiments, multi-task learning includes at least two of anomaly detection tasks, energy consumption optimization tasks, behavior prediction tasks, state estimation tasks, and instruction understanding tasks.

[0116] Anomaly detection tasks primarily rely on multimodal fusion data to detect anomalies in the operating status of home appliances. For example, anomaly detection can be performed by extracting sensor data and the status data of home appliances from the multimodal fusion data.

[0117] The main purpose of energy optimization is to control home appliances to achieve the corresponding tasks with the lowest energy consumption, thereby optimizing the energy consumption of home appliances.

[0118] The main purpose of behavior prediction is to predict user behavior using voice and text data from multimodal fusion data, so as to control corresponding home appliances based on the predicted behavior.

[0119] State estimation tasks primarily rely on multimodal fusion data to estimate the operating state of home appliances. For example, sensor data can be extracted from the multimodal fusion data to estimate the state of home appliances, and then the corresponding estimation results can be provided.

[0120] The main purpose of the instruction understanding task is to understand the speech and text data in the multimodal fusion data so that the corresponding home appliances can be controlled according to the understood instructions.

[0121] Step 84: Perform joint optimization based on contrastive loss and multi-task learning loss, and adjust the weights of relevant data in the multimodal data.

[0122] In some embodiments, the contrastive loss and the multi-task learning loss can be weighted and summed to obtain the final total loss, which can then be used to adjust the weights of relevant data in the multimodal data.

[0123] In this embodiment, at least two fusion methods are used to fuse multimodal data containing sensor data, resulting in multimodal fused data with a unified semantic space. This enables effective fusion of multimodal data in the home appliance field, breaking down data silos. Furthermore, precise alignment of different modal data within the unified semantic space effectively enhances the understanding capabilities of a large language model after inputting the multimodal fused data. Consequently, it significantly improves the generalization ability and robustness of the large language model in home appliance scenarios. This leads to a more natural and intelligent human-computer interaction experience during home appliance control, improving user satisfaction and enhancing the accuracy and efficiency of intelligent home appliance control, while reducing misoperation and energy waste.

[0124] See Figure 9 , Figure 9 This is a schematic flowchart of an embodiment of the control method for home appliances provided in this application. The method includes:

[0125] Step 91: Obtain multimodal fusion data.

[0126] The multimodal fusion data is obtained using the method provided in any embodiment of this application.

[0127] Step 92: Input the multimodal fusion data into the large language model to obtain the multi-task decision output by the large language model.

[0128] Among them, the large language model is a large language model that has been trained in advance.

[0129] Step 93: Control the operation of the corresponding home appliances based on multi-task decision-making.

[0130] In some embodiments, a corresponding intelligent agent is provided in the home appliance. This intelligent agent can generate corresponding control commands based on the corresponding task decisions, thereby controlling the operation of the corresponding home appliance.

[0131] In some embodiments, the intelligent agent possesses autonomous decision-making capabilities. For instance, if a multi-task decision determines that home appliances need to be turned on, the intelligent agent can autonomously decide the activation mode of the appliances based on current environmental data. Taking an air conditioner as an example, when the multi-task decision requires turning on the air conditioner, a command to turn on the air conditioner is sent to the intelligent agent. The intelligent agent determines the activation temperature based on the current environmental data, and after determining the activation temperature, controls the air conditioner to start and operates according to that temperature.

[0132] In this embodiment, the fused data is input into the large language model, which effectively enhances the model's understanding ability and improves the accuracy of its multi-task decision-making output. This significantly improves the model's generalization ability and robustness in home appliance scenarios. Consequently, it enables a more natural and intelligent human-computer interaction experience during home appliance control, increasing user satisfaction and improving the accuracy and efficiency of intelligent home appliance control, while reducing misoperation and energy waste.

[0133] See Figure 10 , Figure 10 This is a schematic flowchart of another embodiment of the control method for home appliances provided in this application. The method includes:

[0134] Step 101: Obtain multimodal fusion data.

[0135] In some embodiments, the multimodal fusion data is obtained using the method provided in any embodiment of this application.

[0136] Step 102: Input the multimodal fusion data into the large language model to obtain the multi-task decision output by the large language model.

[0137] Step 103: Perform a security check on the multi-task decision-making process.

[0138] In some embodiments, the security check primarily examines whether the appliance control scheme corresponding to the multi-tasking decision is safe during execution, such as whether there is a risk of execution anomalies.

[0139] Step 104: After the safety check is passed, control the corresponding home appliances to work according to the multi-task decision; or, if the safety check fails, correct the multi-task decision.

[0140] After the safety check is passed, the corresponding home appliances are controlled to operate based on multi-task decision-making. That is, after the safety check is passed, the task decision is sent to the corresponding home appliances. The home appliances are equipped with corresponding intelligent agents, which can generate corresponding control commands based on the corresponding task decisions, and then control the operation of the corresponding home appliances.

[0141] If the safety check fails, the multi-task decision is revised to ensure it passes the safety check. Once the safety check is passed, the corresponding home appliances are controlled to operate based on the revised multi-task decision.

[0142] In this embodiment, the fused data is input into the large language model, which effectively enhances the model's understanding ability and improves the accuracy of its multi-task decision-making output. This significantly improves the model's generalization ability and robustness in home appliance scenarios. Consequently, it enables a more natural and intelligent human-computer interaction experience during home appliance control, increasing user satisfaction and improving the accuracy and efficiency of intelligent home appliance control, while reducing misoperation and energy waste.

[0143] Furthermore, by performing safety checks on multi-task decisions, the task decisions are sent to the corresponding home appliances after the checks are passed, or the multi-task decisions are corrected when the safety checks fail, thereby reducing the occurrence of sending incorrect task decisions to home appliances and improving the control accuracy of home appliances.

[0144] See Figure 11 , Figure 11 This is a schematic flowchart of another embodiment of the control method for home appliances provided in this application. The method includes:

[0145] Step 111: Obtain multimodal fusion data.

[0146] In some embodiments, the multimodal fusion data is obtained using the method provided in any embodiment of this application.

[0147] Step 112: Input the multimodal fusion data into the large language model to obtain the multi-task decision output by the large language model.

[0148] Step 113: Control the operation of the corresponding home appliances based on multi-task decision-making.

[0149] In some embodiments, steps 111 to 113 have the same or similar technical solutions as any embodiment of this application, and will not be described in detail here.

[0150] Step 114: Record the working results of home appliances and collect user feedback information.

[0151] Step 115: Update the large language model based on the work results and user feedback.

[0152] In this embodiment, updating the large language model by combining work results and user feedback can improve the accuracy of the large language model in task reasoning, and make the reasoning of the large language model more in line with user preferences, thereby improving the user experience.

[0153] See Figure 12 , Figure 12 This is a schematic flowchart of another embodiment of the control method for home appliances provided in this application. The method includes:

[0154] Step 121: Obtain multimodal fusion data.

[0155] The multimodal fusion data is obtained using the method provided in any embodiment of this application.

[0156] Step 122: Input the multimodal fusion data into the large language model so that the large language model can determine at least two task types.

[0157] In some embodiments, the large language model undergoes multi-task training during the training process, thus enabling multi-task inference. When multimodal fusion data representation requires the implementation of multiple tasks, the large language model can obtain at least two task types based on the multimodal fusion data.

[0158] Step 123: Output the task decision control of the corresponding home appliances according to at least two task types.

[0159] Step 124: Control the operation of the corresponding home appliances based on multi-task decision-making.

[0160] In this embodiment, the fused data is input into the large language model, which effectively enhances the model's understanding ability and improves the accuracy of its multi-task decision-making output. This significantly improves the model's generalization ability and robustness in home appliance scenarios. Consequently, it enables a more natural and intelligent human-computer interaction experience during home appliance control, increasing user satisfaction and improving the accuracy and efficiency of intelligent home appliance control, while reducing misoperation and energy waste.

[0161] Furthermore, large language models can handle more complex multimodal fusion data, identify at least two task types from the multimodal fusion data, and then generate corresponding task decisions according to the task types to control the operation of home appliances.

[0162] In the above embodiments, the effective integration of multimodal data in the field of household appliances is achieved, breaking data silos; precise alignment of different modal data is realized in a unified semantic space, enhancing the model's understanding ability; significantly improving the generalization ability and robustness of the model in the household appliance scenario; realizing a more natural and intelligent human-computer interaction experience, enhancing user satisfaction; improving the accuracy and efficiency of intelligent control of household appliances, reducing misoperations and energy waste.

[0163] With the rapid development of artificial intelligence and Internet of Things technologies, smart home systems are evolving towards greater intelligence, personalization, and adaptability. However, related smart home systems are difficult to achieve complex cross-device scenario linkages.

[0164] Based on this, the present application proposes to obtain a user instruction; use a large language model to perform scenario recognition and task decomposition on the user instruction to obtain at least one scenario corresponding to the user instruction and the tasks corresponding to each scenario; wherein, each scenario corresponds to at least one household appliance device, and each scenario corresponds to a scenario agent; use the scenario agent to assign a corresponding device agent to each task so that the device agent controls the corresponding household appliance device; wherein, the manner in which each device agent is deployed in the corresponding household appliance device enables the realization of scenario management for a large number of household appliance devices by setting up the scenario agent, and can achieve complex cross-device scenario linkages, enhancing the scenario-based control ability. For specific reference, see any of the following embodiments.

[0165] See Figure 13 , Figure 13 is a schematic flowchart of an embodiment of the household appliance device control method provided by the present application. The method includes:

[0166] Step 131: Obtain a user instruction.

[0167] In some embodiments, the user instruction may be a voice instruction and / or a text instruction.

[0168] In some embodiments, the user instruction may express the user's needs. And the user's needs may require the cooperation of household appliance devices within at least one scenario.

[0169] Step 132: Use a large language model to perform scenario recognition and task decomposition on the user instruction to obtain at least one scenario corresponding to the user instruction and the tasks corresponding to each scenario.

[0170] Among them, each scenario corresponds to at least one household appliance device, and each scenario corresponds to a scenario agent.

[0171] In some embodiments, the large language model can learn scene recognition capabilities by combining historical data. Then, the trained large language model can be used to perform scene recognition and task decomposition on user commands to obtain at least one scene corresponding to the user command, and the task corresponding to each scene.

[0172] In some embodiments, multiple home appliances in different scenarios typically need to work together to fulfill user needs.

[0173] In some embodiments, a large language model is used to perform scene recognition and task decomposition on user commands to obtain at least two scenes corresponding to the user commands, and a task corresponding to each scene; wherein each scene corresponds to at least two home appliances, and each scene corresponds to one scene agent. That is, one scene agent can manage at least two home appliances.

[0174] In some embodiments, the scenarios involved in this application can be set in advance by the user or obtained through self-learning of a large language model.

[0175] Step 133: Use scene intelligence to assign a corresponding device intelligence to each task so that the device intelligence can control the corresponding home appliances.

[0176] Each device intelligence agent is deployed in its corresponding home appliance.

[0177] In this embodiment, corresponding scene intelligent agents can be set for several home appliances corresponding to each scenario, and then the scene intelligent agents can be used as control relay stations. The scene intelligent agents are used to assign corresponding device intelligent agents to each task, so that the device intelligent agents can control the corresponding home appliances.

[0178] In this embodiment, user instructions are acquired; a large language model is used to perform scene recognition and task decomposition on the user instructions to obtain at least one scene corresponding to the user instructions, and a task corresponding to each scene; wherein each scene corresponds to at least one home appliance, and each scene corresponds to a scene intelligent agent; the scene intelligent agent is used to assign a corresponding device intelligent agent to each task, so that the device intelligent agent controls the corresponding home appliance; wherein each device intelligent agent is deployed in the corresponding home appliance, and scene management of a large number of home appliances is achieved by setting scene intelligent agents, which can realize complex cross-device scene linkage and improve scene-based control capabilities.

[0179] See Figure 14 , Figure 14 This is a schematic flowchart of another embodiment of the home appliance control method provided in this application. The method includes:

[0180] Step 141: Obtain user instructions.

[0181] Step 142: Use a large language model to perform scene recognition and task decomposition on user commands to obtain at least one scene corresponding to the user command, and the task corresponding to each scene.

[0182] Each scenario corresponds to at least one home appliance and one scene-specific intelligent agent.

[0183] Step 143: Based on each task, at least two scene agents collaborate to obtain the execution plan for each device agent.

[0184] In some embodiments, the control logic for multiple home appliances in a task may require multi-scenario collaboration. Therefore, at least two scene agents can collaborate to obtain an execution plan for each device agent.

[0185] In some embodiments, in response to a conflict between the execution plans formulated by the first scene agent and the second scene agent, the system coordinates the two scene agents to revise their execution plans. For example, if the water temperature suggested by the health scene agent differs from the user's default preference, the system will weigh the pros and cons and may choose a slightly lower temperature, explaining the reason to the user. For instance, the first scene agent's execution plan might choose to turn the speakers to maximum volume, while the second scene agent's execution plan might choose to turn the TV to mute mode. In this case, a conflict exists between the first and second scene agents, therefore, the system coordinates the two scene agents to revise their execution plans. In some embodiments, this coordination can be achieved through voice prompts, allowing the user to participate in the coordination.

[0186] In some embodiments, in response to a first execution scheme formulated by a first scene agent and a second execution scheme formulated by a second scene agent corresponding to the same target home appliance, a third execution scheme of the device agent corresponding to the target home appliance is obtained by combining the first and second execution schemes. For example, if both the first and second execution schemes are for lighting devices, the third execution scheme of the device agent corresponding to the target home appliance can be obtained by combining the first and second execution schemes according to the scene type, thus obtaining the optimal lighting scheme for the lighting device.

[0187] In some embodiments, in response to a first scenario agent formulating a fourth execution plan, a second scenario agent adjusts the fourth execution plan. For example, if a cooking scenario agent formulates a recommended menu, a health scenario agent can analyze the user's recent dietary records and health data, adjust the recommended menu based on the analysis results, and provide dietary suggestions.

[0188] In some embodiments, in response to a conflict between the execution plans formulated by the first scene agent and the second scene agent, the first scene agent and the second scene agent coordinate to formulate new execution plans. Furthermore, in response to the first execution plan formulated by the first scene agent and the second execution plan formulated by the second scene agent corresponding to the same target home appliance, a third execution plan for the device agent corresponding to the target home appliance is obtained by combining the first execution plan and the second execution plan.

[0189] In some embodiments, in response to a conflict between the execution plans formulated by the first scene agent and the second scene agent, the first scene agent and the second scene agent coordinate to revise the execution plan. Furthermore, in response to the first scene agent formulating a fourth execution plan, the second scene agent is used to adjust the fourth execution plan.

[0190] In some embodiments, in response to a conflict between the execution plans formulated by the first scene agent and the second scene agent, the first scene agent and the second scene agent coordinate to formulate a new execution plan. Furthermore, in response to the first execution plan formulated by the first scene agent and the second execution plan formulated by the second scene agent corresponding to the same target home appliance, the first execution plan and the second execution plan are combined to obtain a third execution plan for the device agent corresponding to the target home appliance. And in response to the first scene agent formulating a fourth execution plan, the second scene agent is used to adjust the fourth execution plan.

[0191] In some embodiments, in response to a first execution plan formulated by a first scene agent and a second execution plan formulated by a second scene agent corresponding to the same target home appliance, a third execution plan for the device agent corresponding to the target home appliance is obtained by combining the first and second execution plans. Furthermore, in response to a fourth execution plan formulated by the first scene agent, the fourth execution plan is adjusted using the second scene agent.

[0192] Step 144: Assign the execution plan to the corresponding device agent so that the device agent can control the corresponding home appliance.

[0193] Each device intelligence agent is deployed in its corresponding home appliance.

[0194] In some embodiments, the scenario includes at least one of the following: living scenario, work / study scenario, entertainment scenario, sleep scenario, cooking scenario, health scenario, water usage scenario, and environmental control scenario.

[0195] In this embodiment, user instructions are acquired; a large language model is used to perform scene recognition and task decomposition on the user instructions to obtain at least one scene corresponding to the user instructions, and a task corresponding to each scene; wherein each scene corresponds to at least one home appliance, and each scene corresponds to a scene intelligent agent; the scene intelligent agent is used to assign a corresponding device intelligent agent to each task, so that the device intelligent agent controls the corresponding home appliance; wherein each device intelligent agent is deployed in the corresponding home appliance, and scene management of a large number of home appliances is achieved by setting scene intelligent agents, which can realize complex cross-device scene linkage and improve scene-based control capabilities.

[0196] Furthermore, based on each task, at least two scene agents collaborate to obtain the execution plan for each device agent, enabling collaborative control of smart homes based on diverse life scenarios and significantly improving the system's intelligence level.

[0197] See Figure 15 , Figure 15 This is a schematic flowchart of another embodiment of the home appliance control method provided in this application. The method includes:

[0198] Step 151: Obtain user instructions.

[0199] In some embodiments, the user instruction may be a voice instruction and / or a text instruction.

[0200] In some embodiments, the user's needs may be expressed in the user instructions.

[0201] Step 152: Use a large language model to perform scene recognition and task decomposition on user commands to obtain the work and study scenarios, living scenarios, and health scenarios corresponding to the user commands.

[0202] Step 153: Assign the first task to the work and study scenario, the second task to the living scenario, and the third task to the health scenario.

[0203] In some embodiments, the large language model can assign a first task to work and study scenarios, a second task to living scenarios, and a third task to health scenarios. That is, different scenarios have their own corresponding tasks. The specific content of these tasks may be controlling corresponding home appliances or devices associated with the user.

[0204] Step 154: Use scene intelligence agents to assign corresponding device intelligence agents to each task, so that the device intelligence agents can control the corresponding home appliances.

[0205] Each device intelligence agent is deployed in its corresponding home appliance.

[0206] In this embodiment, each scene corresponds to a scene intelligent agent.

[0207] In some embodiments, a first scene agent formulates a first execution plan based on a first task, and a corresponding device agent is determined based on the first execution plan. Wherein, if the first task requires control of lights and / or sound, the first execution plan includes light control and / or sound control. That is, the specific execution plan needs to be determined based on the task content.

[0208] A second scenario agent can be used to formulate a second execution plan based on a second task, and the corresponding device agent can be determined based on the second execution plan. Specifically, if the second task requires controlling the air conditioner and / or curtains, then the second execution plan includes air conditioner control and / or curtain control.

[0209] A third-scenario intelligent agent can be used to formulate a third execution plan based on a third task, and the corresponding device intelligent agent can be determined based on the third execution plan. Specifically, if the third task requires air purification and / or exercise reminders, the third execution plan will include exercise reminders and / or air purification control.

[0210] After receiving the corresponding execution plan, the device agent can automatically control the corresponding device to complete the execution plan by combining environmental data.

[0211] In this embodiment, user instructions are acquired; a large language model is used to perform scene recognition and task decomposition on the user instructions to obtain at least one scene corresponding to the user instructions, and a task corresponding to each scene; wherein each scene corresponds to at least one home appliance, and each scene corresponds to a scene intelligent agent; the scene intelligent agent is used to assign a corresponding device intelligent agent to each task, so that the device intelligent agent controls the corresponding home appliance; wherein each device intelligent agent is deployed in the corresponding home appliance, and scene management of a large number of home appliances is achieved by setting scene intelligent agents, which can realize complex cross-device scene linkage and improve scene-based control capabilities.

[0212] Furthermore, based on each task, at least two scene agents collaborate to obtain the execution plan for each device agent, enabling collaborative control of smart homes based on diverse life scenarios and significantly improving the system's intelligence level.

[0213] See Figure 16 , Figure 16 This is a schematic flowchart of another embodiment of the home appliance control method provided in this application. The method includes:

[0214] Step 161: Obtain user instructions.

[0215] Step 162: Use a large language model to perform scene recognition and task decomposition on user commands to obtain the cooking scene, living scene and health scene corresponding to the user commands.

[0216] Step 163: Assign the fourth task to the cooking scene, the fifth task to the living scene, and the sixth task to the health scene.

[0217] Step 164: Use scene intelligence agents to assign corresponding device intelligence agents to each task, so that the device intelligence agents can control the corresponding home appliances.

[0218] Each device intelligence agent is deployed in its corresponding home appliance.

[0219] In some embodiments, a fourth scene agent formulates a fourth execution plan based on a fourth task, and determines the corresponding device agent based on the fourth execution plan. The fourth execution plan includes food recognition control, menu recommendation, and / or oven control.

[0220] A fifth scenario-based intelligent agent can be used to formulate a fifth execution plan based on a fifth task, and the corresponding device intelligent agent can be determined based on the fifth execution plan. The fifth execution plan includes lighting control and / or air conditioning control.

[0221] A sixth scenario-based intelligent agent can be used to formulate a sixth execution plan based on a sixth task, and the corresponding device intelligent agent can be determined based on the sixth execution plan. The sixth execution plan includes data analysis and control and / or dietary recommendations.

[0222] After receiving the corresponding execution plan, the device agent can automatically control the corresponding device to complete the execution plan by combining environmental data.

[0223] In this embodiment, user instructions are acquired; a large language model is used to perform scene recognition and task decomposition on the user instructions to obtain at least one scene corresponding to the user instructions, and a task corresponding to each scene; wherein each scene corresponds to at least one home appliance, and each scene corresponds to a scene intelligent agent; the scene intelligent agent is used to assign a corresponding device intelligent agent to each task, so that the device intelligent agent controls the corresponding home appliance; wherein each device intelligent agent is deployed in the corresponding home appliance, and scene management of a large number of home appliances is achieved by setting scene intelligent agents, which can realize complex cross-device scene linkage and improve scene-based control capabilities.

[0224] Furthermore, based on each task, at least two scene agents collaborate to obtain the execution plan for each device agent, enabling collaborative control of smart homes based on diverse life scenarios and significantly improving the system's intelligence level.

[0225] See Figure 17 , Figure 17 This is a schematic flowchart of another embodiment of the home appliance control method provided in this application. The method includes:

[0226] Step 171: Obtain user instructions.

[0227] Step 172: Use a large language model to perform scene recognition and task decomposition on user commands to obtain the water use scenario, environmental control scenario and health scenario corresponding to the user commands.

[0228] Step 173: Assign the seventh task to the water use scenario, the eighth task to the environmental control scenario, and the ninth task to the health scenario.

[0229] Step 174: Use scene intelligence to assign a corresponding device intelligence to each task so that the device intelligence can control the corresponding home appliances.

[0230] Each device intelligence agent is deployed in its corresponding home appliance.

[0231] In some embodiments, a seventh scene agent formulates a seventh execution plan based on a seventh task, and a corresponding device agent is determined based on the seventh execution plan. The seventh execution plan includes water temperature regulation, water quality regulation, and / or water pressure regulation.

[0232] The eighth scenario agent can be used to formulate an eighth execution plan based on the eighth task, and the corresponding device agent can be determined based on the eighth execution plan. The eighth execution plan includes lighting control and / or air conditioning control.

[0233] The ninth scenario agent can be used to formulate a ninth execution plan based on the ninth task, and the corresponding device agent can be determined based on the ninth execution plan. The ninth execution plan includes data analysis control and / or dietary recommendations.

[0234] After receiving the corresponding execution plan, the device agent can automatically control the corresponding device to complete the execution plan by combining environmental data.

[0235] In this embodiment, user instructions are acquired; a large language model is used to perform scene recognition and task decomposition on the user instructions to obtain at least one scene corresponding to the user instructions, and a task corresponding to each scene; wherein each scene corresponds to at least one home appliance, and each scene corresponds to a scene intelligent agent; the scene intelligent agent is used to assign a corresponding device intelligent agent to each task, so that the device intelligent agent controls the corresponding home appliance; wherein each device intelligent agent is deployed in the corresponding home appliance, and scene management of a large number of home appliances is achieved by setting scene intelligent agents, which can realize complex cross-device scene linkage and improve scene-based control capabilities.

[0236] Furthermore, based on each task, at least two scene agents collaborate to obtain the execution plan for each device agent, enabling collaborative control of smart homes based on diverse life scenarios and significantly improving the system's intelligence level.

[0237] After the scene ends, the scene agent is used to assign a scene end task to the device agent, so that the device agent can control the corresponding home appliance to complete the scene end task.

[0238] See Figure 18 The control system 1000 for home appliances mainly includes the following modules: scene intelligent agent layer 201, scene collaboration layer 202, and user interaction layer 203.

[0239] The scenario-based intelligent agent layer 201 includes multiple intelligent agents based on specific scenarios, each with a unique skill tree. For example:

[0240] Intelligent agents for living scenarios:

[0241] Inputs: light sensor, temperature sensor, human presence sensor, voice commands.

[0242] Skill Tree: Lighting control, temperature regulation, air quality management, furniture layout optimization.

[0243] Outputs: Lighting control commands, air conditioning control commands, curtain control commands, air purifier control commands.

[0244] Intelligent agents for work and study scenarios:

[0245] Inputs: noise sensor, light sensor, computer usage status, calendar data.

[0246] Skills tree: Focus mode management, environmental noise control, work area lighting optimization, schedule management.

[0247] Outputs: Smart speaker control commands, lighting control commands, schedule reminders, and device usage suggestions.

[0248] Entertainment Scene Intelligent Agent:

[0249] Inputs: audio device status, video device status, human posture sensor, voice commands.

[0250] Skill Tree: Audio-visual equipment collaborative control, ambient lighting management, multimedia content recommendation, and interactive game control.

[0251] Outputs: TV / projector control commands, audio system control commands, lighting control commands, and content recommendations.

[0252] Sleep Scene Intelligent Agent:

[0253] Inputs: circadian rhythm data, environmental noise sensor, light sensor, temperature and humidity sensor.

[0254] Skills tree: Sleep environment optimization, biological rhythm synchronization, noise cancellation, comfort adjustment.

[0255] Outputs: Air conditioner control commands, curtain control commands, humidifier control commands, white noise playback control.

[0256] Intelligent agents for cooking scenarios:

[0257] Inputs: cooking equipment status, food recognition camera, temperature sensor, gas sensor.

[0258] Skill Tree: Recipe Recommendation, Cooking Process Optimization, Kitchen Safety Monitoring, Ingredient Management.

[0259] Outputs: Oven control commands, stove control commands, exhaust fan control commands, recipe display.

[0260] Intelligent agents for health scenarios:

[0261] Inputs: wearable device data, air quality sensor data, motion detection sensor data, and dietary records.

[0262] Skills tree: health data analysis, exercise suggestion generation, diet plan development, environmental health optimization.

[0263] Outputs: Health reports, exercise reminders, dietary recommendations, and air purifier control commands.

[0264] The Scene Collaboration Layer 202 is responsible for coordinating the collaboration between multiple scene agents to achieve cross-scene intelligent control.

[0265] User interaction layer 203 provides users with a natural language interaction interface, supporting users to directly express their needs in a given scenario.

[0266] In some embodiments, the control system operates as follows:

[0267] Users can input commands using natural language, such as "I want to work from home".

[0268] The user interaction layer passes instructions to the scene collaboration layer for scene recognition.

[0269] The scene collaboration layer identifies relevant scenes and decomposes and assigns tasks to the corresponding scene agents.

[0270] Each intelligent agent in a given scenario formulates an execution plan based on its own skill set, device status, and environmental information.

[0271] Scene-based intelligent agents exchange information and make collaborative decisions through the scene collaboration layer.

[0272] Each scenario's intelligent agent executes the finalized control commands.

[0273] The execution results are fed back to the user interaction layer through the scenario collaboration layer and finally presented to the user.

[0274] The control system records this interaction and control process for subsequent learning and optimization.

[0275] In some embodiments, the process for a smart home office scenario is as follows:

[0276] The user said, "I want to work from home."

[0277] Scene recognition: The system identifies that this involves work and study scenarios, living scenarios, and health scenarios.

[0278] Task breakdown and allocation:

[0279] Intelligent agent for work and study scenarios: preparing the office environment.

[0280] Intelligent agent for living scenarios: Adjust living spaces to adapt to work needs.

[0281] Intelligent agents for health scenarios: monitoring and optimizing work health status.

[0282] Execution plans for each agent:

[0283] Intelligent agents for work and study scenarios:

[0284] Check the schedule and remind me of important meetings.

[0285] Adjust the brightness and color temperature of the lights in the work area (output: lighting control command).

[0286] Enable noise cancellation function (output: smart speaker control command).

[0287] Intelligent agents for living scenarios:

[0288] Adjust the indoor temperature to the optimal operating temperature (output: air conditioning control command).

[0289] Adjust the curtains according to the lighting conditions (output: curtain control command).

[0290] Intelligent agents for health scenarios:

[0291] Set a sedentary reminder (output: exercise reminder).

[0292] Adjust the air purifier to optimize the working environment (output: air purifier control command).

[0293] Scene collaboration:

[0294] The intelligent agents in work and study scenarios collaborate with those in living scenarios to jointly determine the best lighting solution based on the type of work and natural light conditions.

[0295] The intelligent agent in the health scenario coordinates with the intelligent agent in the work and study scenario to take care of physical health while ensuring work efficiency.

[0296] Results feedback:

[0297] The system informed the user via voice: "Your home office environment is ready. You have a video conference at 3 PM today. I've set up hourly reminders for you to get up and move around. Best wishes for a successful workday!"

[0298] In some embodiments, the process of a smart cooking scenario is as follows:

[0299] The user said, "I want to make a healthy dinner."

[0300] Scene recognition: The system identifies that this involves cooking scenes, health scenes, and living scenes.

[0301] Task breakdown and allocation:

[0302] Cooking scenario intelligence agent: Prepares the cooking environment and suggests menus.

[0303] Health-related intelligent agents: provide healthy eating advice.

[0304] Smart agent for living scenarios: Adjusting the kitchen environment.

[0305] Execution plans for each agent:

[0306] Intelligent agents for cooking scenarios:

[0307] Scan the food inside the refrigerator (input: food recognition camera).

[0308] Menus are recommended based on available ingredients and health needs.

[0309] Preheat oven (output: oven control command).

[0310] Intelligent agents for health scenarios:

[0311] Analyze users' recent dietary records and health data (inputs: wearable device data, dietary records).

[0312] Adjust menu recommendations based on analysis results to ensure nutritional balance (output: dietary recommendations).

[0313] Intelligent agents for living scenarios:

[0314] Adjust kitchen lighting (output: lighting control command).

[0315] Set the kitchen temperature (output: air conditioning control command).

[0316] Scene collaboration:

[0317] The cooking scenario intelligence agent and the health scenario intelligence agent collaborate to create menus that are both delicious and healthy.

[0318] The intelligent agent in the cooking scene coordinates with the intelligent agent in the living scene to create the best cooking environment.

[0319] Results feedback:

[0320] The system informed the user via voice: "A healthy dinner menu has been prepared for you. Considering your recent increase in physical activity, I've added some high-protein foods. The kitchen environment is now optimized for cooking. Would you like me to play a cooking instruction video for you?"

[0321] In some embodiments, the process of a smart bathing scenario is as follows:

[0322] The user said, "I want to take a nice, relaxing shower."

[0323] Scene Recognition: Through speech recognition and semantic analysis, the system identifies that the user's needs involve the "bathing" scenario. The system then activates three main intelligent agents related to this scenario: the water usage scenario agent, the environmental control agent, and the health scenario agent.

[0324] Task decomposition and allocation: The system decomposes the overall task and allocates it to each agent.

[0325] Water-using scenario intelligent agent: responsible for water quality and temperature control.

[0326] Environmental control agent: responsible for optimizing the bathroom environment.

[0327] Health-related intelligent agent: responsible for providing health-related advice.

[0328] Execution plans for each agent: a) Agent in the water usage scenario:

[0329] Check the water heater status and adjust it to a suitable temperature (default 38℃, can be adjusted according to user preference).

[0330] Activate the water softener to ensure soft water.

[0331] Water pressure can be set according to user preferences.

[0332] b) Environmental control agent:

[0333] Adjust the bathroom temperature to a comfortable range (default 26℃).

[0334] Set the humidity to a suitable level (default 60%).

[0335] Adjust the lighting to a soothing mode (warm tone, moderate brightness).

[0336] Control ventilation equipment to ensure air circulation without creating a chill.

[0337] c) Intelligent agents for health scenarios:

[0338] Retrieve the user's recent health data (such as blood pressure, heart rate, etc.).

[0339] Recommended water and bathroom temperatures based on health data.

[0340] Set a shower duration reminder (default 20 minutes) to prevent the adverse health effects of excessively long showers.

[0341] Scene coordination: Each agent submits its execution plan to the scene coordination layer. The coordination layer comprehensively considers the suggestions of each agent and resolves potential conflicts, such as:

[0342] If the water temperature suggested by the health scenario agent differs from the user's default preference, the system will weigh the pros and cons and may choose a slightly lower temperature, explaining the reason to the user.

[0343] The environmental control agent and the water usage scenario agent coordinate to ensure a comfortable transition between bathroom temperature and water temperature.

[0344] Execution of control commands: The collaboration layer formulates the final execution plan and sends control commands to each smart device.

[0345] Water heater: Set the water temperature to 37.5℃.

[0346] Water softener: Activate.

[0347] Bathroom thermostat system: set temperature to 25.5℃.

[0348] Bathroom humidifier: Set the humidity to 58%.

[0349] Lighting system: Set to warm, soft yellow light, 70% brightness.

[0350] Ventilation system: Set to low speed operation.

[0351] Smart speaker: Ready to play the user's favorite light music.

[0352] Feedback: The system uses a voice assistant to inform the user of the preparation status: "Dear [Username], your bathing environment is ready! The water temperature has been adjusted to 37.5℃, slightly lower than your usual preference, taking into account your high level of activity today. The bathroom temperature is set to 25.5℃ and the humidity to 58%, creating a comfortable bathing environment. The lighting has been adjusted to a soothing mode, and I have also prepared some soft music for you. I suggest you keep your bathing time under 20 minutes. Enjoy your bathing time!"

[0353] Continuous monitoring and adjustment: During the user's shower, the system will continuously monitor various parameters.

[0354] If the water temperature fluctuates, the intelligent system will adjust the water heater's output accordingly.

[0355] If the humidity in the bathroom is too high, the environmental control agent will increase the ventilation.

[0356] The intelligent health-focused system tracks bathing time and gently reminds the user via the bathroom speaker when it approaches 20 minutes.

[0357] Scene End and Learning: After the user finishes showering, the system will:

[0358] Turn off the hot water supply, lower the bathroom temperature, and increase ventilation.

[0359] Record all parameters and user feedback (if any) for this bath.

[0360] We use the collected data to optimize algorithms so that we can provide users with more accurate services next time.

[0361] In some embodiments, the process of a smart morning wake-up scenario is as follows:

[0362] User settings: "Wake me up at 7 a.m. tomorrow. I need to be at the company before 8 a.m."

[0363] Scene Recognition and Initial Planning: The system recognizes this as a complex scene involving multiple aspects such as sleep, daily routines, diet, and travel. The system activates the following agents:

[0364] Intelligent agents for sleep scenarios, living scenarios, eating scenarios, travel scenarios, and health scenarios.

[0365] Nighttime preparation phase: a) Sleep scenario agent:

[0366] Analyze users' sleep cycle data.

[0367] Calculate the optimal time to fall asleep based on the requirement of waking up at 7 a.m., and notify the user at the appropriate time.

[0368] b) Intelligent agents in living scenarios:

[0369] Set up smart curtains to gradually open starting at 6:55 AM.

[0370] Adjust the bedroom temperature to the most comfortable temperature for sleep.

[0371] c) Intelligent agents for travel scenarios:

[0372] Check the weather forecast and traffic conditions to estimate your travel time for the next morning.

[0373] Morning wake-up phase (6:55-7:00): a) Sleep scenario intelligent agent:

[0374] Monitor the user's sleep status.

[0375] The wake-up procedure is triggered at the optimal wake-up time (light sleep stage) between 6:55 and 7:00.

[0376] b) Intelligent agents in living scenarios:

[0377] After receiving a signal from the sleep scene intelligent agent, the room brightness is gradually increased.

[0378] Play some soothing morning music.

[0379] c) Intelligent agents for health scenarios:

[0380] Collect users' sleep quality data and prepare morning health reports.

[0381] Morning activity phase (7:00-7:30): a) Smart agent for living environment:

[0382] Adjust the room temperature to a comfortable level.

[0383] Turn on the bathroom lights and heater.

[0384] b) Intelligent agent for dining scenarios:

[0385] Adjust breakfast recommendations based on sleep quality data provided by the health scenario intelligent agent.

[0386] Turn on the smart coffee machine to prepare coffee.

[0387] c) Intelligent agents for travel scenarios:

[0388] We recommend adjusting departure times based on real-time traffic conditions.

[0389] It shares information with intelligent agents in living and dining scenarios and coordinates time arrangements.

[0390] d) Intelligent agents for health scenarios:

[0391] Generate a morning health report, including sleep quality, recommended water intake, etc.

[0392] Share information with intelligent agents for the dining scenario and the living scenario.

[0393] Preparation and Departure Phase (7:30-8:00): a) Intelligent Agent in Living Environment:

[0394] Based on the suggestions of the intelligent agent in the travel scenario, the reminder time for users to prepare to leave is adjusted.

[0395] Make sure all unnecessary electrical appliances are turned off.

[0396] b) Intelligent agent for dining scenarios:

[0397] Adjust the pace of breakfast preparation according to your travel schedule.

[0398] If time is tight, prepare a portable breakfast option.

[0399] c) Intelligent agents for travel scenarios:

[0400] Continuously monitor road conditions and promptly notify users of any changes.

[0401] Preheat vehicles or book ride-hailing services (depending on user habits).

[0402] Scene coordination and execution: The scene coordination layer integrates the inputs from various agents and coordinates potential conflicts.

[0403] If sleep quality is poor, the health-related AI agent might suggest delaying wake-up time, while the travel-related AI agent might suggest leaving earlier based on traffic conditions. The system will weigh the pros and cons and might choose to slightly delay wake-up time while adjusting the breakfast plan to a quick option.

[0404] The intelligent agents for living and dining scenarios coordinate to ensure optimal scheduling of bathroom use and breakfast time.

[0405] User Interaction and Feedback: The system provides information and suggestions to users via voice assistant or mobile app: "Good morning, [username]. It is 7:00 AM. Your sleep quality is good, with a sleep duration of 7 hours and 20 minutes. The weather is sunny today, but traffic congestion is expected during the morning rush hour. It is recommended to leave at 7:40 AM to ensure you arrive at the company before 8:00 AM. Coffee is ready. A quick and healthy breakfast is recommended: whole wheat toast with eggs and avocado. The bathroom is preheated. Have a wonderful day!"

[0406] Continuous optimization: The system records users' actual behaviors (such as actual wake-up time, breakfast choices, departure time, etc.) for future scenario optimization. Each agent uses this data to adjust its decision-making model to provide more accurate services.

[0407] In some embodiments, this application also provides a method for controlling a home appliance, the method comprising: receiving a control command; wherein the control command is obtained by a multi-task decision generated by the method of any of the above embodiments; and operating in accordance with the control command.

[0408] See Figure 19 , Figure 19 This is a schematic diagram of the structure of a home appliance 300 according to an embodiment of the present application. The home appliance 300 includes a communication interface 301 and a processor 302. The processor 302 is coupled to the communication interface 301 and is used to implement the method of any of the above embodiments.

[0409] In summary, the home appliances and their control methods and systems provided in this application acquire user commands; utilize a large language model to perform scene recognition and task decomposition on the user commands to obtain at least one scene corresponding to the user command, and a task corresponding to each scene; wherein each scene corresponds to at least one home appliance, and each scene corresponds to one scene intelligent agent; the scene intelligent agent is used to assign a corresponding device intelligent agent to each task, so that the device intelligent agent controls the corresponding home appliance; wherein each device intelligent agent is deployed in the corresponding home appliance, and scene management of a large number of home appliances is achieved by setting scene intelligent agents, which can realize complex cross-device scene linkage and improve scene-based control capabilities.

[0410] In the several embodiments provided in this application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed.

[0411] If the integrated units in the other embodiments described above are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processing circuit component (processor) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0412] The above description is merely an embodiment of this application and does not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.

Claims

1. A method for controlling household appliances, characterized in that, The method includes: Obtain user instructions; The user command is identified and decomposed into a scene using a large language model to obtain at least one scene corresponding to the user command and a task corresponding to each scene; wherein each scene corresponds to at least one home appliance and each scene corresponds to a scene agent. The scene intelligence agent is used to assign a corresponding device intelligence agent to each task, so that the device intelligence agent controls the corresponding home appliance; wherein, each device intelligence agent is deployed in the corresponding home appliance. The method of assigning a corresponding device agent to each task using a scene-based intelligent agent, so that the device agent controls the corresponding home appliance, includes: Based on each task, at least two of the scene agents collaborate to obtain an execution plan for each of the device agents. The execution plan is assigned to the corresponding device agent, so that the device agent controls the corresponding home appliance; At least two scene agents, including a first scene agent and a second scene agent, are involved in scene collaboration based on each task, including: In response to a conflict in the execution plans formulated between the first scene agent and the second scene agent, the first scene agent and the second scene agent coordinate to revise the execution plan; and / or, In response to the first execution plan formulated by the first scene agent and the second execution plan formulated by the second scene agent corresponding to the same target home appliance, a third execution plan for the device agent corresponding to the target home appliance is obtained by combining the first execution plan and the second execution plan; and / or, In response to the first scenario agent formulating a fourth execution plan, the second scenario agent is used to adjust the fourth execution plan.

2. The method according to claim 1, characterized in that, The scenarios include at least one of the following: living scenarios, work and study scenarios, entertainment scenarios, sleep scenarios, cooking scenarios, health scenarios, water usage scenarios, and environmental control scenarios.

3. The method according to claim 2, characterized in that, The step of using a large language model to perform scene recognition and task decomposition on the user command to obtain at least one scene corresponding to the user command and a task corresponding to each scene includes: The user instructions are used to perform scene recognition and task decomposition using a large language model to obtain the work and study scene, living scene and health scene corresponding to the user instructions; Assign a first task to the work and study scenario, a second task to the living scenario, and a third task to the health scenario.

4. The method according to claim 3, characterized in that, The scene intelligence agent includes a first scene intelligence agent, a second scene intelligence agent, and a third scene intelligence agent. The step of using the scene intelligence agent to assign a corresponding device intelligence agent to each task includes: The first scene agent formulates a first execution plan based on the first task, and determines the corresponding device agent based on the first execution plan; the first execution plan includes lighting control and / or sound control; The second scene agent formulates a second execution plan based on the second task, and determines the corresponding device agent based on the second execution plan; the second execution plan includes air conditioning control and / or curtain control; The third scene agent formulates a third execution plan based on the third task, and determines the corresponding device agent based on the third execution plan; the third execution plan includes exercise reminders and / or air purification control.

5. The method according to claim 2, characterized in that, The step of using a large language model to perform scene recognition and task decomposition on the user command to obtain at least one scene corresponding to the user command and a task corresponding to each scene includes: The user command is used to perform scene recognition and task decomposition by a large language model to obtain the cooking scene, living scene and health scene corresponding to the user command; A fourth task is assigned to the cooking scenario, a fifth task is assigned to the living scenario, and a sixth task is assigned to the health scenario.

6. The method according to claim 5, characterized in that, The scene intelligence agent includes a fourth scene intelligence agent, a fifth scene intelligence agent, and a sixth scene intelligence agent. The step of using the scene intelligence agent to assign a corresponding device intelligence agent to each task includes: The fourth scene agent is used to formulate a fourth execution plan based on the fourth task, and the corresponding device agent is determined based on the fourth execution plan; the fourth execution plan includes food identification control, menu recommendation and / or oven control; The fifth scene agent formulates a fifth execution plan based on the fifth task, and determines the corresponding device agent based on the fifth execution plan; the fifth execution plan includes lighting control and / or air conditioning control; The sixth scene agent is used to formulate a sixth execution plan based on the sixth task, and the corresponding device agent is determined based on the sixth execution plan; the sixth execution plan includes data analysis and control and / or dietary advice.

7. The method according to claim 2, characterized in that, The step of using a large language model to perform scene recognition and task decomposition on the user command to obtain at least one scene corresponding to the user command and a task corresponding to each scene includes: The user command is used to perform scene recognition and task decomposition by a large language model to obtain the water use scenario, environmental control scenario and health scenario corresponding to the user command; A seventh task is assigned to the water use scenario, an eighth task is assigned to the environmental control scenario, and a ninth task is assigned to the health scenario.

8. The method according to claim 7, characterized in that, The scene intelligence agent includes a seventh scene intelligence agent, an eighth scene intelligence agent, and a ninth scene intelligence agent. The step of using the scene intelligence agent to assign a corresponding device intelligence agent to each task includes: The seventh scene agent formulates a seventh execution plan based on the seventh task, and determines the corresponding device agent based on the seventh execution plan; the seventh execution plan includes water temperature regulation, water quality regulation and / or water pressure regulation; The eighth scene agent is used to formulate an eighth execution plan based on the eighth task, and the corresponding device agent is determined based on the eighth execution plan; the eighth execution plan includes lighting control and / or air conditioning control; The ninth scene agent formulates a ninth execution plan based on the ninth task, and determines the corresponding device agent based on the ninth execution plan; the ninth execution plan includes data analysis and control and / or dietary advice.

9. The method according to any one of claims 1-8, characterized in that, After the scene ends, the scene agent assigns a scene end task to the device agent, so that the device agent controls the corresponding home appliance to complete the scene end task.

10. A method for controlling a household appliance, characterized in that, The method includes: Receive control commands; wherein the control commands are obtained by the method described in any one of claims 1-9; Operate according to the control instructions.

11. A household appliance, characterized in that, The home appliance includes a communication interface and a processor, the processor being coupled to the communication interface for implementing the method as described in claim 10.

Citation Information

Patent Citations

  • Family scene linkage smart home system based on voice recognition and control method and control device thereof

    CN112558491A