Household appliance interaction method, wearable device, and home control system

By using wearable devices to collect multimodal data in smart home systems and analyzing it using large language models, the problem of single input modes and limited interaction methods in smart home control is solved, resulting in a more natural and convenient home appliance control experience.

WO2026091651A1PCT designated stage Publication Date: 2026-05-07WUHU MIDEA KITCHEN & BATH APPLIANCES MFG CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
WUHU MIDEA KITCHEN & BATH APPLIANCES MFG CO LTD
Filing Date
2025-06-30
Publication Date
2026-05-07

AI Technical Summary

Technical Problem

Existing smart home control methods suffer from a single input mode and a single interaction method, lacking a deep understanding of user intent and an intuitive presentation of appliance status information.

Method used

By using wearable devices to collect multimodal data, analyzing the multimodal data through a large language model, generating control commands, and using wearable devices as an interaction medium, the diversity of data collection and interaction input methods are improved, thereby enhancing the depth of understanding of multimodal data and the accuracy of control commands.

Benefits of technology

It enables more natural and convenient control of home appliances, improves the user interaction experience, enhances the deep understanding of user intentions, and provides an intuitive display of home appliance status information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025106008_07052026_PF_FP_ABST
    Figure CN2025106008_07052026_PF_FP_ABST
Patent Text Reader

Abstract

A household appliance interaction method, a wearable device, and a home control system. The method comprises: receiving interaction data corresponding to a user; using a large language model to perform intention recognition on the interaction data, and obtaining a target intention; generating dynamic display data on the basis of the target intention, the dynamic display data being a voice, an image and / or a text corresponding to the interaction data; and displaying the dynamic display data. By means of the described method, a household appliance interaction mode can be more natural and convenient, and the user interaction experience is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Interaction methods for home appliances, wearable devices, home control systems

[0001] Cross-referencing of related applications

[0002] This application claims priority to Chinese Patent Application No. 2024115557045, filed in China on October 31, 2024, the entire contents of which are incorporated herein by reference. [Technical Field]

[0003] This application relates to the field of home appliance technology, and in particular to interaction methods for home appliances, wearable devices, and home control systems. [Background Technology]

[0004] With the development of IoT and AI technologies, smart home control systems are becoming increasingly popular.

[0005] However, some problems still exist in the related smart home control methods, such as a single input mode and a single interaction method. [Summary of the Invention]

[0006] The interaction methods, wearable devices, and home control systems for home appliances provided in this application can make the interaction of home appliances more natural and convenient, and improve the user's interactive experience.

[0007] In a first aspect, this application provides an interaction method for home appliances, applied to wearable devices. The method includes: receiving interaction data corresponding to a user; using a large language model to perform intent recognition on the interaction data to obtain a target intent; generating dynamic display data based on the target intent; wherein the dynamic display data is voice, images, and / or text corresponding to the interaction data; and displaying the dynamic display data.

[0008] In some embodiments, the interaction data includes: voice data; and the interaction data is used to perform intent recognition using a large language model to obtain the target intent, including: using a large language model to perform intent recognition on the voice data to obtain the home appliance in the voice data and the target intent for the home appliance.

[0009] In some embodiments, the interaction data includes: voice data and image data. The intention recognition of the interaction data using a large language model to obtain the target intention includes: using a large language model to perform intention recognition of the voice data and image data to obtain the target intention.

[0010] In some embodiments, using a large language model to perform intent recognition on speech data and image data to obtain target intent includes: using a large language model to perform content recognition on speech data to obtain key content; determining the target object corresponding to the key content from image data; and obtaining the target intent based on the key content and the target object.

[0011] In some embodiments, generating dynamic display data based on the target intent includes: acquiring the voice, image, and / or text corresponding to the target object; and generating dynamic display data based on the voice, image, and / or text.

[0012] In some embodiments, displaying dynamic display data includes: obtaining the user's age group; determining the voice playback speed according to the age group; and displaying the dynamic display data according to the voice playback speed.

[0013] In some embodiments, the method further includes: receiving user instructions during the display of dynamic display data; and adjusting the presentation mode and content of the dynamic display data according to the user instructions.

[0014] In some embodiments, the wearable device is smart glasses, which displays dynamic display data, including: overlaying images and / or text from the dynamic display data onto the visible area of ​​the smart glasses.

[0015] In some embodiments, if the dynamically displayed data includes voice, the voice is played during the display process.

[0016] In some embodiments, the method further includes: receiving a user's dish preparation requirements; inputting the dish requirements into a large language model to obtain dish preparation data corresponding to the dish requirements; and displaying the dish preparation data.

[0017] In a second aspect, this application provides a wearable device, which includes a processor, a memory coupled to the processor, and a communication interface. The memory stores at least one computer program, which, when loaded and executed by the processor, is used to implement the method provided in the first aspect.

[0018] Thirdly, this application provides a home control system, which includes: a cloud; a wearable device, which is communicatively connected to the cloud, receives user-corresponding interactive data, and sends interactive data to the cloud; the cloud uses a large language model to perform intent recognition on the interactive data to obtain the target intent, and generates dynamic display data based on the target intent; wherein, the dynamic display data is voice, image, and / or text corresponding to the interactive data; the wearable device receives the dynamic display data and displays it.

[0019] The beneficial effects of the embodiments of this application are as follows: Unlike the prior art, the interaction method, wearable device, and home control system of home appliances provided in this application utilize wearable devices to receive corresponding user interaction data, increasing the interaction input methods of the home control system. Furthermore, it uses a large language model to perform intent recognition on the interaction data, obtains the target intent, and generates dynamic display data based on the target intent, thereby improving the depth of understanding of the interaction data, improving the accuracy of the dynamic display data, and using wearable devices as the display medium for dynamic display data, making the interaction method of home appliances more natural and convenient, and improving the user interaction experience. [Attached Image Description]

[0020] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. Wherein:

[0021] Figure 1 is a structural schematic diagram of an embodiment of the home control system provided in this application;

[0022] Figure 2 is a structural schematic diagram of another embodiment of the home control system provided in this application;

[0023] Figure 3 is a flowchart illustrating an embodiment of the control method for home appliances provided in this application;

[0024] Figure 4 is a flowchart illustrating another embodiment of the control method for home appliances provided in this application;

[0025] Figure 5 is a flowchart illustrating an embodiment of the interaction method for home appliances provided in this application;

[0026] Figure 6 is a flowchart illustrating another embodiment of the interaction method for home appliances provided in this application;

[0027] Figure 7 is a flowchart illustrating another embodiment of the interaction method for home appliances provided in this application;

[0028] Figure 8 is a flowchart of an embodiment of step 72 in Figure 7 provided in this application;

[0029] Figure 9 is a flowchart illustrating another embodiment of the interaction method for home appliances provided in this application;

[0030] Figure 10 is a structural schematic diagram of an embodiment of the wearable device provided in this application;

[0031] Figure 11 is a schematic diagram of an embodiment of the electronic device provided in this application.

Detailed Implementation Methods

[0032] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. It is understood that the specific embodiments described herein are only for explaining this application and not for limiting it. Furthermore, it should be noted that, for ease of description, only the parts related to this application are shown in the accompanying drawings, not all structures. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.

[0033] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0034] With the development of IoT and AI technologies, smart home control systems are becoming increasingly popular.

[0035] However, some problems still exist with the relevant smart home control methods, such as a single input mode, a lack of deep understanding of user intent, and unintuitive presentation of appliance status information.

[0036] Based on this, this application proposes using wearable devices to collect multimodal data, thereby increasing the diversity of data collection and expanding the interactive input methods of the home control system. Furthermore, the cloud utilizes a large language model to analyze the multimodal data, generating control commands to enhance the depth of understanding of the multimodal data and improve the accuracy of the control commands. Additionally, wearable devices are used as an interaction medium to make home appliance control more natural and convenient. See any of the following embodiments for details.

[0037] Referring to Figure 1, Figure 1 is a schematic diagram of the structure of an embodiment of the home control system provided in this application. The home control system 100 includes: a cloud platform 10 and a wearable device 20.

[0038] The wearable device 20 is connected to the cloud 10 and sends collected multimodal data to the cloud 10. The multimodal data includes at least one of voice data, image data and posture data.

[0039] The cloud-based device 10 uses a large language model to analyze multimodal data, generates control commands, sends the control commands to the home appliances corresponding to the multimodal data, and receives the device status feedback from the home appliances; the wearable device 20 receives the device status and displays it.

[0040] In some embodiments, the cloud 10 selects a target device state that matches the user's preferences from the device states and sends the target device state to the wearable device 20.

[0041] In some embodiments, the wearable device 20 is smart glasses, and the multimodal data includes voice data, image data, and posture data; the cloud 10 uses a large language model to analyze the voice data, image data, and posture data to obtain the type of home appliance corresponding to the voice data, the home appliance that matches the type of home appliance in the image data, and generates control instructions for the corresponding home appliance.

[0042] In some embodiments, after receiving and displaying the device status, the smart glasses receive a gesture command and send the gesture command to the cloud 10. The gesture command is related to the device status.

[0043] In some embodiments, after receiving and displaying the device status, the smart glasses receive voice commands and send voice commands to the cloud 10, the voice commands being related to the device status.

[0044] In some embodiments, after receiving and displaying the device status, the smart glasses receive gesture commands and voice commands, and send the gesture commands and voice commands to the cloud 10. The gesture commands and voice commands are related to the device status.

[0045] In some embodiments, the smart glasses display the device status overlaid within the smart glasses' field of view.

[0046] In this embodiment, wearable device 20 is used to collect multimodal data, thereby increasing the diversity of data collection and adding interactive input methods to the home control system. Furthermore, cloud 10 uses a large language model to analyze the multimodal data, generate control commands, improve the depth of understanding of multimodal data, improve the accuracy of control commands, and use wearable device 20 as an interactive medium to make home appliance control more natural and convenient.

[0047] Referring to Figure 2, which is a schematic diagram of another embodiment of the home control system provided in this application, the home control system 100 includes: a cloud platform 10, a home gateway 30, and a wearable device 20.

[0048] The home gateway 30 connects to the cloud 10, the wearable device 20, and home appliances (not shown). It receives multimodal data sent by the wearable device 20 and forwards it to the cloud 10, and receives control commands sent by the cloud 10 and forwards them to the home appliances.

[0049] In some embodiments, the cloud 10 selects a target device state that matches the user's preferences from the device states and sends the target device state to the wearable device 20 through the home gateway 30.

[0050] In some embodiments, the wearable device 20 is smart glasses, and the multimodal data includes voice data, image data, and posture data; the cloud 10 uses a large language model to analyze the voice data, image data, and posture data to obtain the type of home appliance corresponding to the voice data, the home appliance that matches the type of home appliance in the image data, and generates control instructions for the corresponding home appliance.

[0051] In some embodiments, after receiving and displaying the device status, the smart glasses receive gesture commands and send the gesture commands to the cloud 10 through the home gateway 30. The gesture commands are related to the device status.

[0052] In some embodiments, after receiving and displaying the device status, the smart glasses receive voice commands and send the voice commands to the cloud 10 through the home gateway 30. The voice commands are related to the device status.

[0053] In some embodiments, after receiving and displaying the device status, the smart glasses receive gesture commands and voice commands, and send the gesture commands and voice commands to the cloud 10 through the home gateway 30. The gesture commands and voice commands are related to the device status.

[0054] In some embodiments, the smart glasses display the device status overlaid within the smart glasses' field of view.

[0055] In this embodiment, wearable devices are used to collect multimodal data, which increases the diversity of data collection and adds interactive input methods to the home control system. Furthermore, the cloud uses a large language model to analyze the multimodal data, generate control commands, improve the depth of understanding of multimodal data, improve the accuracy of control commands, and use wearable devices as an interaction medium to make home appliance control more natural and convenient.

[0056] Referring to Figure 3, Figure 3 is a schematic flowchart of an embodiment of the control method for home appliances provided in this application. Applied to wearable devices, the method includes:

[0057] Step 31: Collect multimodal data, which includes at least one of speech data, image data, and pose data.

[0058] In some embodiments, multimodal data can be acquired using relevant sensors on a wearable device. For example, voice data can be acquired using a microphone on the wearable device. Image data can be acquired using an image sensor on the wearable device. Attitude data can be acquired using a gyroscope and inertial navigation system on the wearable device. In other embodiments, attitude data can be acquired by combining image data acquired using an image sensor.

[0059] Step 32: Send multimodal data to the cloud so that the cloud can use a large language model to analyze the multimodal data and generate control commands.

[0060] Step 33: Receive and display the device status of home appliances sent from the cloud; the device status is obtained after the home appliances execute control commands.

[0061] In this embodiment, wearable devices are used to collect multimodal data, which increases the diversity of data collection and adds interactive input methods to the home control system. Furthermore, the cloud uses a large language model to analyze the multimodal data, generate control commands, improve the depth of understanding of multimodal data, improve the accuracy of control commands, and use wearable devices as an interaction medium to make home appliance control more natural and convenient.

[0062] Referring to Figure 4, which is a flowchart illustrating another embodiment of the control method for home appliances provided in this application, the method, applied in the cloud, includes:

[0063] Step 41: Receive multimodal data sent by the wearable device, including at least one of voice data, image data, and posture data.

[0064] Step 42: Analyze the multimodal data using a large language model, generate control commands, and send the control commands to the home appliances corresponding to the multimodal data.

[0065] Step 43: Receive device status feedback from home appliances.

[0066] Step 44: Send device status to wearable device.

[0067] In one application scenario, taking wearable devices as smart glasses as an example, the control process for home appliances with multimodal input is as follows:

[0068] 1) Users issue control commands through smart glasses, such as saying "turn on the water heater" and looking at the target water heater.

[0069] 2) Smart glasses collect multimodal data such as voice, vision, and head posture.

[0070] 3) The multimodal data is uploaded to the cloud server after preprocessing.

[0071] 4) The cloud server uses multimodal large model to analyze data and understand user intent by combining contextual information.

[0072] 5) The cloud server generates precise control commands, which are sent to the target water heater through the home gateway.

[0073] 6) The water heater is turned on and the result is displayed. Users can visually view the remaining hot water volume and water temperature through smart glasses.

[0074] In one application scenario, taking smart glasses as an example of wearable devices, the AR-based visualization process for home appliance status is as follows:

[0075] 1) Smart home appliances (such as dishwashers) report their operating status to the home gateway in real time.

[0076] 2) The home gateway uploads the status information to the cloud server.

[0077] 3) The cloud server generates AR visualization content based on user preferences, such as displaying the current step and remaining time of the dishwasher.

[0078] 4) AR content is delivered to smart glasses.

[0079] 5) Smart glasses overlay AR content onto the user's field of vision.

[0080] 6) Users can interact with the AR interface through eye contact, gestures, or voice, such as adjusting the water temperature of a water dispenser.

[0081] In one application scenario, taking smart glasses as an example of a wearable device, the status visualization and control of a water heater are as follows:

[0082] 1) Users can use voice commands to "check the water heater status" and look at the water heater, and the smart glasses will collect relevant multimodal input.

[0083] 2) The cloud server uses a multimodal big data model to understand the user's intent and retrieves the current operating status data of the water heater, including the remaining hot water volume and water temperature.

[0084] 3) The cloud server generates an AR interface based on the user's location and line of sight, which is overlaid in the user's field of vision to intuitively display the various status indicators of the water heater.

[0085] 4) Users can further control the water heater via gestures or voice, such as adjusting the water temperature. The smart glasses send the user's control commands to the home gateway, which in turn transmits them to the water heater device.

[0086] 5) The water heater executes the user's control commands and provides real-time feedback on the latest operating status data. The cloud server updates the content displayed on the AR interface accordingly.

[0087] In one application scenario, taking smart glasses as an example of a wearable device, the status visualization and control of a dishwasher are as follows:

[0088] 1) The user says the voice command “Check the status of the dishwasher” while focusing their gaze on the dishwasher in their home.

[0089] 2) The smart glasses collect multimodal inputs such as voice and gaze, and send them to the cloud server for processing.

[0090] 3) The cloud server's multimodal big data model understands the user's query intent and pulls the current operating status of the dishwasher from the home gateway, including the washing program stage and remaining time.

[0091] 4) The cloud server generates an AR interface based on the user's location and line of sight, which is overlaid in the user's field of vision to intuitively display the dishwasher's various status information.

[0092] 5) Users can further control the dishwasher via gestures or voice, such as adjusting the water temperature or pausing / resuming the washing program. The smart glasses send the user's control commands to the home gateway, which in turn transmits them to the dishwasher.

[0093] 6) The dishwasher executes the user's control commands and provides real-time feedback on the latest operating status data. The cloud server updates the content displayed on the AR interface accordingly.

[0094] In the embodiments described above, multimodal input improves the accuracy of home appliance control and reduces operational ambiguity; context awareness makes the system more intelligent and able to predict user needs; AR status visualization provides intuitive and real-time home appliance information, enhancing the user experience; smart glasses, as an interaction medium, make home appliance control more natural and convenient; and the home control system has good scalability, making it easy to integrate new home appliances and functions, such as water dispensers and televisions.

[0095] In some embodiments, smart home control systems are becoming increasingly popular with the development of IoT and AI technologies. However, some problems still exist with related smart home control methods, such as a single input mode and limited interaction methods.

[0096] Based on this, this application proposes to utilize wearable devices to receive user-corresponding interaction data, thereby increasing the interactive input methods of the home control system. Furthermore, it utilizes a large language model to perform intent recognition on the interaction data, obtaining the target intent, and generating dynamic display data based on the target intent. This enhances the depth of understanding of the interaction data, improves the accuracy of the dynamic display data, and utilizes wearable devices as the display medium for the dynamic display data, making home appliance interaction more natural and convenient, and improving the user's interactive experience. See any of the following embodiments for details.

[0097] Referring to Figure 5, Figure 5 is a flowchart illustrating an embodiment of the interaction method for home appliances provided in this application. Applied to wearable devices, the method includes:

[0098] Step 51: Receive the user's corresponding interaction data.

[0099] The interaction data includes at least one of voice data, image data, and gesture data.

[0100] In some embodiments, the interactive data includes at least a target object. The target object may be a function of a home appliance, a functional component of a home appliance, or the home appliance itself.

[0101] In some embodiments, the target object can be a product found in a home setting, such as a mobile phone, tablet, and / or recipe.

[0102] Step 52: Use a large language model to identify the intent of the interaction data and obtain the target intent.

[0103] In some embodiments, the large language model is a pre-trained model. In some embodiments, the target intent can represent the intention to understand the target object of the home appliance.

[0104] Step 53: Generate dynamic display data based on the target intent; wherein, the dynamic display data is voice, image and / or text corresponding to the interactive data.

[0105] In some embodiments, the target object that needs to be understood can be obtained from the target intent, and then the corresponding voice, image and / or text of the target object can be obtained; dynamic display data can be generated based on the voice, image and / or text.

[0106] For example, by using a large language model to identify the instruction manual for a home appliance, the system can obtain voice, images, and / or text related to the instruction manual based on the appliance, and then create dynamic display data.

[0107] For example, if a large language model identifies that the user needs to know how a dish is made, then the system can obtain speech, images, and / or text related to the dish's preparation to create dynamic display data.

[0108] In some embodiments, the image can be a video image and / or a single frame image. That is, the dynamically displayed data can consist of video and multiple images.

[0109] Step 54: Display dynamic data.

[0110] Wearable devices are used to display dynamic data, allowing users to quickly and intuitively understand relevant content through these devices.

[0111] In some embodiments, the dynamically displayed data includes images and / or voice, which can utilize wearable devices to display images and / or voice, thereby addressing the disadvantages of users who are illiterate or insensitive to text and improving the user experience.

[0112] In this embodiment, wearable devices are used to receive user-corresponding interactive data, increasing the interactive input methods of the home control system. Furthermore, a large language model is used to perform intent recognition on the interactive data to obtain the target intent, and dynamic display data is generated based on the target intent. This enhances the deep understanding of the interactive data, improves the accuracy of the dynamic display data, and uses wearable devices as the display medium for dynamic display data, making the interaction of home appliances more natural and convenient, and improving the user's interactive experience.

[0113] Referring to Figure 6, Figure 6 is a flowchart illustrating another embodiment of the interaction method for home appliances provided in this application. Applied to wearable devices, the method includes:

[0114] Step 61: Receive the user's voice data.

[0115] In some embodiments, the voice data includes at least a target object. The target object can be a function and / or a functional component of a home appliance. For example, if the corresponding text in the voice data is: "I want to know what functions a dishwasher has," then the target object is the dishwasher.

[0116] In some embodiments, the target object can be a product found in a home setting, such as a mobile phone, tablet, and / or recipe.

[0117] Step 62: Use a large language model to perform intent recognition on the speech data to obtain the home appliances in the speech data and the target intent for the home appliances.

[0118] In some embodiments, a large language model is used to convert speech data into text, resulting in text data. Then, intent recognition is performed on the text data to identify the home appliance in the speech data and the target intent regarding that appliance. For example, if the corresponding text content in the speech data is: "I want to know what functions a dishwasher has," then the home appliance in the speech data is a dishwasher, and the target intent is to know the functions of the dishwasher.

[0119] Step 63: Generate dynamic display data based on the target intent; wherein, the dynamic display data is voice, image and / or text corresponding to the interactive data.

[0120] In some embodiments, the target object to be understood can be obtained from the target intent, which is a dishwasher. Then, the voice, images and / or text corresponding to all functions of the dishwasher can be obtained. Dynamic display data can be generated based on the voice, images and / or text.

[0121] Step 64: Display dynamic data.

[0122] In this embodiment, a wearable device is used to interact with the user and receive the user's voice data; a large language model is used to perform intent recognition on the interaction data to obtain the target intent; dynamic display data is generated based on the target intent, and the wearable device is used as the display medium for the dynamic display data, making the interaction of home appliances more natural and convenient, and improving the user interaction experience.

[0123] Referring to Figure 7, which is a flowchart illustrating another embodiment of the interaction method for home appliances provided in this application, the method, applied to wearable devices, includes:

[0124] Step 71: Receive the user's voice data and image data.

[0125] In some embodiments, image data can be acquired by a home appliance within the field of view of an image sensor on a wearable device.

[0126] For example, image data can be collected simultaneously when the user's voice data is collected.

[0127] Step 72: Use a large language model to perform intent recognition on the speech and image data to obtain the target intent.

[0128] In some embodiments, a large language model can be used to simultaneously perform intent recognition on speech data and image data, and then the target intent can be obtained by combining the content in the speech data and the target object in the image data.

[0129] In some embodiments, referring to FIG8, step 72 may be the following process:

[0130] Step 721: Use a large language model to perform content recognition on the speech data to obtain key content.

[0131] For example, the corresponding text content in the speech data is: "What is the use of this?" Using a large language model to perform content recognition on the speech data, the key content is: "I want to know the function of something." At this point, it's unknown what that "something" actually is.

[0132] Step 722: Identify the target object corresponding to the key content from the image data.

[0133] Furthermore, after extracting key content from the speech data, the large language model determines the target object corresponding to the key content from the image data. For example, if the large language model identifies a button on a range hood from the image data, then that button can be used as the target object.

[0134] Step 723: Obtain the target intent based on the key content and target audience.

[0135] In some embodiments, when the key content of the voice data is: wanting to know the function of something, and the target object in the image data is the button on the range hood, then the target intent is to know the function of the button on the range hood.

[0136] Step 73: Generate dynamic display data based on the target intent; wherein, the dynamic display data is voice, image and / or text corresponding to the interactive data.

[0137] For example, if the user's intention is to know the functions of the buttons on a range hood, dynamic display data can be generated based on this intention. This dynamic display data can be a dynamic instruction manual. Through the display on a wearable device, the user can learn the functions of the range hood's buttons.

[0138] Step 74: Display dynamic data.

[0139] In some embodiments, when displaying dynamic display data using a wearable device, the user's age group can be obtained; the voice playback speed can be determined according to the age group; and the dynamic display data can be displayed according to the voice playback speed. For example, the higher the age group, the older the user, and the slower the voice playback speed. For example, if the user is elderly, the relevant content of the dynamic display data will be played according to the voice playback speed corresponding to the elderly stage.

[0140] In this embodiment, a wearable device is used to interact with the user, receiving voice and image data between the user and the home appliance; and a large language model is used to perform intent recognition on the voice and image data to obtain the target intent; dynamic display data is generated based on the target intent, and the wearable device is used as the display medium for the dynamic display data, making the interaction with the home appliance more natural and convenient, and improving the user's interaction experience.

[0141] In this application, the wearable device can be one of a smart bracelet, a smartwatch, and / or smart glasses. When displaying dynamic display data, the display is performed according to the data type of the dynamic display data.

[0142] In some embodiments, the wearable device is smart glasses, which can overlay images and / or text from dynamically displayed data onto the visible area of ​​the smart glasses. For example, if the dynamically displayed data includes images, the images are overlaid onto the visible area of ​​the smart glasses during display.

[0143] For example, if the dynamically displayed data includes text, the text will be overlaid within the visible area of ​​the smart glasses during the display.

[0144] For example, if the dynamically displayed data includes text and images, then during the display, the text and images will be superimposed on the visible area of ​​the smart glasses in a corresponding arrangement.

[0145] If the dynamically displayed data also includes audio, then audio playback can be performed during the display process.

[0146] In some embodiments, during the display of dynamically displayed data, user instructions are received; and the presentation method and content of the dynamically displayed data are adjusted according to the user instructions. For example, the display speed is adjusted, or the accent of the playback is adjusted, such as changing from Mandarin to a dialect or from Mandarin to another language.

[0147] Referring to Figure 9, Figure 9 is a flowchart illustrating another embodiment of the interaction method for home appliances provided in this application. The method includes:

[0148] Step 91: Receive the user's request for food preparation.

[0149] The request to prepare the dish can be sent to the wearable device in the form of voice, image, or text. That is, the wearable device receives voice data, image data, and / or text data containing the request to prepare the dish.

[0150] Step 92: Input the dish requirements into the large language model to obtain the dish preparation data corresponding to the dish requirements.

[0151] Step 93: Display the dish preparation data.

[0152] During the display of dish preparation data, the system receives user instructions and adjusts the presentation method and content of the dish preparation data according to the user instructions.

[0153] In this embodiment, a wearable device is used to interact with the user, receive the user's requirements for making dishes, and use the large language model to display the dish making data corresponding to the dish requirements. This solves the problems of lack of interactivity and intuitiveness in recipes in related technologies, allowing users to understand the operation steps more intuitively and improving the user experience.

[0154] In one application scenario, taking smart glasses as an example, consider a dynamic instruction manual for a water heater based on these smart glasses. An elderly user wears the smart glasses, looks at a function button on the water heater, and issues a voice command in their dialect, "I want to know what this button does." The smart glasses capture the user's voice and visual information and transmit it to a multimodal large model (large language model). The multimodal large model recognizes the elderly user's dialect, understands the user's intention to understand the function of the button, and generates a dynamic instruction manual containing multiple modalities, including large-font text, high-contrast images, slow-speed speech, and simplified AR animations. The dynamic instruction manual is presented to the user through the smart glasses' AR display function. For example, the function of the button is displayed on the screen in large-font text and slow-speed speech, and the operation steps are overlaid on the water heater image in the form of simplified AR animations.

[0155] In one application scenario, taking wearable devices like smart glasses as an example, a multimodal braised pork recipe is generated based on a multimodal large model. An elderly user selects the braised pork dish they want to learn via their mobile phone. The system transmits the recipe information to the multimodal large model (large language model). The multimodal large model generates a braised pork recipe containing detailed text steps, clear images, slow-speed audio explanations, and segmented video demonstrations, among other modalities. The user can watch the video demonstration of the recipe, listen to the slow-speed audio explanations, or browse the detailed text steps and clear images on their mobile phone. The system can adjust the presentation and content of the recipe based on the user's learning progress and feedback, such as pausing the video or repeating the explanation, providing more detailed explanations or demonstrations when the user encounters difficulties.

[0156] Referring to Figure 10, which is a schematic diagram of the structure of an embodiment of the wearable device provided in this application, the wearable device 20 includes a processor 201, a memory 202 coupled to the processor 201, and a communication interface 203. The memory 202 stores at least one computer program, which, when loaded and executed by the processor, is used to implement the method of any of the above embodiments.

[0157] Referring to FIG11, FIG11 is a schematic diagram of an embodiment of an electronic device provided in this application. The electronic device 300 includes a processor 301, a memory 302 coupled to the processor 301, and a communication interface 303. The memory 302 stores at least one computer program. When the at least one computer program is loaded and executed by the processor, it is used to implement the method of any of the above embodiments.

[0158] The electronic device 300 can be any of the aforementioned cloud devices, wearable devices, and home gateways.

[0159] In some embodiments, the wearable device in the home control system 100 described above is communicatively connected to the cloud, receives interaction data between the user and home appliances, and sends the interaction data to the cloud. The cloud uses a large language model to perform intent recognition on the interaction data, obtains the target intent, and generates dynamic display data based on the target intent; wherein, the dynamic display data is voice, images, and / or text corresponding to the interaction data; and the dynamic display data is displayed. That is, the home control system 100 described above can implement the method of any embodiment of this application.

[0160] In summary, the home appliance interaction method, wearable device, and home control system provided in this application utilize the wearable device to receive corresponding user interaction data, increasing the interactive input methods of the home control system. Furthermore, it uses a large language model to perform intent recognition on the interaction data, obtains the target intent, and generates dynamic display data based on the target intent, thereby improving the depth of understanding of the interaction data, enhancing the accuracy of the dynamic display data, and using the wearable device as a display medium for the dynamic display data, making the home appliance interaction method more natural and convenient, and improving the user interaction experience.

[0161] In the several embodiments provided in this application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed.

[0162] If the integrated units in the other embodiments described above are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processing circuit component (processor) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0163] The above description is merely an embodiment of this application and does not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.

Claims

1. A method for interacting with home appliances, characterized in that, Applied to wearable devices, the method includes: Receive user interaction data; The target intent is obtained by using a large language model to identify the interaction data. Dynamic display data is generated based on the target intent; wherein, the dynamic display data is voice, image and / or text corresponding to the interactive data; The dynamically displayed data is shown.

2. The method according to claim 1, characterized in that, The interaction data includes: voice data, and the process of using a large language model to perform intent recognition on the interaction data to obtain the target intent includes: The large language model is used to perform intent recognition on the speech data to obtain the home appliance in the speech data and the target intent for the home appliance.

3. The method according to claim 1, characterized in that, The interactive data includes: voice data and image data. The step of using a large language model to perform intent recognition on the interactive data to obtain the target intent includes: The target intent is obtained by performing intent recognition on the speech data and image data using the large language model.

4. The method according to claim 3, characterized in that, The step of using the large language model to perform intent recognition on the speech data and the image data to obtain the target intent includes: The large language model is used to perform content recognition on the speech data to obtain key content; Determine the target object corresponding to the key content from the image data; The target intent is obtained based on the key content and the target object.

5. The method according to claim 4, characterized in that, The step of generating dynamic display data based on the target intent includes: Obtain the voice, image, and / or text corresponding to the target object; The dynamic display data is generated based on the voice, the image, and / or the text.

6. The method according to any one of claims 1-5, characterized in that, The display of the dynamic display data includes: Obtain the user's age range; The audio playback speed is determined according to the stated age group; The dynamic display data is shown according to the stated audio playback speed.

7. The method according to any one of claims 1-6, characterized in that, The method further includes: During the display of the dynamic data, user instructions are received; Adjust the presentation method and content of the dynamically displayed data according to the user's instructions.

8. The method according to any one of claims 1-7, characterized in that, The wearable device is smart glasses, and the display of the dynamic display data includes: The images and / or text in the dynamically displayed data are superimposed on the images within the visible range of the smart glasses.

9. The method according to claim 8, characterized in that, If the dynamic display data includes voice, the voice will be played during the display process.

10. The method according to any one of claims 1-9, characterized in that, The method further includes: Receive users' requests for food preparation; The dish requirements are input into the large language model to obtain the dish preparation data corresponding to the dish requirements; The preparation data for the dish is displayed.

11. A wearable device, characterized in that, The wearable device includes a processor and a memory and a communication interface coupled to the processor. The memory stores at least one computer program, which, when loaded and executed by the processor, is used to implement the method as described in any one of claims 1-10.

12. A home control system, characterized in that, The home control system includes: Cloud; Wearable devices are connected to the cloud, receive user-related interactive data, and send the interactive data to the cloud. The cloud-based system uses a large language model to perform intent recognition on the interactive data to obtain the target intent, and generates dynamic display data based on the target intent; wherein, the dynamic display data is voice, image and / or text corresponding to the interactive data; The wearable device receives and displays the dynamic display data.

Citation Information

Patent Citations

  • Method for realizing voice control over smart home based on wearable equipment

    CN107644644A

  • Voice household appliance interaction method, interaction system and computer equipment

    CN113763942A

  • Dialogue system intention recognition method and tool based on large language model

    CN116955618A

  • Multi-modal interaction system and method based on smart watch

    CN117633703A

  • Human-computer interaction method and device, electronic equipment and computer storage medium

    CN118567602A