Smart home control method and device, electronic equipment and storage medium
By recognizing user intent through a large language model and combining it with room device function information, the problems of fuzzy semantic adaptation and response latency in smart home control are solved, enabling efficient multi-device collaborative control.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-24
- Publication Date
- 2026-03-31
AI Technical Summary
Existing smart home control technologies struggle to adapt to fuzzy semantics when faced with dynamically changing environments and user preferences. Furthermore, smart assistants based on large language models suffer from excessively high response latency under concurrent control of multiple devices, failing to meet real-time control requirements.
A large language model is used to identify the intent type of user scenario needs, and corresponding reasoning methods are designed according to explicit or fuzzy intents. Room equipment function information is obtained, equipment execution parameters are determined, and latency is reduced and control accuracy is improved in fuzzy scenarios through staged reasoning.
It achieves accurate parsing of users' ambiguous intentions and collaborative control of multiple devices, reduces inference latency, and improves the intelligence and real-time response capability of smart home control.
Smart Images

Figure CN121763794A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to smart home control methods, devices, electronic devices and storage media. Background Technology
[0002] Smart home scene generation refers to the system automatically generating and executing a set of ordered control actions based on user commands, environmental perception parameters, and device capabilities to meet specific user intentions or living scenario needs. As a core supporting technology for improving living quality, smart home control systems have gradually evolved from remote control of single devices to multi-device collaboration and scenario-based services, often involving the joint control of multiple rooms and various types of devices (lighting, air conditioning, curtains, music, etc.). Its core value lies in achieving automated adjustment, energy optimization, and convenient interaction of the home environment through device interconnection and intelligent scheduling, meeting users' needs for comfortable, efficient, and safe living scenarios.
[0003] Currently, smart home control technology has penetrated into multiple home subsystems such as lighting, HVAC, and security, becoming a core application area in the field of smart living. However, as user needs upgrade from basic control to personalized and adaptive services, the adaptability and intelligence level of existing control technologies face new challenges. Summary of the Invention
[0004] The main objective of this invention is to provide a smart home control method, device, electronic device, and storage medium, aiming to achieve both high inference accuracy and low inference latency, thereby improving the intelligence of smart home control. The technical solution is as follows: In a first aspect, embodiments of this application provide a smart home control method applied to an electronic device, the method comprising: In response to receiving user scenario requirement information for a target space, obtain room equipment function information within the target space; A large language model is used to identify the intent type of the scenario requirement information; If the intent type is an explicit intent, then the large language model is used to determine the device execution parameters of the smart home devices in the target space based on the scenario requirement information and the room device function information; If the intent type is a fuzzy intent, then the scene knowledge that matches the scene requirement information is obtained, and the large language model is used to determine the device execution parameters of the smart home devices in the target space based on the scene requirement information, the scene knowledge and the room device function information. The device executes parameters to control the corresponding smart home devices within the target space.
[0005] Secondly, embodiments of this invention provide a smart home control device, comprising: The demand acquisition unit is used to acquire room equipment function information within the target space in response to receiving scenario demand information from a user for a target space. An intent recognition unit is used to identify the intent type of the scenario requirement information using a large language model. The first reasoning unit is used to determine the device execution parameters of the smart home devices in the target space based on the scenario requirement information and the room device function information if the intent type is an explicit intent. The second reasoning unit is used to obtain scene knowledge that matches the scene requirement information if the intent type is a fuzzy intent, and to use the large language model to determine the device execution parameters of the smart home devices in the target space based on the scene requirement information, the scene knowledge and the room device function information. The control unit is used to control the corresponding smart home devices in the target space using the device execution parameters.
[0006] Thirdly, embodiments of this application provide an electronic device, the electronic device including one or more processors and one or more memories, the one or more memories storing at least one computer program, the computer program being loaded and executed by the one or more processors to implement the smart home control method.
[0007] Fourthly, embodiments of this application provide a storage medium storing a computer program, which, when executed by a processor, implements the steps of the method described above.
[0008] Fifthly, embodiments of this application provide a computer program product, comprising: a computer program that, when executed by a processor of a heating, ventilation, and air conditioning (HVAC) device, enables the processor to at least implement the method described in the first aspect.
[0009] In the embodiments of this invention, in response to receiving user scenario requirement information for a target space, the system obtains room equipment function information within the target space, identifies the user's intent type for the scenario requirement information for the target space through the natural language understanding capability of a large language model, and then automatically infers device execution parameters that can satisfy the user's intent based on two types: fuzzy instructions and explicit instructions. For explicit instructions, device execution parameters are directly generated by combining the room equipment function information within the space, making the inference chain clear. By supplementing scenario knowledge for fuzzy scenarios, the control accuracy of fuzzy scenarios is improved. Attached Figure Description
[0010] To more clearly illustrate the technical solutions in the embodiments of this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0011] Figure 1 A schematic diagram of a smart home Internet of Things scenario provided in an embodiment of this invention application; Figure 2 A flowchart illustrating a smart home control method provided in an embodiment of this invention application; Figure 3 This is a flowchart illustrating a smart home control method provided in an embodiment of this invention. Figure 4 This is a flowchart illustrating a smart home control method provided in an embodiment of this invention. Figure 5 This is a flowchart illustrating a smart home control method provided in an embodiment of this invention. Figure 6 This is a flowchart illustrating a smart home control method provided in an embodiment of this invention. Figure 7 This is a schematic diagram of the architecture of a smart home control method provided in an embodiment of this invention application; Figure 8 This is a schematic diagram of the architecture of a smart home control method provided in an embodiment of this invention application; Figure 9 This is a schematic diagram of the structure of a smart home control device provided in an embodiment of this invention application; Figure 10 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this invention. Detailed Implementation
[0012] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this specification, and not all embodiments. Based on the embodiments in this specification, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this specification.
[0013] In the description of this specification, it should be understood that the terms "first," "second," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance. In the description of this specification, it should be noted that, unless otherwise expressly specified and limited, "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or devices. Those skilled in the art can understand the specific meaning of the above terms in this specification based on the specific circumstances. Furthermore, in the description of this specification, unless otherwise stated, "multiple" means two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, and B alone. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship.
[0014] Please see Figure 1 , Figure 1 This invention provides a schematic diagram of a smart home control method scenario. A smart home scenario refers to the interconnection of various devices in the home, such as lights, air conditioners, curtains, and music players, through Internet of Things (IoT) technology, enabling centralized control and collaborative operation. Users can conveniently manage these devices via voice or an app, making the home environment more comfortable, efficient, and intelligent. Exemplarily, the smart home control method in this invention can be applied to a smart home control device, or specifically executed by its control module or controller. The smart home control device can be a mobile phone, watch, computer, television, speaker, etc., and is not specifically limited. It is understood that the smart home control method provided in this application is not limited to residential homes, but can also be applied to various types of spaces such as offices, hotels, apartments, and elderly care facilities. The system can be embedded in multimodal interactive terminals such as voice assistants, home control screens, and mobile applications, achieving automated scene generation and execution through natural language input.
[0015] Among related technologies, existing smart home control technologies still have significant limitations in practical applications: On the one hand, traditional rule-based control schemes are based on fixed trigger-condition-action paradigms, relying on manually preset logic to achieve device collaboration. Not only is the control logic rigid and lacks scalability, making it difficult to adapt to dynamically changing environments and user preferences, but it also lacks the ability to resolve fuzzy semantics, only responding to explicit commands, thus limiting the intelligent interactive experience. On the other hand, while intelligent assistants based on large language models improve semantic understanding capabilities, they struggle to balance reasoning accuracy and latency. The high accuracy brought by deep reasoning is often accompanied by a surge in computing resource consumption, especially in scenarios with concurrent control of multiple devices. The generation of multiple thought chains leads to excessively high response latency, failing to meet real-time control requirements and hindering the practical application of the technology.
[0016] To address the aforementioned issues, this invention provides a smart home control method. Large Language Modeling (LLM) is used for smart home control, and a clear and simple inference architecture is designed for LLM to reduce inference latency. In response to receiving user scenario requirements for a target space, the method acquires the functional information of room devices within the target space. Then, LLM distinguishes the intent type of the user's input scenario requirements and designs corresponding inference criteria for different intent types. If the intent type is explicit, LLM determines the device execution parameters of the smart home devices in the target space based on the scenario requirements and room device functional information. If the intent type is fuzzy, scenario knowledge matching the scenario requirements is acquired to supplement the user's fuzzy intent. Then, LLM determines the device execution parameters of the smart home devices in the target space based on the scenario requirements, scenario knowledge, and room device functional information, thereby improving the control accuracy of fuzzy scenarios.
[0017] The smart home control method provided in this specification will be described in detail below with reference to specific embodiments.
[0018] Please see Figure 2 This is a flowchart illustrating a smart home control method provided in an embodiment of the present invention. Figure 2 As shown, the smart home control method provided in the embodiments of this invention may include the following steps S101-S105.
[0019] S101, in response to receiving user scenario requirement information for the target space, obtain room equipment function information within the target space; In one embodiment, natural language input is received from a user terminal. This input represents the user's scenario requirements for a target space. The user can input commands via a voice assistant, mobile application, or home control screen. The system also supports secondary confirmation or correction based on user feedback, thereby achieving interactive control. The target space is the space the user wants to control or the space the user is currently in. For example, the user can remotely input requirements to control smart home devices, or the user may be at home and need to control smart home devices.
[0020] Optionally, upon receiving the scenario requirement information, a valid device verification is performed. This step confirms whether there are devices supporting scene object model capabilities, or smart home devices supporting smart home control, in the user's target space. If a device is found to be offline, disconnected, or lacking scene object model capabilities, the system will directly return the corresponding error code to prevent the generation of unexecutable action instructions during the downstream inference stage. If valid devices exist, the room device function information within the target space is obtained. In one embodiment, the room device function information includes rooms within the target space that support scene object model capabilities, the devices within those rooms, and the function information of the devices.
[0021] For example, different target spaces may be equipped with different smart home devices. For instance, smart home devices in an office setting may include a conference tablet, projector, fresh air system, lights, etc.; smart home devices in a classroom setting may include a blackboard light, air conditioner, multimedia player, smart doorbell, etc.
[0022] S102, use a large language model to identify the intent type of the scenario requirement information; In one embodiment, the large language model is a natural language processing model with a large number of parameters built based on deep learning technology, capable of understanding, generating, and reasoning about human language. After receiving scene requirement information, the scene requirement information can be transmitted to the large language model, which then performs intent parsing. To improve the accuracy and efficiency of reasoning, the scene requirement information is first categorized into intent types. Intent types include explicit intents and fuzzy intents. Explicit intents refer to instructions in the scene requirement information that are clearly stated, specific, and directly executable. For example, "Turn on the living room light at 3 PM" or "Turn on the conference tablet and light in the conference room when there are people in the meeting room." Fuzzy intents refer to instruction types where the requirement is expressed as an abstract, scenario-based, or state-based description. This type of intent cannot be directly mapped to device control logic and requires the system to combine scene knowledge, device information, user preferences, and other supplementary data for reasoning and parsing. For example, "I want to relax in my bedroom," "Prepare for the afternoon departmental meeting," or "Generate a leisure scene." The natural language understanding capability of the large language model can be used to determine the intent type corresponding to the scene requirement information. For example, a prompt can be preset in the large language model to require it to perform intent recognition.
[0023] It should be noted that the above-mentioned large language model can adopt GPT series models, deep inference models (such as DeepSeek-R1), as well as DeepSeek-v3 and Qwen3-235b, etc. For the sake of both latency and accuracy, GPT-4.1 is used as the base model in this embodiment.
[0024] S103, if the intent type is an explicit intent, then the large language model is used to determine the device execution parameters of the smart home devices in the target space based on the scenario requirement information and the room device function information; In one embodiment, if the intent type is an explicit intent, such as "turn on the living room light at 3 p.m." or "turn on the conference tablet and light in the conference room when there are people in the conference room", only the relevant room equipment function information will be passed in when inferring the specific device execution parameters of the task. The large language model can determine the room (such as the living room) that can meet the needs of the scenario based on the scenario requirement information, determine the equipment in the room (such as the living room light), and then determine the specific device execution parameters (such as turning it on after a person is detected).
[0025] S104, if the intent type is a fuzzy intent, then obtain the scene knowledge that matches the scene requirement information, and use the big language model to determine the device execution parameters of the smart home devices in the target space based on the scene requirement information, the scene knowledge and the room device function information; In one embodiment, the fuzzy intent, such as "generate a leisure scene," not only inputs device function knowledge but also the matched scene knowledge (including device function knowledge, color knowledge, and device action configuration within the scene). Specifically, retrieving scene knowledge matching the scene requirement information can be done using information retrieval algorithms such as BM25 (Best Matching 25) to perform text matching between the scene requirement information and preset scenes, thereby determining the matched scene knowledge. For example, "I want to relax in my bedroom" can match a relaxation scene. The scene knowledge can be configured with scene configuration templates for various spaces, such as a relaxation scene corresponding to "dim the lights, close the curtains, and play music." Then, the device execution parameters are inferred by combining the room device function information within the target space. For example, if the current user's desired room is a bedroom, then it is necessary to determine from the room device function information whether a bedroom exists and whether there are devices in the bedroom capable of "dim the lights, close the curtains, and play music." The finally inferred device execution parameters include at least the device's triggering conditions and specific execution actions, with the specific execution actions precisely refined to the device function or parameter level, such as on / off, color, brightness, temperature, wind speed, and volume.
[0026] By replacing traditional rule-based script logic with natural language understanding capabilities, the system can accurately interpret users' ambiguous semantics and complex intentions. For example, when a user issues non-deterministic commands such as "make the living room brighter" or "enter rest mode," the system can automatically identify the potential target and associated devices, enabling collaborative control of multiple devices and rooms. This breaks through the constraints of traditional scene arrangement, which is limited by trigger conditions and fixed templates.
[0027] S105, the device executes the parameters to control the corresponding smart home devices in the target space.
[0028] In one embodiment, after determining the device execution parameters, the execution parameters of each device are sent to the corresponding smart home device or the device master control gateway, and the device is driven to perform preset actions, so as to realize the implementation of user scenario requirements.
[0029] In one feasible implementation, this applies to smart home devices (such as smart speakers, smart locks, and Wi-Fi air conditioners) that have independent network connectivity and are directly connected to the system platform. The system directly sends standardized execution parameters to the communication interface of the corresponding device through the device communication link list in the device management module. For example, if the target space is a home, the execution parameter "set the living room air conditioner to 26℃" is directly sent to the Wi-Fi communication module of the living room air conditioner, and the device immediately executes the temperature adjustment action upon receiving it.
[0030] In another feasible implementation, this is applicable to device clusters based on wireless bus networking (such as zoned blackboard lights in a classroom or device groups for meeting scenarios in an office). The system first packages the execution parameters of all devices in the space and sends them to the smart home central control gateway in the target space. The gateway then distributes, generates, and executes instructions based on the device address codes. For example, if the target space is a classroom, the execution parameters of multiple devices, such as the blackboard light zone brightness adjustment, air conditioning group temperature control, and multimedia volume settings corresponding to the "teaching mode," are uniformly sent to the classroom central control gateway. The gateway then distributes these parameters to the blackboard light driver module, air conditioning control module, and multimedia player, respectively, to achieve multi-device collaborative operation.
[0031] Optionally, after sending the device execution parameters to the corresponding smart home devices, the user can receive a prompt message such as "Scene configuration completed" on the user's end.
[0032] In the embodiments of this invention, in response to receiving scene requirement information from a user for a target space, the system obtains the functional information of room equipment within the target space. A large language model is used to distinguish the intent type of the user's input scene requirement information, and corresponding reasoning methods are designed for different intent types. If the intent type is explicit, the large language model determines the device execution parameters of the smart home devices within the target space based on the scene requirement information and the room equipment functional information. If the intent type is ambiguous, scene knowledge matching the scene requirement information is obtained to supplement the user's ambiguous intent. Then, based on the scene requirement information, scene knowledge, and room equipment functional information, the device execution parameters of the smart home devices within the target space are determined, thereby improving the control accuracy of ambiguous scenes.
[0033] Please see Figure 3 This is a flowchart illustrating a smart home control method provided in an embodiment of this invention. Figure 3 As shown, the smart home control method of the present invention embodiment may include the following steps S201-S202.
[0034] S201, if the intent type is an explicit intent, then the large language model is used to determine the target execution room in the target space based on the scenario requirement information and the room equipment function information; In one embodiment, the target execution room is inferred by combining scene requirement information and room equipment function information. The large language model first performs a structured decomposition of the scene requirement information with explicit intent, extracting direct or indirect related information related to space. In a feasible implementation, if the instruction explicitly includes a room name (such as "living room" or "third-floor conference room"), the model directly marks that room as the target execution room.
[0035] In another feasible implementation, if the instruction does not explicitly mention a room, but contains implicit spatial association elements such as device type and functional scenario (e.g., "detect PM2.5" or "turn on the air purifier"), the model matches the device type with the "device-room" attribution relationship in the room's device function information through semantic mapping, and initially narrows down the range of candidate rooms.
[0036] For example, for a request with a clear intent such as "turn on the air purifier when PM2.5 is detected to be too high", the model infers from the room's device function information that the device that can detect PM2.5 (such as the central control screen) is in the living room and the air purifier is on the balcony. Finally, the model infers and outputs that the target execution room is "living room|balcony".
[0037] S202, determine the device execution parameters of the target execution room based on the scenario requirement information and the room device function information.
[0038] In one embodiment, based on a determined target execution room, and combining scene requirement information with room equipment function information, standardized and executable equipment execution parameters are generated through precise mapping reasoning using a large language model. Specifically, the model and functional parameters of the execution equipment in the target execution room are determined. The room equipment function information includes adjustable parameters for different equipment types (such as temperature, mode, and fan speed for air conditioners, and brightness, color temperature, and zoning for lights), parameter value ranges (such as temperature 16℃-30℃, brightness 0%-100%), and default parameter values.
[0039] If the scenario requirement information explicitly specifies parameter values (such as "26℃" or "80% brightness"), they can be directly mapped to the corresponding functional parameters of the device. If the requirement is not explicitly specified (such as "turn on the air conditioner"), the default parameter values can be used, or the user's historical preference parameters can be associated (such as defaulting to cooling mode if the user frequently uses it), or environmental perception parameters can be obtained. The functional parameters can be determined by combining the user's historical preference parameters and environmental perception parameters.
[0040] Optionally, in one embodiment, determining the device execution parameters of the target execution room based on the scenario requirement information and the room device function information includes: S2021, The task planning module of the large language model is used to generate the scene planning task corresponding to the scene requirement information; In one embodiment, the large language model comprises two parts: a task planning module and a task reasoning module, used to implement phased, multi-step reasoning. The task planning module transforms complete natural language requests into executable task frameworks, that is, it generates scenario planning tasks based on scenario requirement information, providing semantic constraints and reasoning scope for subsequent refined reasoning of device execution parameters. The task reasoning module performs reasoning based on the tasks generated by the task planning module.
[0041] Scene planning tasks can be one or more. For simple instructions (such as "turn on the living room lights"), a single scene planning task, "turn on the living room lights," can be generated. For complex instructions (such as "turn on the air conditioner ten minutes before I get home at 6 p.m. on a weekday afternoon"), the complex requirement is broken down into multiple logically related short chain subtasks, such as identifying whether it is 6 p.m. on a weekday afternoon and the task of "turn on the air conditioner."
[0042] By breaking down the thought process of a long-chain large language model into structured short-chain sub-tasks, the inference latency is significantly reduced, while maintaining the logical consistency and interpretability of scene generation, thus achieving a dynamic balance between inference speed and generation quality.
[0043] Optionally, the scenario planning task may include trigger condition subtasks, device action subtasks, and message notification subtasks. The task planning module using a large language model simultaneously infers the trigger condition subtasks, device action subtasks, and message notification subtasks corresponding to the scenario requirement information.
[0044] For example, when a user issues the command "Turn on the air conditioner ten minutes before I get home at 6 PM on a weekday afternoon," the system can split the task into two short chain subtasks: the trigger condition subtask "Ten minutes before I get home at 6 PM on a weekday afternoon" and the device action subtask "Turn on the air conditioner." Furthermore, since the user hasn't specified a notification requirement for this scenario, the notification subtask can be identified as "None," or it can be expanded by the large language model to generate the notification subtask "Notify the user of completed scenario settings."
[0045] S2022, using the task reasoning module of the large language model, the device execution parameters of the target execution room are determined based on the scenario planning task and the room equipment function information.
[0046] In one embodiment, a task reasoning module using a large language model is employed to infer the device execution parameters corresponding to the scene planning task by combining the room equipment function information of the target space. If there are multiple scene planning tasks, the system infers how to implement each scene planning task based on the room equipment function information, that is, it generates the device execution parameters required to complete these scene planning tasks.
[0047] Optionally, in one embodiment, the scene planning task may include a trigger condition subtask, a device action subtask, and a message notification subtask. For example, for the trigger condition subtask "I will return home 10 minutes earlier than 6 PM on a weekday," a device supporting time-triggered execution is matched, generating parameters such as "detect current time, frequency 1 time / minute, judgment condition is 5:50 PM on a weekday." For the execution action subtask, the target execution room is first determined to be the "living room" (equipped with a smart air conditioner), and then the air conditioner is matched, generating device execution parameters such as "power on, 26℃, automatic mode." For example, the device supporting time-triggered execution can be an air conditioner. Furthermore, the device may also support other functions, such as the central control screen being used to detect the presence / absence status and humidity and temperature, for example, "When the humidity in the living room exceeds 60%, turn on the air conditioner's dehumidification mode," where the central control screen triggers the action when the humidity exceeds 60%.
[0048] In one feasible implementation, in order to better match the rules of subtask division, the room device function information can be divided according to two different device types: triggering task and executing task. That is, the room device function information includes room-device information-function mapping that supports conditional events and room-device information-function mapping that supports device action events.
[0049] In this embodiment of the invention, when the intent type is explicit, a large language model is used to determine the target execution room within the target space based on scene requirement information and room equipment function information, and then to determine the device execution parameters of the target execution room based on the scene requirement information and room equipment function information. By controlling the large language model to adopt a phased inference chain, first performing the execution room inference stage, and then performing specific parameter inference on the device execution parameters within the target execution room, the inference latency is significantly reduced, while maintaining the logical consistency and interpretability of scene generation.
[0050] Please see Figure 4 This is a flowchart illustrating a smart home control method provided in an embodiment of this invention. Figure 4 As shown, the method may include the following steps S301-S303.
[0051] S301, if the intent type is a vague intent, then obtain the scene knowledge that matches the scene requirement information; Understandably, deep reasoning-based inference models (such as DeepSeek-R1 and O3) typically output complete thought chains during the generation process. While they possess strong interpretability for complex tasks, their reasoning text often explicitly exposes the business logic and inference guidance details in the prompt, making it difficult to present directly to end users. Furthermore, the output latency of these deep reasoning models during the inference phase is usually between 1 and 2 minutes, which is excessively time-consuming for home control tasks with high real-time requirements. If only non-reasoning models are used instead, problems such as failing to understand user semantics, missing actions, or incomplete scene coverage can easily arise when facing ambiguous goals (e.g., "make the room more comfortable" or "enter sleep mode").
[0052] Based on this, in one embodiment, for fuzzy intents, the large model first determines the execution room in the scene knowledge that satisfies the scene requirements based on the scene requirement information. Then, it combines the user's current room device function information to confirm whether the execution room indicated in the scene knowledge exists. If it does, the parameters configured in the scene knowledge that satisfy the scene requirements are directly reused, that is, the room, device, and specific device parameters to be controlled are determined. Users do not need to manually configure complex rules; the system can automatically generate an executable plan based on natural language.
[0053] Optionally, in one embodiment, the scene knowledge includes at least associated rooms, device functions, environmental perception parameters, and device action configurations corresponding to the environmental perception parameters. Environmental perception parameters include, but are not limited to, temperature, humidity, light intensity, and PM2.5 concentration. Different environmental perception parameters may result in different device action configurations in the scene knowledge. For example, a relaxation scene in cold winter weather might control the heating to be turned on, while a relaxation scene in hot summer weather might control the air conditioning to be turned on. Associated rooms refer to the set of rooms required to realize the scene; for example, the associated room for a "relaxation scene" is the "master bedroom." Device functions refer to the types of equipment and core capabilities required to realize the scene within the associated rooms. This is a device-dimensional constraint of the scene knowledge, clarifying "which equipment is needed" and "what capabilities the equipment must possess" to fulfill the scene requirements. For example, the device functions for a "tea room" include "brightness / color temperature adjustment of ambient lighting" and "constant temperature heating of the smart tea table." Device action configurations refer to the specific execution parameters and operating logic set for each device based on the associated rooms and device functions. For example, ambient lighting with "brightness 40% and color temperature 3000K", smart tea table with "water temperature 85℃ and water level 80%", and sunshade curtain with "opening ratio 60%".
[0054] Optionally, in one embodiment, using the large language model to obtain scene knowledge matching the scene requirement information includes: S3011, The large language model is used to match the scene requirement information with the scene tags corresponding to each pre-stored scene to obtain the target scene tag that matches the scene requirement information; In one embodiment, a library of typical scenes covering multiple target spaces (home, office, classroom) is pre-generated, and standardized scene labels are configured for each pre-stored scene.
[0055] For example, the scene labels for relaxation scenarios could be relaxation, ease and comfort, and soothing before bed; the scene labels for meeting scenarios could be meeting room meetings or business negotiations; the scene labels for lunch break scenarios could be lunch break, quiet and dark, and rest and relaxation. The scene labels for self-study scenarios could be self-study in a study room, quiet and bright, and focused learning.
[0056] The BM25 word segmentation retrieval method directly matches the fuzzy intent scenario requirements input by the user with the tags of each pre-stored scenario. For example, if the user's requirement is "I want to relax in my bedroom," the tag "relax" is directly matched as the target scenario tag. Since the user specified the target execution room as the bedroom, the search can be limited to checking if there are any devices and their parameters that meet the relaxation scenario within the bedroom. If the user's requirement is "I want to take a nap in the office," the tag "lunch break" is directly matched as the target scenario tag. Since the scenario requirement information specifies the office, the target execution room is limited to the office, and the scenario knowledge is applied within that office.
[0057] S3012, Obtain scene knowledge of the pre-stored scene corresponding to the target scene label.
[0058] In one embodiment, scenario knowledge corresponding to the target scenario tag is retrieved from the system scenario knowledge database as supplementary knowledge for this scenario reasoning. Scenario knowledge is a collection of "device combinations, functional parameters, and execution logic" required to realize the scenario. Its core function is to transform vague intentions into a feasible device control framework, bridging the information gap between abstract requirements and concrete execution. For example, the scenario knowledge corresponding to a relaxation scenario could be: curtains are fully closed, smart lights are adjusted to 30% brightness and 2700K color temperature, smart air conditioner is turned on in 24℃ automatic mode, and music devices play soft music at 40% volume.
[0059] S302, using the large language model to obtain the scene execution room in the scene knowledge, and determining whether there is a target execution room in the target space that matches the scene execution room based on the room device function information; S303, if so, then the target execution device of the target execution room and the device execution parameters of the target execution device are determined based on the scenario knowledge.
[0060] In one embodiment, the scene execution room is obtained from the scene knowledge. After obtaining the scene execution room based on the scene knowledge, it is determined whether a scene execution room exists based on the room device function information of the target space. The scene execution room existing in the target space or a room similar to a scene execution room is determined as the target execution room. Then, based on the device execution parameters stored in the scene knowledge corresponding to the matched scene execution room, the device execution parameters of the current target execution device are determined. It is understood that when generating the device execution parameters, it is also necessary to combine the room device function information to determine whether the execution device configured in the scene knowledge exists in the target execution room and whether it supports the control parameters in the device action configuration.
[0061] For example, for a request with a vague intent such as "help me create an afternoon tea scene", the model will match the scene knowledge of the afternoon tea scene, and then obtain the associated rooms in the scene knowledge. For example, if the scene knowledge recommends a tea room, restaurant, living room, and balcony, then these rooms will be used as the scene execution rooms.
[0062] For example, if the recommended execution rooms in the scenario knowledge are tea room, dining room, living room, and balcony, and the actual rooms in the user's home are tea room and scenic balcony, then the inference output target execution room is "tea room|scenic balcony". Further, based on the relevant knowledge about tea room and scenic balcony in the afternoon tea scenario stored in the scenario knowledge, the target execution devices and corresponding device execution parameters in the current user's room are determined. For example, the scenario knowledge presets the core device combination for the afternoon tea scenario (ambient lighting, smart tea table, sunshade curtain, etc.), recommended parameters (lighting brightness 40%, tea table constant temperature 85℃), and execution logic (adjust the balcony environment first, then start the tea room equipment), providing a standardized execution framework for fuzzy intents. This is then adapted to the user's actual space / devices (e.g., if the user only has a "scenic balcony" and no ordinary balcony, it automatically replaces it with the scenic balcony's sunshade curtain and fresh air system), avoiding the generation of invalid parameters for which the user has no supported devices. If the target execution room contains smart ambient lighting, a smart tea table, a smart aromatherapy diffuser, or a smart sunshade curtain, then the corresponding device execution parameters stored in the scenario knowledge are configured accordingly. Compared to large language models that directly apply long-chain reasoning, this method solidifies the optimal execution plan through scenario knowledge, which not only ensures the rationality of the actions but also avoids the model "imagining" parameters that do not conform to the actual device.
[0063] Understandably, when the intent type is fuzzy, for fuzzy scenarios that match the scene knowledge, inference is skipped, and the device configuration indicated in the scene knowledge is directly used as the device's specified parameters. For scenarios that do not match the scene knowledge, such as when there are no available devices in the target execution room or when the device's supported functional parameters do not meet the configuration in the scene knowledge, the large language model will still infer for these execution devices.
[0064] In one embodiment, determining the target execution device in the target execution room and the device execution parameters of the target execution device based on the scenario knowledge includes the following steps S3031-S3032: S3031, The task planning module of the large language model is used to generate the scene planning task corresponding to the scene requirement information; In one embodiment, a task planning module using a large language model generates scene planning tasks corresponding to scene requirement information. This transforms abstract requirements into executable subtasks, avoiding logical confusion in subsequent parameter reasoning. For example, if the scene requirement information is "I want the environment to look warm and inviting," the resulting scene planning task could be "Create a warm and inviting scene."
[0065] Optionally, the scenario planning task may include trigger condition subtasks, device action subtasks, and message notification subtasks. The task planning module using a large language model simultaneously infers the trigger condition subtasks, device action subtasks, and message notification subtasks corresponding to the scenario requirement information.
[0066] For example, if the scenario requirement information is "create a warm and cozy atmosphere throughout the house", based on the above-mentioned pre-divided sub-tasks, we can obtain the following sub-tasks: trigger condition: none (user actively triggers), device action: adjust the color temperature and brightness of lights in each room, turn off strong light devices, turn on soft background music, and message notification: none.
[0067] S3032, using the task reasoning module of the large language model, the target execution device corresponding to the scene planning task and the device execution parameters of the target execution device are determined based on the scene knowledge and the target execution room.
[0068] In one embodiment, in the task reasoning module, the target execution room determined by the scene knowledge and the scene knowledge are combined to continue reasoning about the target execution device required for the scene planning task and the device execution parameters of the target execution device.
[0069] Optionally, environmental perception parameters and / or user dialogue history can also be acquired. Based on scene knowledge, scene planning tasks, room equipment function information, environmental perception parameters, and / or user dialogue history, the device execution parameters for the target execution room can be determined. User dialogue history refers to the user's historical interaction records with the system, including information such as the user's past scene needs and device parameter preferences, used to improve the personalized adaptation of device execution parameters. Environmental perception parameters refer to real-time environmental data collected by sensor devices within the target space, including but not limited to temperature, humidity, light intensity, and PM2.5 concentration.
[0070] Please see Figure 5This is a flowchart illustrating a smart home control method provided in an embodiment of the present invention. In one embodiment, the task reasoning module employing the large language model determines the target execution device corresponding to the scene planning task and the device execution parameters of the target execution device based on the scene knowledge and the target execution room, including the following steps S30321-S30322: S30321, using the task reasoning module of the large language model, the target execution device corresponding to the scene planning task is determined based on the scene knowledge and the target execution room; Specifically, the task reasoning module can first determine the target execution device corresponding to the scene planning task based on scene knowledge and the target execution room. For example, if the scene planning task includes "light color temperature adjustment" and "kettle constant temperature heating," it retrieves the list of valid devices in the target execution room and filters out devices with the corresponding functions, such as all the lights and the kettle. Given that the scene knowledge requires the color temperature to be adjusted to 2700K, it then filters out devices from all the lights that can be adjusted to 2700K as the target execution device. Similarly, if the scene knowledge requires adjusting the color temperature of a table lamp, then a table lamp can be selected from all the lights as the target execution device.
[0071] S30322, If the scenario knowledge includes the device action configuration corresponding to the target execution device, then the device action configuration is used as the device execution parameter of the target execution device; In one embodiment, if the scene knowledge contains the device action configuration (i.e., preset standard parameters) corresponding to the current target execution device, then the configuration is directly called as the device execution parameter. Specifically, the standard parameters in the scene knowledge are bound to the target execution device. For example, if the afternoon tea scene knowledge presets "tea room ambient light brightness 40%, color temperature 3000K", then the parameter is directly assigned to the smart ambient light of the tea room.
[0072] S30323, if the scene knowledge does not include the device action configuration corresponding to the target execution device, then the device execution parameters of the target execution device are generated based on the environmental perception parameters of the target space and the scene planning task.
[0073] In one embodiment, if the scene knowledge does not contain the device action configuration corresponding to the target execution device, parameters are dynamically generated based on the environmental perception parameters of the target space and the scene planning task. For example, if the user's requirement is "I need to make tea and adjust it to a cozy atmosphere," and the scene knowledge does not indicate that a cozy atmosphere requires "kettle constant temperature heating," but the identified user task includes an execution task for the target execution device "kettle," and the scene knowledge does not store the control parameters of the kettle, then the device execution parameters such as the temperature and water volume used by the kettle can be inferred automatically. For example, under the "cozy atmosphere" requirement, if the current light intensity is 200 lux, the inferred light color temperature is 2700K.
[0074] For example, if the scene knowledge only indicates that the ambient light is turned on at 40% brightness, but no color temperature is set, the large language model can further infer the color temperature. It should be noted that environmental perception parameters are not necessarily required during the large language model's inference process. The large language model can determine whether to refer to environmental perception parameters based on the actual situation. For example, for "kettle constant temperature heating," the model directly infers that the kettle temperature needs to be boiling to meet the requirements for brewing tea; in this case, environmental perception parameters are not needed.
[0075] S304, if no scene knowledge matching the scene requirement information is obtained, then identify the keywords in the scene requirement information and obtain the preset prompt words corresponding to the keywords; In one embodiment, for cases where there is no matching scene knowledge for an ambiguous intent, the large language model performs semantic parsing and core keyword extraction on the ambiguous intent command without matching scene knowledge, obtaining preset prompt words corresponding to the keywords. These preset prompt words are used to guide the large language model to recognize the ambiguous intent. For example, meaningless interjections (such as "I want," "help me," "a little") are removed, and core descriptive words are retained, such as "I want to do my homework in the living room, it feels a little glaring," where the extracted keywords can be "do homework" and "glaring."
[0076] For example, a set of fuzzy keywords can be pre-defined, and a preset prompt can be designed for each keyword in the set. For instance, the keyword "stuffy" is an ambiguous instruction, and the preset prompt could be a manually interpreted and annotated message: prioritize the activation of the fresh air system (external circulation mode) or window openers (opening / closing 20%-30%), and if no such equipment is available, adjust the air conditioner to ventilation mode. Another example is the keyword "doing homework," where the preset prompt could be: ensure sufficient and soft lighting, with light brightness at 60%-80% and a color temperature of 4000K-5000K (natural light color temperature), avoid glare, and potentially disable entertainment devices or activate noise reduction mode.
[0077] S305, based on the scenario requirement information, the preset prompt words, and the room device function information, determine the device execution parameters of the smart home devices in the target space.
[0078] In one embodiment, by adding preset prompt words to the large language model, it can be made easier for the large language model to understand ambiguous intentions. Then, by combining scene requirement information and room device function information, it can infer the device execution parameters of smart home devices in the target space, eliminating the need for multiple rounds of thought chain reasoning and shortening the parameter generation time. Based on the instruction, the large language model can first filter the device type specified by the prompt word from the available devices (e.g., the keyword "stuffy" prioritizes fresh air systems, window openers, and air conditioners); match device functions to generate basic parameters: based on the functions of the device, basic parameters are generated according to the parameter rules given by the prompt word (e.g., fresh air systems are set to external circulation and medium fan speed; air conditioners are set to air supply mode); if the device specified by the prompt word does not exist, it is adjusted according to the alternative schemes of the prompt word (e.g., if there is no fresh air system, it is switched to air conditioner air supply mode); if the device does not support the parameters of the prompt word (e.g., air conditioners without air supply mode are adjusted to cooling 26℃ low fan speed).
[0079] In this embodiment of the invention, when the intent type is fuzzy, scene knowledge matching the scene requirement information is obtained. A large language model is used to determine the recommended scene execution room from the scene knowledge. Then, the room's device function information is combined to determine whether there is a target execution room in the target space that matches the scene execution room. If so, the target execution device and the device execution parameters of the target execution room are determined based on the scene knowledge. This solves the problem that pure reasoning schemes based on large language models require multiple rounds of thought chain generation when processing fuzzy intents, resulting in high inference latency and difficulty in meeting the real-time control requirements of smart homes. This method significantly improves inference efficiency through a lightweight process of scene tag matching and scene knowledge retrieval.
[0080] In one embodiment, the inference precision scaling is controlled during model inference. The device information hierarchy in the room device function information includes multiple different levels, each representing a different output precision. When controlling the output device execution parameters of the large language model, the device information hierarchy of the target execution device can be determined based on the current computing resources. For example, if computing resources are sufficient, a smaller level of device information is output, i.e., more specific device information is output; if computing resources are sufficient, a higher level of device information is output.
[0081] In one embodiment, the device information hierarchy in the room device function information, from largest to smallest, includes device category, device type, and device name. Please refer to [link / reference]. Figure 6This is a flowchart illustrating a smart home control method provided in an embodiment of this invention. Figure 6 As shown, the method may include the following steps S401-S402.
[0082] S401, Based on the room equipment function information, obtain all equipment names under the target equipment model to obtain the target equipment name; Specifically, taking lights as an example, the equipment category is "lights," the equipment model includes spotlights / ceiling lights, and the equipment names include spotlight 1 / spotlight 2 / ceiling light 1 / ceiling light 2. To control the output dimension of the large language model, the output precision of the equipment execution parameters is controlled so that the target execution device is the target equipment model; broader equipment categories and finer-grained equipment names are not output. That is, the equipment execution parameters include the target execution room, the target equipment model, and the corresponding function value parameters for the target equipment model.
[0083] Since only the target device model is output during inference, and it is necessary to specifically map it to each device, a post-processing module can be added after the large language model. The post-processing module determines the target device name corresponding to the target device model based on the device information hierarchy stored in the room device function information.
[0084] For example, if the operation controlling the lights directly outputs "all lights" as the device model, without outputting a specific light model, then it covers all lights within the "lights" device category. This is because device name granularity and different light models often involve outputting more than 10 devices. These redundant token outputs significantly increase inference latency and also lead to a decrease in inference accuracy. Therefore, we limit the inference granularity to device models and all lights. In the large language model, we set up a post-processing module to expand the precision of this type of output, matching the device model to the corresponding unique device category. Since the device model is all lights, we can query all device models under this device category and the names of the devices belonging to all device models to obtain the target device name. By using this method of scaling inference precision, we achieve the same inference effect while ensuring inference accuracy and latency.
[0085] Optionally, before obtaining the target device name by acquiring all device names under the target device model based on the room device function information, the process further includes: S4011, based on the room equipment function information, filter abnormal execution parameters from the target execution room, the target execution device, and the function value parameters corresponding to the target execution device.
[0086] In one embodiment, the post-processing module first verifies the target execution room, target execution device, and corresponding function value parameters output by the model, filtering out abnormal execution parameters to avoid situations such as ungeneralized rooms, incorrect device information (e.g., incorrect device model), and function value parameters exceeding the device's control range due to model illusion. Abnormal execution parameters can be one or more of the target execution room, target execution device, and function value parameters. Abnormal execution situations include the model outputting repeated device execution actions, and different function value parameters for the same type of device in the same room. Room device function information is used to provide auxiliary reference; for example, if the output device model does not exist in the target execution room, it is necessary to check the room device function information.
[0087] Optionally, after filtering abnormal execution parameters, a unified format alignment can be performed to ensure that the device execution actions passed to downstream projects are accurate and concise.
[0088] S402, map the function value parameters to each target device name, generate refined device execution parameters, and control the corresponding smart home devices in the target space based on the refined device execution parameters.
[0089] In one embodiment, the device execution parameters also include functional value parameters, such as brightness 80%, color temperature 2000K, color blue, and saturation 50%. In a feasible implementation, the functional value parameters can be directly assigned to the smart home device corresponding to each determined target device name to obtain refined device execution parameters.
[0090] Optionally, each smart home device model may have control limitations (such as parameter value range and control amplitude for each adjustment). In another feasible implementation, the functional value parameters output by the model are mapped to the closest control range to obtain refined device execution parameters. For example, if the color temperature control amplitude of ceiling light 1 is 100K, and the output functional value parameter is to be adjusted to 2050K, then the parameter may be mapped to 2100K or 2000K.
[0091] Optionally, after generating the device execution parameters, the following may also be included: The large language model is used to generate a scene name based on the scene requirement information, and the device execution parameters are stored in correspondence with the scene name.
[0092] In one embodiment, the large language model can also generate a concise scene name for the current scene requirement information, facilitating the storage of the scene and its subsequent invocation by the user. For example, if the scene requirement information is "Turn on the speakers if someone is in the living room," the main actions are identified, resulting in the scene name "Someone is playing." Simultaneously, the generated device execution parameters are stored in relation to this scene name to obtain newly added scene knowledge.
[0093] In the embodiments of this invention, after the device execution parameters are output by model inference, the device execution parameters are post-processed. Based on the room device function information, all device names under the target device model are obtained to obtain the target device name. The function value parameters are mapped to each target device name to generate refined device execution parameters. Based on the refined device execution parameters, the corresponding smart home devices in the target space are controlled to realize the scaling of device execution parameters, thereby refining the device execution parameters to specific smart home devices.
[0094] Please see Figure 7 This is a schematic diagram of the architecture of a smart home control method provided in an embodiment of this invention. The large language model includes a task planning module and a task reasoning module, which implement phased short-chain reasoning of the large language model. Upon receiving scene requirement information sent by the user, the task planning module distinguishes the intent type of the scene requirement information and obtains scene knowledge from the scene knowledge base. Furthermore, the task planning module also obtains the room equipment function information of the target space, thereby reasoning out scene planning tasks based on different intent types. The purpose of the scene planning task is to transform the scene requirement information into an executable task, and then the task reasoning module infers the specific parameters required to implement the scene planning task.
[0095] Please see Figure 8 Furthermore, the large language model also includes a preprocessing module and a post-processing module. The preprocessing module acquires user historical preference parameters and environmental awareness parameters, and integrates them to obtain room equipment function information. Simultaneously, scene knowledge can also be acquired by the preprocessing module. Afterwards, the preprocessing module sends scene requirement information, scene knowledge, room equipment function information, user historical preference parameters, and environmental awareness parameters together to the task planning module. The post-processing module verifies and maps the output device execution parameters. Specifically, the device execution parameters include the target execution room, the target device model, and the corresponding function value parameters for the target device model. Abnormal execution parameters are filtered, and the target device model is mapped to the corresponding target device name. The function value parameters are then configured for each target device name. After configuring the device execution parameters, the post-processing module can provide feedback to the user, such as "Scene configuration complete."
[0096] In one embodiment, the scene planning task may include a trigger condition subtask, a device action subtask, and a message notification subtask. Corresponding to the trigger condition and device action subtasks, the preprocessing module can pre-generate two types of room and device function information: room-device information-function mapping that supports condition events and room-device information-function mapping that supports device action events. Thus, when inferring the trigger condition subtask, rooms and devices can be directly filtered from the room-device information-function mapping that supports condition events; similarly, when inferring the device subtask, rooms and devices can be directly filtered from the room-device information-function mapping that supports device action events.
[0097] To verify the impact of the key modules and model configurations of this invention on the overall system performance, an ablation experiment was designed to evaluate the effectiveness and robustness of the proposed method. (Note: The existing architecture (benchmark) uses the GPT-4.1 model and the staged short-chain inference architecture described above.)
[0098] In the module combination verification experiment, the system performance was compared between removing the task planning module and directly performing task execution module inference, and adding the planning module, to verify the independent contribution of the planning module in the overall inference chain. The results show that the absence of the task planning module does not significantly affect the total system latency, but it will lead to a decrease in the total accuracy of about 14%, indicating that the staged short-chain inference architecture has a significant benefit to the system performance.
[0099] The following will be combined with the appendix Figure 9 This application provides a detailed description of the smart home control device provided in its embodiments. It should be noted that the appendix... Figure 9 The smart home control device described herein is used to execute the instructions. Figures 1-8 The methods shown in the embodiments are illustrated for ease of explanation, showing only the parts relevant to the embodiments of this application. For specific technical details not disclosed, please refer to this specification. Figures 1-8 The example shown.
[0100] Please see Figure 9 This illustration shows a schematic diagram of a smart home control device provided in an exemplary embodiment of this application. The smart home control device can be implemented as all or part of a device through software, hardware, or a combination of both. The device 1 includes a demand acquisition unit 11, an intent recognition unit 12, a first reasoning unit 13, a second reasoning unit 14, and a control unit 15.
[0101] The requirement acquisition unit 11 is used to acquire room equipment function information in the target space in response to receiving the user's scenario requirement information for the target space; Intent recognition unit 12 is used to identify the intent type of the scenario requirement information using a large language model; The first reasoning unit 13 is used to determine the device execution parameters of the smart home devices in the target space based on the scene requirement information and the room device function information if the intent type is an explicit intent; The second reasoning unit 14 is used to obtain scene knowledge that matches the scene requirement information if the intent type is a fuzzy intent, and to determine the device execution parameters of the smart home devices in the target space based on the scene requirement information, the scene knowledge and the room device function information. Control unit 15 is used to control the corresponding smart home devices in the target space using the device execution parameters.
[0102] Optionally, the first reasoning unit 13 is specifically used to determine the target execution room in the target space based on the scene requirement information and the room equipment function information if the intent type is an explicit intent; The device execution parameters for the target execution room are determined based on the scenario requirement information and the room equipment function information.
[0103] Optionally, the first reasoning unit 13 is specifically used to generate a scene planning task corresponding to the scene requirement information using the task planning module of the large language model; The task reasoning module of the large language model is used to determine the equipment execution parameters of the target execution room based on the scenario planning task and the room equipment function information.
[0104] Optionally, the second reasoning unit 14 is specifically used to obtain scene knowledge that matches the scene requirement information if the intent type is a vague intent; The large language model is used to obtain the scene execution room from the scene knowledge, and the room device function information is used to determine whether there is a target execution room in the target space that matches the scene execution room. If so, the target execution device of the target execution room and the device execution parameters of the target execution device are determined based on the scenario knowledge.
[0105] Optionally, the second reasoning unit 14 is specifically used to generate a scene planning task corresponding to the scene requirement information using the task planning module of the large language model; The task reasoning module of the large language model uses the scene knowledge and the target execution room to determine the target execution device and the device execution parameters of the target execution device corresponding to the scene planning task.
[0106] Optionally, the second reasoning unit 14 is specifically used to employ the task reasoning module of the large language model to determine the target execution device corresponding to the scene planning task based on the scene knowledge and the target execution room; If the scenario knowledge includes the device action configuration corresponding to the target execution device, then the device action configuration is used as the device execution parameter of the target execution device; If the scenario knowledge does not include the device action configuration corresponding to the target execution device, then the device execution parameters of the target execution device are generated based on the environmental perception parameters of the target space and the scenario planning task.
[0107] Optionally, the second reasoning unit 14 is specifically used to use the large language model to match the scene requirement information with the scene tags corresponding to each pre-stored scene, so as to obtain the target scene tag that matches the scene requirement information. Obtain scene knowledge of the pre-stored scene corresponding to the target scene label.
[0108] Optionally, the scenario knowledge includes at least the associated room, device functions, and device action configuration.
[0109] Optionally, the second reasoning unit 14 is further configured to identify keywords in the scenario requirement information and obtain preset prompt words corresponding to the keywords if no scenario knowledge matching the scenario requirement information is obtained. Based on the scenario requirement information, the preset prompt words, and the room device function information, the device execution parameters of the smart home devices in the target space are determined.
[0110] Optionally, the device information hierarchy in the room device function information includes multiple different levels; the device execution parameters of the smart home devices in the target space include the target execution room, the target execution device, and the corresponding function value parameters of the target execution device; the device information hierarchy of the target execution device is determined based on the current computing resources.
[0111] Optionally, the equipment information hierarchy in the room equipment function information includes, from largest to smallest, equipment category, equipment model, and equipment name; the target execution device is the target equipment model; the control unit 15 is specifically used to obtain all equipment names under the target equipment model based on the room equipment function information to obtain the target equipment name; Optionally, the control unit 15 is specifically used to map the function value parameters to each of the target device names, generate refined device execution parameters, and control the corresponding smart home devices in the target space based on the refined device execution parameters.
[0112] Optionally, the control unit 15 is further configured to filter abnormal execution parameters from the target execution room, the target execution device, and the corresponding function value parameters of the target execution device based on the room device function information.
[0113] Optionally, the control unit 15 is further configured to generate a scene name based on the scene requirement information using the large language model, and store the device execution parameters corresponding to the scene name.
[0114] Please refer to Figure 10 This diagram illustrates a structural block diagram of an electronic device provided in an exemplary embodiment of this application. The server in this application may include one or more components such as a processor 110, a memory 120, an input device 130, an output device 140, and a bus 150. The processor 110, memory 120, input device 130, and output device 140 may be connected via the bus 150.
[0115] Processor 110 may include one or more processing cores. Processor 110 connects to various parts of the server using various interfaces and lines, and executes various functions of terminal 100 and processes data by running or executing instructions, programs, code sets, or instruction sets stored in memory 120, and by calling data stored in memory 120. Optionally, processor 110 may be implemented using at least one hardware form of Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), or Programmable Logic Array (PLA). Processor 110 may integrate one or more of the following: Central Processing Unit (CPU), Graphics Processing Unit (GPU), and modem. The CPU primarily handles the operating system, user page, and applications; the GPU is responsible for rendering and drawing the displayed content; and the modem handles wireless communication. It is understood that the modem may also not be integrated into processor 110 and may be implemented separately using a communication chip.
[0116] The memory 120 may include random access memory (RAM) or read-only memory (ROM). Optionally, the memory 120 may include non-transitory computer-readable storage medium. The memory 120 may be used to store instructions, programs, code, code sets, or instruction sets. The memory 120 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for implementing at least one function (such as touch function, sound playback function, image playback function, etc.), instructions for implementing the various method embodiments described above, etc. The operating system may be the Android system, including systems deeply developed based on the Android system, the iOS system developed by Apple Inc., including systems deeply developed based on the iOS system, or other systems.
[0117] The memory 120 can be divided into operating system space and user space. The operating system runs in the operating system space, while native and third-party applications run in user space. To ensure that different third-party applications can achieve good running performance, the operating system allocates corresponding system resources for each application. However, different application scenarios within the same third-party application have different requirements for system resources. For example, in local resource loading scenarios, third-party applications have high requirements for disk read speed; in animation rendering scenarios, third-party applications have high requirements for GPU performance. Since the operating system and third-party applications are independent of each other, the operating system often cannot promptly perceive the current application scenario of a third-party application, resulting in the operating system's inability to adapt system resources accordingly.
[0118] In order for the operating system to distinguish the specific application scenarios of third-party applications, it is necessary to establish data communication between the third-party applications and the operating system. This would allow the operating system to obtain the current scenario information of the third-party applications at any time, and then perform targeted system resource adaptation based on the current scenario.
[0119] The input device 130 is used to receive input instructions or data, and includes, but is not limited to, a keyboard, mouse, camera, microphone, or touch device. The output device 140 is used to output instructions or data, and includes, but is not limited to, a display device and a speaker. In one example, the input device 130 and the output device 140 can be combined, and the input device 130 and the output device 140 can be a touch display screen.
[0120] The touch display screen can be designed as a full-screen, curved screen, or irregularly shaped screen. It can also be designed as a combination of a full-screen and a curved screen, or a combination of an irregularly shaped screen and a curved screen; however, this application does not limit the specific design in this regard.
[0121] In addition, those skilled in the art will understand that the structure of the terminal shown in the above figures does not constitute a limitation on the terminal. The terminal may include more or fewer components than shown, or combine certain components, or have different component arrangements. For example, the terminal may also include radio frequency circuits, input units, sensors, audio circuits, Wireless Fidelity (WiFi) modules, power supplies, Bluetooth modules, etc., which will not be described in detail here.
[0122] exist Figure 10 In the computer device shown, the processor 110 can be used to call computer applications stored in the memory 120 and specifically perform the following operations: In response to receiving user scenario requirement information for a target space, obtain room equipment function information within the target space; A large language model is used to identify the intent type of the scenario requirement information; If the intent type is an explicit intent, then the large language model is used to determine the device execution parameters of the smart home devices in the target space based on the scenario requirement information and the room device function information; If the intent type is a fuzzy intent, then the scene knowledge that matches the scene requirement information is obtained, and the large language model is used to determine the device execution parameters of the smart home devices in the target space based on the scene requirement information, the scene knowledge and the room device function information. The device executes parameters to control the corresponding smart home devices within the target space.
[0123] In one embodiment, when the processor 110 executes the operation of determining the device execution parameters in the target space based on the scene requirement information and the room device function information using the large language model if the intent type is an explicit intent, it specifically performs the following operations: If the intent type is an explicit intent, then the large language model is used to determine the target execution room within the target space based on the scenario requirement information and the room equipment function information; The device execution parameters for the target execution room are determined based on the scenario requirement information and the room equipment function information.
[0124] In one embodiment, when the processor 110 determines the device execution parameters of the target execution room based on the scene requirement information and the room device function information, it specifically performs the following operations: The task planning module of the large language model is used to generate the scene planning task corresponding to the scene requirement information. The task reasoning module of the large language model is used to determine the equipment execution parameters of the target execution room based on the scenario planning task and the room equipment function information.
[0125] In one embodiment, when the processor 110 executes the following operations: if the intent type is an ambiguous intent, it acquires scene knowledge matching the scene requirement information and uses the large language model to determine the device execution parameters in the target space based on the scene requirement information, the scene knowledge, and the room device function information: If the intent type is a vague intent, then obtain the scene knowledge that matches the scene requirement information; The large language model is used to obtain the scene execution room from the scene knowledge, and the room device function information is used to determine whether there is a target execution room in the target space that matches the scene execution room. If so, the target execution device of the target execution room and the device execution parameters of the target execution device are determined based on the scenario knowledge.
[0126] In one embodiment, when the processor 110 executes the determination of the target execution device in the target execution room and the device execution parameters of the target execution device based on the scenario knowledge, it specifically performs the following operations: The task planning module of the large language model is used to generate the scene planning task corresponding to the scene requirement information. The task reasoning module of the large language model uses the scene knowledge and the target execution room to determine the target execution device and the device execution parameters of the target execution device corresponding to the scene planning task.
[0127] In one embodiment, when the processor 110 executes the task reasoning module using the large language model to determine the target execution device corresponding to the scene planning task and the device execution parameters of the target execution device based on the scene knowledge and the target execution room, it specifically performs the following operations: The task reasoning module of the large language model determines the target execution device corresponding to the scene planning task based on the scene knowledge and the target execution room. If the scenario knowledge includes the device action configuration corresponding to the target execution device, then the device action configuration is used as the device execution parameter of the target execution device; If the scenario knowledge does not include the device action configuration corresponding to the target execution device, then the device execution parameters of the target execution device are generated based on the environmental perception parameters of the target space and the scenario planning task.
[0128] In one embodiment, when the processor 110 uses the large language model to obtain scene knowledge that matches the scene requirement information, it specifically performs the following operations: The large language model is used to match the scene requirement information with the scene tags corresponding to each pre-stored scene to obtain the target scene tag that matches the scene requirement information. Obtain scene knowledge of the pre-stored scene corresponding to the target scene label.
[0129] In one embodiment, the scene knowledge includes at least the associated room, device functions, and device action configuration.
[0130] In one embodiment, after executing the operation of obtaining scene knowledge matching the scene requirement information if the intent type is a vague intent, the processor 110 further performs the following operations: if no scene knowledge matching the scene requirement information is obtained, the processor 110 identifies the keywords in the scene requirement information and obtains the preset prompt words corresponding to the keywords. The large language model is used to determine the device execution parameters of smart home devices in the target space based on the scene requirement information, the preset prompt words, and the room device function information.
[0131] In one embodiment, the device information hierarchy in the room device function information includes multiple different levels; the device execution parameters of the smart home device in the target space include the target execution room, the target execution device, and the function value parameters corresponding to the target execution device; the device information hierarchy of the target execution device is determined based on the current computing resources.
[0132] In one embodiment, the device information hierarchy in the room device function information, from largest to smallest, includes device category, device model, and device name; the target execution device is the target device model; when the processor 110 controls the corresponding smart home device in the target space using the device execution parameters, it specifically performs the following operations: Based on the room equipment function information, obtain all equipment names under the target equipment model to obtain the target equipment name; The functional value parameters are mapped to each target device name to generate refined device execution parameters, and the corresponding smart home devices in the target space are controlled based on the refined device execution parameters.
[0133] In one embodiment, before the processor 110 controls the corresponding smart home device in the target space using the device execution parameters, it also performs the following operations: Based on the room equipment function information, abnormal execution parameters are filtered from the target execution room, the target execution device, and the corresponding function value parameters of the target execution device.
[0134] In one embodiment, the processor 110 is further configured to use the large language model to generate a scene name based on the scene requirement information, and store the device execution parameters corresponding to the scene name.
[0135] This invention application also provides a storage medium storing a computer program, which, when executed by a processor, implements the above-described functionality. Figures 1-8 The method described in the illustrated embodiment can be found in the following document for a detailed execution process. Figures 1-8 The specific details of the illustrated embodiments will not be elaborated here.
[0136] Additionally, embodiments of this specification provide a computer program product comprising a computer program that, when executed by a processor of an electronic device, enables the processor to at least perform the functions described above. Figures 1 to 8 The method provided in the illustrated embodiment.
[0137] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. The storage medium can be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc.
[0138] The above-disclosed embodiments are merely preferred embodiments of this specification and should not be construed as limiting the scope of this specification. Therefore, any equivalent variations made in accordance with the claims of this specification shall still fall within the scope of this specification.
Claims
1. A smart home control method, characterized by, The method comprises: in response to receiving scene requirement information of a target space by a user, obtaining room device function information in the target space; using a large language model to identify the intent type of the scene requirement information; if the intent type is an explicit intent, using the large language model to determine device execution parameters of smart home devices in the target space based on the scene requirement information and the room device function information; if the intent type is a fuzzy intent, obtaining scene knowledge matching the scene requirement information, and using the large language model to determine device execution parameters of smart home devices in the target space based on the scene requirement information, the scene knowledge and the room device function information; using the device execution parameters to control the corresponding smart home devices in the target space.
2. The method of claim 1, wherein, If the intent type is an explicit intent, using the large language model to determine device execution parameters in the target space based on the scene requirement information and the room device function information, comprising: if the intent type is an explicit intent, using the large language model to determine a target execution room in the target space based on the scene requirement information and the room device function information; determining device execution parameters of the target execution room based on the scene requirement information and the room device function information.
3. The method of claim 2, wherein, Determining device execution parameters of the target execution room based on the scene requirement information and the room device function information, comprising: using a task planning module of the large language model to generate a scene planning task corresponding to the scene requirement information; using a task reasoning module of the large language model to determine device execution parameters of the target execution room based on the scene planning task and the room device function information.
4. The method of claim 1, wherein, If the intent type is a fuzzy intent, obtaining scene knowledge matching the scene requirement information, and using the large language model to determine device execution parameters in the target space based on the scene requirement information, the scene knowledge and the room device function information, comprising: if the intent type is a fuzzy intent, obtaining scene knowledge matching the scene requirement information; using the large language model to obtain a scene execution room in the scene knowledge, and determining whether there is a target execution room matching the scene execution room in the target space based on the room device function information; if yes, determining a target execution device of the target execution room and device execution parameters of the target execution device based on the scene knowledge.
5. The method of claim 4, wherein, Determining a target execution device of the target execution room and device execution parameters of the target execution device based on the scene knowledge, comprising: using a task planning module of the large language model to generate a scene planning task corresponding to the scene requirement information; using a task reasoning module of the large language model to determine a target execution device corresponding to the scene planning task and device execution parameters of the target execution device based on the scene knowledge and the target execution room.
6. The method of claim 5, wherein, The task reasoning module adopting the large language model determines a target execution device corresponding to the scene planning task and a device execution parameter of the target execution device based on the scene knowledge and the target execution room. The task reasoning module adopting the large language model determines a target execution device corresponding to the scene planning task based on the scene knowledge and the target execution room. If the device action configuration corresponding to the target execution device is included in the scene knowledge, the device action configuration is taken as the device execution parameter of the target execution device. If the device action configuration corresponding to the target execution device is not included in the scene knowledge, the device execution parameter of the target execution device is generated based on the environment perception parameter of the target space and the scene planning task.
7. The method of claim 1 or 4, wherein, The scene knowledge matched with the scene demand information includes: The scene demand information is matched with a scene tag corresponding to each pre-stored scene to obtain a target scene tag matched with the scene demand information. The scene knowledge of the pre-stored scene corresponding to the target scene tag is obtained.
8. The method of any one of claims 1 or 4-6, wherein, The scene knowledge at least includes an associated room, a device function, an environment perception parameter, and a device action configuration corresponding to the environment perception parameter.
9. The method of claim 4, wherein, If the intent type is an ambiguous intent, after the scene knowledge matched with the scene demand information is obtained, the method further includes: If the scene knowledge matched with the scene demand information is not obtained, a keyword in the scene demand information is recognized, and a preset prompt word corresponding to the keyword is obtained. The device execution parameter of the smart home device in the target space is determined based on the scene demand information, the preset prompt word, and the room device function information.
10. The method of claim 1 or 9, wherein, The device information hierarchy in the room device function information includes multiple different levels; the device execution parameter of the smart home device in the target space includes a target execution room, a target execution device, and a function value parameter corresponding to the target execution device; and the device information hierarchy of the target execution device is determined based on current computing resources.
11. The method of claim 10, wherein, The device information hierarchy in the room device function information includes, from large to small, a device category, a device model, and a device name; and the target execution device is a target device model. The method of controlling the corresponding smart home device in the target space by using the device execution parameter includes: All device names under the target device model are obtained based on the room device function information to obtain target device names. The function value parameter is mapped to each target device name to generate refined device execution parameters, and the corresponding smart home device in the target space is controlled based on the refined device execution parameters.
12. The method of claim 10, wherein, Before the method of controlling the corresponding smart home device in the target space by using the device execution parameter, the method further includes: Abnormal execution parameters are filtered from the target execution room, the target execution device, and the function value parameter corresponding to the target execution device based on the room device function information.
13. The method of claim 1, wherein, The method further includes: A scene name is generated based on the scene demand information by using the large language model, and the device execution parameter is stored corresponding to the scene name.
14. A smart home control device, characterized in that, The device comprises: a demand acquisition unit configured to acquire room device function information in a target space in response to receiving scene demand information of the target space by a user; an intention recognition unit configured to recognize an intention type of the scene demand information by using a large language model; a first inference unit configured to determine device execution parameters of smart home devices in the target space based on the scene demand information and the room device function information by using the large language model if the intention type is an explicit intention; a second inference unit configured to acquire scene knowledge matched with the scene demand information, and determine device execution parameters of smart home devices in the target space based on the scene demand information, the scene knowledge and the room device function information by using the large language model if the intention type is a fuzzy intention; a control unit configured to control corresponding smart home devices in the target space by using the device execution parameters.
15. An electronic device, comprising: The electronic device comprises one or more processors and one or more memories, and the one or more memories store at least one computer program, which is loaded and executed by the one or more processors to implement the method of any one of claims 1 to 13.
16. A storage medium, characterized by The storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the method in any one of claims 1 to 13 are implemented.
17. A computer program product, characterised in that, The computer program, when executed by a processor of an electronic device, causes the processor to perform the steps of the method in any one of claims 1 to 13. The computer program, when executed by a processor of an electronic device, causes the processor to perform the steps of the method in any one of claims 1 to 13.