Instruction processing method and device for smart home control and electronic device
Patent Information
- Application Number
- CN202611270640.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-20
- Publication Date
- 2026-09-25
AI Technical Summary
[0003]本申请提供了一种面向智能家居控制的指令处理方法、装置及电子设备,以解决现有技术难以兼顾智能家居控制场景中低时延与复杂语义理解的问题
本申请实施例提供的该方法,通过本地智能体对用户指令进行识别,并根据控制类型、风险等级和歧义等级的综合判定结果进行分流处理,满足执行条件时直接由本地生成并下发设备控制指令,不满足执行条件时委派云端大模型生成控制策略后由本地进行设备可行性校验,使得简单控制指令能够通过本地路径快速响应,降低了对云端算力的依赖和网络传输时延,同时复杂指令能够借助云端大模型的语义理解能力完成处理,兼顾了智能家居控制场景下的低时延与复杂任务处理能力。
Smart Images

Figure CN122824533A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of smart home control technology, and in particular to an instruction processing method, device and electronic device for smart home control. Background Technology
[0002] In the smart home field, natural language control has become a significant trend in human-computer interaction. Currently, mainstream technical solutions fall into two main categories: one relies on a large, general-purpose cloud model to perform unified semantic understanding and command generation for user instructions, leveraging its powerful generalization capabilities to handle complex expressions; the other deploys a lightweight model or keyword matching system locally to provide rapid responses to simple commands, reducing latency and network dependence. However, cloud-based solutions are significantly affected by network fluctuations and have high inference costs; while local solutions offer rapid responses, their capacity to handle ambiguous expressions and complex scenarios such as multi-device collaboration is clearly insufficient due to limitations in model size. Summary of the Invention
[0003] This application provides an instruction processing method, device, and electronic device for smart home control, in order to solve the problem that existing technologies cannot simultaneously achieve low latency and complex semantic understanding in smart home control scenarios.
[0004] Firstly, this application provides an instruction processing method for smart home control, comprising: Receive user instructions, and generate context information for the current task based on the user instructions, user information, and device information; the device information refers to the device information of the home-controlled device. The user command is identified to obtain an identification result; the identification result includes control type, risk level, ambiguity level, target device, and action type. When the recognition result meets the preset execution conditions, a first device control command is generated based on the target device and the action type, and the first device control command is sent to the controlled device that matches the user command; When the recognition result does not meet the preset execution conditions, a task package is generated based on the context information and sent to the cloud big model so that the cloud big model can generate a control strategy based on the task package; Receive the control strategy returned by the cloud-based big data model, and obtain the device capability map and device status information; Based on the context information, equipment capability map, and equipment status information, the control strategy is verified for equipment feasibility. After the device feasibility verification is passed, a second device control command corresponding to the control strategy is issued to the controlled device that matches the user command.
[0005] Secondly, this application provides an instruction processing device for smart home control, comprising: The generation unit is used to receive user instructions and generate context information for the current task based on the user instructions, user information, and device information; the device information is the device information of the home-controlled device. The identification unit is used to identify the user command and obtain the identification result; the identification result includes control type, risk level, ambiguity level, target device, and action type. The first issuing unit is used to generate a first device control command based on the target device and action type when the recognition result meets the preset execution conditions, and to issue the first device control command to the controlled device that matches the user command. The sending unit is used to generate a task package based on the context information when the recognition result does not meet the preset execution conditions, and send the task package to the cloud big model so that the cloud big model can generate a control strategy based on the task package; The acquisition unit is used to receive the control strategy returned by the cloud-based big data model, and to acquire the device capability map and device status information; The verification unit is used to verify the feasibility of the control strategy based on the context information, the device capability map, and the device status information. The second issuing unit is also used to issue a second device control command corresponding to the control strategy to the controlled device that matches the user command after the device feasibility verification is passed.
[0006] Thirdly, this application provides an electronic device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, it implements the method as described in any one of the first aspects.
[0007] The technical solutions provided in this application have the following advantages compared with the prior art: The method provided in this application identifies user commands through a local intelligent agent and performs triage processing based on a comprehensive judgment result of control type, risk level, and ambiguity level. When the execution conditions are met, device control commands are directly generated and issued locally. When the execution conditions are not met, the cloud-based large model is delegated to generate a control strategy, and then the local system performs device feasibility verification. This allows simple control commands to respond quickly through the local path, reducing reliance on cloud computing power and network transmission latency. At the same time, complex commands can be processed with the help of the semantic understanding capabilities of the cloud-based large model, thus balancing low latency and complex task processing capabilities in smart home control scenarios. Attached Figure Description
[0008] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0009] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0010] One or more embodiments are illustrated by way of example with reference numerals in the accompanying drawings. These illustrations do not constitute a limitation on the embodiments. Elements with the same reference numerals in the drawings are denoted as similar elements. Unless otherwise stated, the figures in the drawings are not to be limited by scale.
[0011] Figure 1 This is a flowchart illustrating an instruction processing method for smart home control, provided as an embodiment of this application.
[0012] Figure 2 This is a schematic diagram illustrating an instruction processing flow for smart home control according to an exemplary embodiment.
[0013] Figure 3 This is a block diagram illustrating an instruction processing device for smart home control according to an exemplary embodiment.
[0014] Figure 4 This is a block diagram illustrating an apparatus for an instruction processing method for smart home control, according to an exemplary embodiment.
[0015] Figure Labels 301 - Generation unit; 302 - Identification unit; 303 - First sending unit; 304 - Sending unit; 305 - Acquisition unit; 306 - Verification unit; 307 - Second sending unit; 400 - Device; 402 - Processing component; 404 - Memory; 406 - Power component; 408 - Multimedia component; 410 - Audio component; 412 - I / O interface; 414 - Sensor component; 416 - Communication component; 420 - Processor. Detailed Implementation
[0016] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0017] The following disclosure provides numerous different embodiments or examples for implementing various structures of this application. To simplify the disclosure, specific examples of components and arrangements are described below. These are merely examples and are not intended to limit the scope of this application. Furthermore, reference numerals and / or letters may be repeated in different examples. Such repetition is for simplification and clarity and does not in itself indicate a relationship between the various embodiments and / or arrangements discussed.
[0018] For ease of description, spatial relative terms may be used in the text to describe the relative position or movement of one element or feature relative to another element or feature, as shown in the figure. These relative terms include, for example, "inside," "outside," "middle," "outer," "below," "below," "above," "front," "back," etc. Such spatial relative terms are intended to include different orientations of the device in use or operation, other than those depicted in the figure. For example, if the device in the figure undergoes a positional flip, orientation change, or change of motion, these directional indications will change accordingly. For instance, an element described as "below other elements or features" or "below other elements or features" will subsequently be oriented "above other elements or features" or "above other elements or features." Therefore, the example term "below" can include both upper and lower orientations. The device may be otherwise oriented (rotated 90 degrees or in other directions), and the spatial relative descriptors used in the text will be interpreted accordingly.
[0019] To address the technical challenge of balancing low latency and complex semantic understanding in smart home control scenarios in existing technologies, this application provides an instruction processing method and apparatus for smart home control, which improves the security and stability of the smart home control link.
[0020] Figure 1 A flowchart illustrating an instruction processing method for smart home control provided in this application embodiment is shown below. Figure 1 As shown, it should be noted that this method can be deployed on a home central control screen or edge gateway device. The instruction processing method for smart home control in this application embodiment is applied to an instruction processing device for smart home control.
[0021] It should be noted that, in some embodiments of this disclosure, the local intelligent agent is an execution proxy deployed on a home control screen or edge gateway device. It is used to uniformly receive user commands, maintain session context, determine whether to delegate processing to the cloud-based large language model, and drive local device execution after obtaining the processing results. The local intelligent agent is an intermediate layer providing local execution capabilities for the cloud-based large language model, handling the actual distribution of skill calls, tool execution, script execution, file writing, and device control links. In other words, the local intelligent agent is an execution-oriented intelligent proxy located on the home's local side. It is responsible for deciding which commands can be executed directly locally and which commands need to be delegated to the cloud. It is also responsible for verifying the device feasibility and execution gate of the control policies returned from the cloud, ensuring that only control commands that pass all verifications are distributed to the controlled device.
[0022] like Figure 1 As shown, the method may include the following steps: Step 101: Receive user instructions and generate context information for the current task based on user instructions, user information, and device information.
[0023] The device information refers to the device information of the home-controlled devices.
[0024] In some embodiments of this application, user information includes user identity information and historical execution records. User identity is used to identify the operator currently issuing user commands, and can be obtained through voiceprint recognition or facial recognition. Historical execution records are used to characterize the user's operation history with home-controlled devices over a past period, including operation time, operating device, operation type, and execution result.
[0025] Understandably, user identity provides the basis for subsequent gate permission verification, allowing different users to be granted different device operation permissions (e.g., children are not allowed to operate the door lock); historical execution records provide a reference for ambiguity resolution and contextual understanding. For example, when a user says "turn it up a little brighter", the local intelligent agent can infer the actual target of the current instruction based on the device object and parameter value of the most recent dimming operation in the historical execution record.
[0026] In some embodiments of this application, step 101 may specifically include the following sub-steps: Step a1: Receive user instructions and generate task identifier.
[0027] Specifically, after receiving voice or text input from the user, the local agent generates a unique task identifier for that request, which is used for subsequent log tracking and state association. User commands can be any of the following: voice input, text input, touch interaction input, or multimodal fusion input.
[0028] As an example of this application, when a user says "turn on the living room light", the local intelligent agent collects the voice signal through a microphone array, converts it into text through Automatic Speech Recognition (ASR), and generates a task identifier corresponding to the request.
[0029] Step a2: Read the session summary, device online status, user identity, and historical execution records to form the context information of the current task.
[0030] Specifically, the local agent reads the session summary of the most recent N rounds of interaction from local storage or memory, including historical user commands, system responses, and execution results; reads the online status of all currently controlled home devices, such as whether the devices are powered on or connected to the network; reads the user identity information of the user who issued the command, such as the user ID determined by voiceprint recognition or facial recognition; and reads historical execution records related to the user or device, such as the time and result of the most recent operation on the target device. The local agent summarizes the above information to form the context information of the current task.
[0031] Understandably, contextual information provides the data foundation for subsequent rapid judgment, delegation decision-making, and execution verification, enabling the local intelligent agent to make accurate judgments with a full understanding of the current device environment and user background.
[0032] Step 102: Recognize the user command and obtain the recognition result.
[0033] The identification results include control type, risk level, ambiguity level, target device, and action type.
[0034] In this embodiment, the control type is used to characterize whether the user instruction belongs to a control request for home-controlled devices.
[0035] As an example, the local agent categorizes user commands by intent to determine the control type, which includes two values: control request and non-control request. When a user command expresses an intent to operate a device, such as "turn on the living room light" or "set the air conditioner to 26 degrees," its control type is determined to be a control request. When a user command pertains to casual conversation, information retrieval, or other non-device operation content, such as "What's the weather like today?" or "Tell me a joke," its control type is determined to be a non-control request.
[0036] Understandably, by clearly distinguishing between control requests and non-control requests, it is possible to avoid sending non-control requests to the device control link by mistake, and to prevent the system from making incorrect device responses to users' non-operational intentions.
[0037] In this embodiment of the application, the risk level is used to characterize the potential harm to personal safety, property safety or user privacy caused by the control operation corresponding to the execution of user instructions.
[0038] As an example, the local agent uses the target device type and action type extracted from the slot as an index to query the pre-defined mapping relationship in the preset risk rule base to obtain the risk level. The preset risk rule base contains preset risk level mapping relationships corresponding to different combinations of device types and action types. For example, the combination of device type "light" and action type "switch" has a low risk level; the combination of device type "door lock" and action type "open" has a high risk level; the combination of device type "gas valve" and action type "close" has a high risk level; and the combination of device type "camera" and action type "close" has a high risk level.
[0039] Understandably, by separating high-risk control operations from low-risk control operations, high-risk operations are ensured to undergo more rigorous verification before entering the execution chain, thereby reducing the risk of erroneous execution.
[0040] In some embodiments of this application, the ambiguity level is used to characterize the completeness and clarity of the semantic expression of a user instruction.
[0041] As an example, the local agent calculates the ambiguity score based on the completeness of the slot information extracted from the slot, and compares the ambiguity score with a preset ambiguity threshold to obtain the ambiguity level, which includes two values: low ambiguity and high ambiguity.
[0042] Specifically, the local intelligent agent counts the number of entities extracted and the total number of entities required by the request template corresponding to the user instruction, and calculates the ratio of the two as the ambiguity score. When the score is 1 or close to 1, it indicates that the expression of the user instruction is complete and clear, and the ambiguity level is low ambiguity; when the score is 0 or close to 0, it indicates that the user instruction lacks necessary control elements (such as missing target device or unclear action type), and the ambiguity level is high ambiguity.
[0043] For example, in the instruction "turn on the living room light", the extracted entities include the room entity "living room", the device entity "light", and the action entity "turn on". The required number of entities is 3, and the ambiguity score is 3 / 3=1.0, with an ambiguity level of low ambiguity. In the instruction "adjust", the extracted entities only include the action entity "adjust", and lack the device entity and room entity. The ambiguity score is low, with an ambiguity level of high ambiguity.
[0044] Understandably, the determination of ambiguity level is the third threshold for the routing decision. When a user command has semantic gaps or is ambiguous, it is delegated to the cloud-based big model for completion and disambiguation, thus avoiding possible misoperations that may occur if executed directly locally.
[0045] In some embodiments of this application, the target device is used to characterize the home controllable device to be controlled by the user command.
[0046] As an example, the local agent extracts slots from user commands, identifying keywords or phrases related to device entities from the text sequence of the user commands to obtain the target device. For example, for the user command "turn on the living room light," the target device obtained by slot extraction is "living room light"; for the user command "turn off the bedroom air conditioner," the target device obtained by slot extraction is "bedroom air conditioner." When the user command does not explicitly specify a device (e.g., "turn off that"), slot extraction may not be able to obtain a valid target device. In this case, the target device field in the recognition result is empty or has a low confidence level, triggering a high ambiguity level, and thus the command is delegated to a large cloud model for disambiguation processing.
[0047] In some embodiments of this application, the action type is used to characterize the specific operation that the user command requests to be performed on the target device.
[0048] As an example, the local agent extracts slots from user commands and identifies action-related keywords or phrases from the text sequence of user commands to obtain the action type.
[0049] For example, for the user command "turn on the living room light", the action type extracted by the slot is "turn on"; for the user command "set the living room air conditioner temperature to 26 degrees", the action type extracted by the slot is "temperature adjustment", and the parameter "26 degrees" can be further extracted. Common action types include, but are not limited to: turn on, turn off, brighten, dim, temperature adjustment, mode switching, start, pause, stop, etc.
[0050] It is understandable that the target device and the action type together constitute the basic elements of the device control command. The action type is also used to determine the risk level. Different action types for the same target device may correspond to different risk levels (for example, "querying the door lock status" is low risk, while "unlocking the door lock" is high risk).
[0051] In some embodiments of this application, step 102 may specifically include the following sub-steps: Step b1: Extract the slot from the user command to obtain slot information; the slot information includes the target device and action type.
[0052] Specifically, the local agent uses a sequence labeling model to extract slots from user commands. As an example, a Conditional Random Field (CRF) model can be used to sequence label the text sequence of user commands, identifying key entities within it. Slot information includes at least the target device (e.g., "living room light," "door lock") and the action type (e.g., "turn on," "turn off," "dimming"). For the user command "turn on the living room light," the extracted target device is "living room light," and the action type is "turn on." For the user command "adjust the living room air conditioner temperature to 26 degrees," the extracted target device is "living room air conditioner," and the action type is "temperature adjustment," while the parameter "26 degrees" can also be extracted. It is understandable that the quality of slot extraction directly affects the accuracy of all subsequent decisions and executions.
[0053] Step b2: Classify the user commands by intent to obtain the control type.
[0054] Specifically, the local intelligent agent uses an intent classification model to categorize user commands and determine whether the command belongs to the home control category. Control types include two categories: control requests and non-control requests.
[0055] As an example, a lightweight language model or a Support Vector Machine (SVM) classifier can be used to binary classify user commands: if classified as a control request, it means the user wants to control a home appliance; if classified as a non-control request, it means the user may be chatting, querying information, or performing other non-device control operations. Explicit intent classification avoids mistakenly sending non-control requests into the device control chain.
[0056] Step b3: Select the risk level that matches the slot information from the preset risk rule base.
[0057] The preset risk rule base includes a mapping relationship between slot information and risk level.
[0058] Specifically, the local intelligent agent uses the extracted slot information (including at least the device type and action type) as an index to query the local preset risk rule base and obtain the corresponding risk level.
[0059] The preset risk rule base stores the risk level mapping relationship corresponding to different combinations of device types and action types. For example, the combination of device type "light" and action type "switch" has a low risk level; the combination of device type "door lock" and action type "open" has a high risk level; the combination of device type "gas valve" and action type "close" has a high risk level; and the combination of device type "camera" and action type "close" has a high risk level.
[0060] Understandably, by using a pre-built risk rule base to quickly match risk levels, low-latency risk assessment can be achieved, avoiding the latency overhead of sending each request to the cloud for risk judgment.
[0061] Step b4: Calculate the ambiguity score based on the completeness of the slot information, compare the ambiguity score with the preset ambiguity threshold, and obtain the ambiguity level.
[0062] Specifically, the local agent counts the number of entities extracted and the total number of entities required by the request template corresponding to the user command, and calculates the ratio of the number of extracted entities to the total number of entities as the ambiguity score. When the ambiguity score is greater than a preset ambiguity threshold, the ambiguity level is indicated as low ambiguity; when the ambiguity score is not greater than the preset ambiguity threshold, the ambiguity level is indicated as high ambiguity.
[0063] For example, in the user command "turn on the living room light", the extracted entities include the room entity "living room", the device entity "light", and the action entity "turn on". The total number of entities required is 3, and the ambiguity score is 3 / 3=1.0, which is greater than the preset ambiguity threshold of 0.8. Therefore, the ambiguity level is low ambiguity. In the user command "adjust", the extracted entities may only include the action entity "adjust", lacking the device entity and the room entity. The ambiguity score is low, so the ambiguity level is high ambiguity.
[0064] It should be noted that by coordinating four sub-steps—slot extraction, intent classification, risk rule matching, and ambiguity assessment—multi-dimensional and rapid identification of user commands is achieved. Compared to solutions that directly send all requests to a large cloud model for unified parsing, this application uses a local intelligent agent to complete control-type identification, risk classification, and ambiguity classification on the edge. This allows simple, clear, and low-risk control commands to be quickly identified and enter the local execution path without experiencing the full latency of cloud inference. Simultaneously, high-risk, highly ambiguous, or non-control-type requests can also be accurately identified and trigger the delegation process. This multi-dimensional identification mechanism provides a reliable basis for subsequent traffic allocation decisions, balancing response speed and processing accuracy.
[0065] Step 103: When the recognition result meets the preset execution conditions, a first device control command is generated based on the target device and action type, and the first device control command is sent to the controlled device that matches the user command.
[0066] In this embodiment, when the control type is indicated as a control request, the risk level is lower than a preset risk threshold, and the ambiguity level is lower than a preset ambiguity threshold, the local agent determines that the identification result meets the preset execution conditions. At this time, the local agent directly extracts the target device and action type from the identification result, encodes the device identifier and action type of the target device into a device-recognizable instruction format according to the requirements of the underlying device communication protocol, generates a first device control instruction, and confirms that the target device is currently online based on the device online status in the context information, and then sends the first device control instruction to the corresponding controlled device through the local control link.
[0067] It should be noted that by directly processing and issuing simple control commands that meet the conditions of low risk and low ambiguity on the edge, the network transmission latency and computing costs caused by cloud inference for each control request are avoided. This enables high-frequency basic control scenarios such as "turning on the living room light" to receive an instant response, achieving the technical effect of control acceleration.
[0068] In some embodiments of this application, step 103 specifically includes: when the control type is a control request, the risk level is lower than a preset risk threshold, and the ambiguity level is lower than a preset ambiguity threshold, the local agent determines that the recognition result meets the preset execution conditions.
[0069] Specifically, the preset execution conditions include three dimensions: the control type is a control request, the risk level is lower than a preset risk threshold, and the ambiguity level is lower than a preset ambiguity threshold. Only when all three conditions are met simultaneously will the local agent determine that the user instruction is suitable for direct execution locally. If any one of the three conditions is not met, the local agent determines that the user instruction is not suitable for direct local execution and needs to be delegated to a large cloud model for processing.
[0070] When the recognition result meets the preset execution conditions, the local agent directly generates a first device control command based on the target device and action type extracted from the slot in step 102. For example, if the user command is "turn on the living room light", the target device extracted from the slot is "living room light", and the action type is "turn on", the local agent generates a corresponding device control command (such as "living room light_on") and sends the command to the controlled device corresponding to the living room light through the local control link. As an example, after generating the first device control command, the local agent can also confirm that the target device is currently online based on the device online status in the context information obtained in step 101 before sending the command to improve the execution success rate.
[0071] Understandably, by directly processing and issuing low-risk, low-ambiguity control commands that meet preset execution conditions on the device side, this application avoids simple control requests undergoing a complete cloud inference process, significantly reducing response latency. For high-frequency, simple control scenarios such as "turning on the living room lights," users can receive immediate control feedback, improving the smart home user experience. Meanwhile, commands that do not meet preset execution conditions (including non-control requests, high-risk requests, or highly ambiguous requests) are delegated to the cloud-based large model, ensuring that complex scenarios are adequately handled.
[0072] Step 104: When the recognition result does not meet the preset execution conditions, a task package is generated based on the context information and sent to the cloud big model so that the cloud big model can generate a control strategy based on the task package.
[0073] In this embodiment, the task package may include the original text of the user instruction, the recognition result, the candidate device set, the device capability constraints, and the session summary. The task package is sent to the cloud big model via the Internet, so that the cloud big model can complete semantic understanding and task orchestration within the actual capability boundary of the local device defined by the device capability constraints, and generate a control strategy corresponding to the task package.
[0074] Understandably, for non-control requests, high-risk requests, or highly ambiguous requests, processing is delegated to a large cloud model. This model leverages its powerful semantic understanding capabilities to parse complex expressions, scenario-based descriptions, and compound instructions. Meanwhile, the device capability constraints in the task package limit the model's reasoning to the local device's capabilities, preventing the model from generating unexecutable control policies that are detached from the actual device capabilities. This expands the system's scenario understanding capabilities while ensuring the feasibility of the control policies.
[0075] In some embodiments of this application, the cloud-based big data model is deployed on a cloud server to receive task packages sent by the local agent. Within the candidate device set and executable actions defined by the device capability constraints in the task package, the cloud-based big data model performs semantic understanding and generates control strategies. The control strategy output by the cloud-based big data model is a structured control plan, where each control action is mapped to a target device in the candidate device set. In other words, in this application, the cloud-based big data model serves only as a conditional semantic understanding tool delegated by the local agent. Its inference boundaries and output format are constrained by the constraints imposed by the local agent through the task package, rather than freely generating text responses in the form of an open-ended dialogue model.
[0076] Understandably, by limiting the function of the cloud-based large model to "performing conditional reasoning and outputting structured plans within the boundaries of local device capabilities," it distinguishes itself from general conversational large models. This ensures that the cloud-based large model can leverage its powerful semantic understanding capabilities to process complex instructions, while the constraints in the task package strictly limit its output to the range that can be executed locally, thus avoiding unconstrained reasoning that is detached from the actual capabilities of the local device.
[0077] In some embodiments of this application, step 104 specifically includes: when the control type is a non-control request, or the risk level is not lower than a preset risk threshold, or the ambiguity level is not lower than a preset ambiguity threshold, the local intelligent agent determines that the recognition result does not meet the preset execution conditions.
[0078] Specifically, if any one of the three conditions in step 103 is not met, the local agent determines that the processing needs to be delegated to the cloud-based large model. Specific situations include, but are not limited to: the user instruction is identified as a non-control request (such as the user engaging in casual conversation or information retrieval); the user instruction involves high-risk devices (such as door locks, gas valves, cameras, etc.), and its risk level is not lower than a preset risk threshold; the user instruction is ambiguous, and the ambiguity level is not lower than a preset ambiguity threshold (such as the lack of a target device or the unclear action type); the user instruction contains a complex intent (such as "turn on the living room light and turn off the bedroom light") or a scenario-based description (such as "adjust the living room to a suitable state for watching movies").
[0079] When a cloud-based large model needs to be delegated, the local agent constructs a task package based on contextual information. The task package includes at least the original text of the user instruction, the recognition result, a set of candidate devices, device capability constraints, and a session summary. The local agent sends the task package to the cloud-based large model service via the internet, enabling the cloud-based large model to complete semantic understanding and task orchestration under constrained contextual conditions. Device capability constraints are used to limit the inference boundaries of the cloud-based large model, preventing it from generating control plans that are detached from the actual capabilities of the local devices.
[0080] In other embodiments of this application, the cloud-based large model is not limited to a single model. Structured control strategies can also be generated collaboratively by multiple specialized models. For example, one model might be responsible for intent parsing, another for device mapping and parameter generation, and a third for execution sequence orchestration, as long as a structured control result is ultimately formed that accepts gate verification by the local intelligent agent. It is understood that this multi-specialized model collaboration approach allows for the selection of the optimal model combination based on different task types, reducing the computational cost of a single large model while ensuring processing quality.
[0081] Understandably, by constructing a task package that includes recognition results, a set of candidate devices, and device capability constraints, this application enables the cloud-based large model to perform inference with a full understanding of the actual situation of local devices. This avoids the problem of inconsistent generated results and executable conditions caused by the cloud-based large model performing unconstrained inference outside the local control boundary. Simultaneously, the delegation mechanism enables the local agent to transfer tasks to the general large model. That is, while retaining the local control boundary, it utilizes the general large model's understanding of richer user instructions to process scene descriptions, complex constraints, and even non-home-related expressions, and then converges the executable parts back to the local control link, thereby improving the system's scene expansion capabilities.
[0082] Step 105: Receive the control strategy returned by the cloud-based big data model, and obtain the device capability map and device status information.
[0083] In this embodiment, the control strategy is a structured control plan generated by a cloud-based large model based on task packages. The device capability map is a directed graph structure that is pre-built and stored locally to describe the semantic relationships between user nodes, space nodes, device nodes, capability nodes and status nodes in the home. The device status information includes the online status and current operating parameters of each home-controlled device.
[0084] It is understandable that control strategy, equipment capability map, and equipment status information together constitute the three necessary inputs for subsequent equipment feasibility verification. Control strategy provides the set of control actions to be verified, equipment capability map provides the functional boundaries and semantic relationships of each device, and equipment status information provides the real-time operating status of each device.
[0085] In some embodiments of this application, step 105 specifically includes the following sub-steps: Step c1: Receive the control strategy returned by the large model in the cloud.
[0086] Specifically, after receiving the task package sent by the local agent, the cloud-based big model performs semantic understanding and task orchestration based on the original text of the user instructions, recognition results, candidate device set, device capability constraints and conversation summary in the task package, generates a control policy and returns the control policy to the local agent.
[0087] Understandably, the control strategies generated by the cloud-based large model under the constraints of the task package are limited to the capabilities of the local device, thereby reducing the generation of unexecutable instructions from the source.
[0088] Step c2: Obtain the equipment capability map and equipment status information.
[0089] Specifically, the local intelligent agent reads the device capability map from local memory or storage and obtains the real-time status information of each controlled home device. The device capability map adopts a directed graph structure to describe the semantic relationships between users, spaces, devices, capabilities, and states in the home. Device status information includes real-time data such as the online / offline status of each device and current operating parameters (e.g., current brightness, current temperature).
[0090] In some embodiments of this application, the data for constructing the device capability map mainly comes from three parts: standardized device object models (including attributes, services, and event definitions) issued by the IoT platform or edge gateway; spatial topology information (room division, device aliases, scene orchestration) customized by the user in the smart home application (App); and device runtime logs. It is understood that the device capability map remains active in the local gateway or central control memory, and achieves real-time synchronization with the physical device status by subscribing to the edge gateway's message bus, thereby ensuring that the status information in the map is consistent with the actual physical world.
[0091] Step 106: Verify the feasibility of the control strategy based on contextual information, equipment capability map, and equipment status information.
[0092] In this embodiment, the local intelligent agent performs the following three-layer verification for each control action in the control strategy: verifying whether the target device in the control action exists in the device capability map; if it does not exist, the verification fails; verifying whether the target device has the ability to execute the action type in the control action; if the capability node of the target device does not contain the action type, the verification fails; and verifying whether the action parameters in the control action are within the allowable parameter range of the target device; if they exceed the allowable range, the verification fails.
[0093] For any one or more of the above checks that fail, the local agent removes the control action from the control strategy and only retains the control actions that pass all three layers of checks and enters the subsequent execution gate check stage.
[0094] It should be noted that by verifying device existence, capability matching, and parameter range, control actions that are unexecutable due to device non-existence, unsupported capabilities, or out-of-bounds parameters in the control strategy generated by the cloud-based large model are filtered out in advance. This avoids the cloud model directly sending unexecutable inference results to the device control link due to a lack of understanding of the actual situation of the local device. While retaining the semantic understanding capability of the cloud-based large model, a security filtering mechanism for cloud output is formed through feasibility verification on the local end.
[0095] In some embodiments of this application, step 106 may specifically include the following sub-steps: Step d1: Perform a verification operation for each control action in the control strategy.
[0096] Specifically, the control policy returned by the cloud-based large model may include one or more control actions, such as "turn on the main living room light, close the living room curtains, and adjust the living room light strip to a warm color." For each control action in the control policy, the local agent performs the following three layers of verification: The first layer verifies whether the target device in the control strategy exists in the device capability map. Specifically, the local agent uses the target device name in the control action as an index to query whether a corresponding device node exists in the device capability map. If it exists, the verification passes; if it does not exist, it means that the device is not in the home device list, and the verification fails.
[0097] The second layer verifies whether the target device has the capability to execute the action type specified in the control strategy. Specifically, the local agent searches the device capability graph for the set of capability nodes connected to the target device's device node through the "capable" edge, and determines whether the action type is in that set. For example, if the target device is a "living room light," its capability nodes include "switch" and "dimming." If the control strategy requires "adjusting color temperature," then the device does not have this capability, and the verification fails.
[0098] The third layer verifies whether the action parameters in the control strategy are within the allowable parameter range of the target device. Specifically, the local agent obtains the capability parameter range information of the target device node and determines whether the action parameters carried in the control strategy are within the allowable range. For example, if the target device is a "living room light," its dimming capability has an allowable parameter range of 0% to 100%. If the control strategy requires "brightness adjusted to 150%," then this parameter exceeds the allowable range, and the verification fails.
[0099] Step d2: If any one or more of the above verification operations fail, the corresponding control action will be removed from the control strategy.
[0100] Specifically, for control actions that fail any of the three layers of verification, the local agent removes the control action from the control strategy and prevents it from proceeding to the subsequent execution gate verification stage. Control actions that pass all three layers of verification are retained and proceed to the next step. Through device feasibility verification, this application can pre-filter out control actions in the cloud-generated large model that are unexecutable due to device non-existence, capability mismatch, or parameter out-of-bounds errors, thus avoiding execution failures or erroneous operations caused by device problems during the actual device control execution phase.
[0101] It should be noted that a device feasibility verification mechanism is used to achieve strong consistency verification between the results generated in the cloud and the actual capabilities of the local device. Compared to the approach of directly issuing control commands by using the output of the large cloud model, this application uses a device capability map to perform three-layer filtering on the control strategy returned by the cloud: existence, capability matching, and parameter range. This effectively avoids the problem of generating unexecutable commands due to the cloud model's lack of understanding of the actual situation of the local device, thus improving the security and stability of the control link.
[0102] Step 107: After the equipment feasibility verification is passed, a second equipment control command corresponding to the control strategy is issued to the controlled equipment that matches the user command.
[0103] In this embodiment, after the control strategy passes the device feasibility verification, the local intelligent agent calls the execution gate to perform execution verification on the control strategy. The execution verification includes at least one of permission verification, device status verification, time window verification, and risk level threshold verification. After the execution verification passes, the local intelligent agent generates a corresponding second device control command according to the target device, action type, and action parameters of each control action in the control strategy, and sends the second device control command to the corresponding controlled device.
[0104] In some embodiments of this application, step 107 may specifically include the following sub-steps: Step e1: After the equipment feasibility verification is passed, identify whether the control strategy includes control operations for preset high-risk equipment types.
[0105] Specifically, the preset high-risk device types include, but are not limited to, door locks, gas valves, cameras, and main power switches. These preset high-risk device types can be pre-configured by the system or customized by the user in the smart home application. The local smart agent iterates through each control action retained after the device feasibility verification passes, determining whether its target device belongs to a preset high-risk device type. If so, it is determined that the control strategy includes a high-risk control operation, requiring additional high-risk verification; otherwise, it proceeds directly to step e3.
[0106] Step e2: When the control strategy includes control operations on preset high-risk device types, perform at least one of the following high-risk verification operations: The first method involves initiating interactive secondary confirmation via voice broadcast or a pop-up window on the central control screen, and receiving confirmation commands from the user. Specifically, the local intelligent agent initiates a confirmation request to the user through voice synthesis broadcast or by displaying a pop-up window on the central control screen, such as "You are attempting to open the door lock, please confirm whether to proceed?", and opens a preset listening time window (e.g., 5 seconds). The verification is only confirmed as successful upon receiving a clear affirmative response from the user (e.g., "Confirm," "Yes," "Open"). If the user does not provide an affirmative response or provides a negative response within the time window, the verification is confirmed as unsuccessful.
[0107] The second method involves obtaining the voiceprint features of the user's command when the command is a voice command. These features are then compared with a registered voiceprint database to obtain the voiceprint verification result. Specifically, the local agent extracts the voiceprint feature vector from the original user's voice command and compares it with the voiceprint features in a pre-registered family member voiceprint database. If the similarity exceeds a preset voiceprint threshold, the voiceprint verification is successful; otherwise, it fails. Voiceprint verification can identify whether the user issuing the command is an adult or an administrator. For the voiceprints of children or unregistered users, high-risk control operations can be automatically rejected.
[0108] The third method involves acquiring a facial image from the central control screen's camera and comparing it with a registered facial database to obtain a facial verification result. Specifically, the local intelligent agent acquires the current user's facial image through the central control screen's camera, extracts facial feature vectors, and compares them with features from a pre-registered family member facial database. If the similarity exceeds a preset facial threshold, the facial verification result is successful; otherwise, it fails.
[0109] The fourth method involves pushing an authorization request to the bound terminal and receiving authorization confirmation information from the bound terminal. Specifically, the local smart agent pushes an authorization request to the smart home application on the homeowner's bound mobile terminal (e.g., a mobile phone), requiring the homeowner to remotely confirm. After the homeowner confirms via the application or enters a Personal Identification Number (PIN), the bound terminal returns authorization confirmation information to the local smart agent. It is understandable that the out-of-band authorization mechanism provides multi-factor authentication protection for extremely high-risk operations (such as disarming security systems).
[0110] As another high-risk verification method, the gate also performs scenario and time window constraint judgments: verifying whether the current control behavior conforms to preset security logic rules. For example, during preset late-night periods (such as 10:00 PM to 6:00 AM the next day), it refuses to deactivate the security system via voice broadcast or refuses to unlock the door via voice command. Understandably, this time and physical environment logic verification, as an additional security verification layer, can provide extra protection when users may issue high-risk commands due to fatigue or inattention.
[0111] As an alternative or supplement, the execution gate can also employ geofence verification or historical behavior deviation detection for high-risk verification: geofence verification obtains the current location information of the mobile terminal or central control screen that issued the user command to determine whether the current user is within the preset home geofence range. If the user is outside the geofence range, high-risk control operations are refused. Historical behavior deviation detection compares the control actions in the current user command with the user's historical operation records to calculate the deviation. If the deviation exceeds a preset deviation threshold, a secondary confirmation process is triggered. It is understandable that geofence verification can effectively prevent the risk of unauthorized remote access, and historical behavior deviation detection can identify abnormal operation patterns, further enhancing the security protection capabilities of the execution gate.
[0112] Step e3, after the high-risk verification operation passes, executes the step of issuing a second device control command corresponding to the control strategy to the controlled device that matches the user command.
[0113] Specifically, when the control strategy does not include high-risk control operations, or when the control strategy includes high-risk control operations and all of the above-mentioned high-risk verification operations pass, the local intelligent agent calls the device control link to convert the control strategy into a second device control command that the device can recognize and sends it to the corresponding controlled device.
[0114] In some embodiments of this application, after issuing the second device control command, the local intelligent agent obtains the device control execution result and the latest device status returned by the controlled device; it writes the execution result, failure reason, and latest device status back to the log audit module and the session state of the current task for reuse in subsequent requests. It is understood that by writing the execution result and the latest device status back to the session state in real time, subsequent user commands can directly reuse the latest device status information without re-querying the device, further improving system response efficiency; simultaneously, the execution result and failure reason recorded in the log audit module can be used for subsequent troubleshooting and behavior analysis.
[0115] It should be noted that by adding multiple independent verification methods for high-risk device control operations before execution, a multi-dimensional and multi-level execution gate protection mechanism is formed. Interactive secondary confirmation requires users to explicitly express authorization; biometric verification (voiceprint and face) ensures that the operation initiator is the legitimate user; out-of-band authorization provides a remote confirmation channel for extremely high-risk scenarios; time windows and scenario logic judgments provide rule-based dynamic protection; geofencing and historical behavior deviation detection further enrich the verification dimensions of the execution gate. The combined effect of these multi-layered verification mechanisms significantly reduces the probability of erroneous execution in high-risk control scenarios such as accidental door lock opening, gas valve misoperation, and unauthorized camera shutdown, ensuring the execution security of the smart home system and the personal and property safety of users.
[0116] In some embodiments of this application, the device capability map adopts a directed graph structure, which includes user nodes, space nodes, device nodes, capability nodes, and status nodes.
[0117] As an example, the specific structure of the equipment capability map is as follows: Device nodes and spatial nodes are connected by edges that represent the location of the device. For example, the device node "living room light" is connected to the spatial node "living room" through the edge "located in", indicating that the living room light is located in the living room.
[0118] The device node and the state node are connected by an edge that represents the current real-time state of the device. For example, the device node "living room light" is connected to the state node "on" or "brightness 80%" through the "current state" edge, which represents the current real-time state of the living room light.
[0119] The device node and the user node are connected by an edge that represents the device's operation permission. For example, the device node "Door Lock" is connected to the user node "Homeowner" through the "Operation Permission" edge, indicating that the homeowner has the operation permission to operate the door lock.
[0120] Device nodes and capability nodes are connected by edges that characterize the functions of the device. For example, the device node "living room light" is connected to the capability nodes "switch" and "dimmer" through the "capable" edge, indicating that the living room light has both switching and dimming functions.
[0121] It is understood that the device capability graph is not limited to the directed graph structure described above. In other embodiments of this application, the device capability graph can also be stored in the form of a device type mapping table, a device ontology graph, a knowledge graph, a rule base, a configuration file, or a database, as long as it can provide the positional relationship between the device and space, the functional boundaries of the device, the current real-time status of the device, and the user's operation permission information for the device. Different implementation forms can be flexibly selected according to the computing power and storage resources of the local edge device. For example, resource-constrained devices can use lightweight configuration files or rule bases, while resource-rich devices can use a complete graph database to support more complex semantic queries.
[0122] Understandably, through the aforementioned directed graph structure, the device capability graph unifies the user, space, device, capability, and status information in the home as nodes and the connections between nodes, providing a structured data query basis for the local intelligent agent when performing feasibility verification, enabling the local intelligent agent to quickly query the location, current status, permission ownership, and functional boundaries of the target device.
[0123] In some embodiments of this application, the mechanism for updating the association between the device capability map and the real-time device status includes the following sub-steps: Step f1: Access the controlled device through the edge gateway.
[0124] The edge gateway is used to receive status change messages reported by the controlled devices and publish the status change messages to the message bus.
[0125] Specifically, the local intelligent agent accesses all controlled home devices through an edge gateway deployed locally in the home. The edge gateway connects to each controlled device via a wireless communication protocol (such as Zigbee, Wi-Fi, or Bluetooth) to receive status change messages reported by each device and publish these messages to the edge gateway's message bus. The message bus can use either the Message Queuing Telemetry Transport (MQTT) protocol or the WebSocket protocol.
[0126] Step f2: When a state change occurs in a home-controlled device, receive a state change event pushed by the edge gateway.
[0127] Specifically, the local agent pre-subscribes to device state change topics on the edge gateway's message bus. When a physical device's state changes (e.g., a user manually turns off a light by pressing a physical switch, the device goes offline due to a malfunction, or sensor values change), the gateway pushes the state change event to the message bus in real time, and the local agent receives the event by subscribing. The state change event includes at least the device identifier and the changed state value.
[0128] Step f3: Based on the state change event, update the edge between the device node and the state node that represents the current real-time state of the device.
[0129] Specifically, after receiving a state change event, the local agent locates the corresponding device node in the device capability graph based on the device identifier in the event, and updates the edge between the device node and the state node, which represents the current real-time state of the device, to the changed state value. For example, when a user manually turns off the living room light, the edge gateway receives the light's state change message ("living room light, state = off"), publishes the message to the MQTT message bus, and after the local agent subscribes to the event, it updates the current state edge between the device node "living room light" and the state node in the device capability graph to "off".
[0130] Understandably, through the aforementioned event subscription and listening mechanism, the device capability graph remains "active" in the local gateway or central control memory, achieving real-time synchronization between the graph and the physical device status. When the cloud-based big data model returns a control policy, the local agent can query the graph in real time to obtain the current real status of each target device, thereby ensuring that the status information used for verification is strictly consistent with the actual situation in the physical world. If the cloud-based big data model requests "turn on the living room light," but the graph shows that the current status of the light node is "offline," the execution gate will directly intercept the instruction and inform the user that the device is unavailable, thus achieving strong consistency verification between the physical state and the AI control instructions.
[0131] In some embodiments of this application, the rapid identification process involved in the aforementioned method steps can be implemented using a hybrid architecture of a lightweight language model and a rule engine. As an example, the lightweight language model can be an edge model with approximately 1B of parameters (such as MobileBERT or ALBERT), which performs low-latency inference at the edge by adding a text classification header (for intent and risk classification) and a sequence labeling header (for extracting room, device, and action slots) on top of it.
[0132] As another example, an algorithm based on traditional Natural Language Understanding (NLU) and rule matching can also be used: First, the input text is segmented and key entities are extracted using a Term Frequency-Inverse Document Frequency (TF-IDF) or Conditional Random Field (CRF) model to obtain slot information; then, based on the extracted features, a Support Vector Machine (SVM) or Lightweight Decision Tree is used to determine whether it represents a control intent; finally, the extracted entities are input into a pre-set local rule graph to calculate risk values, and ambiguity scores are calculated based on slot completeness. Both implementation methods can achieve the recognition function in step 102 above, and those skilled in the art can flexibly choose according to the computing resources of the edge device.
[0133] In some embodiments of this application, the control strategy returned by the cloud-based large model is a structured control plan. The structured control plan may include fields such as target room, target device, action type, parameter range, execution order, and whether user confirmation is required.
[0134] Understandably, by using a structured output format, large cloud models cannot directly generate free-flowing chat-style text to send to the device control link, effectively reducing the risk of mis-controlling devices due to model illusions or abnormal output formats.
[0135] In some embodiments of this application, the execution gate in step 107 can be triggered either after the device feasibility verification in step 106 or deployed as a pre-verification in the local fast control channel in step 103. Specifically, in the local fast control channel, the execution gate can also perform permission verification and device online status verification on the first device control command generated locally, thereby ensuring that all device control commands processed by the solution of this application (regardless of whether they come from the local fast path or the cloud-delegated path) can pass the unified execution gate security verification, achieving full-link security coverage.
[0136] In some embodiments of this application, the preset risk rule base can be constructed as follows: During the initialization phase of the smart home system, the system obtains the object model definitions of each device through the Internet of Things platform, and assigns a corresponding risk level value to each combination based on the combination of device type and action type. For example, the mapping relationship between device type "light" and action type "switch" is "low risk"; the mapping relationship between device type "door lock" and action type "open" is "high risk"; the mapping relationship between device type "camera" and action type "close" is "high risk"; and the mapping relationship between device type "gas valve" and action type "close" is "high risk". Users can customize the above default mapping relationships in the smart home application to meet the personalized security needs of different families.
[0137] In some embodiments of this application, the specific formats of the first and second device control instructions depend on the underlying device communication protocol. As an example, for devices supporting the Zigbee protocol, the device control instruction may be a data frame encoding the target device, action type, and parameters into a Zigbee cluster command format; for devices supporting the Wi-Fi protocol, the device control instruction may be a JSON format command sent via Hypertext Transfer Protocol (HTTP) or Message Queuing Telemetry Transport Protocol. It is understood that the local agent automatically adapts the corresponding instruction format based on the device type and protocol information recorded in the device capability map, without requiring manual configuration by the user.
[0138] In some embodiments of this application, the motion parameter verification in step 106 can be differentiated according to the specific device type. For example, for lighting devices, the motion parameter range verification includes whether the brightness value is within the range of 0% to 100%; for air conditioning devices, the motion parameter range verification includes whether the set temperature is within the range of 16℃ to 30℃ and whether the working mode is a valid value among "cooling", "heating", "air supply", and "dehumidification"; for curtain devices, the motion parameter range verification includes whether the opening degree is within the range of 0% to 100%. It is understood that different types of devices have different parameter spaces. By performing parameter verification using the parameter range information of each device pre-stored in the device capability map, it is possible to prevent device abnormalities or damage caused by parameters exceeding the limits.
[0139] In some embodiments of this application, such as Figure 2As shown, after a user inputs a voice or text command to the local AI agent, the local AI agent first generates a task identifier and reads the session summary, device online status, user identity, and historical execution records to form the context information of the current task. Then, it calls the fast judgment module to extract slots, classify intents, match risk rules, and evaluate ambiguity for the user command, obtaining recognition results including control type, risk level, ambiguity level, target device, and action type. When the recognition result meets the preset execution conditions, the local AI agent directly enters the local fast control channel, generates the first device control command based on the target device and action type, and sends it to the controlled device. When the recognition result does not meet the preset execution conditions, the local AI agent constructs a delegation task package including the original text, recognition result, candidate device set, device capability constraints, and session summary, and sends it to the cloud-based big model. This allows the cloud-based big model to generate a structured control strategy within the boundaries defined by the device capability constraints and return it to the local AI agent. After receiving the results from the cloud, the local intelligent agent performs constraint convergence verification (i.e., device feasibility verification) on the structured control strategy by combining the device capability map and real-time device status. It filters out control actions that do not exist, have mismatched capabilities, or have parameters that exceed limits. Control actions that pass the verification further enter the execution gate. After passing permission verification, device status verification, time window verification, and risk level threshold verification, the second device control command is sent to the controlled device and the execution status is written back. If the execution gate verification fails, the command is rejected and a supplementary interaction process is triggered, requiring the user to confirm or supplement information before making a judgment.
[0140] According to the instruction processing method for smart home control proposed in this application, the local intelligent agent identifies user instructions and performs triage processing based on a comprehensive judgment result of control type, risk level, and ambiguity level. When the execution conditions are met, the device control instruction is directly generated and issued locally. When the execution conditions are not met, the cloud-based large model is delegated to generate a control strategy, and the local system performs device feasibility verification. This allows simple control instructions to respond quickly through the local path, reducing dependence on cloud computing power and network transmission latency. At the same time, complex instructions can be processed with the help of the semantic understanding capabilities of the cloud-based large model, balancing low latency and complex task processing capabilities in smart home control scenarios. By performing a three-layer feasibility verification of the control strategy returned by the cloud—device existence, capability matching, and parameter range—and issuing the execution only after the verification is passed, the risk of erroneous execution caused by the cloud-generated control strategy being out of sync with the actual capabilities of the local device is avoided, thus improving the security and stability of the smart home control link.
[0141] Figure 3 This is a block diagram of an instruction processing device for smart home control, according to an exemplary embodiment. (Refer to...) Figure 3The device includes a generation unit 301, an identification unit 302, a first distribution unit 303, a sending unit 304, an acquisition unit 305, a verification unit 306, and a second distribution unit 307.
[0142] The generation unit 301 is used to receive user instructions and generate context information for the current task based on user instructions, user information, and device information; the device information is the device information of the home-controlled device. The identification unit 302 is used to identify user commands and obtain identification results; the identification results include control type, risk level, ambiguity level, target device, and action type. The first issuing unit 303 is used to generate a first device control command based on the target device and action type when the recognition result meets the preset execution conditions, and to issue the first device control command to the controlled device that matches the user command. The sending unit 304 is used to generate a task package based on context information when the recognition result does not meet the preset execution conditions, and send the task package to the cloud big model so that the cloud big model can generate a control strategy based on the task package. The acquisition unit 305 is used to receive the control strategy returned by the cloud-based large model and to acquire the device capability map and device status information; Verification unit 306 is used to verify the feasibility of the control strategy based on context information, equipment capability map and equipment status information. The second issuing unit 307 is also used to issue a second device control command corresponding to the control strategy to the controlled device that matches the user command after the device feasibility verification is passed.
[0143] In some embodiments, the identification unit 302 may specifically be used for: Slot extraction is performed on user commands to obtain slot information; slot information includes the target device and action type. User commands are classified by intent to obtain control type; Select a risk level that matches the slot information from the preset risk rule base; the preset risk rule base includes the mapping relationship between slot information and risk levels; The ambiguity score is calculated based on the completeness of the slot information. The ambiguity score is then compared with a preset ambiguity threshold to obtain the ambiguity level.
[0144] In some embodiments, the preset execution conditions include: the control type is a control request, the risk level is lower than a preset risk threshold, and the ambiguity level is lower than a preset ambiguity threshold.
[0145] In some embodiments, the task package includes the raw text of the user instruction, the recognition result, the candidate device set, the device capability constraints, and the session summary.
[0146] In some embodiments, the verification unit 306 may specifically be used for: For each control action in the control strategy, perform the following verification operation: Verify whether the target device in the control strategy exists in the device capability map; Verify whether the target device has the capability to execute the action types specified in the control strategy; Verify whether the action parameters in the control strategy are within the allowable parameter range of the target device; If any one or more of the above verification operations fail, the corresponding control action will be removed from the control strategy.
[0147] In some embodiments, the verification unit 306 may also be used for: When a control policy is found to include a control operation targeting a preset high-risk device type, at least one of the following high-risk verification operations is performed: Initiate interactive secondary confirmation via voice broadcast or pop-up window on the central control screen, and receive confirmation instructions from user feedback; When the user command is a voice command, the voiceprint feature of the user command is obtained, and the voiceprint feature is compared with the registered voiceprint database to obtain the voiceprint verification result. The system acquires facial images captured by the central control screen camera, compares these images with the registered facial database, and obtains the facial verification result. Push an authorization request to the bound terminal and receive authorization confirmation information returned by the bound terminal; If all high-risk verification operations pass, the next step is to issue a second device control command corresponding to the control strategy to the controlled device that matches the user command.
[0148] In some embodiments, the device capability graph adopts a directed graph structure, which includes user nodes, spatial nodes, device nodes, capability nodes, and state nodes. The device node and the spatial node are connected by edges that represent the location of the device; The device node and the status node are connected by an edge that represents the current real-time status of the device; The device node and the user node are connected by an edge that represents the device's operating permissions; The device node and the capability node are connected by edges that characterize the functions of the device.
[0149] In some embodiments, the apparatus further includes an updating unit, which may be used for: The controlled device is accessed through an edge gateway; the edge gateway is used to receive status change messages reported by the controlled device and publish the status change messages to the message bus. When a home-controlled device experiences a state change, it receives a state change event pushed by the edge gateway. Based on the state change event, update the edge between the device node and the state node that represents the current real-time state of the device.
[0150] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.
[0151] According to the instruction processing device for smart home control proposed in this application embodiment, the local intelligent agent identifies user instructions and performs triage processing based on a comprehensive judgment result of control type, risk level, and ambiguity level. When the execution conditions are met, the device control instruction is directly generated and issued locally. When the execution conditions are not met, the cloud-based large model is delegated to generate a control strategy, and the device feasibility is verified locally. This allows simple control instructions to respond quickly through the local path, reducing dependence on cloud computing power and network transmission latency. At the same time, complex instructions can be processed with the help of the semantic understanding capability of the cloud-based large model, taking into account both low latency and complex task processing capabilities in smart home control scenarios. By performing a three-layer feasibility verification of the control strategy returned by the cloud—device existence, capability matching, and parameter range—and issuing the execution only after the verification is passed, the risk of erroneous execution caused by the control strategy generated by the cloud being out of sync with the actual capabilities of the local device is avoided, thus improving the security and stability of the smart home control link.
[0152] Figure 4 This is a block diagram illustrating an apparatus for a command processing method for smart home control, according to an exemplary embodiment. For example, apparatus 400 may be an electronic device, such as a mobile phone, computer, digital broadcasting terminal, messaging device, tablet device, personal digital assistant, etc.
[0153] Reference Figure 4 The device 400 may include one or more of the following components: processing component 402, memory 404, power component 406, multimedia component 408, audio component 410, input / output I / O interface 412, sensor component 414, and communication component 416.
[0154] Processing component 402 typically controls the overall operation of device 400, such as operations associated with display, telephone calls, data communication, camera operation, and recording. Processing component 402 may include one or more processors 420 to execute instructions to perform all or part of the steps of the methods described above. Furthermore, processing component 402 may include one or more modules to facilitate interaction between processing component 402 and other components. For example, processing component 402 may include a multimedia module to facilitate interaction between multimedia component 408 and processing component 402.
[0155] Memory 404 is configured to store various types of data to support the operation of device 400. Examples of such data include instructions for any application or method operating on device 400, contact data, phonebook data, messages, pictures, videos, etc. Memory 404 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0156] The power supply component 406 provides power to the various components of the device 400. The power supply component 406 may include a power management system, one or more power sources, and other components associated with generating, managing, and distributing power to the device 400.
[0157] Multimedia component 408 includes a screen that provides an output interface between the device 400 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touchscreen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors may sense not only the boundaries of the touch or swipe action but also the duration and pressure associated with the touch or swipe operation. In some embodiments, multimedia component 408 includes a front-facing camera and / or a rear-facing camera. When the device 400 is in an operating mode, such as a shooting mode or a video mode, the front-facing camera and / or the rear-facing camera may receive external multimedia data. Each front-facing camera and rear-facing camera may be a fixed optical lens system or have focal length and optical zoom capabilities.
[0158] Audio component 410 is configured to output and / or input audio signals. For example, audio component 410 includes a microphone (MIC) configured to receive external audio signals when device 400 is in an operating mode, such as call mode, recording mode, and voice recognition mode. The received audio signals may be further stored in memory 404 or transmitted via communication component 416. In some embodiments, audio component 410 also includes a speaker for outputting audio signals.
[0159] I / O interface 412 provides an interface between processing component 402 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include, but are not limited to, home buttons, volume buttons, power buttons, and lock buttons.
[0160] Sensor assembly 414 includes one or more sensors for providing status assessments of various aspects of device 400. For example, sensor assembly 414 may detect the on / off state of device 400, the relative positioning of components such as the display and keypad of device 400, changes in the position of device 400 or a component of device 400, the presence or absence of user contact with device 400, the orientation or acceleration / deceleration of device 400, and temperature changes of device 400. Sensor assembly 414 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. Sensor assembly 414 may also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, sensor assembly 414 may also include an accelerometer, a gyroscope, a magnetometer, a pressure sensor, or a temperature sensor.
[0161] Communication component 416 is configured to facilitate wired or wireless communication between device 400 and other devices. Device 400 can access wireless networks based on communication standards, such as WiFi, 2G, or 4G, or combinations thereof. In one exemplary embodiment, communication component 416 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, communication component 416 also includes a near-field communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on radio frequency identification (RFID) technology, Infrared Data Association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0162] In an exemplary embodiment, the apparatus 400 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the methods described above.
[0163] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 404 including instructions, which can be executed by a processor 420 of the device 400 to perform the above-described method. For example, the non-transitory computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.
[0164] In an exemplary embodiment, a computer program product is also provided, including a computer program that implements the above-described method when executed by the processor 420 of the device 400.
[0165] It should be understood that the terminology used herein is for the purpose of describing particular exemplary embodiments only and is not intended to be limiting. Unless the context clearly indicates otherwise, the singular forms “a,” “an,” and “described” as used herein may also include the plural forms. The terms “comprising,” “including,” “containing,” and “having” are inclusive and therefore indicate the presence of the stated features, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, elements, components, and / or combinations thereof. The method steps, processes, and operations described herein are not construed as requiring them to be performed in a particular order described or illustrated unless the order of performance is explicitly indicated. It should also be understood that additional or alternative steps may be used.
[0166] Although terms such as first, second, third, etc., may be used in this document to describe multiple elements, components, regions, layers, and / or segments, these elements, components, regions, layers, and / or segments should not be limited by these terms. These terms may be used only to distinguish one element, component, region, layer, or segment from another. Unless the context clearly indicates otherwise, terms such as "first," "second," and other numerical terms used herein do not imply order or sequence. Therefore, the first element, component, region, layer, or segment discussed below may be referred to as the second element, component, region, layer, or segment without departing from the teachings of the exemplary embodiments.
[0167] The above description is merely a specific embodiment of this application, enabling those skilled in the art to understand or implement this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features claimed herein.
Claims
1. A method for processing instructions for smart home control, characterized in that, Applied to local intelligent agents, including: Receive user instructions, and generate context information for the current task based on the user instructions, user information, and device information; the device information refers to the device information of the home-controlled device. The user command is identified to obtain an identification result; the identification result includes control type, risk level, ambiguity level, target device, and action type. When the recognition result meets the preset execution conditions, a first device control command is generated based on the target device and the action type, and the first device control command is sent to the controlled device that matches the user command; When the recognition result does not meet the preset execution conditions, a task package is generated based on the context information and sent to the cloud big model so that the cloud big model can generate a control strategy based on the task package; Receive the control strategy returned by the cloud-based big data model, and obtain the device capability map and device status information; Based on the context information, equipment capability map, and equipment status information, the control strategy is verified for equipment feasibility. After the device feasibility verification is passed, a second device control command corresponding to the control strategy is issued to the controlled device that matches the user command.
2. The method according to claim 1, characterized in that, The process of recognizing the user instruction and obtaining the recognition result includes: The user command is processed to extract slot information; the slot information includes the target device and the action type. The user instructions are classified according to intent to obtain the control type; The risk level that matches the slot information is selected from the preset risk rule base; the preset risk rule base includes the mapping relationship between slot information and risk level; The ambiguity score is calculated based on the completeness of the slot information, and the ambiguity score is compared with a preset ambiguity threshold to obtain the ambiguity level.
3. The method according to claim 1, characterized in that, The preset execution conditions include: the control type is a control request, the risk level is lower than a preset risk threshold, and the ambiguity level is lower than a preset ambiguity threshold.
4. The method according to claim 1, characterized in that, The task package includes the original text of the user instruction, the recognition result, a set of candidate devices, device capability constraints, and a session summary.
5. The method according to claim 1, characterized in that, The step of verifying the equipment feasibility of the control strategy based on the context information, equipment capability map, and equipment status information includes: For each control action in the control strategy, perform the following verification operation: Verify whether the target device in the control strategy exists in the device capability map; Verify whether the target device has the capability to execute the action type in the control strategy; Verify whether the action parameters in the control strategy are within the allowable parameter range of the target device; If any one or more verification operations fail, the corresponding control action will be removed from the control strategy.
6. The method according to claim 1, characterized in that, After the equipment feasibility verification is passed, the following is also included: In response to the detection that the control strategy includes a control operation targeting a preset high-risk device type, at least one of the following high-risk verification operations is performed: Initiate interactive secondary confirmation via voice broadcast or pop-up window on the central control screen, and receive confirmation instructions from user feedback; When the user command is a voice command, the voiceprint features of the user command are obtained, and the voiceprint features are compared with the registered voiceprint database to obtain the voiceprint verification result. Acquire a facial image captured by the central control screen camera, compare the facial image with the registered facial database, and obtain the facial verification result; The system pushes an authorization request to the bound terminal and receives authorization confirmation information returned by the bound terminal. If all the high-risk verification operations pass, the step of issuing a second device control command corresponding to the control strategy to the controlled device that matches the user command is executed.
7. The method according to claim 1, characterized in that, The device capability graph adopts a directed graph structure, which includes user nodes, space nodes, device nodes, capability nodes, and status nodes. The device node and the spatial node are connected by an edge that represents the location of the device; The device node and the status node are connected by an edge that represents the current real-time status of the device. The device node and the user node are connected by an edge that represents the device's operating permissions; The device node and the capability node are connected by an edge that represents the functions of the device.
8. The method according to claim 7, characterized in that, Also includes: The controlled device is accessed via an edge gateway; The edge gateway is used to receive status change messages reported by the controlled device and publish the status change messages to the message bus; When a home-controlled device undergoes a state change, it receives a state change event pushed by the edge gateway; Based on the state change event, update the edge between the device node and the state node that represents the current real-time state of the device.
9. An instruction processing device for smart home control, characterized in that, include: The generation unit is used to receive user instructions and generate context information for the current task based on the user instructions, user information, and device information. The device information refers to the device information of home-controlled devices. The identification unit is used to identify the user command and obtain the identification result; the identification result includes control type, risk level, ambiguity level, target device, and action type. The first issuing unit is used to generate a first device control command based on the target device and action type when the recognition result meets the preset execution conditions, and to issue the first device control command to the controlled device that matches the user command. The sending unit is used to generate a task package based on the context information when the recognition result does not meet the preset execution conditions, and send the task package to the cloud big model so that the cloud big model can generate a control strategy based on the task package; The acquisition unit is used to receive the control strategy returned by the cloud-based big data model, and to acquire the device capability map and device status information; The verification unit is used to verify the feasibility of the control strategy based on the context information, the device capability map, and the device status information. The second issuing unit is also used to issue a second device control command corresponding to the control strategy to the controlled device that matches the user command after the device feasibility verification is passed.
10. An electronic device, characterized in that, include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, implements the method as described in any one of claims 1 to 8.