A context reasoning and control method and system based on a home agent
Patent Information
- Application Number
- CN202610759988.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-29
- Publication Date
- 2026-09-29
AI Technical Summary
其一,缺乏情境理解能力:现有系统无法理解用户的深层意图以及非结构化的环境状态
[0009]通过如上所提供的基于家庭智能体的情境推理与控制方案,本申请实施例构建了一个从“环境感知-意图理解-动作编排-闭环执行”的全链路家庭智能控制体系。首先,通过采集多源异构感知数据生成情境快照,打破了单一传感器或设备的信息孤岛,实现了对家庭复杂物理环境与用户状态的全面、精准捕获。其次,将情境快照与家庭习惯图谱结合提取双层意图标签,使系统突破了传统机械式声控或触控指令的局限,具备了挖掘用户深层真实需求的推理能力,从而能够提供高度个性化的主动式服务。再次,将宏观的服务目标细粒度拆解为原子动作序列,保障了跨设备协同编排的灵活性与执行的有序性。最后,引入了执行后数据实时反馈的闭环机制,使系统能够根据终端执行的真实物理效果动态刷新情境快照。这种“感知-决策-执行-反馈”的闭环设计赋予了家庭智能体极强的自适应纠偏与动态演进能力,大幅提升了全场景服务的鲁棒性与用户体验的连贯性。
Smart Images

Figure CN122837265A_ABST
Abstract
Description
Technical Field
[0001] This application generally relates to the interdisciplinary field of smart home and artificial intelligence. More specifically, this application relates to a situational reasoning and control method and system based on a home intelligent agent. Background Technology
[0002] The current smart home industry is still in the "connectivity and control" stage. Users typically need to issue explicit commands to devices via mobile apps or voice assistants (such as "turn on the living room lights"), and the devices can only passively execute these commands. This traditional interaction paradigm has the following fundamental limitations in practical applications: First, there is a lack of contextual understanding: existing systems cannot understand users' deeper intentions and unstructured environmental states. For example, the system cannot autonomously recognize the complex life scenario of "the child is asleep," and therefore cannot automatically dim the lights or turn off the TV to meet the user's implicit needs.
[0003] Secondly, there are privacy and latency bottlenecks: existing advanced semantic understanding relies heavily on large cloud models, which means that highly private data such as audio and video within the home must be uploaded to the cloud for processing, posing a serious risk of data leakage; at the same time, network latency caused by cloud communication also seriously affects the real-time interactive experience of smart homes.
[0004] Third, services are fragmented: existing smart devices (such as cameras, sensors, home appliances, etc.) often operate independently, with isolated underlying data and a lack of a unified cross-modal semantic understanding framework, which makes it impossible for devices to form collaborative, proactive, and forward-looking whole-house smart services.
[0005] In view of this, there is an urgent need to provide a situational reasoning and control scheme based on home intelligent agents to solve the above problems and enable smart home systems to leap from the traditional passive execution of commands to proactive perception and service. Summary of the Invention
[0006] In order to at least address one or more of the technical problems mentioned above, this application proposes a situational reasoning and control scheme based on family intelligent agents in several aspects.
[0007] In a first aspect, this application provides a contextual reasoning and control method based on a family intelligent agent, comprising: collecting multi-source heterogeneous sensing data within the family and generating a contextual snapshot based on the multi-source heterogeneous sensing data; obtaining a two-layer intent label based on the contextual snapshot and a pre-established family habit map; generating a service target set based on the two-layer intent label and decomposing the service target set into atomic action sequences; distributing the atomic action sequences to corresponding terminal devices for execution, and collecting multi-source heterogeneous sensing data after execution in real time to feed back to the step of generating a contextual snapshot containing visual, voice, environmental, and device status features.
[0008] In a second aspect, this application provides a contextual reasoning and control system based on a home intelligent agent, employing the contextual reasoning and control method based on a home intelligent agent as described in any embodiment of the first aspect. The system includes: a data perception and snapshot generation module, used to collect multi-source heterogeneous perception data within the home and generate a contextual snapshot based on the multi-source heterogeneous perception data; an intent graph analysis module, used to obtain two-layer intent labels based on the contextual snapshot and a pre-established family habit graph; a target decomposition and sequence generation module, used to generate a service target set based on the two-layer intent labels and decompose the service target set into atomic action sequences; and a distribution execution and feedback closed-loop module, used to distribute the atomic action sequences to corresponding terminal devices for execution and collect multi-source heterogeneous perception data after execution in real time to feed back to the step of generating a contextual snapshot containing visual, voice, environmental, and device status features.
[0009] Based on the scenario reasoning and control scheme provided above for home intelligent agents, this application embodiment constructs a full-link home intelligent control system from "environmental perception - intent understanding - action orchestration - closed-loop execution". First, by collecting multi-source heterogeneous perception data to generate scenario snapshots, the information silos of single sensors or devices are broken, achieving comprehensive and accurate capture of the complex physical environment of the home and the user's state. Second, by combining scenario snapshots with a family habit map to extract dual-layer intent tags, the system overcomes the limitations of traditional mechanical voice or touch commands, possessing the reasoning ability to uncover the user's deep, real needs, thus providing highly personalized proactive services. Third, the macro-level service goals are broken down into fine-grained atomic action sequences, ensuring the flexibility of cross-device collaborative orchestration and the orderly execution. Finally, a closed-loop mechanism for real-time data feedback after execution is introduced, enabling the system to dynamically refresh the scenario snapshot based on the actual physical effects of terminal execution. This closed-loop design of "perception-decision-execution-feedback" endows the home intelligent agent with strong adaptive correction and dynamic evolution capabilities, significantly improving the robustness of full-scenario services and the consistency of user experience.
[0010] Furthermore, in some embodiments, a deep intent recognition mechanism is constructed. First, a negative space feature vector is introduced. Traditional smart homes rely solely on currently occurring explicit states for judgment, while this solution, by comparing family habit maps, keenly captures missing elements in the current situation that should have occurred but did not (such as absence of personnel, interruption of habitual behavior, etc.). This mechanism of reasoning using reverse features greatly broadens the system's perception dimension, enabling it to gain insight into users' implicit needs and abnormal situations. Second, a complete temporal logical chain is established from tracing historical causes to predicting future trends. By retrieving antecedent causes such as time and events from the contextual knowledge base and combining them with the current snapshot and negative space features for joint matching, the system not only clarifies why the current state occurred but also accurately predicts the user's next action. Finally, the reasoning results are scientifically decoupled into surface and deep dual-layer intent labels. Surface intent addresses immediate direct needs, while deep intent, based on causes and predictions, directly addresses the user's fundamental purpose. This endows the home intelligent agent with a high degree of anthropomorphic empathy, forward-looking predictive ability, and the advantage of providing coherent, proactive, and in-depth butler-style services.
[0011] Furthermore, in some embodiments, a dynamic optimization and generation mechanism for service targets, guided by core user values and strictly controlled within security boundaries, is constructed. First, an absolutely feasible and secure candidate target solution space is built, eliminating the generation of contradictory instructions at the source and ensuring the physical security of hardware execution and system stability. Second, in the service target optimization phase, the complex intelligent decision-making process is transformed into a rigorous mathematical optimization problem, enabling precise evaluation of each candidate target set within the vast candidate target solution space. Finally, using an optimal calculation formula that includes utility satisfaction and a hard constraint masking function, not only can the service target set best matching the dual-layer intent label be selected, but it can also ensure that this target set maximizes user value experience while absolutely not violating family safety boundaries or causing unnecessary disturbance to the user. This endows the family intelligent agent with strong artificial intelligence decision-making capabilities for dynamic programming of globally optimal solutions under complex constraints.
[0012] Furthermore, in some embodiments, an ordered sequence of atomic actions is constructed. First, the macroscopic service target set is precisely quantified and mapped to a fine-grained set of atomic actions, and a conflict score calculation formula is introduced. This enables the pre-identification and interception of high-risk conflict operations, and employs a multi-dimensional flexible resolution strategy to effectively avoid hardware resource contention and physical deadlock caused by multi-task concurrency, greatly improving the security and stability of the system during concurrent processing. Second, by rigorously analyzing the prerequisite state set and supply state set of each legal atomic action, the causal constraints between actions are automatically derived, thereby constructing a rigorous set of logically dependent edges. Combined with a priority-based topological sorting algorithm formula, scattered actions can be arranged into an absolutely reasonable and ordered sequence of atomic actions. This eliminates logical paradoxes caused by disordered control and ensures the temporal continuity of complex cross-device collaborative actions. Attached Figure Description
[0013] The above and other objects, features, and advantages of exemplary embodiments of this application will become readily understood by reading the following detailed description with reference to the accompanying drawings. In the drawings, several embodiments of this application are illustrated by way of example and not limitation, and the same or corresponding reference numerals denote the same or corresponding parts, wherein: Figure 1 An exemplary flowchart of the situational reasoning and control method based on a home intelligent agent according to an embodiment of this application is shown; Figure 2 An exemplary flowchart illustrating the process of obtaining a two-layer intent label according to an embodiment of this application is shown; Figure 3 An exemplary flowchart illustrating the generation of service target sets based on two-layer intent tags according to an embodiment of this application is shown; Figure 4 An exemplary flowchart illustrating the construction of the candidate target set solution space based on context snapshots in an embodiment of this application is shown; Figure 5 An exemplary flowchart illustrating conflict resolution of corresponding atomic actions according to an embodiment of this application is shown; Figure 6 An exemplary flowchart illustrating the generation and output of atomic action sequences according to an embodiment of this application is shown; Figure 7 An exemplary structural block diagram of a home-based intelligent agent-based situational reasoning and control system according to an embodiment of this application is shown. Detailed Implementation
[0014] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0015] It should be understood that the terms "comprising" and "including" used in the specification and claims of this application indicate the presence of the described features, integrals, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.
[0016] It should also be understood that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the application. As used in this specification and claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in this specification and claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes such combinations.
[0017] Figure 1 An exemplary flowchart of a situational reasoning and control method 100 based on a family intelligent agent according to an embodiment of this application is shown.
[0018] like Figure 1 As shown, in step S110, multi-source heterogeneous sensing data within the home is collected, and a contextual snapshot is generated based on the multi-source heterogeneous sensing data.
[0019] In the embodiments of this application, the home smart gateway continuously subscribes to the data streams of all sensing devices (cameras, microphone arrays, various sensors) within the home via the HarmonyOS soft bus.
[0020] In embodiments of this application, the context snapshot includes visual features, voice features, environmental features, and device status features.
[0021] In the embodiments of this application, during the generation of a contextual snapshot: First, continuous video frames are analyzed to extract motion intentions representing human movement vectors and speeds, and sparse keyframes are simultaneously filtered to obtain visual features. The audio stream is processed in real time to extract text features containing semantic text content, acoustic features containing voiceprints and intonation variations, and spatial sound source localization features based on microphone array directivity, thus obtaining speech features. Environmental features are obtained by real-time monitoring of the numerical fluctuation trends of environmental sensors. Device state features are obtained by real-time capture of discrete state switching events of smart home devices. Then, visual features, motion intentions, speech features, environmental features, and device state features belonging to the same time window are bound together to obtain a contextual snapshot.
[0022] Specifically, by analyzing consecutive video frames and calculating the offset of pixels between adjacent frames using optical flow, motion intent is generated. This image can represent the vector (direction) and velocity of human movement, thus identifying whether the user is "walking towards the sofa" or "running towards the kitchen." Simultaneously, redundant background is removed using algorithms, retaining only sparse keyframes containing key motion changes, preserving the most important visual information with minimal computational overhead.
[0023] Specifically, the audio stream is subjected to speech recognition (ASR) to extract text features, voiceprints reflecting the speaker's biometrics, rhythm and pitch reflecting emotions and tone, and the phase difference (TDOA) of sound waves arriving at each sensor is obtained using a microphone array. The specific coordinates of the sound source in three-dimensional space are calculated using geometric algorithms.
[0024] Specifically, the system continuously monitors the numerical sequences of environmental sensors (such as temperature, humidity, light intensity, and air quality). Instead of recording a single value, the system calculates the slope change (rising, falling, or remaining stable) over a specific period to determine the dynamic changes in the environmental state.
[0025] Specifically, it captures the binary state (such as on / off, locked / unlocked) or multi-state switching signals of smart terminals in real time. This data reflects the immediate results of the interaction between the physical device and the user.
[0026] Specifically, after acquiring the aforementioned heterogeneous features, the system performs time-series alignment of the data using a preset time window threshold T (e.g., 1 second). Visual displacement, voice information, environmental fluctuations, and device feedback collected within the same timestamp range (i.e., within period T) are mapped into contextual snapshots.
[0027] After completing step S110, in step S120, a two-layer intent label is obtained based on the contextual snapshot and the pre-established family habit map.
[0028] In the embodiments of this application, the family habit graph includes family member nodes, spatiotemporal scene nodes, behavioral pattern edges, and collaborative relationship edges. Specifically, family member nodes store the identity identifiers of family members (e.g., elderly, children, father) and behavioral habit vectors representing their personal characteristics. Spatiotemporal scene nodes include specific time periods (e.g., early morning, late night), spatial areas (e.g., bedroom, kitchen), and corresponding typical activity types. Behavioral pattern edges connect family member nodes and spatiotemporal scene nodes, representing the habitual behavioral logic of specific individuals in specific time and space. Collaborative relationship edges connect multiple family member nodes, representing collaborative behavioral patterns involving multiple people (e.g., dining together, watching movies together).
[0029] In the embodiments of this application, in the process of obtaining two-layer intent labels based on contextual snapshots and pre-established family habit maps, firstly, the contextual snapshots are compared with the pre-established family habit maps, and a negative space feature vector is generated based on the comparison results. Then, the two-layer intent labels are obtained by matching the contextual snapshots with the negative space feature vectors.
[0030] In the embodiments of this application, the comparison results represent the missing expected elements in the current context. Through comparison, situations that "should theoretically occur but did not actually occur" can be identified. The missing expected elements in the current context include: missing members, missing collaboration, and missing states. Missing members refer to family members who should have appeared in a specific time period and spatial area (based on map prediction) but were not actually sensed. Missing collaboration refers to multi-person collaborative behavior that should have occurred in a specific time and space but did not. Missing states refer to intermediate transition states that should exist in a specific standard behavioral sequence but did not actually exist.
[0031] By extracting these missing elements and encoding them into vectors, a negative space feature vector that can represent "abnormal or implicit needs" is obtained.
[0032] In the embodiments of this application, the specific details regarding obtaining the two-layer intent label based on contextual snapshot and negative space feature vector matching can be found in [reference needed]. Figure 2 .
[0033] Figure 2 An exemplary flowchart illustrating how to obtain a two-layer intent label according to an embodiment of this application is shown.
[0034] like Figure 2As shown, in step S210, the context snapshot is input as a query vector into the pre-built context knowledge base to retrieve historical contexts with similarity meeting a preset threshold and their corresponding antecedents. In step S220, antecedent labels and corresponding confidence levels are generated based on the antecedents. In step S230, the current context snapshot and negative space feature vector are concatenated to obtain joint query features, and corresponding historical behavior trajectories are matched in the pre-built user behavior knowledge base based on the joint query features. In step S240, subsequent behavior sequences are matched based on historical behavior trajectories, and the next behavior prediction result is generated based on the subsequent behavior sequence. In step S250, a two-layer intent label is output based on the context snapshot, the antecedent label, and the one-step behavior prediction result.
[0035] In the embodiments of this application, the pre-existing triggers include time-triggered conditions (such as regular work and rest patterns), event-triggered conditions (such as pre-action guidance), and state-triggered conditions (such as sudden environmental changes or device-reported state changes).
[0036] In the embodiments of this application, during step S220, the preceding triggers retrieved in step S210 are first classified and mapped to a preset trigger semantic library. Specifically, this includes time-based trigger mapping, event-based trigger mapping, and state-based trigger mapping. Time-based trigger mapping converts specific time points (e.g., "18:00") into periodic semantic labels. For example, based on a family habit map, if this time point coincides with a user's dinner time, it is mapped as a "periodic sleep-wake cycle trigger." Event-based trigger mapping converts discrete sensor events (e.g., "detecting the door opening") into logical trigger labels. For example, if "door opening" is detected and the user is carrying a "shopping bag," it is mapped as a "returning home from shopping trigger." State-based trigger mapping converts changes in environmental values (e.g., "increased PM2.5 levels") into environmentally driven labels, such as "sudden environmental quality change trigger."
[0037] Then, a multi-factor fusion algorithm is used to calculate the corresponding confidence score Conf for each generated trigger label. The calculation formula is as follows: Conf=∑(A×Sim+B×Freq+C×Rel), where Sim is the context similarity weight, representing the vector cosine similarity between the current context snapshot and the retrieved historical context. The higher the similarity, the greater the reference value of the current trigger. Freq is the habit recurrence rate, which is the statistical probability that the preceding trigger leads to the occurrence of a specific intention in the family habit map. For example, if there is a 90% probability that "preparing dinner" will occur after "18:00" in historical data, then Freq will have a higher value. Rel is the perceptual reliability factor, which is the confidence coefficient determined based on the accuracy of the sensor that triggered the trigger and the timeliness of data updates. For example, the weight of high-precision visual recognition triggers is greater than that of infrared sensing triggers. A, B, and C are the dynamic adjustment coefficients of each factor, and satisfy A+B+C=1.
[0038] In the embodiments of this application, during step S240, the joint query features generated in step S230 are used to retrieve multiple historical behavior trajectories with similarity higher than a preset threshold from the user behavior knowledge base. For each successfully matched historical trajectory, an ordered sequence of actions of length L is extracted from the spatiotemporal node corresponding to the current context snapshot as the starting anchor point, forming a candidate subsequent action chain set. In this process, the system not only extracts the action type (e.g., "move to the kitchen"), but also simultaneously extracts the duration of the action and the associated equipment state change parameters to ensure that the candidate sequence has complete temporal characteristics. Statistical analysis is performed on the candidate subsequent action chain set to calculate the conditional probability of transitioning from the current state to various possible subsequent actions. Based on the conditional probability, the action with the highest probability value is selected as the prediction result for the next action.
[0039] In the embodiments of this application, the dual-layer intent label includes a surface intent label and a deep intent label.
[0040] In embodiments of this application, the surface intent label is generated based on a context snapshot map. It represents the user's currently performing, observable behavioral intent (e.g., "the user is opening the refrigerator").
[0041] Specifically, the multimodal features from the context snapshot are input into a pre-built action semantic mapping model. Using the encoder of the large multimodal model, the snapshot feature vector is projected onto a predefined atomic action vocabulary space. This vocabulary contains all observable physical action labels in the home scene (e.g., "open the door," "pick up a water glass," "sit on the sofa"). The cosine similarity between the current snapshot features and the action label vectors in the vocabulary is calculated. The system selects the label with the highest similarity score as the surface intent. For example, if visual features identify physical contact between a hand and a refrigerator handle, and the acceleration vector is upward, the system directly maps and outputs the surface intent label "open the refrigerator."
[0042] In the embodiments of this application, the deep intent label is generated based on the inference of the trigger label and the prediction result of the next behavior. It represents the real demand logic behind the user's behavior (e.g., "based on the exercise trigger and hydration prediction, the user intends to find a cold drink").
[0043] Specifically, the trigger label provides the context of the behavior's need. For example, if the trigger label is "State Trigger: Room temperature too high (>28℃)," then the deep intent must point to the category of "adjusting the perceived temperature." The next behavior prediction result provides the endpoint goal of the behavior. For example, if the predicted behavior is "turn on the air conditioner" or "move to the fan," then the specific means of executing the goal are further locked. Through logistic regression or causal graph reasoning, the optimal path that can explain both the "trigger" and the "predicted behavior" is retrieved from the knowledge base. For example, input: trigger label = "Time Trigger: 18:00 (dinner time)"; next behavior prediction = "take out the ingredients and place them on the stove." The "time point requirement" and the "ingredient preparation action" are combined, eliminating distracting items such as simply "organizing the refrigerator." Output: deep intent label Ideep = "preparing to cook dinner."
[0044] When generating deep intents, the negative space feature vector N is referenced. If it is identified that "the person who should be collaborating in the kitchen is missing", the deep intent label may be corrected or supplemented to "prepare dinner independently" or "cook after waiting for the members to arrive", thereby improving the granularity of intent understanding.
[0045] After step S120 is completed, in step S130, a service target set is generated based on the two-layer intent label, and the service target set is decomposed into atomic action sequences.
[0046] For details regarding the specific process of generating a service target set based on two-layer intent tags in the embodiments of this application, please refer to [link / reference]. Figure 3 .
[0047] Figure 3 An exemplary flowchart illustrating the generation of service target sets based on two-layer intent tags according to an embodiment of this application is shown.
[0048] like Figure 3 As shown, in step S310, the user's core value vector in the current context is determined based on user profiles or historical behavior assessments. In step S320, a preset set of safety and interference thresholds is loaded, which defines the negative hard constraint boundaries that cannot be triggered in the current context. In step S330, a candidate target set solution space is constructed based on the context snapshot. In step S340, based on the two-layer intent tags, context snapshot, user core value vector, and safety and interference threshold set, optimization calculations are performed in the candidate target set solution space to obtain the service target set.
[0049] Specifically, the user core value vector represents the user's preference weights for efficiency, comfort, safety, quietness, or energy saving. The set of safety and interference thresholds defines hard constraint boundaries, such as "prohibiting voice interruptions during user calls" and "prohibiting high-brightness lighting at night."
[0050] In the embodiments of this application, the specific process involved in constructing the candidate target set solution space based on context snapshots can be found in [reference needed]. Figure 4 .
[0051] Figure 4 An exemplary flowchart illustrating an embodiment of this application is shown, demonstrating the construction of a candidate target set solution space based on a context snapshot.
[0052] like Figure 4 As shown, in step S410, the scenario snapshot is parsed, and all online and non-exclusively used terminal devices within the current household are traversed to generate an available device matrix. In step S420, a preset device capability registry is queried to extract the discrete control commands and continuous adjustment parameters supported by each device in the available device matrix, which are used as atomic operation units. In step S430, the atomic operation units are combined along the time and device dimensions to generate an initial set containing multiple candidate execution strategies. In step S440, candidate execution strategies with physical conflicts or logical mutual exclusions in the initial set are removed, and the remaining set of candidate execution strategies is defined as the candidate target set solution space.
[0053] In the embodiments of this application, during step S410, the online status of each smart terminal is polled or subscribed to through underlying communication protocols (such as Zigbee, Matter, Wi-Fi). The occupancy semantics of the devices need to be identified. If a device (such as a mobile phone or speaker) is currently explicitly occupied by a user (e.g., during a call or screen mirroring), it is marked as "exclusively occupied." All devices in the "online and not exclusively occupied" state are aggregated to generate an available device matrix. Each row of the matrix represents a physical entity device, and the column vector contains its device ID, location tag (kitchen / living room), current battery level, and communication link quality.
[0054] In the embodiments of this application, during step S420, the adjustable semantic primitives of each device in the available device matrix are extracted by accessing a cloud-based or local device capability model library. Boolean (e.g., "on / off") or enumeration (e.g., "mode: sweeping / mopping / recharging") control commands are extracted. Numerical range parameters, including brightness (0%-100%), volume, color temperature, air conditioning temperature, curtain opening / closing degree, etc., are extracted, and their step values are set according to the device hardware precision. Each independently callable command is encapsulated into an atomic operation unit. For example, the unit set corresponding to the main kitchen light is: {u on / off ,u brightness ,u color_temp}
[0055] In the embodiments of this application, during the execution of step S440, it is checked whether there is a physical paradox within the same time segment of the same device. For example, the instruction set strictly prohibits the simultaneous execution of atomic operations such as "forward rotation (opening the curtains)" and "reverse rotation (closing the curtains)" on the same motor. Filtering is performed based on a semantic rule base. For example, if the current scenario is "preparing to cook," the system will automatically remove branches containing the strategy "the robot vacuum cleaner starts cleaning the kitchen" because it is logically mutually exclusive with "people are walking in the kitchen." All legal, conflict-free, and physically consistent operation combinations retained after pruning are defined as the candidate target solution space, serving as the basis for subsequent optimization calculations.
[0056] In the embodiments of this application, the following calculation formula is used in the process of performing optimization calculations in the solution space of the candidate target set to obtain the service target set: G represents the service target set, and g represents the candidate target set. Let be the solution space for the candidate target set. As the first dynamic adjustment factor, This is a two-layer intent label, where V is the user's core value vector and C is the context snapshot. To determine the utility satisfaction of the candidate target set g with respect to the user's core value vector under the constraint of context snapshot C, Let be the hard constraint masking function, and s be the set of security and interference thresholds.
[0057] In the embodiments of this application, during the process of decomposing a service target set into a sequence of atomic actions, firstly, the service target set is mapped to a set of atomic actions according to a preset decomposition rule. Next, by analyzing the matching degree between the set of execution resources required by each atomic action within the set of atomic actions and the currently occupied resource set, conflict resolution is performed on the corresponding atomic actions to obtain multiple legal atomic actions. Then, an atomic action sequence is generated based on the logical dependencies between the legal atomic actions.
[0058] In the embodiments of this application, the following calculation formula is used in the process of mapping the service target set to the atomic action set: A=G / U D and A represent the set of atomic actions, G represents the set of service targets, U represents the smallest unit of execution (such as triggering a reminder or adjusting a status), and / represents the decomposition operator. For mapping operators, D is used to match the set of execution capabilities and bind unit requirements to specific devices and parameters (such as "kitchen light + set brightness + 80%)", where D is the set of execution capabilities.
[0059] In the embodiments of this application, the specific process of conflict resolution for the corresponding atomic actions can be found in [reference needed]. Figure 5 .
[0060] Figure 5 An exemplary flowchart illustrating conflict resolution of corresponding atomic actions according to an embodiment of this application is shown.
[0061] like Figure 5 As shown, in step S510, the conflict score of each atomic action within the atomic action set is calculated based on the conflict score calculation formula. In step S520, it is determined whether the conflict score is greater than a security threshold. In response to a conflict score not being greater than the security threshold, in step S530, the corresponding atomic action is determined to be a legal atomic action, and its execution parameters are retained. In response to a conflict score greater than the security threshold, in step S540, the corresponding atomic action is determined to be an illegal atomic action, and at least one resolution strategy is triggered based on the attributes and resource conflict type of the corresponding atomic action to eliminate the illegal atomic action or convert the illegal atomic action into a legal atomic action.
[0062] In the embodiments of this application, the conflict score calculation formula is: Conflict=(R task ∩R current )×W, Conflict is the conflict score of atomic actions, R task R is the set of execution resources required for an atomic action. current Let W be the currently occupied resource set, and W be the interference weight matrix.
[0063] In the embodiments of this application, the resolution strategies include interactive replacement, timing delay, action cancellation, and degraded execution.
[0064] In the embodiments of this application, when a conflict occurs in the perception channel (such as audio or visual space) and the action has a high information notification attribute, an interactive replacement is performed. Interactive replacement refers to retrieving a multimodal mapping table and converting the output modality of the original action into another non-conflicting modality. For example, if it is detected that the user is making a voice call, the originally planned action of "the smart speaker plays a voice reminder (e.g., the washing machine is finished)" will trigger an interactive replacement. After the interactive replacement, the speaker's audio output is blocked, and instead, the user's smartwatch provides a physical vibration reminder, or a text notification silently pops up on the smart tablet in front of the user.
[0065] In the embodiments of this application, when the conflicting resource has the characteristic of "temporary occupation" (such as the device running a short-term task), and the real-time requirements of the action are not extremely high, the execution timing is delayed. During the execution timing delay, the atomic action is removed from the current execution queue and attached to a listening trigger. This trigger continuously monitors the release signal of the conflicting resource. For example, if a user is watching a movie in the living room (audio and spatial resources are occupied), and it is predicted that the action of "starting the robot vacuum cleaner to clean the living room" should be executed, the execution timing is delayed, the robot vacuum cleaner enters a "waiting sequence," and the system continuously monitors the TV status. Once the TV is turned off or switched to standby mode, the system immediately triggers the delayed cleaning action.
[0066] In the embodiments of this application, when the priority weight of an action belongs to the "comfort" or "ambience" category (low priority), and the conflict weight is extremely high, and there is no suitable alternative channel, the action is canceled. In this case, the atomic action is directly removed, and a log entry of "not executed due to interference avoidance" is recorded. For example, when the deep intent is identified as "the user is preparing to fall asleep," the original plan was to "play soothing sleep-aiding background music." However, it is detected that another white noise device is already running in the bedroom (exclusive resource usage). In this case, the background music is determined to be an unnecessary action, and its execution is directly canceled to avoid auditory interference caused by the superposition of multiple audio sources.
[0067] In the embodiments of this application, when the conflict stems from excessively large parameter magnitudes (such as excessively bright light or loud sound), and the action falls into the category of "necessary but not urgent," execution is downgraded. During the downgraded execution process, a parameter correction operator is invoked to forcibly compress the execution parameters of the action to below a safe threshold. For example, at 2:00 AM, a user gets out of bed to go to the kitchen (triggering the "turn on the lights" intent), but the system detects that another member is sleeping in the bedroom (Wlight=0.9). The system downgrades the originally planned 80% brightness "full brightness mode" to a 5% brightness "low-light night light mode," satisfying the current user's obstacle avoidance needs while ensuring that the sleeping member is not awakened by the bright light.
[0068] In the embodiments of this application, the specific process involved in generating a sequence of atomic actions based on the logical dependencies between each legal atomic action can be found in [reference needed]. Figure 6 .
[0069] Figure 6 An exemplary flowchart illustrating the generation and output of atomic action sequences according to an embodiment of this application is shown.
[0070] like Figure 6 As shown, in step S610, the set of prerequisite states and the set of supply states corresponding to each legal atomic action are extracted. In step S620, it is determined whether the intersection of the set of prerequisite states corresponding to the i-th legal atomic action and the set of supply states corresponding to the j-th legal atomic action is empty. In response to the fact that the intersection of the set of prerequisite states corresponding to the i-th legal atomic action and the set of supply states corresponding to the j-th legal atomic action is empty, in step S630, it is determined that the i-th legal atomic action and the j-th legal atomic action are logically independent actions that do not restrict each other. In response to the fact that the intersection of the set of prerequisite states corresponding to the i-th legal atomic action and the set of supply states corresponding to the j-th legal atomic action is not empty, in step S640, it is determined that the i-th legal atomic action is a predecessor dependency node of the j-th legal atomic action, and a corresponding directed constraint edge is generated. All directed constraint edges are aggregated to form a dependency edge set. In step S650, an atomic action sequence is generated based on the dependency edge set.
[0071] In the embodiments of this application, the prerequisite state set consists of necessary but not sufficient conditions that the physical environment or device must meet before a legal atomic action can be performed. For example, the prerequisite states for the action "push ingredient reminders to a smartwatch" include: {device online: smartwatch, ambient light: 30 lux (ensuring the user can see), user location: kitchen}.
[0072] In the embodiments of this application, the supply state set represents the deterministic change to the physical environment or device state that occurs after a valid atomic action is successfully executed. For example, the supply state for the action "turn on the main kitchen light" includes: {ambient light: 500 lux, light fixture state: ON}.
[0073] In the embodiments of this application, the i-th legal atomic action A is determined. i For the j-th legal atomic action A j When generating a predecessor dependency node, a path is generated from A. i Pointing to A j Directed constrained edge e i,j All detected directed constraint edges are aggregated to obtain a set of dependent edges.
[0074] In the embodiments of this application, the following calculation formula is used in the process of generating atomic action sequences: Order=TopologicalSort(Edges,Priority(t)), where Order is the atomic action sequence, TopologicalSort is the weighted topological sorting operator, Edges is the set of dependent edges, and Priority(t) is the task priority weight.
[0075] By using the above calculation formula, the dependency edge set generated by the preceding steps is deeply integrated with the priority weight W of each action, thereby enabling the priority execution of safe tasks while ensuring logical correctness.
[0076] After step S130 is completed, in step S140, the atomic action sequence is distributed to the corresponding terminal device for execution, and the multi-source heterogeneous sensing data after execution is collected in real time to feed back to the step of generating a context snapshot containing visual, voice, environmental and device status features.
[0077] In the embodiments of this application, during the process of distributing atomic action sequences to corresponding terminal devices for execution, execution feedback and multi-source heterogeneous sensing data within the home are continuously monitored.
[0078] In the embodiments of this application, it is determined whether the currently executing atomic action has failed based on execution feedback. If the currently executing atomic action fails, a fault tolerance mechanism is triggered, and the physical channel is automatically switched for retry execution according to a preset alternative dictionary. If the currently executing atomic action does not fail, the next atomic action in the atomic action sequence continues execution.
[0079] Specifically, after issuing an atomic action command, a dynamic timer based on the command type is started. The execution status of the action is determined by receiving ACK response signals from the device, status query command feedback, or indirect feedback from associated sensors (such as current sensors and illuminance meters). If a command timeout occurs, the device returns an error code, or the expected state is inconsistent with the sensor's perceived state (e.g., a "turn on the light" command has been issued but the ambient illuminance has not increased), the current atomic action is considered to have failed. In response to the failure, the system immediately searches a preset dictionary. This dictionary defines functional peer mappings across physical entities. For example, if the original plan was to broadcast a voice reminder via a "kitchen smart speaker," but execution failed due to network fluctuations or device offline, the system searches the dictionary and finds that the alternative channels for the "voice broadcast function" in the current context are the "user's smartwatch" or the "living room TV screen." The system then modifies the physical identifier (Device ID) of the atomic action, redirects the control command to the alternative device, and initiates a retry to ensure that critical information is not lost.
[0080] In the embodiments of this application, the presence of a contextual mutation is determined based on multi-source heterogeneous sensing data within the home. In response to the presence of a contextual mutation, a reconstruction mechanism is triggered, discarding any remaining unexecuted atomic actions in the atomic action sequence and regenerating the atomic action sequence. In response to the absence of a contextual mutation, the remaining unexecuted atomic actions in the atomic action sequence continue to be executed.
[0081] Specifically, the system continuously retrieves multi-source heterogeneous sensor data streams within the home, processing them in real time, including data from human motion sensors (PIR), visual recognition streams (such as facial recognition), door and window magnetic status, device manual operation signals, and third-party API interface data (such as sudden weather changes or unexpected phone calls). A context drift calculation operator is introduced to compare the current instantaneous context vector with the baseline context vector during task planning. When the offset between the two exceeds a preset mutation threshold, a context mutation alarm is triggered. For example, during the "preparing dinner service" process, if it detects that "the user suddenly leaves the kitchen and enters the study to answer a phone call" or "a non-family member (visitor) is detected entering the target space," the system will immediately terminate the current linear execution process and initiate a reconfiguration mechanism once a context mutation is determined.
[0082] In summary, through the contextual reasoning and control scheme based on home intelligent agents provided above, this application embodiment constructs a full-link home intelligent control system from "environmental perception - intent understanding - action orchestration - closed-loop execution". First, by collecting multi-source heterogeneous sensing data to generate contextual snapshots, the information silos of single sensors or devices are broken down, achieving comprehensive and accurate capture of the complex physical environment of the home and the user's state. Second, by combining contextual snapshots with a family habit map to extract dual-layer intent tags, the system overcomes the limitations of traditional mechanical voice or touch commands, possessing the reasoning ability to uncover the user's deep, real needs, thereby providing highly personalized proactive services. Third, by decomposing macroscopic service goals into fine-grained atomic action sequences, the flexibility of cross-device collaborative orchestration and the orderliness of execution are ensured. Finally, a closed-loop mechanism for real-time data feedback after execution is introduced, enabling the system to dynamically refresh the contextual snapshot based on the actual physical effects of terminal execution. This closed-loop design of "perception-decision-execution-feedback" endows the home intelligent agent with strong adaptive correction and dynamic evolution capabilities, greatly improving the robustness of full-scenario services and the consistency of user experience.
[0083] Furthermore, in some embodiments, a deep intent recognition mechanism is constructed. First, a negative space feature vector is introduced. Traditional smart homes rely solely on currently occurring explicit states for judgment, while this solution, by comparing family habit maps, keenly captures missing elements in the current situation that should have occurred but did not (such as absence of personnel, interruption of habitual behavior, etc.). This mechanism of reasoning using reverse features greatly broadens the system's perception dimension, enabling it to gain insight into users' implicit needs and abnormal situations. Second, a complete temporal logical chain is established from tracing historical causes to predicting future trends. By retrieving antecedent causes such as time and events from the contextual knowledge base and combining them with the current snapshot and negative space features for joint matching, the system not only clarifies why the current state occurred but also accurately predicts the user's next action. Finally, the reasoning results are scientifically decoupled into surface and deep dual-layer intent labels. Surface intent addresses immediate direct needs, while deep intent, based on causes and predictions, directly addresses the user's fundamental purpose. This endows the home intelligent agent with a high degree of anthropomorphic empathy, forward-looking predictive ability, and the advantage of providing coherent, proactive, and in-depth butler-style services.
[0084] Furthermore, in some embodiments, a dynamic optimization and generation mechanism for service targets, guided by core user values and strictly controlled within security boundaries, is constructed. First, an absolutely feasible and secure candidate target solution space is built, eliminating the generation of contradictory instructions at the source and ensuring the physical security of hardware execution and system stability. Second, in the service target optimization phase, the complex intelligent decision-making process is transformed into a rigorous mathematical optimization problem, enabling precise evaluation of each candidate target set within the vast candidate target solution space. Finally, using an optimal calculation formula that includes utility satisfaction and a hard constraint masking function, not only can the service target set best matching the dual-layer intent label be selected, but it can also ensure that this target set maximizes user value experience while absolutely not violating family safety boundaries or causing unnecessary disturbance to the user. This endows the family intelligent agent with strong artificial intelligence decision-making capabilities for dynamic programming of globally optimal solutions under complex constraints.
[0085] Furthermore, in some embodiments, an ordered sequence of atomic actions is constructed. First, the macroscopic service target set is precisely quantified and mapped to a fine-grained set of atomic actions, and a conflict score calculation formula is introduced. This enables the pre-identification and interception of high-risk conflict operations, and employs a multi-dimensional flexible resolution strategy to effectively avoid hardware resource contention and physical deadlock caused by multi-task concurrency, greatly improving the security and stability of the system during concurrent processing. Second, by rigorously analyzing the prerequisite state set and supply state set of each legal atomic action, the causal constraints between actions are automatically derived, thereby constructing a rigorous set of logically dependent edges. Combined with a priority-based topological sorting algorithm formula, scattered actions can be arranged into an absolutely reasonable and ordered sequence of atomic actions. This eliminates logical paradoxes caused by disordered control and ensures the temporal continuity of complex cross-device collaborative actions.
[0086] This application also provides a situational reasoning and control system based on home intelligent agents. It can use the aforementioned situational reasoning and control method 100 based on home intelligent agents to perform situational reasoning and control based on home intelligent agents, or it can use other methods to perform situational reasoning and control based on home intelligent agents. This application does not limit this.
[0087] Figure 7 An exemplary structural block diagram of a home-based intelligent agent-based situational reasoning and control system according to an embodiment of this application is shown.
[0088] like Figure 7 As shown, the system 700 includes a data perception and snapshot generation module 710, an intent graph analysis module 720, a target decomposition and sequence generation module 730, and a distribution execution and feedback closed-loop module 740.
[0089] Specifically, the data sensing and snapshot generation module 710 is used to collect multi-source heterogeneous sensing data within the home and generate a contextual snapshot based on the multi-source heterogeneous sensing data.
[0090] Specifically, the intent graph analysis module 720 is used to obtain two-layer intent labels based on contextual snapshots and pre-established family habit graphs.
[0091] Specifically, the target decomposition and sequence generation module 730 is used to generate a service target set based on the two-layer intent label and decompose the service target set into atomic action sequences.
[0092] Specifically, the distribution execution and feedback closed-loop module 740 is used to distribute the atomic action sequence to the corresponding terminal device for execution, and collect multi-source heterogeneous perception data after execution in real time to feed back to the step of generating a context snapshot containing visual, voice, environmental and device status features.
[0093] When system 700 uses the aforementioned family agent-based situational reasoning and control method 100 to perform family agent-based situational reasoning and control, the aforementioned step S110 is executed through the data perception and snapshot generation module 710, the aforementioned step S120 is executed through the intent graph analysis module 720, the aforementioned step S130 is executed through the target decomposition and sequence generation module 730, and the aforementioned step S140 is executed through the distribution execution and feedback closed-loop module 740. The specific execution process can be found above and will not be repeated here.
[0094] While numerous embodiments of this application have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Many modifications, alterations, and alternatives will arise for those skilled in the art without departing from the spirit and intent of this application. It should be understood that various alternatives to the embodiments of this application described herein may be employed in the practice of this application. The appended claims are intended to define the scope of protection of this application and therefore cover equivalents or alternatives within the scope of these claims.
Claims
1. A situational reasoning and control method based on family intelligent agents, characterized in that, include: Collect multi-source heterogeneous sensing data within the home and generate contextual snapshots based on the multi-source heterogeneous sensing data; Based on the aforementioned contextual snapshot and the pre-established family habit map, a two-layer intent label is obtained; A service target set is generated based on the two-layer intent tags, and the service target set is decomposed into atomic action sequences; The atomic action sequence is distributed to the corresponding terminal device for execution, and multi-source heterogeneous sensing data after execution is collected in real time to feed back to the step of generating a context snapshot containing visual, voice, environmental and device status features.
2. The situational reasoning and control method based on family intelligent agents according to claim 1, characterized in that, The family habit graph includes family member nodes, spatiotemporal scene nodes, behavioral pattern edges, and collaborative relationship edges; In the process of obtaining two-layer intent labels based on the aforementioned contextual snapshot and the pre-established family habit map, the following steps are performed: The contextual snapshot is compared with a pre-established family habit map, and a negative space feature vector is generated based on the comparison results; Two-layer intent labels are obtained by matching contextual snapshots with negative space feature vectors; The comparison results are the expected elements that are missing in the current context. The expected elements that are missing in the current context include: missing members, missing collaborations, and missing states.
3. The situational reasoning and control method based on family intelligent agents according to claim 2, characterized in that, In the process of obtaining two-layer intent labels based on contextual snapshots and negative space feature vector matching, the following steps are performed: The context snapshot is used as a query vector to input into a pre-built context knowledge base, and historical contexts with similarity satisfying a preset threshold and the corresponding antecedents of the historical contexts are retrieved. The antecedents include time-triggered conditions, event-triggered conditions and state-triggered conditions. Generate trigger labels and corresponding confidence levels based on the aforementioned pre-existing triggers; The current context snapshot is concatenated with the negative space feature vector to obtain a joint query feature. Based on the joint query feature, the corresponding historical behavior trajectory is matched in the pre-built user behavior knowledge base. Matching subsequent behavior sequences based on historical behavior trajectories, and generating prediction results for the next behavior based on the subsequent behavior sequences; Based on the context snapshot, the trigger label, and the one-step behavior prediction result, a two-layer intent label is output. The dual-layer intent label includes a surface intent label and a deep intent label; The surface intent label is generated based on the context snapshot mapping; The deep intent label is generated based on the trigger label and the prediction result of the next behavior.
4. The situational reasoning and control method based on family intelligent agents according to claim 1, characterized in that, In the process of generating the service target set based on the two-layer intent tags, the following steps are performed: Based on user profiles or historical behavior assessments, determine the core value vector of users in the current context; Load a preset set of safety and interference thresholds, which defines the negative hard constraint boundary that cannot be triggered in the current situation; Construct a candidate target set solution space based on context snapshots; Based on the aforementioned two-layer intent tags, context snapshots, user core value vectors, and security and interference threshold sets, optimization calculations are performed in the candidate target set solution space to obtain the service target set. The following calculation formula is used in the process of performing optimization calculations in the solution space of the candidate target set to obtain the service target set: G represents the service target set, and g represents the candidate target set. Let be the solution space for the candidate target set. As the first dynamic adjustment factor, This is a two-layer intent label, where V is the user's core value vector and C is the context snapshot. To determine the utility satisfaction of the candidate target set g with respect to the user's core value vector under the constraint of context snapshot C, Let be the hard constraint masking function, and s be the set of security and interference thresholds.
5. The situational reasoning and control method based on family intelligent agents according to claim 4, characterized in that, In the process of constructing the solution space of the candidate target set based on context snapshots, the following steps are performed: The scenario snapshot is parsed, and the terminal devices that are currently online and not exclusively occupied within the household are traversed to generate an available device matrix; Query the preset device capability registry, extract the discrete control commands and continuous adjustment parameters supported by each device in the available device matrix, and use them as atomic operation units; The atomic operation units are combined along the time dimension and the device dimension to generate an initial set containing multiple alternative execution strategies; Eliminate candidate execution strategies from the initial set that have physical conflicts or logical mutual exclusions in device functions, and define the remaining set of candidate execution strategies as the candidate target set solution space.
6. The situational reasoning and control method based on family intelligent agents according to claim 1 or 4, characterized in that, In the process of decomposing the service target set into atomic action sequences, the following steps are performed: The service target set is mapped to an atomic action set according to a preset decomposition rule. The following calculation formula is used in the process of mapping the service target set to the atomic action set: A = G / U D and A represent the set of atomic actions, G represents the set of service targets, U represents the smallest unit of execution, and / represents the decomposition operator. Let D be the mapping operator, and D be the set of execution-side capabilities. By analyzing the matching degree between the set of execution resources required by each atomic action in the atomic action set and the set of currently occupied resources, conflict resolution is performed on the corresponding atomic actions to obtain multiple legal atomic actions; A sequence of atomic actions is generated based on the logical dependencies between each legal atomic action.
7. The situational reasoning and control method based on family intelligent agents according to claim 6, characterized in that, In the process of conflict resolution of the corresponding atomic actions, the following steps are performed: The conflict score of each atomic action within the set of atomic actions is calculated based on the conflict score calculation formula, whereby: Conflict=(R task ∩R current )×W, Conflict is the conflict score of atomic actions, R task R is the set of execution resources required for an atomic action. current Let W be the currently occupied resource set, and W be the interference weight matrix. Determine whether the conflict score is greater than a security threshold; If the conflict score is not greater than the security threshold, the corresponding atomic action is determined to be a legal atomic action, and its execution parameters are retained. In response to a conflict score greater than the security threshold, the corresponding atomic action is determined to be an illegal atomic action. At least one resolution strategy is triggered based on the attributes of the corresponding atomic action and the type of resource conflict, in order to eliminate the illegal atomic action or convert the illegal atomic action into a legal atomic action. The resolution strategies include interactive replacement, timing delay, action cancellation, and degraded execution.
8. The situational reasoning and control method based on family intelligent agents according to claim 6, characterized in that, In the process of generating a sequence of atomic actions based on the logical dependencies between each legal atomic action, the following steps are performed: Extract the set of prerequisite states and the set of supply states corresponding to each legal atomic action; Determine whether the intersection of the set of prerequisite states corresponding to the i-th legal atomic action and the set of supply states corresponding to the j-th legal atomic action is an empty set; In response to the fact that the intersection of the premise state set corresponding to the i-th legal atomic action and the supply state set corresponding to the j-th legal atomic action is an empty set, the i-th legal atomic action and the j-th legal atomic action are determined to be logically independent actions that do not restrict each other. In response to the fact that the intersection of the premise state set corresponding to the i-th legal atomic action and the supply state set corresponding to the j-th legal atomic action is not empty, the i-th legal atomic action is determined to be the predecessor dependency node of the j-th legal atomic action, and the corresponding directed constraint edge is generated. All the directed constraint edges are aggregated to form the dependency edge set. Generate atomic action sequences based on dependency edge sets; In the process of generating the atomic action sequence, the following calculation formula is used: Order=TopologicalSort(Edges,Priority(t)), where Order is the atomic action sequence, TopologicalSort is the weighted topological sorting operator, Edges is the set of dependent edges, and Priority(t) is the task priority weight.
9. The situational reasoning and control method based on family intelligent agents according to claim 6, characterized in that, During the process of distributing the atomic action sequence to the corresponding terminal device for execution, the execution feedback and multi-source heterogeneous sensing data within the home are continuously monitored. Among them, it is determined whether the currently executed atomic action has failed based on the execution feedback; In response to the failure of the currently executing atomic action, the fault tolerance mechanism is triggered, and the physical channel is automatically switched according to the preset alternative dictionary for retry execution; If the currently executing atomic action does not fail, continue executing the next atomic action in the sequence of atomic actions; Determine whether there are sudden changes in the context based on multi-source heterogeneous sensing data within the family; In response to a sudden change in the situation, a reconstruction mechanism is triggered, discarding any remaining atomic actions in the atomic action sequence that have not yet been executed, and regenerating the atomic action sequence. In response to the absence of a situational change, continue executing the remaining atomic actions in the atomic action sequence that have not yet been executed.
10. A situational reasoning and control system based on a family intelligent agent, characterized in that, The system employs the scenario reasoning and control method based on family agents as described in any one of claims 1-9, wherein the system comprises: The data sensing and snapshot generation module is used to collect multi-source heterogeneous sensing data within the home and generate contextual snapshots based on the multi-source heterogeneous sensing data. The intent graph analysis module is used to obtain two-layer intent labels based on the context snapshot and the pre-established family habit graph; The target decomposition and sequence generation module is used to generate a service target set based on the two-layer intent label, and decompose the service target set into atomic action sequences; The distribution execution and feedback closed-loop module is used to distribute the atomic action sequence to the corresponding terminal device for execution, and collect multi-source heterogeneous perception data after execution in real time to feed back to the step of generating a context snapshot containing visual, voice, environmental and device status features.