Multi-user-oriented smart home resource conflict negotiation and distribution method

By constructing causal models and counterfactual inferences, smart home systems achieve more accurate and fair resource allocation in multi-user scenarios, solving the problems of rigid traditional decision-making logic and data association errors, and dynamically adapting to changes in the home environment.

CN120979867APending Publication Date: 2025-11-18NANJING FORESTRY UNIV
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202511310365.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-15
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Existing smart home systems struggle to effectively handle device resource conflicts in multi-user scenarios. Their rigid traditional decision-making logic or reliance on data correlation leads to erroneous decisions, lacking fairness and dynamic adaptability.

Method used

We construct causal models to understand user behavior through causal relationships, generate candidate decision strategies, and achieve fair and dynamic resource allocation through counterfactual inference and Pareto optimal selection.

Benefits of technology

It enables decisions that better reflect the user's true intentions, improves the accuracy and fairness of resource allocation, and can dynamically adapt to changes in the home environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120979867A_ABST
    Figure CN120979867A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-user-oriented smart home resource conflict negotiation and allocation method, which comprises the following steps of: detecting equipment use conflicts caused by at least two users by constructing a structural causal model for representing a causal relationship of a plurality of variables in a smart home environment, generating a candidate decision strategy set comprising resource isolation and alternative compensation, and allocating the candidate decision strategy set to the smart home environment; carrying out anti-fact inference by utilizing a causal model, quantifying the causal effect of each strategy on the user state, and selecting an optimal strategy for execution based on a minimum negative effect and a Pareto optimal principle; besides, the method ensures that the system dynamically adapts to user habit changes through online monitoring of prediction errors and correction of the causal model, and compared with the prior art, the method improves the decision accuracy through causal reasoning, realizes fair and personalized user resource allocation, and improves the user experience and long-term effectiveness of the smart home system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of smart home, artificial intelligence and multi-user interaction technology, and in particular to a method for negotiating and allocating resource conflicts in smart homes for multiple users. Background Technology

[0002] With the deep integration of the Internet of Things (IoT) and artificial intelligence (AI) technologies, smart home systems have evolved from single-device control into a complex ecosystem involving multiple devices and users. In this development process, effectively handling conflicting control requests from multiple family members simultaneously for shared devices (such as ambient lighting, central air conditioning, and public audio systems) has become a key technological bottleneck for improving user experience and system intelligence. To address this challenge, existing technologies often employ arbitration mechanisms based on static priorities, such as pre-setting the homeowner's authority to be higher than other family members. While simple and direct, this decision-making logic is rigid and cannot adapt to the dynamically changing priorities within a home setting. Another approach is to predefine "cinema mode" or "guest mode" based on scene patterns to coordinate device states. While these solutions improve automation to some extent, their scene coverage is limited, making it difficult to handle sudden or personalized conflicts outside of pre-defined modes.

[0003] To overcome these limitations, current mainstream research is shifting towards applying machine learning techniques to build predictive models by analyzing massive amounts of historical interaction data. These methods can learn the correlation between user habits and environmental states, thus predicting the most likely user intent when a conflict occurs. However, these predictive models have a fundamental flaw: they cannot distinguish between "correlation" and "causation." When the real causal situation changes, these models are prone to making erroneous decisions that contradict the user's true intent due to reliance on spurious correlations. Furthermore, existing predictive models often employ utility maximization or voting mechanisms in their optimization objectives, which may lead to the long-term suppression of the reasonable needs of a minority of users by the conventional preferences of the majority, lacking consideration for decision-making fairness.

[0004] Therefore, there is an urgent need in this field for a resource conflict resolution method that can go beyond superficial data correlations, deeply understand the causal logic behind user behavior, and make forward-looking inferences about the potential consequences of different decisions. Summary of the Invention

[0005] The purpose of this section is to outline some aspects of embodiments of the present invention and to briefly describe some preferred embodiments. Simplifications or omissions may be made in this section, as well as in the abstract and title of this application, to avoid obscuring the purpose of these documents; however, such simplifications or omissions should not be construed as limiting the scope of the invention.

[0006] In view of the aforementioned existing problems, this invention is proposed. Therefore, this invention provides a method for negotiating and allocating resource conflicts in smart homes for multiple users, in order to solve the problems mentioned in the background art.

[0007] To address the aforementioned technical problems, this invention provides the following technical solution: a method for negotiating and allocating resource conflicts in a multi-user smart home system, comprising:

[0008] Construct a causal model that characterizes the causal relationships among multiple variables in a smart home environment;

[0009] When a conflict in the use of smart home device resources caused by at least two users is detected, a set of candidate decision strategies is generated.

[0010] Using the causal model, counterfactual inference is performed on each group of candidate decision strategies to determine the causal effect of the strategy on the state of at least one user.

[0011] Based on the causal effect, an optimal decision strategy is selected and executed from the candidate decision strategies.

[0012] As a preferred embodiment of the intelligent home resource conflict negotiation and allocation method for multiple users described in this invention, the causal model is a structural causal model, which includes a directed acyclic graph representing the causal relationship between the multiple variables, and a set of structural equations describing how each non-root node variable is determined by its direct cause variable.

[0013] As a preferred embodiment of the multi-user smart home resource conflict negotiation and allocation method described in this invention, the method further includes a step of maintaining the causal model online, which includes:

[0014] Monitor the prediction error distribution of the causal model, and when it is determined that the distribution has undergone a statistical change, relearn the local structure in the causal graph that is related to the error.

[0015] As a preferred embodiment of the intelligent home resource conflict negotiation and allocation method for multiple users described in this invention, the method for generating the set of candidate decision strategies includes, in addition to the control request issued when resource usage conflicts occur, the method also includes: using a pre-trained generative model based on the current conflict state to generate decision strategies for resource isolation or providing alternative compensation.

[0016] As a preferred embodiment of the multi-user smart home resource conflict negotiation and allocation method described in this invention, the counterfactual inference for each group of candidate decision strategies includes:

[0017] Based on the causal model and the current observation at the time of the conflict, the unmodeled random factors that caused the observation are inferred.

[0018] An intervention is applied to the causal model to force the state of the smart home device to be set to the value defined by the candidate decision strategy, resulting in an intervened causal model.

[0019] By combining the inferred unmodeled random factors with the applied intervention, the counterfactual prediction of the user's state is calculated, and the causal effect is determined.

[0020] As a preferred embodiment of the multi-user smart home resource conflict negotiation and allocation method described in this invention, the method for inferring the unmodeled random factors includes:

[0021] Using a pre-trained encoder neural network that takes the current observation as input, the output is a parameter describing the posterior probability distribution of the unmodeled random factors through a single forward propagation computation.

[0022] As a preferred embodiment of the multi-user smart home resource conflict negotiation and allocation method described in this invention, the method further includes, before determining the causal effect:

[0023] By using inverse reinforcement learning, a user well-being function is learned from the user's historical behavior data. This function is used to map the user's state to a scalarized well-being value.

[0024] The causal effect is obtained by quantifying the difference in well-being values ​​between the counterfactual predicted user state and the currently observed user state, based on the user well-being function.

[0025] As a preferred embodiment of the intelligent home resource conflict negotiation and allocation method for multiple users described in this invention, the selection of the optimal decision strategy is determined based on minimizing the negative causal effects that any set of candidate decision strategies may cause to any single user.

[0026] As a preferred embodiment of the multi-user smart home resource conflict negotiation and allocation method described in this invention, the step of selecting the optimal decision strategy includes:

[0027] For each set of candidate decision strategies, construct a decision vector containing multiple evaluation dimensions;

[0028] Determine a Pareto optimal policy set from all candidate decision policies, and select the optimal decision policy from this policy set.

[0029] As a preferred embodiment of the multi-user smart home resource conflict negotiation and allocation method described in this invention, the method further includes, after executing the optimal decision-making strategy:

[0030] Observe the actual user state after the decision is implemented and compare it with the determined counterfactual predicted user state to determine the prediction error;

[0031] When it is determined that there is a systematic bias in the prediction error, the causal model is corrected based on the bias.

[0032] Compared with existing technologies, the beneficial effects of the invention are:

[0033] 1. This invention, by constructing and utilizing a causal model for counterfactual inference, transcends the decision-making paradigm of traditional machine learning models based on data correlation. It enables the model to gain a deeper understanding of the real causal links behind user behavior, rather than merely learning the surface correlations of data. As a result, when faced with changes in the situation or new types of conflicts that have never been seen before, it can make more accurate decisions that are more in line with the user's true intentions.

[0034] 2. By introducing a fairness decision-making criterion based on minimizing negative causal effects and combining it with Pareto optimal sets for multi-objective optimization, the drawback of traditional methods that simply maximize overall utility may ignore individual feelings is changed. This effectively avoids the systematic and long-term suppression of the reasonable needs of a few users, and achieves a fairer and more humane resource allocation result under advanced negotiation intelligence.

[0035] 3. Furthermore, the causal model proposed in this invention has an online maintenance and self-correction mechanism. That is, by continuously monitoring the deviation between the actual effect after the decision is implemented and the counterfactual prediction, the smart home system can automatically identify and correct its understanding of the causal relationship of the home environment, thereby dynamically adapting to changes in the habits, preferences and even the family structure of family members, ensuring that the system's decision-making ability has long-term effectiveness. Attached Figure Description

[0036] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. Wherein:

[0037] Figure 1 This is a flowchart illustrating the overall process of a multi-user smart home resource conflict negotiation and allocation method according to an embodiment of the present invention. Detailed Implementation

[0038] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the protection scope of the present invention.

[0039] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.

[0040] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that is mutually exclusive with other embodiments.

[0041] This invention is described in detail with reference to the schematic diagrams. When detailing the embodiments of this invention, for ease of explanation, the cross-sectional views illustrating the device structure may be partially enlarged, not adhering to the usual scale. Furthermore, the schematic diagrams are merely examples and should not be construed as limiting the scope of protection of this invention. In actual fabrication, the three-dimensional spatial dimensions of length, width, and depth should be included.

[0042] Furthermore, in the description of this invention, it should be noted that the terms "upper," "lower," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. These terms are used solely for the convenience of describing the invention and for simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on the invention. In addition, the terms "first," "second," or "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.

[0043] Unless otherwise explicitly specified and limited, the terms "installation," "connection," and "joining" in this invention should be interpreted broadly. For example, they can refer to fixed connections, detachable connections, or integral connections; similarly, they can refer to mechanical connections, electrical connections, or direct connections, or indirect connections through an intermediate medium, or internal connections between two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.

[0044] Example 1

[0045] Reference Figure 1This is the first embodiment of the present invention, which provides a method for negotiating and allocating smart home resource conflicts for multiple users, including:

[0046] S1. Construct a causal model that represents the causal relationships between multiple variables in a smart home environment;

[0047] It should be noted that the goal of this step is to build a causal model that can accurately understand the internal workings of a smart home system, rather than simply fitting the surface correlations of the data.

[0048] Specifically, the causal model in this invention is preferably a structural causal model (SCM), whose construction process is divided into an offline initialization stage and an online adaptive maintenance stage.

[0049] Furthermore, for the offline initialization phase, a basic causal model is constructed using historical data accumulated during the initial deployment of the smart home system.

[0050] Specifically, the smart home system first collects heterogeneous, timestamped data from various deployed sensors and devices. These data sources include, but are not limited to:

[0051] Visual data, namely, the user's location, posture, and activity category captured by an indoor camera (output as a probability distribution by a pre-trained action recognition model).

[0052] Audio data, namely, voice commands collected by the microphone array, ambient noise levels, and emotion vectors output by the voice emotion recognition model;

[0053] Environmental sensor data, namely, continuous readings of temperature, humidity, illuminance, CO2 concentration, etc.

[0054] Device log data includes the on / off status, set parameters (such as light brightness and air conditioning temperature), and energy consumption of all smart devices.

[0055] User interaction data refers to explicit control commands issued by users through apps, voice assistants, and physical panels.

[0056] Furthermore, based on the collected data sources, a set of variables constituting the smart home system is defined. Each variable in this set is operationalized. For example, a user's "attention level" is no longer a vague concept, but is operationalized into a heart rate variability (HRV) index from wearable devices (such as smartwatches) and the gaze direction entropy value captured by the camera. This entropy value can be calculated by a fusion model, and the output of the calculation is a scalar in the range of [0,1].

[0057] Specifically, in one feasible implementation of the present invention, the fusion model can be a multilayer perceptron (MLP). Its input layer receives heart rate variability index and gaze direction entropy value, and performs feature crossover and nonlinear transformation through one or two hidden layers containing nonlinear activation functions (such as ReLU). Finally, it passes through an output layer with a sigmoid activation function to map the result to the [0,1] interval. Meanwhile, the training of this model can be completed in a supervised learning manner by collecting a small-scale dataset with user subjective attention annotations (for example, by periodically asking the user "Are you currently focused?" in a wearable device and collecting the feedback results). Its loss function can be the mean squared error (MSE) or cross-entropy loss.

[0058] Furthermore, in order to learn the causal structure among multiple variables, the present invention constructs a Directed Acyclic Graph (DAG), denoted as […]. Where the node set V is the set of variables defined above, and the directed edge set E represents the direct causal relationship between the variables;

[0059] It should be noted that the learning of this DAG can be carried out offline. It utilizes historical time-series data accumulated from the long-term operation of the smart home system and employs advanced causal discovery algorithms (such as the NOTEARS algorithm based on continuous optimization or the PC algorithm based on constraints) for structure learning. The most likely causal graph is inferred by analyzing the conditional independence in the data.

[0060] Specifically, in one feasible implementation of the present invention, the NOTEARS algorithm is used to transform the combinatorial search space of the DAG into a differentiable continuous optimization problem, and gradient descent is used to find the optimal weighted adjacency matrix. The core idea is to establish the necessary and sufficient conditions for the establishment of the DAG: ,in, Let d represent the Hadamard product, where d is the number of variables in the directed acyclic graph. This condition is added as a penalty term to the loss function of gradient descent, thereby efficiently learning the causal structure among large-scale variables.

[0061] Furthermore, after constructing the DAG, for each non-root node... By fitting its structural equation, we obtain:

[0062]

[0063] It should be noted that, considering the complex nonlinear relationships in smart homes, Parameterization can be performed using deep neural networks, for example, for variables that depend on time-series inputs (such as predicting the indoor temperature at the next moment). It can be a recurrent neural network (RNN) or a long short-term memory network (LSTM), while for variables with static relationships, then... It is a multilayer perceptron (MLP); and, in the present invention, it is not assumed that... It is not simple Gaussian noise, but rather a probability distribution is learned for it;

[0064] Preferably, we employ the concept of a variational autoencoder (VAE), where for each structural equation, we simultaneously train an encoder network. The encoder network transmits the observations (i.e., ) and its parent node As input, the output describes the noise. posterior distribution The encoder network can quickly infer "unmodeled causes" (i.e., noise) from "effects" and "partial causes" when making causal inferences, using parameters such as the mean and variance of a Gaussian distribution. (Specific implementation method)

[0065] Furthermore, for the online adaptive maintenance phase, although a basic causal model has been established through the aforementioned offline initialization phase, since the actual home environment is dynamically changing, in order to ensure the long-term effectiveness of the model, the present invention also includes an online maintenance and self-correction mechanism to ensure that the model can continuously evolve with changes in the home environment and members' habits.

[0066] Specifically, during the operation of the smart home system, a baseline window (e.g., the past week) of residual distribution is maintained by continuously calculating the residual between the predicted value and the actual observed value of the variable in each causal model. At the same time, the residual distribution of a sliding window (e.g., the past hour) is compared with the residual distribution of the maintained baseline window.

[0067] Preferably, this comparison process uses a non-parametric two-sample Kolmogorov-Smirnov test (KS test) to determine whether there is a significant difference between the two distributions. The smart home system only triggers a causal model update operation when the p-value of the KS test is lower than a preset threshold α (e.g., 0.05). For example, assuming that the residuals follow a normal distribution, the mean of the baseline window is μ=0, and the variance is σ²=1, then the p-value calculated by the sliding window is 0.03<0.05, which triggers an update.

[0068] Specifically, when an update is triggered, the smart home system does not blindly retrain the entire model. Instead, it locates the variable that produces the maximum distribution drift. Then, within its Markov Blanket (i.e., its parent nodes, child nodes, and other parent nodes of the child nodes), the system performs a local, Bayesian-scored causal structure search. During this search, the system attempts to add, delete, or reverse edges in the local subgraph and calculates the Bayesian Information Criterion (BIC) score for each new structure. Since the BIC score considers both the goodness of fit of the causal model to the new data and the complexity of the causal model, the system selects the new local structure that yields the optimal BIC score to replace the old structure of the causal model. For example, for a local subgraph containing 5 nodes, the search space can be implemented using enumeration (computational complexity O(log n)). The BIC formula can be: Where L is the likelihood function, k is the number of parameters, and n is the number of samples. Training uses the Adam optimizer with a learning rate of 0.001 and a batch size of 64. When the structure of the causal model is adjusted, only the structural equations related to that adjustment are considered. It will be retrained or fine-tuned, rather than completely retrained;

[0069] It should be noted that this retraining or fine-tuning process uses the latest data buffer and adapts quickly to emerging causal patterns with low computational cost.

[0070] S2. When a conflict in the use of smart home device resources caused by at least two users is detected, a set of candidate decision strategies is generated.

[0071] It should be noted that the core task of this step is to quickly generate a rich and high-quality decision space when a user-initiated conflict over the use of smart home device resources is detected. This provides a basis for subsequent counterfactual inferences and optimal choices. Furthermore, this process is not limited to satisfying the direct requests of both parties in the conflict, but also aims to explore solutions that can improve the overall well-being of users.

[0072] Furthermore, the smart home system uses an instruction intent buffer (all user control instructions (regardless of the interface through which they are issued) to enter this buffer before execution) to monitor all user control instructions to shared home devices in real time. The system parses the instructions in the buffer in real time to extract the target home device, control parameters, expected values, and initiating user. When multiple instructions for the same home device and the same parameter appear in the buffer, and the expected value parsed by the instruction meets the conflict determination rules, a conflict is triggered. The rules specifically include a numerical parameter conflict rule and a categorical parameter conflict rule.

[0073] Specifically, numerical parameter conflict rules refer to situations where the expected values ​​of the instructions are in opposite directions (i.e., one increases and the other decreases), or the absolute value of the difference exceeds a preset sensitivity threshold.

[0074] Specifically, categorical parameter conflict rules refer to instructions whose expected values ​​are mutually exclusive categories (such as "cooling" and "heating" for air conditioning equipment).

[0075] Furthermore, as soon as a conflict is triggered, the smart home system immediately captures a snapshot of the current panoramic state, that is, the current values ​​of all variables in the causal model in step S1. Subsequently, the system constructs a high-dimensional conflict state input tensor, which contains the following parts:

[0076] The first part is a core description of the conflict, which includes the conflicting device ID, parameter name, ID of each conflicting user and their specific request value;

[0077] The second part is user context embedding, that is, for each conflicting user, its relevant state variables (such as activity, location, emotion, physiological indicators) are extracted and mapped into dense vector representations through a pre-trained embedding layer;

[0078] The third part is the environment and device context, that is, all other variables (environment readings, other device states) in the panoramic state snapshot, except for the user state, are also vectorized;

[0079] The fourth part is the device capability coding, which encodes the functional attributes of conflicting devices and other potentially available devices in the environment. For example, a multi-zone light strip will be encoded as an entity that supports multiple independent control points such as zone 1 and zone 2.

[0080] It should be noted that, in order to ensure the diversity of the strategy set, the present invention also adopts a hybrid strategy generation mechanism, which combines deterministic strategies and generative strategies, wherein the two strategies are executed in parallel by the system's strategy generation decision engine.

[0081] Specifically, for deterministic strategies, the original requests from both sides of the conflict are directly used as candidate decision strategies. In particular, for numerical parameters, such as continuously adjustable parameters (e.g., brightness, temperature, volume), a series of compromise values ​​are generated. These compromise values ​​can be the mean, median, and weighted average (where the weights can be dynamically adjusted based on historical user priorities). Furthermore, a small rule base can be embedded in the system to match specific conflict patterns. For example, if the conflict is about audio playback, a time-division multiplexing strategy is generated, i.e., "play user A's song for 10 minutes first, then play user B's song for 10 minutes."

[0082] It should be noted that for deterministic strategies, there is also a no-operation, which is to maintain the equipment state before the conflict occurred as a reference benchmark for evaluating the effectiveness of all equipment's proactive intervention.

[0083] Specifically, for generative strategies, this invention employs a pre-trained generative model conditioned on the current conflict state input tensor to accomplish this task. This pre-trained generative model is a conditional sequence generation model based on the Transformer architecture, which consists of an encoder and an autoregressive decoder.

[0084] Specifically, the encoder is responsible for receiving and deeply understanding the conflict state input tensor mentioned above, and compressing it into a hidden state representation with contextual information; while the decoder, based on this hidden state representation, generates a symbol sequence describing the decision-making strategy token by token.

[0085] Specifically, to make the symbol sequence parseable, this invention also defines a microformat language. The vocabulary of this language includes the IDs of all controllable devices, the names of all controllable parameters, special operators (such as ":=", ", ";) and a set of discretized parameter values. A complete symbol sequence consists of one or more control clauses separated by semicolons. The format of each clause is: parameter name:= parameter value.

[0086] In addition, to enable the pre-trained generative model to generate structured, executable complex policies, we also defined two policies: a resource isolation policy and an alternative compensation policy.

[0087] Specifically, the resource isolation strategy enables the pre-trained generative model to understand the functional decomposition of the device. For example, for a smart light strip that supports zoned control, when user A wants high brightness in the reading area and user B wants a dimly lit viewing atmosphere in the sofa area, the pre-trained generative model will not simply choose an intermediate brightness, but will output a strategy:

[0088] P1:{LivingRoom_Light_Zone1_Brightness:=100%, LivingRoom_Light_Zone2_Brightness:=10%};

[0089] This strategy means setting the brightness of the light strip in the current reading area to 100% and the brightness of the light strip in the current sofa area to 10%.

[0090] Specifically, alternative compensation strategies enable pre-trained generative models to consider solutions across devices and domains. When directly fulfilling one party's request would have a significant negative impact on the other party, the generative model will attempt to provide compensation. For example, user A wants to set the air conditioner temperature very low, but user B (possibly an elderly person) is very sensitive to this. After analyzing the user's physiological indicators contained in the conflict state input tensor, the generative model generates the following strategy:

[0091] P2:{AC_Temperature:=25℃,User_B_SmartBlanket:='ON',User_A_SmartFan:='ON'};

[0092] The strategy involves setting the air conditioner to 25°C, turning on the smart electric blanket for user B, and turning on the smart fan for user A. It is worth noting that this strategy does not fully meet user A's request for low temperature, but it provides a localized cooling solution by turning on the smart fan, while turning on user B's smart electric blanket to compensate for the discomfort that the slightly lower room temperature may cause.

[0093] It should be noted that, through the above strategy, a single conflicting resource can be decomposed into multiple independent sub-resources, thereby simultaneously satisfying the needs of both parties.

[0094] Furthermore, for the training dataset of this generative model, it is only necessary to adapt a resource conflict scenario simulator to generate massive and diverse synthetic conflict data. Then, for each simulated conflict, an expert system can be used to label one or more high-quality solutions. At the same time, the generative model will take the conflict state input tensor as input and learn to predict the corresponding expert solution sequence. Moreover, the generative model adopts the standard cross-entropy loss function during the training process to maximize the probability of the model generating the correct token.

[0095] Specifically, in one feasible implementation of the present invention, an expert system is used to label one or more high-quality solutions. The working principle of the resource conflict scenario simulator includes: based on multiple preset virtual user profiles (e.g., including parameters such as preferences, work and rest schedules, and physiological sensitivity), a series of behavioral goals (e.g., "start work" or "prepare to rest") are generated for each virtual user on the simulation timeline. These behavioral goals are then mapped to a specific sequence of resource requests to smart home devices. Finally, structured conflict scenario data is generated by detecting and outputting resource requests that overlap in time and have mutually exclusive parameters. The expert system embeds a set of heuristic rules for labeling solutions. This set of rules is arranged by priority, such as: safety and health rules (prioritizing the elderly's temperature needs), fairness rules (avoiding extreme negative values ​​for the welfare of either party), resource optimization rules (prioritizing lower energy consumption solutions), and innovation rules (exploring the possibility of alternative compensation). By applying this heuristic rule set, one or more high-quality labeling strategies can be automatically generated for each conflict scenario generated by the simulator.

[0096] Furthermore, when a conflict occurs, the conflict state input tensor is fed into the trained model encoder. In the decoding stage, instead of using the traditional greedy search, we use a bundle search or kernel sampling method. The advantage of this is that the generative model can explore multiple possibilities at each generation step and finally output a list of k (e.g., k=5) high-probability and distinct policies.

[0097] Specifically, by performing deduplication and feasibility filtering on the above-mentioned hybrid strategy generation mechanism (e.g., checking whether the strategy violates the physical limitations of the device), a set of candidate decision strategies can be obtained.

[0098] S3. Using a causal model, counterfactual inference is performed on each group of candidate decision strategies to determine the causal effect of the strategy on the state of at least one user.

[0099] It should be noted that this step mainly utilizes the causal model built in step S1 through the smart home system to "imagine" the potential consequences that each set of candidate decision strategies generated in step S2 will lead to once it is executed. This is the counterfactual inference process, which is executed independently for each strategy in the candidate decision strategy set.

[0100] Furthermore, using the constructed causal model and real observation data at the time of conflict, the background factors that cause all unobservable factors are inferred in reverse. In the context of the structural causal model, these background factors are modeled as exogenous noise variables, and the posterior estimate of the specific realization value of the exogenous noise variable at time t is output.

[0101] It should be noted that the posterior estimate can be understood as the set of all particular features at the current moment, such as user A being particularly tired today (which the sensor did not directly detect), or the cold wind blowing in from the window. If these particular feature sets are not attributed, then any prediction will be based on the average situation, which cannot explain the individual differences in specific conflict scenarios, thus leading to the failure of counterfactual inference.

[0102] Furthermore, since the calculated posterior distribution is usually difficult to process, this invention employs an efficient variational inference scheme, namely, approximating the posterior distribution through a dedicated "context encoder" neural network.

[0103] Specifically, the context encoder is trained together with the structural equation of the causal model to form a causal variational autoencoder, whose input is the entire observation vector (including all sensor readings, device status, user behavior, etc.).

[0104] Specifically, in one feasible implementation of the present invention, the input layer dimension of the neural network is the size of the observation vector, and each of the two hidden layers has 256 neurons activated by ReLU. The neural network uses the ELBO loss function, the Adam optimizer, and the learning rate is 0.001.

[0105] Specifically, when a conflict occurs, the actual observation data is fed into the causal variational autoencoder. Each time the neural network passes through a forward propagation operation, it outputs parameters describing the posterior distribution of each exogenous noise. For example, if the posterior distribution is assumed to be Gaussian, the neural network will output the mean and log-variance of each exogenous noise. At the same time, by sampling from these parameterized posterior distributions, a specific noise vector can be obtained. This noise vector is the smart home system's quantitative understanding of the "hidden background" of the current world.

[0106] Furthermore, for each set of candidate decision strategies generated in step S2, we create a causal model after intervention.

[0107] It's important to note that creating a post-intervention causal model is called a "do" operation in causal science. For example, for a candidate decision strategy = {set the air conditioner temperature to 22℃}, do(AC_Temp=22) means forcibly setting the air conditioner temperature to 22℃ and cutting off all factors that might have previously affected the air conditioner temperature (such as user voice commands, automatic temperature control logic, etc.). In a causal model, this do operation signifies a minor, procedural modification to the model, with the following steps:

[0108] Locate the target variable and identify the equipment variables directly manipulated in the candidate decision-making strategies;

[0109] Replace the structural equation by removing the original structural equation of the variable from the original causal model;

[0110] In the post-intervention causal model, an assignment statement is used instead of an intervention.

[0111] Furthermore, by constructing the causal model after intervention and the noise vector obtained above, we can calculate how the user's state will evolve under the current candidate decision-making strategy.

[0112] Specifically, the calculation process is as follows:

[0113] All root node variables The value is set to its observed value in the observation vector;

[0114] Following the topological sorting of the causal graph, calculate the counterfactual value of each endogenous variable one by one from parent node to child node, and for any root node variable... Its counterfactual value The calculation formula can be expressed as: ,in, It is the counterfactual value that its parent node has already calculated. These are the corresponding components in the inferred noise vector;

[0115] It should be noted that this calculation process continues until all leaf nodes related to the user state have been calculated, and finally a complete set of user state counterfactual prediction vectors can be obtained.

[0116] Furthermore, in order to make the decision-making of candidate strategies based on evidence, it is necessary to transform the multidimensional user state vector (such as {comfort: 0.8, focus: 0.3}) into a single, comparable scalar—wellness value. This invention learns a personalized wellness function for each user k through inverse reinforcement learning. Inverse reinforcement learning can infer the intrinsic reward function that best explains the choice by observing the user's actual choices in thousands of historical situations. That is, the wellness function, which represents the user's unspoken preferences.

[0117] Specifically, in one feasible implementation of the present invention, maximum entropy inverse reinforcement learning is employed. Within this framework, the state of the smart home system (i.e., the vector of some or all observable variables defined in step S1) is defined as state s in a Markov decision process (MDP). The device control operation (or inaction) actually performed by the user in any state is regarded as the action a chosen by the user. Then, the user's historical interaction data constitutes a series of expert trajectories. The goal of this reinforcement learning algorithm is to find the reward function (i.e., the welfare function) with the largest entropy among all reward functions (i.e., welfare functions) that can explain the expert trajectory. This welfare function is usually parameterized as a linear or nonlinear function of state features, whose parameters are learned through an optimization algorithm (gradient descent) to maximize the probability of the expert trajectory under this policy.

[0118] Specifically, the net causal effect of candidate decision-making strategies on user k. Defined as:

[0119] in, It is the predicted observed state of user k when the conflict occurs. It is the actual observed state of user k at the time of the conflict. As a personalized well-being function, the difference can be perfectly quantified by calculating: "If this candidate decision strategy is adopted, how much better (or worse) will user k's subjective feeling be than it is now";

[0120] It should be noted that after performing the above operations on all candidate decision strategies, the final output is a causal effect matrix. Each row of the causal effect matrix corresponds to a candidate decision strategy, and each column corresponds to a user. Each element in the matrix is ​​a net causal effect, which accurately and interpretably reveals the potential impact of the decision on the user's well-being.

[0121] S4. Based on causal effects, select and execute an optimal decision strategy from the candidate decision strategies;

[0122] It should be noted that this step aims to transform the causal effect matrix obtained above into an executable device control instruction that is most equitable for all users.

[0123] Furthermore, a corresponding decision vector is established for each candidate decision strategy, and the dimension of the decision vector consists of the net causal effect of all relevant users;

[0124] In addition, to achieve more comprehensive decision-making, the decision vector can also include other non-user well-being assessment dimensions, such as the additional energy consumption expected by the candidate decision strategy, the potential damage to equipment (such as frequent start-stop of the compressor) caused by the candidate decision strategy, and the complexity of the candidate decision strategy or the time required to execute it.

[0125] Specifically, in this way, each candidate decision strategy is mapped to a point in the multi-objective optimization space;

[0126] It should be noted that traditional decision-making methods (in this invention, maximizing total well-being) may severely harm the interests of an individual for the sake of a small overall improvement, thus leading to the problem that "majority approval" is greater than "minority approval". To avoid this problem, this invention introduces the Pareto Optimality principle for preliminary screening.

[0127] Specifically, in the set of candidate decision strategies, if there are no other candidate decision strategies that can improve the well-being of at least one user without harming the well-being of any other user, then that candidate decision strategy is Pareto optimal.

[0128] Specifically, the screening process is as follows:

[0129] By traversing all candidate decision strategy pairs, assuming that... If candidate decision strategies The causal effect is greater than or equal to the candidate decision strategy across all users. And is strictly greater than in at least one dimension Then it is called Dominate That is, all strategies dominated by other candidate strategies are removed from the candidate strategy set, and the remaining candidate strategies constitute a Pareto optimal set. Also known as the Pareto front;

[0130] It should be noted that all candidate decision strategies in this optimal set are efficient, and there is no absolute superiority or inferiority among them. The improvement of any candidate decision strategy will inevitably come at the cost of sacrificing a certain dimension of another candidate decision strategy. Through this screening process, the decision range is greatly narrowed, and it is ensured that the final selected candidate decision strategy will not be a "bad" strategy with obvious room for improvement.

[0131] Furthermore, since the Pareto optimal set may contain multiple decision strategies, a decision criterion is needed to select one of the strategies. Based on this, the present invention adopts a fairness principle based on Rawlsianism, namely, maximizing the minimum gain. Its goal is to prioritize the worst-performing individuals, that is, to minimize the negative causal effects that any single user may cause.

[0132] Furthermore, for each candidate decision strategy in the Pareto optimal set, calculate its minimum causal effect among all users. :

[0133] Here, the minimum causal effect represents the effect of implementing the candidate decision strategy. The changes in the well-being of the "least satisfied" user; and the optimal decision-making strategy. That is the strategy that maximizes the minimum causal effect:

[0134]

[0135] Furthermore, once the optimal decision-making strategy is determined, the smart home system's decision engine will parse it into a set of underlying device control commands (e.g., via the MQTT protocol or HTTP API) and send them to the corresponding smart home devices for execution.

[0136] Specifically, within a short time window after the strategy is implemented (e.g., 5 minutes), the system will continuously monitor the actual status of all users and calculate the actual user well-being after the decision is implemented. This actual well-being is then compared with the counterfactual well-being predicted by the optimal decision-making strategy. By comparison, the counterfactual prediction error is obtained. :

[0137]

[0138] In addition, the system will continuously track the distribution of this prediction error. If a systematic bias is found in the prediction error (for example, for user A, the mean of the prediction error continues to deviate significantly from the normal value), it indicates that there are erroneous or outdated causal assumptions in the causal model (for example, the model underestimates the negative impact of air conditioner noise on user A's concentration). When such a systematic bias is detected, the system will trigger the online maintenance mechanism in step S1.

[0139] It should be noted that through this feedback mechanism, the method of the present invention can not only make immediate decisions, but also learn and evolve from the consequences of each decision, so as to continuously deepen the understanding of the causal relationship of the family environment, thereby achieving sustainable intelligence.

[0140] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. The solutions in the embodiments of this application can be implemented using various computer languages, such as the object-oriented programming language Java and the interpreted scripting language JavaScript.

[0141] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0142] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0143] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0144] Although preferred embodiments of this application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of this application.

[0145] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.

Claims

1. A method for negotiating and allocating resource conflicts in a multi-user smart home system, characterized in that: include: Construct a causal model that characterizes the causal relationships among multiple variables in a smart home environment; When a conflict in the use of smart home device resources caused by at least two users is detected, a set of candidate decision strategies is generated. Using the causal model, counterfactual inference is performed on each group of candidate decision strategies to determine the causal effect of the strategy on the state of at least one user. Based on the causal effect, an optimal decision strategy is selected and executed from the candidate decision strategies.

2. The method for negotiating and allocating smart home resource conflicts for multiple users as described in claim 1, characterized in that, The causal model is a structural causal model, which includes a directed acyclic graph representing the causal relationships among the multiple variables, and a set of structural equations describing how each non-root variable is determined by its direct cause variable.

3. The method for negotiating and allocating smart home resource conflicts for multiple users as described in claim 2, characterized in that, The method further includes a step of maintaining the causal model online, which includes: Monitor the prediction error distribution of the causal model, and when it is determined that the distribution has undergone a statistical change, relearn the local structure in the causal graph that is related to the error.

4. The method for negotiating and allocating smart home resource conflicts for multiple users as described in claim 1, characterized in that, The method for generating the set of candidate decision strategies, in addition to including the control request issued when there is a resource usage conflict, also includes: using a pre-trained generative model conditioned on the current conflict state to generate decision strategies for resource isolation or providing alternative compensation.

5. The method for negotiating and allocating smart home resource conflicts for multiple users as described in claim 1 or 2, characterized in that, The counterfactual inference for each group of candidate decision strategies includes: Based on the causal model and the current observation at the time of the conflict, the unmodeled random factors that caused the observation are inferred. An intervention is applied to the causal model to force the state of the smart home device to be set to the value defined by the candidate decision strategy, resulting in an intervened causal model. By combining the inferred unmodeled random factors with the applied intervention, the counterfactual prediction of the user's state is calculated, and the causal effect is determined.

6. The method for negotiating and allocating smart home resource conflicts for multiple users as described in claim 5, characterized in that, Methods for inferring the unmodeled random factors include: Using a pre-trained encoder neural network that takes the current observation as input, the output is a parameter describing the posterior probability distribution of the unmodeled random factors through a single forward propagation computation.

7. The method for negotiating and allocating smart home resource conflicts for multiple users as described in claim 5, characterized in that, Before determining the causal effect, the following is also included: By using inverse reinforcement learning, a user well-being function is learned from the user's historical behavior data. This function is used to map the user's state to a scalarized well-being value. The causal effect is obtained by quantifying the difference in well-being values ​​between the counterfactual predicted user state and the currently observed user state, based on the user well-being function.

8. The method for negotiating and allocating smart home resource conflicts for multiple users as described in claim 1, characterized in that, The selection of the optimal decision-making strategy is based on minimizing the negative causal effects that any set of candidate decision-making strategies may have on any single user.

9. The method for negotiating and allocating smart home resource conflicts for multiple users as described in claim 1 or 8, characterized in that, The steps for selecting the optimal decision-making strategy include: For each set of candidate decision strategies, construct a decision vector containing multiple evaluation dimensions; Determine a Pareto optimal policy set from all candidate decision policies, and select the optimal decision policy from this policy set.

10. The method for negotiating and allocating smart home resource conflicts for multiple users as described in claim 1, characterized in that, After executing the optimal decision-making strategy, the following is also included: Observe the actual user state after the decision is implemented and compare it with the determined counterfactual predicted user state to determine the prediction error; When it is determined that there is a systematic bias in the prediction error, the causal model is corrected based on the bias.

Citation Information

Cited By

  • Interaction quantity causal contribution prediction method and device, and storage medium

    CN121685017A

  • Household scene control method, ternary architecture system and electronic equipment

    CN121918437A

  • A home scene control method, a ternary architecture system and an electronic device

    CN121918437B