Smart home scenarized audio linkage control system based on multi-mode perception

By integrating the conditional probability model of device electromagnetic radiation signals and network traffic data, a behavioral inconsistency indication signal is generated, which solves the problem of insufficient response accuracy in smart home audio linkage control and achieves higher robustness and accuracy.

CN121523081AActive Publication Date: 2026-02-13GUANGZHOU BLUE LIGHT ELECTRONICSTECH CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202511847363.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-09
Publication Date
2026-02-13
Estimated Expiration
2045-12-09

AI Technical Summary

Technical Problem

Existing technologies for audio linkage control in smart homes lack consideration of the macro-environmental context of device events, resulting in insufficient response accuracy, especially limited robustness when handling ambiguous commands.

Method used

By acquiring device electromagnetic radiation signals and network traffic metadata, and using a conditional probability model to fuse the device's instantaneous operating status and the macroscopic activity patterns of the environment, behavioral inconsistency indication signals are generated, enabling proactive early warning and precise interactive decision-making.

Benefits of technology

It improves the response accuracy and robustness of audio linkage control, reduces the false alarm rate, and can make the operation selection with the highest causal consistency under ambiguous user commands.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121523081A_ABST
    Figure CN121523081A_ABST
Patent Text Reader

Abstract

The invention provides a smart home scenarized audio linkage control system based on multi-modal perception, relates to the technical field of electric digital data processing, and realizes improvement of smart home audio linkage control capability by deeply fusing a microscopic equipment event and a macroscopic environment mode and making a decision based on causal consistency between the microscopic equipment event and the macroscopic environment mode. By constructing a multi-modal fusion decision-making mechanism based on a conditional probability model, an indication signal for quantitatively representing'behavior-scene 'consistency can be generated. The rationality of the equipment behavior is continuously evaluated, so that the accuracy of identifying the abnormity of the real equipment is improved, and the false alarm and the missing alarm are reduced; in the man-machine interaction process, the ability of the system to understand and analyze the fuzzy instruction is enhanced, so that the control decision is more in line with the real intention of the user in a specific scene, and the interaction robustness is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of electronic digital data processing technology, specifically to a smart home scene-based audio linkage control system based on multimodal perception. Background Technology

[0002] With the continuous evolution of smart home technology, audio interaction has become a core hub for communication between people and their environment. Current technological development is shifting from passive responses that execute explicit voice commands to intelligent systems capable of understanding deeper scenarios and providing proactive services. This requires control systems not only to interpret the literal meaning of commands but also to possess the ability to perceive and understand the macroscopic state of the environment, thereby enabling audio-linked control to seamlessly integrate into the user's daily rhythms. However, existing technologies still face some challenges in achieving highly intelligent audio-linked control.

[0003] 1. Existing technologies analyze instantaneous device events in isolation when making decisions, lacking consideration of the macro-environmental context in which these events occur. This approach makes it difficult for the system to effectively distinguish between normal household activities and genuine device malfunctions, resulting in issues with the accuracy of proactive early warning services.

[0004] 2. There are also certain limitations at the decision-making logic level of audio interaction. For example, the technical solution with publication number CN107357547A, while providing an effective way to dynamically transmit and adjust audio parameters while the audio device is in operation, focuses on the data transmission implementation method and does not address how to generate more intelligent and reliable control decisions based on multi-dimensional environmental information. Especially when processing ambiguous user commands, the lack of a deep understanding of the current scenario limits its robustness in interpreting the user's true intent. Summary of the Invention

[0005] The purpose of this invention is to provide a smart home scenario-based audio linkage control system based on multimodal perception to solve the problems mentioned in the background art.

[0006] To achieve the above objectives, the present invention provides the following technical solution: A smart home scenario-based audio linkage control system based on multimodal perception, specifically including: The first data acquisition module is configured to acquire first device event data characterizing the instantaneous operating state of at least one electronic device, the first device event data being generated based on the analysis of electromagnetic radiation signals generated by the electronic device; The second data acquisition module is configured to acquire second macro-state data that characterizes the macro-activity pattern of the environment in which the electronic device is located. The second macro-state data is generated based on long-term statistical analysis of network traffic metadata generated by multiple electronic devices in the environment. The fusion decision module is configured to: based on a preset conditional probability model, fuse the first device event data and the second macroscopic state data to calculate the conditional probability of the instantaneous operating state occurring under the macroscopic activity mode; The indicator signal generation module is configured to: generate a behavioral inconsistency indicator signal based on the conditional probability to characterize the degree of causal consistency between the instantaneous operating state and the macroscopic activity pattern; The warning generation module is configured to generate an active warning signal when the value of the behavior inconsistency indication signal exceeds a preset warning threshold, provided that the first device event data is not triggered by a user instruction. The interactive decision module is configured to: upon receiving a user's ambiguous voice command, generate a control command for parsing the ambiguous voice command based on the behavior inconsistency indication signal, so as to prioritize the candidate operation with the highest causal consistency with the macroscopic activity pattern of the environment.

[0007] Compared with existing technologies, the beneficial effects of this invention are as follows: By proposing the fusion of two distinct types of data: one is primary device event data, obtained through high-frequency analysis of device electromagnetic radiation signals, which accurately characterizes the instantaneous operating state of the device; the other is secondary macroscopic state data, abstracted through long-term statistical analysis of network traffic metadata in the environment, which stably characterizes the rhythms of family life. By introducing a conditional probability model, the rationality of a specific instantaneous device event occurring under a particular macroscopic activity pattern can be quantitatively calculated. This calculation result is shaped into a "behavioral inconsistency indication signal," and the generation of this core indication signal provides a basis for system decision-making. On the one hand, it enables the system to accurately identify non-user-triggered device behaviors that contradict the current macroscopic pattern, thereby achieving proactive early warning with a low false alarm rate.

[0008] On the other hand, when processing users' ambiguous voice commands, this indication signal can provide the system with scene logic support, helping it to prioritize the operation with the highest causal consistency with the current environment among multiple candidate operations, thereby improving the robustness and accuracy of interactive decision-making and making up for the shortcomings of technologies such as CN107357547A in decision-making intelligence. Attached Figure Description

[0009] Fig. 1 This is a schematic diagram of the overall system technical route of the present invention; Fig. 2This is a technical roadmap of the first data acquisition module, the second data acquisition module, and the fusion decision module of the present invention; Fig. 3 This is a technical roadmap for the indicator signal generation module, early warning generation module, and interactive decision-making module of the present invention. Detailed Implementation

[0010] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0011] It is understood that the terms “first,” “second,” etc., used in this application may be used herein to describe various elements, but unless otherwise stated, these elements are not limited by these terms. These terms are used only to distinguish one element from another.

[0012] Example 1: Please see Figs. 1-3 The present invention provides a technical solution: A smart home scenario-based audio linkage control system based on multimodal perception, used to monitor and control at least one electronic device and its surrounding environment, including: The first data acquisition module is configured to acquire first device event data characterizing the instantaneous operating state of at least one electronic device, the first device event data being generated based on the analysis of electromagnetic radiation signals generated by the electronic device; The second data acquisition module is configured to acquire second macro-state data that characterizes the macro-activity pattern of the environment in which the electronic device is located. The second macro-state data is generated based on long-term statistical analysis of network traffic metadata generated by multiple electronic devices in the environment. The fusion decision module is configured to: based on a preset conditional probability model, fuse first device event data and second macroscopic state data to calculate the conditional probability of an instantaneous operating state occurring under the macroscopic activity mode; The indicator signal generation module is configured to generate, based on conditional probability, a behavioral inconsistency indicator signal to characterize the degree of causal consistency between the instantaneous operating state and the macroscopic activity pattern.

[0013] for Fig. 1 Explanation: Fig. 1The smart home scenario on the left represents the operating environment of this invention, which includes various electronic devices. The technical roadmap derived from this scenario clearly reveals the core processing flow of the system. The roadmap starts at "multimodal data acquisition," which corresponds to the process by which the first data acquisition module acquires first device event data representing the instantaneous operating state of the devices, and the second data acquisition module acquires second macroscopic state data representing the macroscopic activity patterns of the environment. Subsequently, the process enters the "fusion analysis and inconsistency quantification" stage. This stage embodies the core technical steps of the fusion decision module fusing the above two types of data based on a preset conditional probability model, and the indicator signal generation module calculating and generating a behavioral inconsistency indicator signal representing the degree of causal consistency. This core indicator signal is applied to two parallel decision output paths: the "scene anomaly proactive warning" path corresponds to the function of the warning generation module generating a proactive warning signal when preset conditions are met; while the "fuzzy interaction precise decision" path corresponds to the interaction decision module using this signal to parse fuzzy voice commands to achieve more robust control command generation.

[0014] Further explanation: The steps for obtaining the first device event data specifically include: By constructing an electromagnetic leakage feature spectrum library, the electromagnetic radiation signals collected in real time are matched with the pre-stored features in the electromagnetic leakage feature spectrum library to identify the device identity and instantaneous operating status of electronic equipment. The steps for obtaining the second macroscopic state data specifically include: By calculating the information entropy of network traffic metadata and using a hidden Markov model to model the time series of information entropy, the macroscopic activity patterns of the environment can be identified. The steps for calculating conditional probability include: Conditional probabilities are calculated using a dynamically learned Bayesian inference model.

[0015] Further explanation: The electromagnetic leakage feature spectrum library is further limited to a pre-trained machine learning classification model; and the matching operation is further limited to the model inference process. The model inference process specifically includes: receiving real-time signal features extracted from the electromagnetic radiation signal, outputting a multi-dimensional micro-event feature vector characterizing the instantaneous operating state, and determining the device identity and the instantaneous operating state based on the multi-dimensional micro-event feature vector through a classification layer; The step of modeling the time series of the information entropy using a hidden Markov model is further configured to: output a macro-state probability distribution vector representing the macro-activity pattern, and determine the state with the maximum probability value in the macro-state probability distribution vector as the current macro-activity pattern; The step of calculating conditional probabilities is further configured as follows: the dynamically learned Bayesian inference model is a vectorized Bayesian network, which is configured to perform the fusion calculation of the micro-event feature vector and the macro-state probability distribution vector.

[0016] Further explanation: The fusion decision module is further configured to include a layered fusion processing architecture, which includes: The logical rationality gating unit is configured to: perform preliminary rationality verification on the events represented by the micro-event feature vectors based on the macro-state probability distribution vector, and generate a rationality adjustment coefficient; and, The dynamic weighted fusion unit is configured to receive micro-event feature vectors, macro-state probability distribution vectors, and rationality adjustment coefficients, and calculate conditional probabilities through a dynamically adjusted weighted network.

[0017] Further explanation: Before performing calculations, the dynamic weighted fusion unit is also configured as follows: Obtain a short-term activity context vector representing a recent sequence of events in the environment in which the electronic device is located; and obtain a periodic time context vector representing the current time. The weighted network in the dynamically weighted fusion unit is dynamically adjusted, and its internal weights are adjusted in real time based on the short-term activity context vector and the periodic time context vector.

[0018] The following is a detailed description of the implementation of the above content: This invention aims to address the challenge of accurately and reliably understanding and predicting user behavior and device status in complex smart home environments. Existing technologies analyze single types of data in isolation or perform simple data overlays, making it difficult to distinguish between normal user habits and genuine abnormal events. For example, the plausibility of an event like "the TV is turned on" late at night is closely related to whether the environment at that time is characterized by "family members being awake and active" or "family members being asleep." By deeply integrating microscopic instantaneous device events with macroscopic environmental activity patterns, and introducing an understanding of the event's temporal context, a foundation for intelligent control capable of logical reasoning is constructed.

[0019] 1) Specific implementation of the first data acquisition module: In this embodiment, the first data acquisition module aims to solve the technical problem of how to extract information that can accurately characterize the instantaneous operating state of the device from the original electromagnetic radiation signal. Its core idea originates from the time-frequency analysis theory in the field of signal processing. The technology used in this embodiment can simultaneously provide the distribution information of the signal in time and frequency, making it suitable for capturing and characterizing transient events with short durations and frequency variations. Its specific implementation process is decomposed into the following logical steps: An electromagnetic radiation signal segment of a preset time length is acquired; a wavelet packet transform is then performed on this segment. This transform decomposes the original signal into a series of orthogonal subspaces with different frequencies and time resolutions, yielding a set of wavelet packet decomposition coefficients. The wavelet packet transform can adaptively and more finely divide the entire frequency band of a signal, and compared to the standard wavelet transform, it can more effectively extract weak features hidden in each frequency band of the signal. In this preferred embodiment, the wavelet packet transform specifically uses "Daubechies-6" as the mother wavelet function and decomposes the signal to the 5th layer. The wavelet packet transform has good tight support and regularity in both the time and frequency domains, making it suitable for analyzing such non-stationary electromagnetic signals. The 5th layer decomposition in this embodiment achieves a technical balance between the frequency resolution of the signal features and the computational complexity, effectively distinguishing the characteristic frequency bands generated by different electronic devices.

[0020] Based on this set of wavelet packet decomposition coefficients, an energy distribution feature set is calculated. Specifically, the energy value of the coefficients within each frequency band subspace is calculated, and the energy values ​​of all frequency bands are combined into an energy feature sequence. This energy feature sequence is then matched against a pre-trained event classification model. The input to the event classification model is the energy feature sequence, and the output is a structured micro-event feature vector.

[0021] 2) Specific Implementation of the Second Data Acquisition Module: The second data acquisition module aims to solve the problem of how to abstract macroscopic patterns that can stably represent the overall activity rhythm of the environment from continuous and seemingly chaotic network traffic metadata. Its implementation is based on an improved Hidden Markov Model (HMM); the implementation process is as follows: A time series of information entropy is generated by calculating the information entropy of network traffic metadata within a unit time window. This information entropy sequence serves as the observation. This observation sequence is then input into a pre-trained Hidden Markov Model (HMM), whose state transition probability matrix and observation probability matrix are loaded from an external structured data carrier. The parameters of the HMM, namely the state transition probability matrix and observation probability matrix, are obtained by learning from a labeled macroscopic state sequence dataset.

[0022] The macroscopic state sequence dataset comprises two parallel, time-aligned data streams: one is the information entropy time series calculated from network traffic metadata, i.e., the observation sequence; the other is the corresponding baseline true macroscopic state sequence, i.e., the hidden state sequence, obtained through manual annotation or other auxiliary means. Other auxiliary means include indoor positioning and scheduling. The states in the baseline truth macroscopic state sequence originate from a predefined, complete, and mutually exclusive set of states that comprehensively covers the activities of the smart home environment. In a preferred embodiment of the invention, this set specifically includes the following four states: "Sleep Mode," indicating that all family members are in a resting sleep state; "At Home Activity Mode," indicating that at least one member is at home and active; "Away from Home Mode," indicating that all family members have left the environment; and "Visitor Mode," indicating that the activity mode of non-family members is detected. This well-defined set of states constitutes the state space of the HMM.

[0023] The model training process uses the Baum-Welch algorithm to train the HMM parameters. Its core idea is to start with random or preset initial parameters and iteratively adjust the parameters to maximize the probability of a given observation sequence occurring.

[0024] The training process executes the following iterative loop: In the E-step (expectation step), using the current HMM parameters, the expected number of times the model is in each hidden state at each time step, and the expected number of state transitions, are calculated given the observation sequence. In the M-step (maximization step), based on the expectation statistics calculated in the E-step, the state transition probabilities and observation probabilities are re-estimated, allowing the model with the new parameters to better interpret the observation data. This iterative process continues until the model parameters converge to a stable state. The final parameters are then used in the online system.

[0025] This embodiment utilizes the forward algorithm of the Hidden Markov Model (HMM) to calculate the probability of being in each predefined macrostate ("sleep mode", "work mode", "away mode", "visitor mode") at the current moment. These probability values ​​are combined into a macrostate probability distribution vector, denoted as Vstate, and used as the output of this module.

[0026] 3) The integrated decision-making module enables the construction of a hierarchical dynamic reasoning engine that simulates causality, temporal sequence, and logical relationships. Its internal implementation mechanism is described below.

[0027] In this embodiment, all key parameters are defined and normalized to a unified, closed real number range, which in this embodiment is between 0 and 1, to ensure the stability and comparability of the calculations. All configurable operating parameters are read from locally stored spreadsheet files to decouple the algorithm logic from the application strategy.

[0028] The micro-event feature vector, denoted as Vent, is a multi-dimensional vector used to comprehensively represent instantaneous device events. In this embodiment, it consists of three components: Event type encoding, denoted as Etype: It is a vector that has been one-hot encoded, and its dimension is equal to the total number of predefined event types. It is used to uniquely identify the category of the event, including "TV on" and "Microwave oven running".

[0029] Event intensity indicator, denoted as Eintensity: This is a scalar value that represents the normalized energy or amplitude of the detected event signal, reflecting the physical intensity of the event. It is determined by dividing the original signal energy value by a preset maximum possible energy value and then performing a saturation cutoff.

[0030] Event confidence, denoted as Econfidence, is a scalar value that represents the degree of confidence the event classification model has in the current identification result. It is directly provided by the model's output layer.

[0031] It should be noted that the construction and training process of the micro-event feature vector (Vevent) is as follows: A Multilayer Perceptron (MLP) network architecture is chosen. In this embodiment, the MLP model consists of an input layer, two hidden layers, and an output layer. The number of nodes in the input layer is the same as the dimension of the energy feature sequence. Both the first and second hidden layers use the Corrected Linear Unit (ReLU) as the activation function, which introduces non-linearity and enhances the model's expressive power. The structure of the output layer corresponds to the dimension of the micro-event feature vector (Vevent) and uses different activation functions: for the part used to output the event type encoding (Etype), the Softmax function is used to generate the probability distribution of each category; for the parts used to output the event intensity indication (Eintensity) and event confidence (Econfidence), the Sigmoid function is used to ensure that the output values ​​are in the range of 0 to 1.

[0032] The training of the above model relies on an event-signal labeled dataset with specific attributes. This dataset is obtained through the following process: in a controlled environment, various target electronic devices are sequentially operated to generate all predefined events. During this process, a sensing module with key performance attributes is used to synchronously acquire electromagnetic radiation signals. The key attribute requirements for this sensing module are: a sampling frequency of at least 100kHz to ensure the capture of high-frequency electromagnetic leakage characteristics; and a signal-to-noise ratio of at least 40dB to guarantee signal quality.

[0033] Each acquired raw signal is labeled with its corresponding ground truth labels for "device identity" and "instantaneous operating status." Subsequently, all signal samples undergo wavelet packet transform and energy feature extraction, generating a large number of "energy feature sequence-event label" data pairs, forming the training dataset. This dataset is divided into training, validation, and test sets in a predetermined ratio of 8:7:7. The model training process aims to find an optimal set of network weights and bias parameters through iterative optimization. This process is guided by a loss function. In this embodiment, the loss function is a composite loss, a weighted combination of cross-entropy loss for classification and mean squared error loss for regression. The optimization process is performed using an adaptive moment estimation (Adam) optimizer. In each iteration, a small batch of data is extracted from the training set, and the network is used for forward propagation to obtain the predicted output. Then, the loss value between the predicted output and the true label is calculated. Finally, the gradient of the loss with respect to each parameter is calculated using the backpropagation algorithm, and the optimizer updates the parameters based on this gradient. This process is repeated until the loss on the validation set no longer decreases significantly. After training, the obtained network weights and bias parameters are fixed, forming an executable event classification model. When running online, the model receives real-time energy feature sequences and directly outputs a structured micro-event feature vector, Vevent.

[0034] This embodiment further discloses a set of key hyperparameters used in the training process of the event classification model. The number of neurons in both hidden layers of the MLP model is set to 128. This value was determined based on experimental verification, which shows a good technical balance between ensuring the model has sufficient ability to capture complex nonlinear relationships of input features and avoiding overfitting due to excessive model complexity. During training, the initial learning rate of the Adam optimizer is set to an optimal value of 0.01%. This value aims to ensure stable convergence of the model in the early stages of training, avoiding oscillations caused by an excessively large learning rate. The batch size is set to 64, which is to balance the gradient estimation accuracy of a single parameter update with the computational efficiency of the training process. The entire training process continues until the loss on the validation set no longer decreases for 10 consecutive epochs, at which point it terminates early to prevent overfitting.

[0035] 4) Macrostate probability distribution vector Vstate: This is a multidimensional vector whose dimension is equal to the total number of predefined macrostates ("Sleep mode", "Work mode", "Away mode", "Visitor mode"). Each element in the vector represents the probability of the current environment being in the corresponding macrostate, and the sum of all elements is 1.

[0036] The rationality constraint matrix, denoted as M1, is a two-dimensional matrix where rows correspond to event types and columns correspond to macroscopic states. Each element M(i,j) in the matrix is ​​a preset value between 0 and 1, representing the inherent rationality or logical compatibility of event i occurring in state j. In this embodiment, the rationality value of the "oven on" event in the "sleep" state is preset to be low. The rationality constraint matrix is ​​loaded from an external data carrier. The steps for constructing the rationality constraint matrix M1 are as follows: Define all "event types" (rows of the matrix) and all "macro states" (columns of the matrix). Then, define a logical compatibility classification system. In this embodiment, the system includes three categories: "logically strong association" (characterizing that an event occurring in a certain state is expected and common), "logically neutral" (characterizing that an event has no obvious association with a state), and "logically weak association or conflict," characterizing that an event occurring in a certain state is rare or has a potential logical conflict with the definition of that state. For each logical compatibility category, assign a unique value normalized to the range of 0 to 1. This mapping rule aims to transform abstract logical relationships into computable parameters. In this embodiment, the rule is set as follows: The "logically strong association" category is mapped to a high reasonableness benchmark value. In this embodiment, the high reasonableness benchmark value is set to 0.95.

[0037] The "Logical Neutrality" category is mapped to a neutral rationality benchmark value, which is set to 0.5 in this embodiment.

[0038] The "Logically Weakly Related or Conflicting" category is mapped to a low reasonableness benchmark value. This parameter effectively suppresses the weighting of illogical combinations of events. In this embodiment, the low reasonableness benchmark value is set to 0.05.

[0039] Based on the above definitions and rules, each element M(i,j) in the matrix is ​​filled through a logical judgment process. For each combination of "event type i - macro state j", a judgment is made to determine its logical compatibility category, and the corresponding quantified value is entered. For the combination of "event: playing lullaby" and "state: sleep mode", it is determined to be "logically strongly related", so 0.95 is entered. For the combination of "event: coffee machine starts" and "state: away from home mode", it is determined to be "logically weakly related or conflicting", so 0.05 is entered.

[0040] The rationality adjustment coefficient, denoted as Ch, is a single scalar value calculated by the logical rationality gating unit and is used to adjust the weights of subsequent fusion calculations.

[0041] 5) Short-term activity context vector, denoted as Vactivity: This is a multi-dimensional vector used to capture environmental activity patterns over a period of time before an event occurs. It is determined by maintaining a fixed-length queue of recent micro-event feature vectors. When calculating this short-term activity context vector, the "event intensity indication" and "event confidence" components of all vectors in the queue are subjected to time-decay weighted averaging to form a compact vector that characterizes the frequency, intensity, and determinism of recent activities.

[0042] The time-decay weighted average uses an exponential decay model. The calculation logic is as follows: for each historical event feature vector in the event queue, a weight coefficient is calculated based on its time step from the current time. This weight coefficient is calculated by exponentially powering a preset decay benchmark parameter by that time step. The "event intensity indicator" and "event confidence" components of each historical event vector are multiplied by their corresponding weight coefficients. All weighted "event intensity indicator" components are summed, and all weighted "event confidence" components are summed to obtain the final short-term activity context vector. In this embodiment, the decay benchmark parameter is set to 0.9, which means that for each time step forward in the queue, its influence on the current context decays to 90% of its original value.

[0043] 6) Periodic Time Context Vector, denoted as Vtemporal: This is a multi-dimensional vector used to encode current time information into periodic features that are easy for the model to understand. It is determined by obtaining the current hour and day of the week values. The hour is transformed using sine and cosine functions, mapping it to a two-dimensional vector to address the issue that 11 PM and midnight are numerically different but actually adjacent. Similarly, the day of the week value is transformed in a similar way. These transformed results are then concatenated to form the time context vector.

[0044] 7) The layered fusion processing architecture logic of this embodiment is as follows: Obtain the micro-event feature vector Vevent and the macro-state probability distribution vector Vstate of the current input.

[0045] Parse the event type index from Vevent. Using this event type index, extract the corresponding row from the rationality constraint matrix M1. This row vector represents the inherent rationality of the event under all macro-states. Calculate the dot product of this row vector and the macro-state probability distribution vector Vstate. The physical meaning of this operation is to sum the inherent rationality of the event under each state using the actual probability of the current environment being in each state, obtaining a comprehensive "expected rationality value" in the current context. Output this "expected rationality value" as the rationality adjustment coefficient Ch. The closer the rationality adjustment coefficient Ch is to 1, the more rational the event is under the current macro-environment; the closer it is to 0, the less rational it is.

[0046] The dynamically weighted fusion unit is the core of the final conditional probability calculation. It constructs a computational network capable of dynamically adjusting its internal information flow based on multiple contexts. This solves the problem of fixed weights and inability to adapt to changing scenarios in conventional fusion methods. Its workflow is as follows: Obtain all necessary input vectors, including: micro-event feature vector Vevent, macro-state probability distribution vector Vstate, short-term activity context vector Vactivity, periodic time context vector Vtemporal, and rationality adjustment coefficient Ch. Calculate the context modulation weights. This step internally contains two parallel sub-computations: Sub-computation 1: Activity context modulation: The short-term activity context vector Vactivity is linearly combined with a set of preset "activity influence weights" loaded from an external data carrier. A "activity modulation gate value" is then generated using a non-linear activation function. In this embodiment, the non-linear activation function is the Sigmoid function, which maps any real number to the interval between 0 and 1. Sub-computation 2: Temporal context modulation: The periodic temporal context vector Vtemporal is linearly combined with another set of "temporal influence weights" and passed through a nonlinear activation function to generate a "temporal modulation gating value".

[0047] Each component of the micro-event feature vector Vevent is multiplied element-wise with the "activity modulation gate value".

[0048] Each component of the macroscopic state probability distribution vector Vstate is multiplied element-wise with the "time modulation gate value".

[0049] If current activities are frequent, the characteristics of the event itself will be amplified; if the current state is sleep, which is characterized by a macroscopic state that is sensitive to time, the impact of time information will be amplified.

[0050] The two gated and weighted vectors are concatenated to form a longer fused feature vector. This fused feature vector is then input into the final linear layer, where it undergoes a dot product operation with the final output weights to obtain a preliminary fused score. This preliminary fused score is then multiplied by a rationality adjustment coefficient Ch. The adjusted score is then passed through a final non-linear sigmoid function to ensure that the output value lies within the probability interval of 0 to 1; this final output value is the conditional probability.

[0051] It should be further explained that all internal weight parameters of the dynamic weighted fusion unit are treated as a unified, end-to-end machine learning model for training. The model is trained on a "context-decision" dataset containing complete input and expected output. Each sample in the dataset is a data tuple containing: a micro-event feature vector Vevent, a macro-state probability distribution vector Vstate, a short-term activity context vector Vactivity, a periodic time context vector Vtemporal, and an "expected conditional probability" label serving as the baseline truth. In this embodiment, this label is represented as a binary classification label (1 representing "normal / expected", 0 representing "abnormal / unexpected"), obtained through manual review and annotation of historical data.

[0052] The dynamically weighted fusion unit is treated as a differentiable computational graph. The goal of training is to adjust all its internal weight parameters so that, for a given input tuple, the model's final output is as close as possible to the ground truth label. The training process employs a supervised learning paradigm similar to the aforementioned MLP model: a binary cross-entropy loss function is used to quantify the difference between the predicted probability and the ground truth label. The Adam optimizer is also used, and the gradient of the loss with respect to all trainable parameters, including the "activity impact weights," "time impact weights," and "final output weights," is calculated using backpropagation and iteratively updated until the model reaches optimal performance on the validation set. This embodiment further discloses its key training hyperparameters. The initial learning rate of the Adam optimizer is set to 0.01% to achieve stable convergence. The batch size is set to 32, a value chosen because the dimensionality of each sample in the "context-decision" dataset is high, and a smaller batch size helps to perform effective training with limited computational resources. The total number of training epochs is capped at 200, and the same strategy of terminating early when the validation set performance shows no improvement for 10 consecutive epochs is adopted to ensure that the model achieves optimal generalization performance.

[0053] Through the above-described layered fusion processing architecture, the present invention achieves the following beneficial effects: The rationality gating unit filters out a large number of logically illogical event-state combinations, improving the accuracy and robustness of the judgment. The dynamic weighted fusion unit can adjust the focus on different information in real time based on short-term event history and periodic time patterns, enabling the model to understand "habits" and thus distinguish true "anomalies." By using feature vectors instead of discrete labels, key information such as event intensity and state uncertainty is preserved, making the foundation of fusion decision-making more solid.

[0054] 8) The calculation process of this invention based on the first data acquisition module, the second data acquisition module, and the fusion decision module in the embodiment is as follows: The first and second data acquisition modules are started in parallel. The first data acquisition module continuously processes electromagnetic radiation signals and generates a micro-event feature vector (Vevent) in real time through wavelet packet transform and classification models. The second data acquisition module continuously processes network traffic metadata and generates a macro-state probability distribution vector (Vstate) in real time through information entropy calculation and HMM forward algorithm. Whenever a new micro-event feature vector is generated, the calculation of the fusion decision module is triggered. The fusion decision module updates and calculates the current short-term activity context vector (Vactivity) and periodic time context vector (Vtemporal). Subsequently, Vevent and Vstate are input into the logical rationality gating unit to calculate the rationality adjustment coefficient (Ch). All five key inputs (Vevent, Vstate, Vactivity, Vtemporal, Ch) are sent to the dynamic weighted fusion unit. The dynamic weighted fusion unit performs its internal multi-step calculation, including context modulation, gating weighting, feature aggregation, and final adjustment. The final output represents the conditional probability value of the instantaneous operating state under the current macro-activity mode and multiple contexts. This conditional probability value can be used as a decision basis by the subsequent indication signal generation module and interactive decision module. It should be understood that the above embodiments based on the layered fusion processing architecture are only one of the preferred implementations of the present invention. Without departing from the core spirit of the present invention, the "conditional probability model" can also be implemented through other technical paths.

[0055] In another alternative embodiment, the fusion decision module can be implemented as an adaptive inference engine based on a discretized conditional probability table. Specifically, the first data acquisition module can use Fast Fourier Transform (FFT) and pre-stored spectral templates to perform cross-correlation calculations to output discrete event labels. The second data acquisition module can divide the environment into discrete macro states such as "high activity," "medium activity," and "low activity" by setting a set of threshold rules based on the number of active network devices. In this alternative embodiment, the core of the fusion decision module is a two-dimensional lookup table, with row indices representing event labels and column indices representing macro states. The values ​​stored in the table are the conditional probabilities obtained through long-term data statistics. This module continuously updates the table through an online learning mechanism, which is implemented by maintaining a frequency counter with time decay for each "event-state" combination. Whenever a new "event-state" pair is observed, the value of the corresponding counter increases by one unit, and the values ​​of all counters are multiplied by a decay factor less than 1. The conditional probability value is dynamically calculated by dividing the value of a specific counter by the total value of its corresponding state (column).

[0056] The following further elaboration addresses the above: The final output of the fusion decision module is a conditional probability value located within the closed interval [0,1]. The technical meaning of this value is defined as follows: When the conditional probability value is closer to 1, it indicates that the system of this invention is more likely to determine the reasonableness, predictability, or logical compatibility of the occurrence of the "micro-event" in the current "macro-state" and multiple additional contexts (time, recent activities). Conversely, when the conditional probability value is closer to 0, it indicates that the reasonableness of the event occurring in the current situation is lower, and the possibility of it constituting a potential "unexpected" or "abnormal" event that requires attention is higher.

[0057] The final output conditional probability value is the result of the nonlinear effects of multiple key input parameters processed through a hierarchical fusion architecture. The trends and underlying logic of the influence of each parameter are as follows: The impact of event intensity indicator Eintensity and event confidence Econfidence on the micro-event feature vector Vevent: With other input parameters remaining constant, an increase in Eintensity or Econfidence will lead to a non-linear increase in the conditional probability value of the final output. In the dynamic weighted fusion unit of this invention, the Vevent participates in the subsequent weighted calculation as a whole. Events with stronger physical intensity (higher Eintensity) or more deterministic classification models (higher Econfidence) have feature vectors with a larger norm in the fusion calculation, thus having a greater impact in interactions with other vectors (including dot products and weighted summations), tending to output more deterministic results (closer to 0 or 1). This enhancement effect is positively correlated when the event itself is deemed reasonable. The impact of the rationality adjustment coefficient: There is a direct positive multiplicative correlation between the parameter and the conditional probability value of the final output. The rationality adjustment coefficient is an adjustment factor that is directly multiplied by the preliminary fusion score in the later stages of the dynamic weighted fusion unit calculation. A low rationality adjustment coefficient value (representing a logical conflict between the event and the current macroscopic state) will directly and forcefully suppress the final conditional probability value, causing it to approach 0. A "logical veto gate" is implemented at the algorithm level, ensuring that any event that fundamentally contradicts the macroscopic background will have its final rationality score decisively suppressed.

[0058] The impact of the short-term activity context vector Vactivity: This parameter's effect on the final output is context-dependent and non-linear. Vactivity is not directly added to or multiplied by the final output; instead, it controls the weighting of information flow by adjusting the "activity modulation gate" within the dynamically weighted fusion unit. When Vactivity represents frequent recent activity (multiple device events occurring consecutively), the corresponding gate value increases, amplifying the impact of Vevent itself in the fusion computation. This indicates that the system becomes more sensitive to the "current event." Conversely, after a long period of silence, the system's response to a single isolated event will be relatively flat. This design maps to the logic of "behavioral inertia" in the real world, where the occurrence of an event is related to its preceding sequence of events.

[0059] The impact of the periodic temporal context vector Vtemporal: Similar to Vactivity, the effect of this parameter is also context-dependent and non-linear. Vtemporal controls the weight of Vstate (macrostate) in the fusion process by adjusting the "time modulation gating value". During late-night hours, the encoded value of Vtemporal amplifies the weights associated with "sleep mode". If the Vstate at this time does indeed point to "sleep mode", the influence of this macrostate on the final decision will be enhanced. This allows the present invention to understand complex anomalous scenarios such as "a seemingly correct event occurring at the wrong time".

[0060] Further explanation: The indicator signal generation module is specifically configured to define the magnitude of the behavior inconsistency indicator signal as being proportional to the negative logarithm of the conditional probability.

[0061] Based on time series of behavioral inconsistency indicator signals, a cumulative inconsistency score is generated through a time accumulation processing step, which represents the cumulative effect of the degree of inconsistency in recent history.

[0062] Further explanation: The warning generation module is defined and configured to generate an active warning signal when the value of the behavior inconsistency indication signal exceeds a preset warning threshold, provided that the first device event data is not triggered by a user command.

[0063] Further explanation: The early warning generation module is further configured to generate proactive early warning signals through a two-stage hierarchical early warning logic, which includes: The instantaneous attention gate is configured to: activate the second stage of the warning logic when the instantaneous value of the behavior inconsistency indication signal exceeds a preset attention threshold; and, The continuous confirmation gate is configured to generate an active warning signal when the accumulated inconsistency score exceeds a preset confirmation threshold, provided that the instantaneous attention gate is activated.

[0064] Further explanation: Define the interaction decision module, which is configured to: when receiving a user's ambiguous voice command, generate control commands for parsing the ambiguous voice command based on the behavior inconsistency indication signal, so as to prioritize the candidate operation with the highest causal consistency with the macroscopic activity pattern of the environment.

[0065] Further explanation: The interactive decision-making module is further configured as follows: For each potential candidate operation, obtain the prior weight representing the inherent priority of the candidate operation; based on the behavioral inconsistency indication signal corresponding to the candidate operation, determine a scene adaptability score; By fusing prior weights and scenario adaptability scores, a final decision score is generated for each potential candidate operation. Calculate the difference in decision scores between the first candidate operation with the highest final decision score and the second candidate operation with the second highest final decision score among the candidate operations. Control commands are generated by selecting steps in interactive mode. The steps are configured as follows: When the difference in decision scores is greater than a preset ambiguity threshold, a control instruction is generated to directly execute the first candidate operation. When the difference in decision scores is not greater than the ambiguity threshold, a voice prompt is generated via the audio output device to request further clarification from the user.

[0066] The following is a detailed description of the implementation of the above content: In complex intelligent environments, audio control decisions based on instantaneous data make it difficult to distinguish between one-off signal glitches and persistent real anomalies, resulting in high false alarm or false negative rates in the warning function. Furthermore, when processing ambiguous user voice commands, it tends to select the "most likely" option and execute it directly, even if the confidence level of this selection is only slightly higher than the second-best option. This decision-making method sacrifices the accuracy of the interaction and the user's sense of control. This invention aims to solve the above-mentioned technical defects caused by the lack of time dimension consideration and adaptive decision-making logic. This embodiment involves the following key parameters.

[0067] The Behavioral Inconsistency Index (BII) is a dimensionless scalar used to quantify the degree of unexpectedness or logical inconsistency in the occurrence of instantaneous equipment events under specific macroscopic activity patterns. A higher BII value indicates a weaker causal relationship between the event and the current environmental state. The determination process involves receiving "conditional probabilities" from the upstream "fusion decision module" as input. The calculation model is as follows: obtain the conditional probability value; perform a logarithmic operation on the conditional probability value with the natural constant as the base; take the arithmetic negation of the result; the resulting value is the BII.

[0068] The cumulative inconsistency score, denoted as AIS, is a dimensionless scalar that characterizes the cumulative intensity and persistence of behavioral inconsistencies within a recent historical time window. AIS exhibits a "memory effect," reflecting whether inconsistencies are isolated events or persistent states. Its determination process is executed by the "time accumulation processing unit." The core computational model of this unit is as follows: Obtain the new BII value calculated at the current time, denoted as BII. current And the AIS value at the previous moment, denoted as AIS previous Obtain the preset "time decay coefficient". The time decay coefficient controls the influence weight of historical information on the current AIS value, and its value ranges from 0 to 1. In this embodiment, the time decay coefficient is set to 0.9. Specifically, the "time decay coefficient" is determined through the following steps: Obtain a benchmark dataset containing two types of typical signals; the first type is the transient interference set: collect at least one hundred independent high BII pulse events triggered by known benign causes (including transient power grid fluctuations and transient electromagnetic radiation from unrelated equipment). Each event is recorded as a time series, characterized by a rapid increase in BII value within a short period of three calculation cycles followed by a drop back to the baseline level.

[0069] Category 2 is the persistent anomaly set: at least one hundred independent BII sequences triggered by known faults or persistent anomalous states (including equipment entering an abnormal operating cycle, continuous non-mandatory operations). Each event is characterized by an elevated BII value that remains significantly above the baseline, including those lasting for more than 10 calculation cycles.

[0070] Define the discrimination evaluation index. Define "peak-valley difference" as the evaluation index. For transient interference signals, the lower the peak value of its AIS sequence, the better; for persistent anomalous signals, the higher the "valley value" of its AIS sequence after stabilization, i.e., the lowest value during the stable period, the better. "Peak-valley difference" is defined as: the difference between the mean of the valley values ​​of the AIS during the stable period of all signals in the persistent anomalous set and the mean of the peak value of the AIS of all signals in the transient interference set. Specifically, the following deterministic iterative calculation process is executed: Initialize a candidate "time decay coefficient" list, starting from 0.80 and increasing in increments of 0.01 to 0.99. For each candidate coefficient value representing the time decay coefficient in the list: apply the candidate coefficient value to the BII sequences of all signals in the reference dataset to calculate the corresponding AIS time series. For the transient interference set, calculate the average of the peak values ​​of all AIS sequences. For the persistent anomaly set, calculate the average of the valley values ​​of all AIS sequences during the stable period; in this embodiment, the stable period begins at the tenth calculation cycle; calculate the "peak-valley difference" under the current candidate coefficient.

[0071] After iterating through all candidate coefficients, the candidate coefficient that maximizes the "peak-valley difference" is selected as the final value of the "time decay coefficient" in this invention. The preferred value range for the "time decay coefficient" is limited to 0.9 to 0.99. This range is determined as follows: when the "time decay coefficient" is 0.9, the corresponding "target event half-life" is seven time steps, which is suitable for scenarios requiring rapid response and rapid forgetting of abnormal events. When the time decay coefficient is 0.99, the corresponding "target event half-life" is 69 time steps, which is suitable for scenarios requiring stable accumulation of long-term, slowly changing trends, such as environmental comfort monitoring. Based on specific application requirements, the desired value is selected within the defined range of "target event half-life".

[0072] AIS previous Multiply the result by the "time decay coefficient" to obtain the decayed historical score. Then, use the BII... current Multiply the result of "1 minus the time decay coefficient" to obtain a weighted current score. Add the decayed historical score to the weighted current score, and the sum is the AIS value at the current moment.

[0073] 9) The two-stage hierarchical early warning mechanism is explained as follows: Directly comparing BII with a single fixed threshold cannot distinguish between instantaneous high-intensity interference (including electromagnetic crosstalk caused by the starting of appliances from neighboring homes) and continuous, moderate-intensity real anomalies (including abnormal start-stop cycles of the refrigerator compressor). The former will trigger unnecessary false alarms, while the latter will be ignored because it does not reach a high threshold. This embodiment sets up the following cognitive process for hierarchical decision-making logic: The first stage, the instantaneous attention gate, aims to quickly capture any potential abnormal signals. It uses the instantaneous BII value as the criterion. Its working logic is as follows: it acquires the current BII value and compares it with a preset "attention threshold." This ensures that no potentially risky signals are missed. Only when the BII value exceeds this "attention threshold" will the process proceed to the second stage; otherwise, the warning process terminates. This effectively filters out a large amount of low-intensity background noise, reducing the computational load of subsequent processing.

[0074] The second stage, the continuous confirmation gate, is initiated after the "instantaneous attention gate" is activated. Its purpose is to confirm whether the abnormal signal is persistent. The AIS value, which has a memory effect, is used as the criterion. Its working logic is as follows: it acquires the current AIS value and compares it with a preset "confirmation threshold." The confirmation threshold is used to trigger the final alarm only when the inconsistency accumulates to a level sufficient to be considered a persistent anomaly. The warning generation module will generate the final proactive warning signal only if and only if the AIS value exceeds the "confirmation threshold."

[0075] It should be further explained that a baseline dataset under normal operating conditions is constructed. Under the condition that there are no known anomalies in the environment, BII and AIS time series data are continuously recorded for N1 calculation cycles to form a "normal baseline dataset".

[0076] Extract all BII values ​​from the "Normal Baseline Dataset". Perform statistical analysis on these BII values ​​to calculate their probability density distribution. Set a target "Initial Attention False Alarm Rate", which defines the acceptable probability that the "Instantaneous Attention Gate" will be erroneously triggered under normal conditions. In this embodiment, the Initial Attention False Alarm Rate is set to 0.1%. This value allows the system to maintain attention on rare normal high-value signals while avoiding being overwhelmed by subsequent calculations due to excessive attention. The "Attention Threshold" is determined to be the 999th percentile of the BII value probability density distribution. That is, during normal operation, 999 out of 1000 BII values ​​are below this attention threshold.

[0077] Furthermore, all AIS values ​​are extracted from the "normal baseline dataset." Statistical analysis is performed on these AIS values ​​to calculate their probability density distribution. A target "final confirmation false alarm rate" is set; this parameter defines the acceptable probability of the system issuing a final false alarm under normal conditions. In this embodiment, the "confirmation threshold" is set to 4. This two-stage mechanism not only detects events but also performs preliminary risk classification by distinguishing between transient and persistent events. It effectively filters out false alarms caused by transient interference while improving the detection sensitivity for persistent and progressive faults, thereby enhancing the accuracy and reliability of the early warning system.

[0078] 10) Adaptive Interaction Mode Selection Mechanism: When parsing ambiguous voice commands, the system calculates the scores of each candidate operation and selects the highest-scoring operation for execution. When the highest score is close to the second-highest score, this strategy represents information loss, ignoring the decision-making uncertainty within the system and potentially executing operations that are not intended by the user. This embodiment has a mechanism that can perceive its own decision-making uncertainty and adaptively switch interaction modes. For each candidate operation, a "final decision score" is calculated by fusing its "prior weight" and "scene adaptability score".

[0079] It should be noted that the determination process for the "prior weight" of each executable candidate operation is as follows: the system designer scores the operation on the following three core factors, with scores ranging from 1 to 10: Safety Impact Factor: Assess the potential safety risks arising from accidental triggering of this operation. The lower the risk, the higher the score. Examples include: "Turning on the night light" scores 10, and "Turning on the gas stove" scores 1.

[0080] Energy Consumption Impact Factor: Assess the typical energy consumption level of this operation. The lower the energy consumption, the higher the score. Examples include: "Adjusting the air conditioner to energy-saving mode" scores 9, and "Turning on a high-power oven" scores 3.

[0081] Convenience Influence Factor: This assesses the likelihood of this operation being used frequently and preferentially in daily life. The higher the likelihood, the higher the score. Examples include: "Playing music" scores 8, and "Starting the robot vacuum cleaner" scores 5.

[0082] A set of "factor weight vectors" is preset, with the following weights: safety impact factor weight 0.5, energy consumption impact factor weight 0.2, and convenience impact factor weight 0.3. The sum of these weight values ​​is 1. The technical consideration is that "safety" is given the highest decision priority.

[0083] Multiply the "Safety Impact Factor" score of this operation by its corresponding factor weight. Multiply the "Energy Consumption Impact Factor" score of this operation by its corresponding factor weight. Multiply the "Convenience Impact Factor" score of this operation by its corresponding factor weight. Add these products together to obtain a total score. Divide this total score by 10 and normalize it. The result is the final "prior weight" of this operation, whose value is strictly limited to between 0 and 1.

[0084] This embodiment further discloses the scoring and weighting determination mechanism. Regarding factor scoring, each operation is scored according to the following objective scoring guidelines: Safety Impact Factor Scoring Guidelines: A score of 1 to 3 corresponds to the possibility of serious personal injury or significant property damage due to accidental operation (including starting gas equipment); a score of 4 to 7 corresponds to the possibility of equipment damage or significant functional abnormalities; a score of 8 to 10 corresponds to almost no safety risk (including playing music or adjusting light brightness).

[0085] Energy Consumption Impact Factor Scoring Guidelines: Mapped based on the rated power of the operating equipment. Equipment with a rated power higher than 1000 watts is scored from 1 to 3; equipment with a rated power between 100 watts and 1000 watts is scored from 4 to 7; and equipment with a rated power lower than 100 watts is scored from 8 to 10.

[0086] Convenience Impact Factor Scoring Guidelines: Based on the average daily usage frequency within the reference user group (including a survey of one hundred target users). Average daily usage less than once is scored 1 to 3; average daily usage between 1 and 5 times is scored 4 to 7; average daily usage more than 5 times is scored 8 to 10.

[0087] The determination of the "factor weight vector" (0.5, 0.2, 0.3) is achieved through the Analytic Hierarchy Process (AHP). The process involves constructing a judgment matrix and comparing each of the three criteria—"safety," "energy consumption," and "convenience"—pair by pair to determine their relative importance. In this embodiment, "safety" is established as absolutely important relative to "convenience," "safety" is absolutely important relative to "energy consumption," and "convenience" is slightly more important than "energy consumption." Based on this judgment matrix, its largest eigenvalue and corresponding normalized eigenvector are calculated, which is (0.5, 0.2, 0.3), thus objectively determining the weights of each factor.

[0088] 11) Among the final decision scores of all candidate operations, find the highest and second-highest scores. Then, perform a subtraction operation, that is, subtract the second-highest score from the highest score to obtain the "decision score difference". This difference is a direct quantification of decision ambiguity: the larger the difference, the more obvious the advantage of the optimal option and the higher the certainty of the decision; the smaller the difference, the less certain the system is about the optimal choice and the higher the ambiguity.

[0089] For interaction mode selection: the calculated "decision score difference" is compared with a preset "ambiguity threshold" read from an external data carrier. The technical consideration of the ambiguity threshold is to define the lower bound of acceptable decision certainty for the system.

[0090] Execution mode: If the "decision score difference" is greater than the "ambiguity threshold", the system determines that the decision is clear and unambiguous. At this time, a control instruction is generated to directly execute the first candidate operation with the highest score.

[0091] Clarification Mode: If the "Decision Score Difference" is not greater than the "Ambiguity Threshold," the system determines that there is high ambiguity. In this case, no action is taken; instead, an interface device drives the audio output device to generate a voice prompt requesting clarification from the user. This includes outputting the voice prompt: "Do you want to turn off the TV or the desk lamp?" The technical effect of this mechanism is that it enables the system to assess its own decision confidence and proactively initiate dialogue to eliminate ambiguity, thus becoming an intelligent collaborator. It improves the robustness, accuracy, and user experience of human-computer interaction, transforming the interaction from a one-way "command-execution" paradigm to a two-way "proposal-confirmation" paradigm.

[0092] To determine the ambiguity threshold: Construct an interactive reference dataset; in an environment where this invention has been deployed, collect all interactive events triggered by ambiguous commands over a period of time. For each event, record the following information: the "final decision score" of all candidate operations calculated by the system, and the user's final feedback, i.e., whether the user accepted the highest-scoring operation recommended by the system, or whether the user corrected it through subsequent commands. Divide the dataset into two groups: "Revision Group": All interaction events where the user revised the system's initial recommendation. "Acceptance Group": All interaction events where the user directly accepted the system's initial recommendation.

[0093] For each event in the "correction group," its "decision score difference" is calculated. For the dataset consisting of all "decision score differences" in the "correction group," a specific percentile is calculated, specifically the 90th percentile. This calculated percentile is determined as the "ambiguity threshold" of this invention. Choosing the 90th percentile of the "correction group" as the threshold ensures that the decision score difference is below this threshold when 90% of users require correction. Therefore, setting this value as the boundary indicates that the system will proactively initiate clarifying dialogues in the vast majority of scenarios where there is indeed high ambiguity (90%), effectively intercepting most potential interaction errors while avoiding unnecessary disturbances in scenarios with low ambiguity.

[0094] 12) This embodiment further performs the following calculation process: Data is continuously collected through the "first data acquisition module" and the "second data acquisition module", and the "fusion decision module" calculates the conditional probability of any device event under the current macroscopic state.

[0095] The "Indication Signal Generation Module" receives conditional probabilities and calculates the instantaneous "Behavioral Inconsistency Index (BII)". Simultaneously, the "Time Accumulation Processing Unit" updates and calculates the current "Cumulative Inconsistency Score (AIS)" based on the current BII value and historical AIS values. For the early warning function, the BII value is compared to the "Instantaneous Attention Gate". If it passes, the AIS value is compared to the "Continuous Confirmation Gate". If it passes again, an active early warning signal is generated.

[0096] For interactive functions, when a vague voice command is received and converted into internal data through a voice input device, a "final decision score" is calculated for each candidate operation. Then, the score difference is calculated and compared with the "ambiguity threshold" to finally decide whether to execute directly or ask a voice question for clarification.

[0097] The above content is further elaborated as follows: The "Behavioral Inconsistency Index (BII)" and the "Cumulative Inconsistency Score (AIS)" together constitute a quantitative assessment system for the anomalies of environmental conditions.

[0098] The Behavioral Inconsistency Index (BII) outputs non-negative real numbers. When the BII value approaches zero, it indicates a higher causal correlation between the instantaneous operating state indicated by "first device event data" and the macroscopic activity pattern indicated by "second macroscopic state data," meaning the event was predictable and the system was in a highly logically consistent state. Conversely, as the BII value increases, it indicates a weaker causal correlation between the instantaneous operating state and the macroscopic activity pattern, an increased "unexpectedness" of the event, and a rising risk of the system deviating from its logically consistent state.

[0099] The Cumulative Inconsistency Score (AIS) output range is a non-negative real number, and its dynamic range is related to the BII (Body Inconsistency Index). When the AIS value approaches zero, it indicates that the system has been logically consistent in its recent history, with no or only minor inconsistencies that can be ignored by time decay. When the AIS value increases, it indicates that one or more significant inconsistencies have occurred in the recent history, and their impact is being amplified and sustained by the time-cumulative processing steps. This reveals a persistent or high-frequency anomalous trend and is a strong indicator that the system may be entering a persistent failure or anomalous mode.

[0100] Conditional probability is the sole input for calculating BII. Its value shows a clear negative correlation with the BII value. The BII value is directly proportional to the negative logarithm of the conditional probability. The logarithmic function is monotonically increasing, but becomes monotonically decreasing when its negative value is taken. Therefore, as the conditional probability changes from near 1 (indicating a high probability of the event) to near 0 (indicating a low probability of the event), its logarithm changes from 0 to negative infinity, and the negative logarithm changes from 0 to positive infinity. The lower the probability of an event, the greater the amount of information it carries.

[0101] The "time decay coefficient" determines the weight of historical AIS values ​​on the current AIS value. Its increase or decrease is clearly positively correlated with the smoothness of instantaneous BII fluctuations by AIS, i.e., the system's "memory length." According to the calculation logic of AIS, the AIS value at the previous moment... previous Multiply by this time decay factor, and the current BII value is BII. current This is then multiplied by "1 minus the time decay coefficient". As the "time decay coefficient" increases (approaching 1), the value of "1 minus the time decay coefficient" decreases. This indicates that when updating AIS, the weight of historical values ​​is amplified, while the influence of the current instantaneous value is weakened. This makes the changes in AIS smoother, effectively filtering out brief BII spikes, thus more stably reflecting the long-term cumulative trend of inconsistencies. This positive correlation enables controllable adjustment of the system's time response characteristics.

[0102] 13) To quantitatively verify the advancements of the proposed "two-stage hierarchical early warning logic" and "adaptive interaction mode selection mechanism" compared to existing technologies, a digital environmental state assessment platform was built. This platform can generate and record device event sequences and macroscopic state sequences in a simulated intelligent environment, and calculate the core output indicators of the proposed method and the control group method accordingly. This experiment aims to quantitatively verify the superiority of the proposed "two-stage hierarchical early warning logic" in distinguishing between various transient disturbances and persistent anomalies. The control group uses traditional techniques, directly comparing the instantaneous "Behavioral Inconsistency Index (BII)" with a fixed, optimized single threshold of 5.0 to trigger an alarm. The proposed method sets the "attention threshold" to 3.5 and the "confirmation threshold" to 4.0. The experiment covers four of the most typical operating conditions to ensure the comprehensiveness of the assessment. See Table 1 below for details. Table 1: Performance comparison of the present invention and the control group in typical early warning scenarios The Comprehensive Early Warning Effectiveness (CWS) metric is designed to provide a single quantitative score for the overall performance of the early warning system across multiple scenarios. Its calculation logic is as follows: each decision outcome is assigned an effectiveness score: a correct alarm (True Positive, TP) earns +2 points, a correct no alarm (True Negative, TN) earns +1 point, a missed alarm (False Negative, FN) earns -2 points, and a false alarm (False Positive, FP) earns -1 point. CWS is the sum of the system's effectiveness scores across all test scenarios. CWS quantitatively characterizes the overall effectiveness of the early warning system. Its scoring rules are as follows: the harm of a missed alarm (FN) is far greater than that of a false alarm (FP), therefore it is given twice the negative weight; while a correct alarm (TP) is more valuable than a normal correct no alarm (TN) because it successfully avoids risk. The higher the value of this metric, the stronger the system's comprehensive early warning effectiveness.

[0103] Scenario 1 (Instantaneous Strong Interference): The instantaneous peak value of BII (5.30) exceeded the single threshold of the control group, causing the control group to generate false alarms. Since this interference is transient, the AIS value of the present invention only accumulated to 1.17, which is far below its "acknowledgment threshold" of 4.0. Therefore, the present invention did not trigger an alarm and successfully suppressed false alarms.

[0104] Scenario 2 (Persistent Weak Anomaly): The BII peak value (3.91) is lower than the threshold of the control group, causing a false negative in the control group. However, since this anomaly is persistent, the AIS value of this invention accumulates steadily and eventually reaches 4.25, successfully exceeding the "confirmation threshold" and thus correctly triggering the alarm.

[0105] The data from these two key scenarios directly demonstrate that the two-stage logic of this invention can effectively decouple the judgment of the "intensity" and "persistence" of an event, and solve the inherent dilemma of "false alarm-false negative" in traditional single threshold methods.

[0106] In both Scenario 3 (no anomaly) and Scenario 4 (transient weak interference), both methods exhibited the correct silent behavior and did not trigger alarms.

[0107] In Scenario 5 (persistent strong anomalies), both methods correctly and quickly triggered alerts. This demonstrates that while enhancing the ability to handle complex scenarios, they did not sacrifice basic performance in simple or routine scenarios, exhibiting high robustness.

[0108] To conduct the final quantitative evaluation, the performance of the two methods in all five scenarios was calculated according to the CWS scoring rules: CWS calculation for the control group: Scenario 1 (FP): -1 point; Scenario 2 (FN): -2 points; Scenario 3 (TN): +1 point; Scenario 4 (TN): +1 point; Scenario 5 (TP): +2 points; The total CWS is: (-1) + (-2) + 1 + 1 + 2 = +1 point.

[0109] The CWS calculation method of this invention is as follows: Scenario 1 (TN): +1 point; Scenario 2 (TP): +2 points; Scenario 3 (TN): +1 point; Scenario 4 (TN): +1 point; Scenario 5 (TP): +2 points; The total CWS is: 1+2+1+1+2=+7 points.

[0110] In extended testing covering all four typical operating conditions, the overall early warning performance score (+7) of this invention was seven times that of the control group (+1). This data demonstrates that the two-stage hierarchical early warning logic of this invention not only performs superiorly in critical scenarios that existing technologies cannot handle, but also maintains the same correctness as existing technologies in other scenarios, thus achieving an overall performance advantage.

[0111] 14) Further demonstrate the process of determining the "confirmation threshold" corresponding to the "cumulative inconsistency score (AIS)" in this invention by using the "Objective Threshold Derivation (OTD)" methodology; The F1 score for early warning decisions was selected as the verifiable performance indicator (VEI). The F1 score is a comprehensive statistical metric used to measure the accuracy of binary classification models, taking into account both precision and recall. The rationale for choosing the F1 score as the VEI lies in the fact that an excellent early warning system must simultaneously achieve both high accuracy (high precision, avoiding false positives) and comprehensiveness (high recall, avoiding missed reports), and the F1 score is the most appropriate quantitative description of this comprehensive performance.

[0112] Based on an expanded dataset containing labeled "normal" and "abnormal" events, the impact of different "confirmation threshold" values ​​on the F1 score of the final warning decision was analyzed. The results show a typical unimodal asymmetric curve relationship between the "confirmation threshold" and the F1 score. Specifically, when the threshold is too low, recall is high but precision is low, resulting in a low F1 score; as the threshold increases, precision increases rapidly, and the F1 score rises accordingly; once the threshold exceeds an optimal point, recall begins to decline sharply, causing the F1 score to decrease again.

[0113] A mathematical identification method for finding the peak point of the function is employed. The optimal "confirmation threshold" is defined as the threshold point that maximizes the VEI "F1 score". This peak point is precisely calculated by differentiating the unimodal curve function and setting its first derivative to zero. The threshold point that maximizes the F1 score is found to be Toptimal≈4.0.

[0114] Setting up a Level 1 intervention: Generating proactive early warning signals (including audio or visual cues). This has limited impact on the user, low computational cost, and its "intervention utility" is to prompt the user to pay attention to potential problems.

[0115] Zero intervention: Silent monitoring at zero cost.

[0116] The primary intervention strategy is precisely mapped to an interval defined by the threshold Toptimal. When the AIS value is below Toptimal, zero intervention is performed; when the AIS value exceeds Toptimal, primary intervention is initiated. This mapping is optimal, and initiating intervention at this threshold ensures that every intervention action is based on a data-proven decision boundary with the highest overall benefit. This maximizes the accuracy of early warnings while minimizing the negative impacts of false alarms (costs) and missed alarms (utility losses).

[0117] By rigorously implementing the OTD methodology, the following practical application range with a solid data foundation was identified for the Cumulative Inconsistency Score (AIS): Safety / monitoring interval: AIS∈[0,4.0]; its upper boundary 4.0 is determined by the unique mathematical optimum Toptimal that maximizes the overall performance index (F1 score) of the early warning system. Zero intervention is performed. The system operates silently within this interval, performing only background data recording and continuous monitoring.

[0118] Warning / Intervention Interval: AIS∈(4.0,+∞); its lower boundary 4.0 is determined by the optimal point Toptimal. Immediately initiate Level 1 intervention, generating an active warning signal. The trigger boundary for this operation, after cost-benefit analysis, has been confirmed as the decision point for achieving optimal warning performance within this technical framework.

[0119] The computational logic involved in this application can be constructed using algorithms such as regression analysis in machine learning, establishing a mathematical model by analyzing the inherent trends and interrelationships of the collected parameters. This process can be implemented using specialized computational tools (such as Python's Scikit-learn library or the R language environment). Throughout all calculations, to eliminate the influence of different physical dimensions and ensure that data is compared and analyzed on the same scale, the input parameters in each formula are dimensionless. The dimensionless techniques used include, but are not limited to, max-min normalization or Z-score standardization.

[0120] The algorithm of this invention is implemented as a Python script. Before executing the core logic, the program first executes a data loading module (e.g., using the widely used pandas library in Python) configured to read the aforementioned spreadsheet file and load its contents into the program's working memory (e.g., a DataFrame data structure). Subsequent algorithm steps will directly query and retrieve the required configuration parameters from this in-memory data structure.

[0121] It should be emphasized that the foregoing embodiments are merely illustrative of preferred implementations of the present invention and are not intended to limit the scope of protection of the present invention. This application also provides a computer-readable storage medium having computer program instructions stored thereon.

Claims

1. A smart home scenario-based audio linkage control system based on multimodal perception, characterized in that, Specifically, it includes: The first data acquisition module is configured to acquire first device event data characterizing the instantaneous operating state of at least one electronic device, the first device event data being generated based on the analysis of electromagnetic radiation signals generated by the electronic device; The second data acquisition module is configured to acquire second macro-state data that characterizes the macro-activity pattern of the environment in which the electronic device is located. The second macro-state data is generated based on long-term statistical analysis of network traffic metadata generated by multiple electronic devices in the environment. The fusion decision module is configured to: based on a preset conditional probability model, fuse the first device event data and the second macroscopic state data to calculate the conditional probability of the instantaneous operating state occurring under the macroscopic activity mode; The indicator signal generation module is configured to: generate a behavioral inconsistency indicator signal based on the conditional probability to characterize the degree of causal consistency between the instantaneous operating state and the macroscopic activity pattern; The warning generation module is configured to generate an active warning signal when the value of the behavior inconsistency indication signal exceeds a preset warning threshold, provided that the first device event data is not triggered by a user instruction. The interactive decision module is configured to: upon receiving a user's ambiguous voice command, generate a control command for parsing the ambiguous voice command based on the behavior inconsistency indication signal, so as to prioritize the candidate operation with the highest causal consistency with the macroscopic activity pattern of the environment.

2. The smart home scenario-based audio linkage control system based on multimodal perception according to claim 1, characterized in that: The steps for obtaining the first device event data specifically include: By constructing an electromagnetic leakage feature spectrum library, the electromagnetic radiation signal collected in real time is matched with the pre-stored features in the electromagnetic leakage feature spectrum library to identify the device identity and instantaneous operating status of the electronic device; The steps for obtaining the second macroscopic state data specifically include: By calculating the information entropy of the network traffic metadata and using a hidden Markov model to model the time series of the information entropy, the macroscopic activity patterns of the environment can be identified. The steps for calculating conditional probability include: The conditional probability is calculated using a dynamically learned Bayesian inference model.

3. The smart home scene-based audio linkage control system based on multimodal perception according to claim 2, characterized in that: The electromagnetic leakage feature spectrum library is further defined as a pre-trained machine learning classification model; and the matching operation is further defined as a model inference process. The model inference process specifically includes: receiving real-time signal features extracted from the electromagnetic radiation signal, outputting a multi-dimensional micro-event feature vector characterizing the instantaneous operating state, and determining the device identity and the instantaneous operating state based on the multi-dimensional micro-event feature vector through a classification layer; The step of modeling the time series of the information entropy using a hidden Markov model is further configured to: output a macro-state probability distribution vector representing the macro-activity pattern, and determine the state with the maximum probability value in the macro-state probability distribution vector as the current macro-activity pattern; The step of calculating conditional probabilities is further configured as follows: the dynamically learned Bayesian inference model is a vectorized Bayesian network, which is configured to perform the fusion calculation of the micro-event feature vector and the macro-state probability distribution vector.

4. The smart home scene-based audio linkage control system based on multimodal perception according to claim 3, characterized in that: The fusion decision module is further configured to include a hierarchical fusion processing architecture, which includes: The logical rationality gating unit is configured to: perform preliminary rationality verification on the events represented by the micro-event feature vector based on the macro-state probability distribution vector, and generate a rationality adjustment coefficient; and, The dynamic weighted fusion unit is configured to receive the micro-event feature vector, the macro-state probability distribution vector, and the rationality adjustment coefficient, and calculate the conditional probability through a dynamically adjusted weighted network.

5. The smart home scene-based audio linkage control system based on multimodal perception according to claim 4, characterized in that: Before performing calculations, the dynamic weighted fusion unit is further configured as follows: Obtain a short-term activity context vector representing a recent sequence of events in the environment in which the electronic device is located; and obtain a periodic time context vector representing the current time; The weighted network in the dynamically weighted fusion unit is dynamically adjusted in real time based on the short-term activity context vector and the periodic time context vector.

6. The smart home scene-based audio linkage control system based on multimodal perception according to claim 5, characterized in that: The indicator signal generation module is specifically configured to define the magnitude of the behavior inconsistency indicator signal as being proportional to the negative logarithm of the conditional probability. Based on the time series of the behavioral inconsistency indication signal, a cumulative inconsistency score is generated through a time accumulation processing step, which represents the cumulative effect of the degree of inconsistency in recent history.

7. The smart home scene-based audio linkage control system based on multimodal perception according to claim 6, characterized in that: The early warning generation module is further configured to generate the active early warning signal through a two-stage hierarchical early warning logic, the two-stage hierarchical early warning logic including: The instantaneous attention gate is configured to: activate the second stage of the warning logic when the instantaneous value of the behavior inconsistency indication signal exceeds a preset attention threshold; and, The continuous confirmation gate is configured to generate the active warning signal when the cumulative inconsistency score exceeds a preset confirmation threshold, provided that the instantaneous attention gate is activated.

8. The smart home scene-based audio linkage control system based on multimodal perception according to claim 7, characterized in that: The interactive decision-making module is further configured as follows: For each potential candidate operation, obtain a priori weight representing the inherent priority of the candidate operation; determine the scene adaptability score based on the behavior inconsistency indication signal corresponding to the candidate operation; By fusing the prior weights and the scene adaptability score, a final decision score is generated for each potential candidate operation. Calculate the difference in decision scores between the first candidate operation with the highest final decision score and the second candidate operation with the second highest final decision score among the candidate operations.

9. The smart home scene-based audio linkage control system based on multimodal perception according to claim 8, characterized in that: When the difference in decision scores is greater than a preset ambiguity threshold, a control instruction is generated to directly execute the first candidate operation. When the difference in decision scores is not greater than the ambiguity threshold, a voice prompt is generated via an audio output device to request further clarification from the user.

Citation Information

Patent Citations

  • Audio control method, audio control device and audio device

    CN107357547A

  • Voice scene recognition method and device, voice control method and equipment and air conditioner

    CN109741747A

  • Intention category identification method and device

    CN111027667A

  • Network traffic classification method and related equipment

    CN120030481A

  • Multi-scene adaptive intelligent waterproof socket remote regulation and control system and method

    CN120540046A