User behavior intention prediction method and system for smart home
By extracting the voiceprint features and semantic features in the user's voice control signals, combining identity tags and historical behavior data, a dynamic analysis network is built, and a personalized problem of user behavior intention prediction in smart homes is solved, achieving more accurate and flexible intention prediction.
Patent Information
- Application Number
- CN202510478191.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-07-18
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing smart home user behavior intention prediction technology fails to fully consider individual language habits differences, resulting in misjudgment of semantic analytical models, and identity tags are not deeply integrated into the semantic understanding process, making it difficult to adapt to the expression habits of different users.
By collecting user voice control signals, extracting voiceprint features to determine identity tags, and combining voice control semantic features for binding expression, a dynamic analysis network based on user identity is built, and the current behavior intention and historical behavior sequence data are deeply integrated to generate the final behavior intention.
It improves the accuracy and robustness of user intention prediction, and can adapt to the expression patterns of different users in complex natural language scenarios, reducing intent misjudgment.
Smart Images

Figure CN120340482A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of smart home, and in the embodiments of this application, it relates to a method and system for predicting user behavior intention for smart home. Background Art
[0002] With the popularization of smart home systems, user behavior intention prediction technology has become the core to enhance the intelligent experience. Existing technologies usually build prediction models based on user historical behavior sequences and environmental data to infer users' potential needs. For example, Patent CN109818839B proposes a personalized behavior prediction method, device and system applied to smart home, which models the temporal correlation of the behavior sequence and the influence of the environmental context by using a recurrent neural network (RNN) and a deep learning classification model respectively by combining user identity, current behavior intention and historical behavior sequence, and finally determines the predicted behavior by fusing the output probabilities of the two sub-models. However, this solution still has significant defects in practical applications: existing methods rely on basic identity recognition technologies such as voice or fingerprint, but when parsing user voice commands, the differences in individual language habits are not fully considered. For example, different users may have different semantic references for the same command (such as "turn on the light") (such as specific room or lighting mode), and a general semantic parsing model is prone to misjudgment of intentions due to ambiguity. And although historical behaviors are matched by user identity in existing technologies, the identity label is not deeply integrated into the semantic understanding process. That is to say, the identity information is only used to filter historical data, but does not dynamically optimize the semantic parsing model, and it is difficult to adapt to the personalized expression habits of different users.
[0003] Therefore, an optimized solution for predicting user behavior intention for smart home is desired. Summary of the Invention
[0004] To solve the above technical problems, this application is proposed. The embodiments of this application provide a method and system for predicting user behavior intention for smart home, which collect user voice control signals and extract user voiceprint features to determine user identity labels, and simultaneously extract user voice control semantic features and perform constrained expression based on user identity labels to determine the current behavior intention. Then, obtain the historical behavior sequence data of the user and perform deep fusion and prediction in combination with the current behavior intention to obtain the user's final behavior intention. In this way, the identity label is deeply integrated into the semantic understanding process and adapts to the expression patterns of different users, thereby assisting in improving the accuracy and robustness of user intention prediction in complex natural language scenarios.
[0005] According to one aspect of this application, there is provided a method for predicting user behavior intention for smart home, which includes:
[0006] Collect and analyze the user voice control signals for smart home to generate the current behavior intention of the user, including: extracting the user voiceprint features from the user voice control signals and determining the user identity label based on the user voiceprint features; extracting the user voice control semantic features from the user voice control signals and determining the current behavior intention by performing a constrained expression on the user voice control semantic features based on the user identity label;
[0007] Obtain the historical behavior sequence data of the user;
[0008] Input the current behavior intention and the historical behavior sequence data into the first user behavior prediction sub-model to generate the first predicted user behavior intention;
[0009] Input the current behavior-environment data pair and the current environment data into the second user behavior prediction sub-model to generate the second predicted user behavior intention, where the current behavior-environment data pair includes the current behavior intention and the environmental information before collecting the user voice control signals;
[0010] Predict the final behavior intention of the user based on the first predicted user behavior intention and the second predicted user behavior intention.
[0011] According to another aspect of the present application, there is provided a user behavior intention prediction system for smart home, which includes:
[0012] A user current behavior intention collection and analysis module, configured to collect and analyze the user voice control signals for smart home to generate the current behavior intention of the user;
[0013] A user historical behavior sequence data acquisition module, configured to obtain the historical behavior sequence data of the user;
[0014] A first predicted user behavior intention generation module, configured to input the current behavior intention and the historical behavior sequence data into the first user behavior prediction sub-model to generate the first predicted user behavior intention;
[0015] A second predicted user behavior intention generation module, configured to input the current behavior-environment data pair and the current environment data into the second user behavior prediction sub-model to generate the second predicted user behavior intention, where the current behavior-environment data pair includes the current behavior intention and the environmental information before collecting the user voice control signals;
[0016] A user final behavior intention prediction module, configured to predict the final behavior intention of the user based on the first predicted user behavior intention and the second predicted user behavior intention.
[0017] Compared with the prior art, a user behavior intention prediction method and system for smart home provided by the present application collect user voice control signals, extract user voiceprint features to determine user identity tags, synchronously extract user voice control semantic features, and perform constrained expression based on the user identity tags to determine the current behavior intention. Then, historical behavior sequence data of the user is obtained and deeply fused with the current behavior intention for prediction to obtain the final behavior intention of the user. In this way, the identity tags are deeply integrated into the semantic understanding process and adapt to the expression patterns of different users, thereby assisting in improving the accuracy and robustness of user intention prediction in complex natural language scenarios. Description of the Drawings
[0018] The above and other objects, features, and advantages of the present application will become more apparent by describing the embodiments of the present application in more detail in conjunction with the accompanying drawings. The drawings are used to provide a further understanding of the embodiments of the present application and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the present application and do not constitute a limitation to the present application. In the drawings, the same reference numerals generally represent the same components or steps.
[0019] Figure 1 It is a flowchart of a user behavior intention prediction method for smart home according to an embodiment of the present application.
[0020] Figure 2 It is a schematic diagram of data flow of a user behavior intention prediction method for smart home according to an embodiment of the present application.
[0021] Figure 3 It is a flowchart of collecting and analyzing user voice control signals for smart home in a user behavior intention prediction method for smart home according to an embodiment of the present application to generate the current behavior intention of the user.
[0022] Figure 4 It is a flowchart of extracting user voiceprint features from user voice control signals and determining user identity tags based on the user voiceprint features in a user behavior intention prediction method for smart home according to an embodiment of the present application.
[0023] Figure 5 It is a flowchart of extracting user voice control semantic features from user voice control signals and performing constrained expression on the user voice control semantic features based on the user identity tags to determine the current behavior intention in a user behavior intention prediction method for smart home according to an embodiment of the present application.
[0024] Figure 6A flowchart for obtaining a user behavior intention representation vector from a sequence of the speech control semantic representation vectors and the user identity label encoding vector in the user speech control signal through a control semantic dynamic analysis network based on user identity constraints in the user behavior intention prediction method for smart home according to an embodiment of the present application.
[0025] Figure 7 A system block diagram of a user behavior intention prediction system for smart home according to an embodiment of the present application. Detailed implementation manners
[0026] Various exemplary embodiments, features and aspects of the present application will be described in detail below with reference to the accompanying drawings. The same reference numerals in the drawings denote elements having the same or similar functions. Although various aspects of the embodiments are shown in the drawings, the drawings are not necessarily drawn to scale unless otherwise specified.
[0027] With the wide use of smart home systems, one of the key technologies to enhance the intelligent experience is user behavior intention prediction. Most current solutions rely on analyzing the user's past behavior patterns and environmental information to build a prediction model to guess the user's potential needs. The patent (CN109818839B) introduces a personalized behavior prediction method and system for a smart home environment. The system uses the user's identity, current behavior intention, and past behavior records, and employs a recurrent neural network (RNN) and a deep learning classification model to process the temporal correlation of behaviors and the influence of the environmental background respectively, and finally combines the results of the two parts to predict the user's behavior. However, when parsing the user's voice commands, existing methods usually do not fully consider the differences in personal language habits. This means that even for the same instruction (such as "turn on the light"), different users may refer to different rooms or lighting settings, which may lead to misunderstandings by a general semantic understanding model. On the other hand, although existing technologies match the user's identity with their historical behaviors, these identity information do not deeply participate in the semantic understanding process. Here, the identity information is mainly used to select an appropriate historical data set, rather than being used to adjust the semantic parsing model in real time to better adapt to the unique expressions of different users. This limits the ability of the model to accurately identify the user's true intention.
[0028] To address the above technical problems, the present application proposes a user behavior intention prediction method for smart home. Figure 1 A flowchart of a user behavior intention prediction method for smart home according to an embodiment of the present application. Figure 2 A data flow diagram of a user behavior intention prediction method for smart home according to an embodiment of the present application. As Figure 1 and Figure 2As shown in the figure, the user behavior intention prediction method for smart home according to the embodiment of the present application includes: S110, collecting and analyzing the user voice control signal for smart home to generate the current behavior intention of the user; S120, obtaining the historical behavior sequence data of the user; S130, inputting the current behavior intention and the historical behavior sequence data into the first user behavior prediction sub-model to generate the first predicted user behavior intention; S140, inputting the current behavior-environment data pair and the current environment data into the second user behavior prediction sub-model to generate the second predicted user behavior intention, where the current behavior-environment data pair includes the current behavior intention and the environmental information before collecting the user voice control signal; S150, predicting the final behavior intention of the user based on the first predicted user behavior intention and the second predicted user behavior intention.
[0029] Figure 3 It is a flowchart of collecting and analyzing the user voice control signal for smart home to generate the current behavior intention of the user in the user behavior intention prediction method for smart home according to the embodiment of the present application. As Figure 3 As shown in the figure, in the above-mentioned user behavior intention prediction method for smart home, the step S110, collecting and analyzing the user voice control signal for smart home to generate the current behavior intention of the user, includes: S111, extracting the user voiceprint feature from the user voice control signal, and determining the user identity label based on the user voiceprint feature; S112, extracting the user voice control semantic feature from the user voice control signal, and determining the current behavior intention by performing a constrained expression on the user voice control semantic feature based on the user identity label.
[0030] In particular, during the process of identifying and predicting the current behavior intention of the user, the known user identity information can help improve the accuracy of semantic understanding. For example, different users may have different vocabulary habits, command sentence patterns, or preference settings. In the semantic understanding stage, the identity label can be used to adjust the semantic parsing model to make it more suitable for the language characteristics and habits of the current user, providing a basis for the intelligent and personalized control of the smart home. That is to say, in the technical solution of the present application, through the accurate identification of identity driven by voiceprint features and the dynamic optimization of voice control semantics, the personalized improvement of intention parsing is realized. Specifically, the user voiceprint feature is deeply fused with the voice semantic representation, and a dynamic analysis network with identity constraints is designed to make the semantic parsing process adapt to the expression patterns of different users, thereby assisting in improving the accuracy and robustness of user intention prediction in complex natural language scenarios.
[0031] Figure 4 It is a flowchart of extracting the user voiceprint feature from the user voice control signal and determining the user identity label based on the user voiceprint feature in the user behavior intention prediction method for smart home according to the embodiment of the present application. As Figure 4As shown, in the embodiment of the present application, the step S111, which extracts the user's voiceprint feature from the user voice control signal and determines the user identity tag based on the user voiceprint feature, includes: S1111, receiving the user voice control signal for the smart home; S1112, extracting the user voiceprint feature from the user voice control signal to obtain the user voiceprint feature vector as the user voiceprint feature; S1113, passing the user voiceprint feature vector through the identity tag recognizer based on the classifier to obtain the user identity tag.
[0032] Specifically, in the step S1111, the user voice control signal for the smart home is received. It should be understood that by receiving the user's voice commands, smart home devices can achieve functions such as turning on and off lights, adjusting the temperature, and controlling household appliances. These operations not only improve the convenience of home life but also bring more intelligent experiences. To improve the accuracy of voice recognition, modern smart home systems usually adopt technologies such as noise suppression, echo cancellation, and multi-microphone arrays to ensure that even in a noisy environment, the user's voice commands can be captured as accurately as possible. Specifically, the reception of the user voice control signal begins with the operation of the hardware device, which is usually completed by a microphone array or a voice sensor. The microphone in the smart home system captures the voice signal emitted by the user and converts it into an electronic signal through the sensor. This signal contains the user's voice characteristics, which are usually sound waves in different frequency ranges. By collecting the sound waves, the basic information about the user's commands can be obtained.
[0033] In an embodiment of the present application, the step S1112, which extracts the user's voiceprint features from the user voice control signal to obtain a user voiceprint feature vector as the user voiceprint feature, includes: passing the user voice control signal through a user voiceprint feature extractor based on TDNN to obtain the user voiceprint feature vector. It should be understood that in the prior art, although the user's identity is matched with historical behaviors, the identity recognition depends on basic technologies (such as fingerprints or simple voice features), and the personalized biometric features hidden in the voice signal are not fully explored, resulting in the problem of the disconnection between identity recognition and semantic parsing. That is to say, different users may have different semantic references (such as specific rooms or lighting modes) for the same instruction (such as "turn on the light"), and traditional general models are difficult to distinguish such personalized differences. Therefore, in the technical solution of the present application, voiceprint features are extracted from the user voice control signal to identify the user identity label. Specifically, the user voiceprint features are extracted from the user voice control signal to obtain a user voiceprint feature vector. In a specific example of the present application, advanced models such as time delay neural network (TDNN) or ECAPA-TDNN can be used to extract the voiceprint features, aiming to capture the user's unique acoustic patterns from the voice signal, such as fundamental frequency, formant distribution, and time-frequency characteristics. For example, ECAPA-TDNN can extract robust speaker embedding vectors from the voice signal through an enhanced channel attention mechanism and multi-layer feature aggregation, effectively distinguishing the voiceprint differences between family members. This process not only improves the accuracy of identity recognition, but also provides an identity coding vector for the subsequent semantic dynamic analysis network, enabling it to integrate the user's historical preferences (such as a user's habit of referring "turn on the light" to the bedroom light) when parsing voice instructions. Finally, the deep fusion of voiceprint features and semantic representations significantly reduces the intention misjudgment rate. Especially in the smart home scenario with multiple users coexisting, it can quickly lock the user's identity based on the voiceprint and accurately infer their behavioral intentions based on personalized semantic rules, thus achieving the leap from "general response" to "personalized service".
[0034] Specifically, in step S1113, the user voiceprint feature vector is passed through an identity label recognizer based on a classifier to obtain the user identity label. It should be understood that although historical behavior data is screened by user identity in the prior art, identity recognition relies on fingerprints or simple voice features, and fails to fully utilize the deep biological features (such as vocal tract structure and pronunciation habits) contained in the voice signal, resulting in a lack of dynamic interaction between the identity label and the semantic model. Therefore, in the technical solution of this application, the voiceprint feature vector is converted into a user identity label through a classifier and further one-hot encoded to solve the shallow coupling problem of identity information and semantic parsing in the prior art. Specifically, the user voiceprint feature vector is passed through an identity label recognizer based on a classifier to obtain the user identity label. Through an identity label recognizer based on a classifier (such as a support vector machine or a deep neural network classifier), the voiceprint feature vector (such as the user voiceprint feature vector extracted by TDNN) can be mapped to a specific user identity label. This process establishes a strong association between acoustic features and user identity from a large number of labeled voiceprint data through supervised learning, so as to accurately distinguish family members or different users. For example, when there are slight differences in the frequency domain distribution or time domain dynamic characteristics of the voiceprint features of user A and user B, the classifier can achieve high-precision identity determination through a non-linear decision boundary.
[0035] Figure 5 FIG. is a flowchart of extracting user voice control semantic features from a user voice control signal and determining the current behavior intention by performing a constrained expression on the user voice control semantic features based on the user identity label in the user behavior intention prediction method for a smart home according to an embodiment of the present application. As Figure 5 shown, in the embodiment of the present application, step S112, extracting user voice control semantic features from the user voice control signal and performing a constrained expression on the user voice control semantic features based on the user identity label to determine the current behavior intention, includes: S1121, performing one-hot encoding on the user identity label to obtain a user identity label encoding vector; S1122, converting the user voice control signal into a Log-Mel spectrogram to obtain a voice control signal Log-Mel spectrogram; S1123, performing speech control semantic feature extraction based on segment parsing on the voice control signal Log-Mel spectrogram to obtain a sequence of speech control semantic representation vectors as the user voice control semantic features; S1124, passing the sequence of speech control semantic representation vectors and the user identity label encoding vector through a control semantic dynamic analysis network based on user identity constraints to obtain a user behavior intention representation vector; S1125, determining the current behavior intention based on the user behavior intention representation vector.
[0036] Specifically, in step S1121, one-hot encoding is performed on the user identity tag to obtain a user identity tag encoding vector. It should be understood that the purpose of one-hot encoding the identity tag is to convert discrete categorical tags into orthogonal representations in a dense vector space, avoiding the spurious order relationships that may be introduced by numerical encoding (for example, the numerical proximity of user 1 and user 2 may mislead the model into thinking that the two are similar). The one-hot encoded identity tag vector (such as 1,0,01,0,0 representing user A and 0,1,00,1,0 representing user B) not only retains the uniqueness of the identity but can also be used as the input to the subsequent semantic dynamic analysis network for cross-modal fusion with the speech semantic representation vector. For example, in the control semantic dynamic analysis network based on user identity constraints, the identity encoding vector dynamically weights semantic features (such as the semantic reference of "turn on the light") through an attention mechanism, enabling the same instruction to generate differentiated intention representations according to the preferences of different users (for example, user A is accustomed to controlling the living room light, while user B prefers the bedroom light). This mechanism that deeply couples identity and semantics effectively solves the problem of intention ambiguity caused by the general model's neglect of individual expression habits, thus significantly improving the personalization level of behavior prediction in multi-user scenarios.
[0037] Specifically, in step S1122, the user voice control signal is converted into a Log-Mel spectrogram to obtain a voice control signal Log-Mel spectrogram. It should be understood that since existing solutions rely on general semantic parsing models to directly process raw speech or shallow features, it is difficult to capture the fine-grained semantic information hidden in the speech signal (such as the user's unique stress, speech rate, or keyword habits), resulting in a contradiction between the semantic ambiguity of voice commands and the user's personalized expression habits. Therefore, the user voice control signal is further converted into a Log-Mel spectrogram to obtain a voice control signal Log-Mel spectrogram, so as to map the time-frequency characteristics of the user voice control signal into a two-dimensional image representation that conforms to the human ear's perception characteristics. For example, the fundamental frequency and formant distribution of the speech are retained in the low-frequency band (such as 0 - 4 kHz), and at the same time, the high-frequency redundant information is compressed through a Mel filter bank, thereby highlighting the semantic key frequency band while reducing noise. The reason for this step is due to the complexity of natural language instructions - users may issue ambiguous instructions (such as the reference of "that" in "turn down that") under different environmental noises, and the frequency-domain representation of the spectrogram can separate the speech content from the interference components, providing a robust input basis for the subsequent model.
[0038] In an embodiment of the present application, step S1123, extracting speech control semantic features based on segment parsing from the Log-Mel spectrogram of the speech control signal to obtain a sequence of speech control semantic representation vectors as the user speech control semantic features, includes: S1123-1, decomposing the Log-Mel spectrogram of the speech control signal into segments to obtain a sequence of local Log-Mel spectrograms of the speech control signal; S1123-2, respectively passing each local Log-Mel spectrogram in the sequence of local Log-Mel spectrograms of the speech control signal through a speech control semantic feature extractor based on the VGGish network to obtain the sequence of speech control semantic representation vectors.
[0039] Specifically, in step S1123-1, the Log-Mel spectrogram of the speech control signal is decomposed into segments to obtain a sequence of local Log-Mel spectrograms of the speech control signal. It should be understood that decomposing the Log-Mel spectrogram of the speech control signal into segments (such as dividing it into a local frame sequence of 200 ms according to a time window) aims to solve the problem of semantic dynamic changes in long-term speech commands. For example, a user may mention multiple operations successively in a piece of speech (such as "turn on the light first, then adjust the air conditioner temperature"). By dividing the local spectrogram sequence, the model can parse the temporal dependence relationship segment by segment, avoiding the loss of key semantics caused by global feature pooling.
[0040] Specifically, in step S1123-2, each local Log-Mel spectrogram in the sequence of local Log-Mel spectrograms of the speech control signal is respectively passed through a speech control semantic feature extractor based on the VGGish network to obtain the sequence of speech control semantic representation vectors. It should be understood that particularly, each local Log-Mel spectrogram in the obtained sequence of local Log-Mel spectrograms of the speech control signal extracts semantic representation information through a pre-trained VGGish network to obtain the sequence of speech control semantic representation vectors, then utilizes the deep feature representation ability learned by this network in the audio event detection task. That is to say, VGGish captures hierarchical semantic patterns from the local spectrogram through stacked convolutional layers. For example, the underlying convolutional kernels identify phoneme-level acoustic units (such as the plosive / f / or the vowel / a: / ), and the high-level features are associated with semantic intents (such as the action category corresponding to "turn on"). The finally generated sequence of speech control semantic representation vectors not only encodes the global semantics of the speech command but also retains the context relevance in the time dimension. For example, in the user command "turn off the bedroom light but don't turn off the air conditioner", the semantics of the local spectrogram segments "turn off the bedroom light" and "don't turn off the air conditioner" are represented and temporally modeled, thereby achieving robust cross-intent prediction in a personalized scenario.
[0041] Figure 6 It is a flowchart for obtaining a user behavior intention representation vector from a user voice control signal in a user behavior intention prediction method for smart home according to an embodiment of the present application, where a sequence of the voice control semantic representation vectors and the user identity tag encoding vector are passed through a control semantic dynamic analysis network based on user identity constraints. As Figure 6As shown, in the embodiment of the present application, in step S1124, the sequence of the voice control semantic representation vectors and the user identity tag encoding vector are passed through a control semantic dynamic analysis network based on user identity constraints to obtain a user behavior intention representation vector, including: S1124-1, performing decision response analysis on each voice control semantic representation vector in the sequence of the user identity tag encoding vector and the voice control semantic representation vectors to obtain a set of implicit encoding vectors of the user identity-voice control semantic decision point state; S1124-2, constructing a Laplacian matrix of the set of implicit encoding vectors of the user identity-voice control semantic decision point state to obtain a user identity-voice control semantic decision point state Laplacian matrix; S1124-3, performing spectral decomposition on the user identity-voice control semantic decision point state Laplacian matrix to obtain a set of core component encoding vectors of the user identity-voice control semantic decision point; S1124-4, performing adaptive fusion on the set of core component encoding vectors of the user identity-voice control semantic decision point to obtain the user behavior intention representation vector. It should be understood that although historical behavior data is screened through identity tags in the prior art, the interaction between identity and semantics is limited to shallow associations (such as rule-based preference matching), and the complex dynamic relationship between the user's personalized expression habits and voice commands cannot be modeled. Therefore, in order to solve the limitation of the static coupling of identity information and semantic features in the prior art, in the technical solution of the present application, the sequence of the voice control semantic representation vectors and the user identity tag encoding vector are further passed through a control semantic dynamic analysis network based on user identity constraints to obtain a user behavior intention representation vector. By introducing a control semantic dynamic analysis network based on user identity constraints, it is possible to break through the traditional method from two dimensions: First, a non-linear decision response unit (such as a deep neural network) is used to capture the implicit association between the identity tag encoding feature and the voice control semantic representation sequence. For example, when the user identity tag encoding vector (such as the one-hot encoding of user A's identity tag) and the voice control semantic representation vector (such as the semantic reference of "turn on the light") are input into the network, the non-linear mapping layer can learn the interaction pattern between the two (such as user A's habit of associating "light" with the bedroom light), thereby constructing a personalized decision boundary in the latent space. Further, the control semantic dynamic analysis network based on user identity constraints explicitly models the graph structure relationship of the identity-semantic fusion feature by constructing a decision point state neighborhood matrix and a Laplacian matrix. The neighborhood matrix quantifies the local association strength between different semantic segments and the user identity (such as the "adjust temperature" command of user B mostly refers to the air conditioner rather than the heater in the historical data), and the spectral decomposition of the Laplacian matrix extracts the low-dimensional topological features of the data manifold (such as the common pattern of the user's specific intention). For example, in the spectral domain, user A's "turn on the light" command may share a low-dimensional embedding with the "draw the curtain" action in his historical behavior, revealing the potential structure of his night-time homecoming behavior chain.This topological modeling enables the semantic dynamic analysis of identity constraints to not only rely on surface statistics but also uncover implicit behavioral logic networks. Ultimately, through an adaptive fusion module (such as the attention mechanism) to dynamically weight the encoded vectors of the core components in the spectral domain, the network can adjust the feature weights according to the current environment and the user's real-time instructions. For example, in a low-light environment, the "turn on the light" instruction of user C may be assigned a higher spectral component weight, while the identity label encoding strengthens its directivity towards the study light. This fusion mechanism overcomes the problem of feature dilution caused by traditional linear weighting or simple concatenation, making the generated user behavior intention representation vector possess both identity sensitivity, semantic integrity, and environmental adaptability, thus achieving a dual improvement in the accuracy and generalization ability of intention prediction in multi-user and multi-scenario applications of smart homes.
[0042] Specifically, in step S1124-1, each voice control semantic representation vector in the sequence of the user identity label encoding vector and the voice control semantic representation vector is subjected to a decision responsiveness analysis to obtain a set of implicit encoding vectors of the user identity-voice control semantic decision point states, which is represented by the user identity-voice control decision responsiveness analysis formula as:
[0043] V2 = {v 21 , v 22 ,..., v2i,..., v 2n}
[0044]
[0045] V 1,2 = {v 1,21 , v 1,22 ,..., v 1,2i ,..., v 1,2n}
[0046] where V2 is the sequence of the voice control semantic representation vectors, v 21 , v 22 , v 2i and v 2n are respectively the 1st, 2nd, ith, and nth voice control semantic representation vectors in the sequence of the voice control semantic representation vectors, V1 is the user identity label encoding vector, is matrix multiplication, W i and b i are respectively the decision response weight matrix and the decision response bias vector corresponding to v 2i , sigmoid is the activation function, v 1,21 , v 1,22 , v 1,2i , v 1,2j and v 1,2nThe 1st, 2nd, i-th, j-th, and n-th user identity-voice control semantic decision point state implicit coding vectors in the set of user identity-voice control semantic decision point state implicit coding vectors, respectively, V 1,2 is the set of user identity-voice control semantic decision point state implicit coding vectors. It should be understood that in traditional smart homes, the combination of identity tags and voice semantic features often remains at a shallow association level. However, the dynamic association between users' personalized expression habits and voice commands has highly non-linear characteristics, and different identity users may have semantic pointing differences for the same command (e.g., "turn on the light" points to different rooms). Existing linear fusion methods are difficult to capture the complex interaction patterns between identity tags and semantic features, resulting in blurred key decision boundaries during the intent parsing process. In addition, voice control commands in natural language scenarios are often accompanied by interference factors such as environmental noise and semantic ambiguity, further exacerbating the difficulty of modeling the dynamic relationship between identity and semantics. Therefore, introducing non-linear decision responsiveness analysis becomes a method to break through the limitations of traditional methods. By constructing a hidden space that can represent the deep interaction between identity tags and semantic features, and through non-linear mapping layers (such as multi-layer perceptrons or attention mechanisms), the cross-modal feature interaction between the user identity tag coding vector and the voice control semantic representation vector can be carried out, forcing the model to learn the dynamic decision response rules between the two. During this process, the identity tag serves as a prior condition to guide the dynamic reconstruction of semantic features. For example, user A's "adjust the temperature" command is mapped to the air conditioner control semantic cluster in the hidden space, while user B's same-name command is associated with the floor heating device. The construction of the hidden space decouples the mixed identity correlation and semantic pointing in the original high-dimensional features, and the originally overlapping decision regions in the Euclidean space (such as the different definitions of "turn off the device" by different users) are transformed into linearly separable low-dimensional subspaces, thus providing a more discriminative intermediate representation for subsequent intent prediction. The set of user identity-voice control semantic decision point state implicit coding vectors generated through decision responsiveness analysis realizes the optimized reconstruction of identity-sensitive semantic representations.
[0047] In an embodiment of the present application, step S1124-2 of constructing the Laplacian matrix of the user identity-voice control semantic decision point state for the set of user identity-voice control semantic decision point state implicit coding vectors includes: S1124-21, calculating the user identity-voice control semantic decision point state class neighborhood matrix based on the set of user identity-voice control semantic decision point state implicit coding vectors; S1124-22, calculating the user identity-voice control semantic decision point state class degree matrix based on the set of user identity-voice control semantic decision point state implicit coding vectors; S1124-23, calculating the Laplacian matrix of the user identity-voice control semantic decision point state based on the user identity-voice control semantic decision point state class neighborhood matrix and the user identity-voice control semantic decision point state class degree matrix.
[0048] Specifically, step S1124-21 of calculating the user identity-voice control semantic decision point state class neighborhood matrix based on the set of user identity-voice control semantic decision point state implicit coding vectors is represented by the user identity-voice control semantic decision point state class neighborhood matrix calculation formula as:
[0049]
[0050] where ||·||2 is the two-norm of the calculation vector, arccosh is the inverse hyperbolic cosine function, A 11 , A n1 , A 1n , A ij and A nnare the eigenvalues at each position in the neighborhood matrix of the user identity-voice control semantic decision point state class. A is the neighborhood matrix of the user identity-voice control semantic decision point state class. It should be understood that although the set of latent encoding vectors of the user identity-voice control semantic decision point state has captured feature interactions through non-linear mapping, its high-dimensional nature still obscures the local topological structure of the data manifold. In the original latent encoding space of the user identity-voice control semantic decision point state, the semantic correlation degree between different decision points lacks quantitative representation, resulting in the subsequent intent prediction model being unable to accurately model the dynamic evolution path of the user's behavior intention in the feature space. Therefore, it is necessary to use mathematical tools to transform the spatial distribution of the latent encoding vectors of the user identity-voice control semantic decision point state into a computable graph structure, so as to explicitly reveal the local association characteristics of the identity-semantic decision points. By constructing a graph structure model representing the local association strength of the identity-semantic decision points to calculate the neighborhood matrix of the user identity-voice control semantic decision point state class, the complex relationship between the high-dimensional latent encoding vectors of the user identity-voice control semantic decision point state can be transformed into discretized adjacency weights, where the matrix element values quantify the semantic correlation strength between two decision points in the feature space (for example, the "turn on the light" instruction of user C and the latent encoding vector of the user identity-voice control semantic decision point state of "activate the night light mode" in his historical behavior have a high adjacency weight). This transformation enables the manifold structure originally hidden in the data distribution (such as sub-clusters formed by user-specific intents, cross-user semantic common patterns) to be explicitly encoded as the connection relationship between graph nodes, providing a structured input for subsequent intent feature decoupling based on spectral analysis. The construction of the neighborhood matrix of the user identity-voice control semantic decision point state class essentially completes the mapping from the continuous feature space to the discrete graph topology space, enabling the generation mechanism of the user's behavior intention to be analyzed through graph theory tools.
[0051] Specifically, in step S1124-22, based on the set of latent encoding vectors of the user identity-voice control semantic decision point state, calculate the degree matrix of the user identity-voice control semantic decision point state, which is expressed by the formula for calculating the degree matrix of the user identity-voice control semantic decision point state:
[0052]
[0053] where is the square of the calculation of the vector one-norm, n is the number of vectors in V minus one, D1, D 1,2 in the middle of the vector, D i and D nThey are the eigenvalues at each position on the diagonal of the user identity-voice control semantic decision point state degree matrix, and D is the user identity-voice control semantic decision point state degree matrix. It should be understood that after constructing the user identity-voice control semantic decision point state neighborhood matrix, the connection strength distribution of the nodes in the graph structure still lacks a normalized measure, resulting in difficulty in accurately characterizing the global characteristics of the data manifold in subsequent spectral analysis. When traditional graph theory methods directly use the adjacency matrix to calculate the user identity-voice control semantic decision point state Laplacian matrix, they do not consider the cumulative effect of the correlation degree between nodes on the graph structure (such as the problem of excessive dominance of high-frequency interaction user nodes), resulting in spectral features being vulnerable to noise interference. In the set of original user identity-voice control semantic decision point state implicit coding vectors, the connection density differences of different decision points are significant (for example, the "turn on the light" instruction of user A is associated with 10 historical operation nodes, while the new user instruction is only associated with 2 nodes). This non-uniform distribution will distort the mathematical representation of the internal structure of the data, and it is necessary to quantitatively calibrate the node connection strength through the user identity-voice control semantic decision point state degree matrix. By constructing a diagonal form of the user identity-voice control semantic decision point state degree matrix, the topological influence of each decision point in the graph structure can be quantitatively measured in a mathematical way. Among them, aggregating and calculating the node connection degree (such as counting the sum of the association weights between the "adjust temperature" instruction node of user B and all other nodes) can form the diagonal matrix element value reflecting the node centrality. The user identity-voice control semantic decision point state degree matrix essentially completes the transformation of the graph structure from the original connection relationship to the probabilistic transition relationship, providing structural stability guarantee for the intention feature extraction based on graph convolution.
[0054] Specifically, in step S1124-23, based on the user identity-voice control semantic decision point state neighborhood matrix and the user identity-voice control semantic decision point state degree matrix, calculate the user identity-voice control semantic decision point state Laplacian matrix, which is expressed by the user identity-voice control semantic decision point state Laplacian matrix calculation formula as:
[0055] L = D - A
[0056] Among them, L is the Laplacian matrix of the user identity-voice control semantic decision point state. It should be understood that after constructing the neighborhood matrix and degree matrix of the user identity-voice control semantic decision point, the local connection relationship and node importance information of the graph structure are still in a separated state, and the global spectral characteristics of the data manifold cannot be directly revealed. The original neighborhood matrix only reflects the direct association strength between nodes, while the degree matrix isolates and represents the connection density of nodes themselves. The two lack a collaborative analysis mechanism under a unified mathematical framework. Therefore, it is necessary to fuse the topological information of the two types of matrices into a linear algebra operator by constructing the Laplacian matrix of the user identity-voice control semantic decision point state, providing a structured mathematical basis for the frequency-domain analysis of graph signals. The construction of the Laplacian matrix of the user identity-voice control semantic decision point state essentially realizes the mathematical mapping of the graph structure from the discrete topological relationship to the continuous spectral domain space. Its symmetric semi-positive definite property ensures that the eigen-decomposition can extract the frequency-domain components of the graph signal (low frequencies correspond to smooth semantic patterns across nodes, and high frequencies reflect local specific perturbations).
[0057] In the embodiment of the present application, in step S1124-3, performing spectral decomposition on the Laplacian matrix of the user identity-voice control semantic decision point state to obtain a set of core component coding vectors of the user identity-voice control semantic decision point, including: S1124-31, performing matrix structure optimization on the Laplacian matrix of the user identity-voice control semantic decision point state to obtain an optimized Laplacian matrix of the user identity-voice control semantic decision point state; S1124-32, performing spectral decomposition on the optimized Laplacian matrix of the user identity-voice control semantic decision point state to obtain the set of core component coding vectors of the user identity-voice control semantic decision point.
[0058] Specifically, in step S1124-31, performing matrix structure optimization on the Laplacian matrix of the user identity-voice control semantic decision point state to obtain an optimized Laplacian matrix of the user identity-voice control semantic decision point state. Specifically, first, pseudo-inverse operation is used to extract the connected component information of the neighborhood matrix of the user identity-voice control semantic decision point state and the degree matrix of the user identity-voice control semantic decision point state, that is, let:
[0059]
[0060] Among them, matrices M1 and M2 can respectively represent the partition matrix and loop matrix of the Laplacian matrix L of the user identity-voice control semantic decision point state.
[0061] Then, the Laplacian matrix of the decision point state is optimized through cut space closure operation and cycle space closure operation, expressed as:
[0062]
[0063] where exp is the exponential function value with the natural constant e as the base, is point addition by position, and L' is the Laplacian matrix that optimizes the user identity - voice control semantic decision point state.
[0064] In this way, by performing a spatial closed operation based on the pseudo - inverse to extract the global structure invariants in the graph structure information, the spectral topology information representation of the user identity - voice control semantic decision point state Laplacian matrix L becomes more robust.
[0065] Specifically, for the step S1124 - 32, spectral decomposition is performed on the optimized user identity - voice control semantic decision point state Laplacian matrix to obtain a set of user identity - voice control semantic decision point core component coding vectors, which is represented by the user identity - voice control spectral decomposition formula as:
[0066]
[0067] where Spectral Decomposition(L') is the operation of performing spectral decomposition on L', U is the user identity - voice control semantic feature matrix, that is, a set of user identity - voice control semantic decision point core component coding vectors, x1, x2, x i and x i are respectively the 1st, 2nd, i - th, and k - th user identity - voice control semantic decision point core component coding vectors in the set of user identity - voice control semantic decision point core component coding vectors, diag(λ1, λ2, …, λ i …, λ k ) is the user identity - voice control semantic eigenvalue diagonal matrix with elements λ1, λ2, λ i and λ k on the diagonal, and λ1, λ2, λ i and λ k are respectively the eigenvalues corresponding to x1, x2, x i and x kThe corresponding eigenvalue, and Λ is the diagonal matrix of user identity - voice control semantic eigenvalue. It should be understood that after obtaining the optimized Laplacian matrix of the user identity - voice control semantic decision point state, its high - dimensional algebraic form is still difficult to directly represent the essential structural characteristics of the data manifold. Although the optimized Laplacian matrix of the user identity - voice control semantic decision point state encodes the complete topological information of the graph structure, its implicit frequency - domain characteristics (low - frequency corresponding to smooth semantic patterns and high - frequency reflecting random perturbations) need to be decoupled and extracted through mathematical tools to meet the requirements of robust intention representation in the smart home scenario. By spectral decomposition, the optimized Laplacian matrix of the user identity - voice control semantic decision point state is transformed into a set of basis functions that can be analyzed in the frequency domain, and a low - dimensional embedding vector representing the core structure of the data manifold can be extracted. Through eigenvalue - eigenvector decomposition, the Laplacian matrix of the user identity - voice control semantic decision point state is mapped to the spectral domain space, where the eigenvector corresponding to the smallest eigenvalue captures the principal components of the global data distribution (such as the "leaving home security" instruction cluster shared by multiple users), and the eigenvectors associated with larger eigenvalues represent local perturbations (such as mis - instructions caused by environmental noise).
[0068] Specifically, in step S1124 - 4, the set of encoded vectors of the core components of the user identity - voice control semantic decision point is adaptively fused to obtain the user behavior intention representation vector, and the adaptive fusion is expressed as:
[0069]
[0070] where AF(U) is the adaptive fusion operation on U, W i and b i are the fusion weight matrix and fusion bias vector corresponding to x i respectively, softmax is the softmax function, is the user identity - voice control semantic scoring weight vector corresponding to x i a i is the user identity - voice control semantic weight value corresponding to x i mask is the masking operation, τ is the preset threshold, w i is the user identity - voice control semantic masking weight value corresponding to x i v fIt is the user behavior intention representation vector. It should be understood that after the spectral decomposition of the core component coding vector of the user identity-voice control semantic decision point, although the set of core component coding vectors of the user identity-voice control semantic decision point contains the key structural information of the data manifold, the contribution degrees of different core component coding vectors of the user identity-voice control semantic decision point to the final intention representation vary significantly due to the dynamic changes in the scenario. Traditional static weighted fusion methods (such as average pooling or fixed-weight linear combination) cannot meet the requirements of the ever-changing smart home scenarios. For example, the "adjust temperature" instruction of the same user may be associated with air conditioner cooling in summer, while it points to floor heating in winter, making it difficult for a fixed fusion strategy to capture context-sensitive semantic associations. In addition, high-frequency noise (such as abnormal instruction fluctuations caused by environmental interference) and low-frequency effective signals (such as the stable behavior patterns of users) may be mixed in the spectral domain core components, and a dynamic mechanism is needed to distinguish and suppress the interference of irrelevant components on the intention representation. By using the attention mechanism or the gating network, the weight distribution of each component vector can be automatically learned according to the real-time input (such as the current environmental data, the semantic focus of the user voice instruction). For example, when recognizing the user's "activate night mode" instruction, the model will assign higher weights to the low-frequency core components such as "dim the lights" and "mute the air conditioner" that are frequently associated in the historical operations, while suppressing the components related to the "morning wake-up mode" that conflict with the current timestamp. This dynamic feature combination strategy not only realizes noise filtering and key signal enhancement, but also adapts to the behavior preference differences of different user groups (such as elderly users preferring stable operation clusters and young users tending to use scenario-based linkage instructions), thereby generating a user behavior intention representation vector with strong discriminative power.
[0071] In an embodiment of the present application, step S1125, determining the current behavior intention based on the user behavior intention representation vector, includes: passing the user behavior intention representation vector through a user behavior intention parser based on a large model to obtain the current behavior intention. It should be understood that by introducing a user behavior intention parser based on a large model, deeper semantic associations and logical structures can be mined from the user behavior intention representation vector. For example, the large model performs global context modeling on the user behavior intention representation vector through the self-attention mechanism, identifying implicit temporal order (such as "first... then..."), conditional logic (such as "turn on the air conditioner if the temperature exceeds 28 degrees"), or referential relationships (such as the specific reference of "that room"), thereby mapping the abstract representation vector into specific executable operations (such as "turn on the bedroom air conditioner and set it to 26 degrees"). The reason for performing this step stems from the polysemy and dynamism of natural language - the intention of the same instruction may be completely different for different users or scenarios, and the large model, with its massive pre-trained knowledge base and fine-tuning ability, can dynamically adapt to user personal habits (such as user A's "brighten" defaulting to increasing the brightness by 30%, while user B prefers 50%). The large model can integrate temporal behavior patterns, making the parsing result of the current behavior intention not only conform to the immediate instruction but also fit the user's long-term habits, thus achieving a leap from "mechanical response" to "contextual intelligent decision-making" in the smart home scenario. Specifically, the current behavior intention is comprehensive information containing multiple factors such as the user's historical behavior pattern, the current voice instruction, and environmental data. Based on the user behavior intention representation vector, the user's behavior intention can be accurately judged, and corresponding actions can be taken, such as turning on electrical appliances, adjusting settings, or triggering scene modes, etc.
[0072] In the above user behavior intention prediction method for smart home, in step S120, historical behavior sequence data of the user is obtained. It should be understood that by analyzing the historical behavior sequence data of the user, the user's preferences, needs and habits can be understood more deeply, so as to make a more intelligent and accurate response when the user issues a new voice control instruction. The historical behavior sequence data can not only help accumulate the user's behavior patterns, but also provide strong support when facing complex environmental changes or diverse voice control instructions. In a smart home, every operation of the user, whether through voice, gesture, touch screen, etc., will leave a behavior record. These records contain the interaction information between the user and the device, reflecting the user's preferences and needs at different times and in different situations. For example, the user frequently adjusts the light brightness within a specific time period, or is more inclined to turn on the air conditioner in a certain specific room. These behavior data provide valuable reference bases to help analyze the user's living habits and behavior patterns. Moreover, with the gradual popularization of smart home, the user's needs become increasingly diverse and have strong personalized characteristics. Simply relying on the current voice input information often makes it difficult to accurately judge the user's true intention. At this time, the historical behavior sequence data can play a huge role. By continuously learning and analyzing the user's historical behavior, the user's possible future behavior can be inferred based on the user's past behavior patterns. For example, if a certain user automatically turns on the kitchen light and coffee machine at 7 o'clock every morning, and turns off the living room light and adjusts the air conditioner temperature at 9 o'clock in the evening, then by learning these historical behavior data, these tasks can be automatically executed at similar time points in the future, and even when the user issues a similar instruction, a response consistent with the past can be made, thus improving the response speed and accuracy of voice control.
[0073] In the above user behavior intention prediction method for smart home, in step S130, the current behavior intention and historical behavior sequence data are input into the first user behavior prediction sub-model to generate the first predicted user behavior intention. It should be understood that the current behavior intention generally refers to the behavior signal or instruction issued by the user at a specific moment. These instructions often reflect the user's needs or goals at the current moment. For example, the intention expressed by the voice control instructions "turn on the air conditioner" or "turn off the bedroom light". The acquisition of the current behavior intention is usually based on the user's immediate input, which can accurately capture the user's specific needs in the short term. However, relying solely on the current behavior intention to predict future behavior may have certain limitations. The user's needs and intentions are often affected by their past behavior patterns, and these behaviors may have a certain time dependence or periodicity. Therefore, relying solely on the currently input data cannot fully reflect the user's behavior characteristics. The historical behavior sequence data refers to the records of the behaviors performed by the user over a period of time in the past. Historical data can usually reveal information such as the user's behavior patterns, preferences, and living habits. For example, whether the user is accustomed to turning on the air conditioner every morning, whether they like to turn off the lights at night, or whether they often adjust the temperature of the home environment at specific times. By analyzing the historical behavior sequence, the user's preferences in different situations can be understood, and the user's future behavior can be inferred based on these patterns. The historical behavior sequence data is often a dataset with a long time span, containing the user's behavior trajectories under various environmental changes and time changes. Therefore, it is crucial for understanding the user's long-term behavior characteristics and preferences. Combining the current behavior intention with the historical behavior sequence data and inputting them into the first user behavior prediction sub-model can make full use of the complementarity of these two types of information. Specifically, in a specific embodiment of the present application, when the historical behavior sequence and the current behavior intention are input into the first user behavior prediction sub-model simultaneously, the complex relationship between the two is modeled by this sub-model. The first user behavior prediction sub-model is constructed based on a deep learning framework, and it can handle the non-linear relationships and temporal dependencies in the input data. The model will combine the temporal information in the historical behavior data to identify the long-term trends of the user's behavior and the immediate responses in specific situations. In this way, the user's next behavior intention can be more accurately inferred, and a prediction result that fits the current situation and the user's historical behavior pattern can be generated. Figure 1 When input into the first user behavior prediction sub-model together, the complex relationship between the two is modeled by this sub-model. The first user behavior prediction sub-model is constructed based on a deep learning framework, and it can handle the non-linear relationships and temporal dependencies in the input data. The model will combine the temporal information in the historical behavior data to identify the long-term trends of the user's behavior and the immediate responses in specific situations. In this way, the user's next behavior intention can be more accurately inferred, and a prediction result that fits the current situation and the user's historical behavior pattern can be generated.
[0074] In the above user behavior intention prediction method for smart home, in step S140, the current behavior-environment data pair and the current environment data are input into the second user behavior prediction sub-model to generate the second predicted user behavior intention. Herein, the current behavior-environment data pair includes the current behavior intention and the environmental information before collecting the user voice control signal. It should be understood that in the current behavior-environment data pair, the environmental information and the behavior intention are intertwined. Usually, the change of the environment will directly affect the user's behavior decision-making, and the expression of the behavior intention is a reaction of the user to the environmental state. Therefore, by inputting these two types of information, the second user behavior prediction sub-model can better capture the interaction relationship between the environmental change and the user's needs, so as to generate a more accurate prediction result.
[0075] In the above user behavior intention prediction method for smart home, in step S150, based on the first predicted user behavior intention and the second predicted user behavior intention, the final user behavior intention is predicted. It should be understood that the first predicted user behavior intention generation module mainly relies on the fusion of the current user behavior intention and the user historical behavior sequence data. Through the first user behavior prediction sub-model, based on large-scale data analysis, a preliminary prediction of the current behavior intention can be obtained. This prediction is not just a simple speculation of the user intention at the current moment, but through pattern recognition and modeling of their historical behaviors, a dynamic and real-time prediction result is obtained, with strong adaptability. On the basis of the generation of the first prediction model, the second predicted user behavior intention generation module further conducts refined analysis by introducing the current environmental data. The uniqueness of this module lies in that it not only considers the current behavior intention of the user, but also incorporates environmental factors into the consideration scope, taking into account richer external factors. The advantage of this model is that it can more accurately simulate the possible behavioral responses of the user in a complex environment, providing more dimensions and background information for the final behavior intention prediction. The user behavior intentions generated by the first prediction and the second prediction are both analyzed and calculated in their respective independent sub-models, focusing on the mining of historical behavior data and the impact of environmental changes respectively. Although these two prediction modules have different focuses, their combination provides a more comprehensive and accurate basis for the final user behavior intention prediction. The prediction of the final user behavior intention is achieved by fusing and weighting the prediction results from these two different sources. This process is not just a simple combination, but an optimized fusion process, which involves the in-depth cooperation of various algorithms and technologies. By fusing the results obtained from the first prediction and the second prediction, the errors and biases that may exist in a single prediction can be eliminated, thereby obtaining a more accurate prediction result that meets the actual needs of the user. For example, when the user historical behavior data is relatively rich and the environmental change is small, the first prediction may occupy a higher weight because the historical behavior data has strong stability and can provide a more accurate prediction for the current behavior intention. While in the case of drastic environmental changes, the second prediction may occupy a higher weight because the environmental factors have a significant impact on the user behavior and can provide more reference value for the user's behavior intention. After such weighted fusion, a final prediction result of the user behavior intention will be output. This prediction result not only considers the user's historical behavior and the current environmental information, but also eliminates the possible biases in different prediction models through a multi-level fusion process. Therefore, the final behavior intention prediction is more in line with the actual needs and expectations of the user, and has strong robustness and adaptability.
[0076] In summary, the user behavior intention prediction method for smart home based on the embodiments of the present application is elucidated. It collects user voice control signals and extracts user voiceprint features to determine user identity tags, synchronously extracts user voice control semantic features, and performs constrained expression based on the user identity tags to determine the current behavior intention. Then, it obtains the historical behavior sequence data of the user, performs deep fusion with the current behavior intention, and predicts to obtain the final user behavior intention. In this way, the identity tags are deeply integrated into the semantic understanding process and adapt to the expression patterns of different users, thereby assisting in improving the accuracy and robustness of user intention prediction in complex natural language scenarios.
[0077] Figure 7 FIG. is a system block diagram of a user behavior intention prediction system for smart home according to an embodiment of the present application. As Figure 7 shown, the user behavior intention prediction system 100 for smart home according to an embodiment of the present application includes: a user current behavior intention collection and analysis module 110, configured to collect and analyze user voice control signals for smart home to generate the current behavior intention of the user; a user historical behavior sequence data acquisition module 120, configured to acquire the historical behavior sequence data of the user; a first predicted user behavior intention generation module 130, configured to input the current behavior intention and the historical behavior sequence data into a first user behavior prediction sub-model to generate a first predicted user behavior intention; a second predicted user behavior intention generation module 140, configured to input the current behavior-environment data pair and the current environment data into a second user behavior prediction sub-model to generate a second predicted user behavior intention, where the current behavior-environment data pair includes the current behavior intention and the environmental information before collecting the user voice control signal; and a user final behavior intention prediction module 150, configured to predict the final user behavior intention based on the first predicted user behavior intention and the second predicted user behavior intention.
[0078] Here, those skilled in the art can understand that the specific operations of each step in the above user behavior intention prediction system for smart home have been described in detail in the description of the user behavior intention prediction method for smart home above, and therefore, the repeated description thereof will be omitted. Figures 1 to 6 of the user behavior intention prediction method for smart home, and thus, the repeated description thereof will be omitted.
[0079] In summary, the user behavior intention prediction system for smart home based on the embodiments of the present application is elucidated. It collects user voice control signals and extracts user voiceprint features to determine user identity tags, synchronously extracts user voice control semantic features, and performs constrained expression based on the user identity tags to determine the current behavior intention. Then, it obtains the historical behavior sequence data of the user, performs deep fusion with the current behavior intention and prediction to obtain the final behavior intention of the user. In this way, the identity tags are deeply integrated into the semantic understanding process and adapt to the expression patterns of different users, thereby assisting in improving the accuracy and robustness of user intention prediction in complex natural language scenarios.
Claims
1. A method for predicting user behavior intention for smart home, characterized in that, Including: Collect and analyze user voice control signals for smart home to generate the current behavior intention of the user, including: extracting the user voiceprint feature from the user voice control signal and determining the user identity label based on the user voiceprint feature; extracting the user voice control semantic feature from the user voice control signal and determining the current behavior intention by performing a constrained expression on the user voice control semantic feature based on the user identity label; Obtain the historical behavior sequence data of the user; Input the current behavior intention and the historical behavior sequence data into the first user behavior prediction sub-model to generate the first predicted user behavior intention; Input the current behavior-environment data pair and the current environment data into the second user behavior prediction sub-model to generate the second predicted user behavior intention, where the current behavior-environment data pair includes the current behavior intention and the environmental information before collecting the user voice control signal; Predict the final behavior intention of the user based on the first predicted user behavior intention and the second predicted user behavior intention.
2. The method for predicting user behavior intention for smart home according to claim 1, wherein Extracting the user voiceprint feature from the user voice control signal and determining the user identity label based on the user voiceprint feature, including: Receiving the user voice control signal for smart home; Extracting the user voiceprint feature from the user voice control signal to obtain the user voiceprint feature vector as the user voiceprint feature; Passing the user voiceprint feature vector through an identity label recognizer based on a classifier to obtain the user identity label.
3. The method for predicting user behavior intention for smart home according to claim 2, wherein Extracting the user voiceprint feature from the user voice control signal to obtain the user voiceprint feature vector as the user voiceprint feature, including: passing the user voice control signal through a user voiceprint feature extractor based on TDNN to obtain the user voiceprint feature vector.
4. The method for predicting user behavior intention for smart home according to claim 3, wherein, Extracting the user voice control semantic feature from the user voice control signal and determining the current behavior intention by performing a constrained expression on the user voice control semantic feature based on the user identity label, including: Performing one-hot encoding on the user identity label to obtain the user identity label encoding vector; Converting the user voice control signal into a Log-Mel spectrogram to obtain the voice control signal Log-Mel spectrogram; Performing voice control semantic feature extraction based on segment parsing on the voice control signal Log-Mel spectrogram to obtain a sequence of voice control semantic representation vectors as the user voice control semantic feature; Passing the sequence of voice control semantic representation vectors and the user identity label encoding vector through a control semantic dynamic analysis network based on user identity constraint to obtain the user behavior intention representation vector; Determining the current behavior intention based on the user behavior intention representation vector.
5. The user behavior intention prediction method for smart home according to claim 4, wherein Performing voice control semantic feature extraction based on segment parsing on the voice control signal Log-Mel spectrogram to obtain a sequence of voice control semantic representation vectors as the user voice control semantic feature, including: Performing segment decomposition on the voice control signal Log-Mel spectrogram to obtain a sequence of local voice control signal Log-Mel spectrograms; Respectively pass each local Log-Mel spectrogram of the sequence of the voice control signal local Log-Mel spectrograms through a voice control semantic feature extractor based on the VGG ish network to obtain a sequence of the voice control semantic representation vectors.
6. The method for predicting user behavior intention for smart home according to claim 5, characterized in that, Pass the sequence of the voice control semantic representation vectors and the user identity label encoding vector through a control semantic dynamic analysis network based on user identity constraints to obtain a user behavior intention representation vector, including: Perform a decision responsiveness analysis on the user identity label encoding vector and each voice control semantic representation vector in the sequence of the voice control semantic representation vectors to obtain a set of user identity-voice control semantic decision point state implicit encoding vectors; Construct a Laplacian matrix of the set of user identity-voice control semantic decision point state implicit encoding vectors to obtain a user identity-voice control semantic decision point state Laplacian matrix; Perform spectral decomposition on the user identity-voice control semantic decision point state Laplacian matrix to obtain a set of user identity-voice control semantic decision point core component encoding vectors; Perform adaptive fusion on the set of user identity-voice control semantic decision point core component encoding vectors to obtain the user behavior intention representation vector.
7. The user behavior intention prediction method for smart home according to claim 6, wherein Construct a user identity-voice control semantic decision point state Laplacian matrix of the set of user identity-voice control semantic decision point state implicit encoding vectors, including: Based on the set of user identity-voice control semantic decision point state implicit encoding vectors, calculate a user identity-voice control semantic decision point state class neighborhood matrix; Based on the set of user identity-voice control semantic decision point state implicit encoding vectors, calculate a user identity-voice control semantic decision point state degree matrix; Based on the user identity-voice control semantic decision point state class neighborhood matrix and the user identity-voice control semantic decision point state degree matrix, calculate the user identity-voice control semantic decision point state Laplacian matrix.
8. The method for predicting user behavior intention for smart home according to claim 7, characterized in that, Perform spectral decomposition on the user identity-voice control semantic decision point state Laplacian matrix to obtain a set of user identity-voice control semantic decision point core component encoding vectors, including: Perform matrix structure optimization on the user identity-voice control semantic decision point state Laplacian matrix to obtain an optimized user identity-voice control semantic decision point state Laplacian matrix; Perform spectral decomposition on the optimized user identity-voice control semantic decision point state Laplacian matrix to obtain the set of user identity-voice control semantic decision point core component encoding vectors.
9. The method for predicting user behavior intention for smart home according to claim 8, wherein, Based on the user behavior intention representation vector, determine the current behavior intention, including: Pass the user behavior intention representation vector through a user behavior intention parser based on a large model to obtain the current behavior intention.
10. A user behavior intention prediction system for smart home, characterized in that, Including: A user current behavior intention acquisition and analysis module, configured to collect and analyze user voice control signals for a smart home to generate the user's current behavior intention; A user historical behavior sequence data acquisition module, configured to acquire the user's historical behavior sequence data; The first predicted user behavior intention generation module is used to input the current behavior intention and historical behavior sequence data into the first user behavior prediction sub-model to generate the first predicted user behavior intention; The second predicted user behavior intention generation module is used to input the current behavior-environment data pair and the current environment data into the second user behavior prediction sub-model to generate the second predicted user behavior intention, where the current behavior-environment data pair includes the current behavior intention and the environmental information before collecting the user voice control signal; The user's final behavior intention prediction module is used to predict the user's final behavior intention based on the first predicted user behavior intention and the second predicted user behavior intention.
Citation Information
Patent Citations
Personalized behavior prediction method, device and system for smart home
CN109818839B
Intelligent water meter control system and method based on Internet of Things
CN119961329A
Production optimization method and system for anylayer core board layer
CN120146303A
Intelligent management system and method for cable production
CN120235557A
Vent travel content intelligent recommendation method and system based on meta universe
CN120316345A
Cited By
Data processing method and system for virtual reality animation playing
CN120599101A
Smart home control method and system based on voice recognition
CN121768389A