An intelligent speaker permission design management method and system

Through artificial intelligence technology and data analysis, a user-situ joint feature vector is built, combined with Gaussian hybrid model and fuzzy matching algorithm, the dynamic and intelligent nature of smart speaker permission management is realized, solving the shortcomings of permission management in the existing technology and improving user experience and security.

CN119203100BActive Publication Date: 2025-06-20GUANGZHOU RUIGAO INTELLIGENT SYST CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411217025.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-02
Publication Date
2025-06-20
Estimated Expiration
2044-09-02

AI Technical Summary

Technical Problem

The permission management of existing smart speakers lacks dynamic and intelligentity, cannot adapt to the complex needs of multiple users and multiple scenarios, poses privacy and security challenges, and cannot provide personalized experience and effective multi-user management.

Method used

Using artificial intelligence technology and data analysis methods, the user-situ joint feature vector is constructed through user audio data and real-time environment data, combined with Gaussian hybrid model and fuzzy matching algorithm to realize dynamic permission configuration and real-time management.

Benefits of technology

It realizes the dynamic and intelligent nature of smart speaker permission management, improves user experience and security, and supports flexible management of multiple users and multiple scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119203100B_ABST
    Figure CN119203100B_ABST
Patent Text Reader

Abstract

The present invention proposes a method and system for intelligent speaker permission design and management. The method includes: collecting user audio data and preprocessing the audio data; constructing a user-situation joint feature vector to adjust the recognition accuracy; if there is a complex speech environment with multiple languages and multiple accents, designing a fuzzy matching and dynamic adjustment model for real-time dynamic management of the intelligent speaker permissions; constructing a behavior monitoring and anomaly detection model to analyze user permissions in real time through feature engineering and time series modeling, and dynamically adjusting the permission level according to the results of anomaly detection and behavior monitoring; proposing a mapping function between emotional state and permission response to dynamically adjust the permission level, and the intelligent speaker immediately applies the new permission configuration; constructing an adaptive permission optimization model based on reinforcement learning to adaptively learn the permission strategy of the intelligent speaker through reinforcement learning for long-term strategy optimization. The present invention meets the wide application requirements of intelligent speakers in diverse scenarios such as home and office.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of permission design management, and particularly relates to a method and system for intelligent speaker permission design management. Background Art

[0002] With the rapid development of the Internet of Things (IoT) technology, intelligent speakers have gradually become indispensable intelligent devices in home and office environments. They provide users with various convenient services such as music playback, information query, and smart home control through voice interaction. However, the wide application of intelligent speakers has also raised a series of problems related to privacy protection, permission management, and multi-user collaboration. Existing intelligent speakers mainly control the usage permissions of devices through simple static permission settings, and this method has many limitations and deficiencies.

[0003] Firstly, the existing static permission management mode lacks dynamics and intelligence and cannot adapt to the complex requirements of multi-users and multi-scenarios. In a home environment, different members may have different usage requirements and permission requirements for the functions of the speaker. For example, children and adults should have different permissions for content access, but the existing static permission settings cannot flexibly handle these changes, often resulting in children accessing inappropriate content or adults being unable to obtain the required functions in a timely manner. Secondly, the traditional permission management based on static rules cannot provide a personalized experience, lacking in-depth understanding and learning of user individual behaviors and usage habits, resulting in a poor user experience. As a device with a very high daily usage frequency, intelligent speakers should have a higher ability to optimize the user experience.

[0004] In addition, with the popularization of intelligent speakers in home, office and other environments, their permission management faces more severe privacy and security challenges. The traditional permission setting method usually requires users to manually adjust permissions through cumbersome operations, which is neither convenient nor error-prone. More seriously, speakers are usually exposed in public environments, and unauthorized users may perform unauthorized operations through voice commands, posing a great security risk. For example, unauthorized users can easily access confidential information, control smart home devices, or even make improper voice purchases, causing property and privacy losses. In the existing technology, most of the solutions to these security problems are after-the-fact remedies, lacking forward-looking protection strategies.

[0005] Finally, the permission management of smart speakers lacks an effective management mechanism in multi-user scenarios. Especially in the scenario where multiple people use it simultaneously, how to distinguish user identities and provide personalized services is a major challenge. The current permission management systems can often only distinguish simple user roles or rely on manual login methods to switch users, which is neither convenient nor reliable in actual use and cannot meet diverse usage requirements. Facing the alternating use of users of different ages and roles, the existing technologies cannot provide an intelligent and dynamic permission management solution, which limits the function of smart speakers and the optimization of user experience.

[0006] Therefore, the current smart speaker permission management technology has significant deficiencies, mainly manifested in the lack of dynamic intelligence, personalized experience, security guarantee, and insufficient support for multi-user scenarios. These problems severely limit the potential of smart speakers and their market applications. Summary of the Invention

[0007] The objective of the present invention is to design a smart speaker permission design management method and system, introduce advanced artificial intelligence technologies and data analysis methods, overcome the defects of the existing technologies, and provide a smart speaker permission management system with higher intelligence, security, and excellent user experience.

[0008] To achieve the above objective, in the first aspect of the present invention, a smart speaker permission design management method is provided, and the method includes:

[0009] S1. Collect user audio data and preprocess the audio data, and construct a Gaussian mixture model based on the preprocessed data to output an audio signature;

[0010] S2. Use the audio signature as an input representing the user's identity characteristics, collect real-time environmental data and preprocess it, construct a user-situation joint feature vector based on the user's identity characteristics and the preprocessed real-time environmental data, add a regularization term to the user-situation joint feature vector to adjust the recognition accuracy, and then perform situation awareness and dynamic permission configuration initialization to obtain an optimized permission level and permission score;

[0011] S3. According to the optimized permission level and permission score, if there is a complex speech environment with multiple languages and accents, design a fuzzy matching and dynamic adjustment model for the real-time dynamic management of smart speaker permissions; wherein, the fuzzy matching and dynamic adjustment model uses a feature weighted aggregation method to calculate the credibility score γ of the speech feature vector C new as follows:

[0012]

[0013] Among them, γ represents the credibility score of the voice input, indicating the credibility of the current voice feature matching degree; M represents the total number of voice features; w j represents the weight of the j-th feature, indicating the importance of this feature when calculating the voice credibility; α j represents the fuzzy adjustment parameter of the j-th feature, controlling the sensitivity of feature matching; C new,j represents the component of the j-th voice feature vector; μ j represents the expected value of the j-th feature, indicating the typical value of this feature under normal circumstances;

[0014] Using the credibility score γ of the voice input and the optimized permission score According to the permission adjustment function Dynamically adjust the user's permission level It is expressed as follows:

[0015]

[0016] Among them, represents the dynamically adjusted permission level, indicating the permission setting adjusted according to the voice credibility and the original permission score; represents the initial permission level output by the second step; η represents the permission adjustment gain coefficient, controlling the permission adjustment range; γ represents the credibility score of the voice input; δ represents the intermediate threshold of the credibility score; represents the optimized permission score; τ represents the benchmark threshold of the permission score; According to the calculated new permission level The smart speaker system immediately adjusts the user's permission configuration and dynamically changes the accessibility of functions;

[0017] S4. According to the new permission level optimized by combining the complex voice environment Combined with real-time user behavior data, construct a behavior monitoring and anomaly detection model to analyze the user's permissions in real time through feature engineering and time series modeling, detect anomalies and monitor user behavior, and dynamically adjust the permission level according to the results of anomaly detection and behavior monitoring

[0018] S5. According to the permission level Combine the user's emotional state to classify the emotional state through the emotion classifier model, generate emotion labels, and according to the emotion state label and the abnormal behavior response strategy, propose a mapping function of emotion state and permission response to dynamically adjust the permission level. According to the permission level calculated by the emotion state and permission response mapping function The smart speaker immediately applies the new permission configuration;

[0019] S6. According to the permission level Build an adaptive permission optimization model based on reinforcement learning, combine the long-term behaviors and usage patterns of users, and perform long-term policy optimization by adaptively learning the permission policies of smart speakers through reinforcement learning.

[0020] Preferably, the preprocessing is to process the collected audio data x(t) to remove noise and perform normalization processing;

[0021] Among them, the probability density function of the Gaussian mixture model is expressed as:

[0022]

[0023] Among them, C represents the extracted Mel-frequency cepstral coefficient feature vector, which is used to represent the voiceprint feature of the audio; represents the parameter set of the Gaussian mixture model; π k represents the mixing weight of the k-th Gaussian component, satisfying and π k ≥0; μ k represents the mean vector of the k-th Gaussian component; Σ k represents the covariance matrix of the k-th Gaussian component; K represents the number of Gaussian components, indicating the number of Gaussian distributions in the model; represents the multivariate normal distribution, which is expressed as:

[0024]

[0025] Among them, d represents the dimension of the feature vector C.

[0026] Preferably, for each new audio input, the user identity is recognized by calculating the log-likelihood estimate of the feature vector of the new audio under each user model, which is expressed as follows:

[0027]

[0028] Among them, logL(C new |S i ) represents the log-likelihood value of the new feature vector under the i-th user model, which is used to measure the matching degree between the audio input and the user model; C new : represents the feature vector of the new audio input; S i represents the audio signature of the i-th user, that is, the GMM parameter set Θ i ; T represents the total number of frames of the new audio input, indicating the number of frames after the audio signal is segmented; P(C new (t)|Θ i ) represents the probability of the feature vector C new (t) under the user model Θ i ).

[0029] Preferably, the real-time environmental data is collected through a variety of sensors, including microphones, cameras, light sensors, temperature and humidity sensors, and accelerometers. The real-time environmental feature vector is denoted as E;

[0030] The user-situation joint feature vector is fused through a weighted multi-layer perceptron model, expressed as follows:

[0031] C e = σ(W2·σ(W1·[E norm ,S i +b1)+b2)

[0032] where C e represents the generated situation feature vector, indicating the behavioral characteristics of the user in the current environment; [E norm ,S i represents the input feature after combining the environmental feature vector and the user audio signature; W1, W2 represent weight matrices with dimensions m×(n + 1) and m×m, respectively, learned through model training; b1, b2 represent bias vectors with dimensions m and m, respectively, learned through model training; σ represents a non-linear activation function used to capture non-linear relationships;

[0033] Combining the situation feature vector C e and the user audio signature S i , the user-situation joint feature vector F = [S i ,C e is generated;

[0034] In the user-situation joint feature vector F = [S i ,C e , a regularization term is added to achieve higher recognition accuracy through weighted combination of user and environmental features, expressed as follows:

[0035] F opt = F + λ·(F⊙W reg )

[0036] where F opt represents the optimized joint feature vector; λ represents the regularization parameter used to balance the influence of the original feature and the regularization term; ⊙ represents the element-wise product operation for element-wise weighting; W reg represents the weight matrix indicating the mutual relationship between features.

[0037] Preferably, the configuration initialization includes:

[0038] Design a permission scoring function based on the adjusted user-situation joint feature vector to calculate the permission score of the user in the current situation Expressed as follows:

[0039]

[0040] Among them, represents the permission score, which is used to reflect the user's permission level in the current context; represents the weight vector, indicating the importance of each feature; β represents the weight adjustment coefficient, which controls the overall weight of the combined features; represents the second norm of the combined features, which is used as a measure of feature importance and adds a constraint on the balance between features; represents the bias term;

[0041] According to the calculated permission score and the preset permission level threshold, perform dynamic permission adjustment. Among them, the permission adjustment function is expressed as follows:

[0042]

[0043] Among them, represents the dynamically adjusted permission level; represents the permission threshold, which is set according to user requirements and system configuration.

[0044] Preferably, in S3, a feedback mechanism for error correction of the fuzzy matching and dynamic adjustment model design is used to dynamically adjust the weights and parameters of the fuzzy matching model. The error correction mechanism is expressed as follows:

[0045] Δw = ∈·(γ target -γ)·C new

[0046] Among them, Δw represents the weight adjustment vector, which is used to correct the weights of the speech features; ∈ represents the learning rate, which is used to control the step size of the weight adjustment; γ target represents the target credibility score, which is usually set to the desired credibility level of the system; γ represents the credibility score of the current speech input; C new represents the feature vector of the current speech input.

[0047] Preferably, based on the weighted behavior feature vector B w construct a monitoring and anomaly detection model based on the long short-term memory network to predict the behavior features at the next moment and identify the abnormal patterns of the behavior. Among them, the update formula of the behavior monitoring and anomaly detection model is expressed as follows:

[0048] h t = f LSTM (B w,t ,h t-1 )

[0049] Among them, h tRepresents the hidden state at the current time t, capturing the time-dependence of the user's behavior; B w,t Represents the weighted behavior feature vector at the current time t; h t-1 Represents the hidden state at the previous time t-1; f LSTM Represents the state update function of the long short-term memory network;

[0050] By calculating the residual between the actual behavior feature and the predicted behavior feature, it is determined whether there is an abnormal behavior; where the abnormal score ξ is expressed as follows:

[0051]

[0052] Among them, ξ represents the abnormal score, indicating the degree of difference between the current behavior and the predicted behavior; B w,t Represents the actual weighted behavior feature vector at the current time t; Represents the weighted behavior feature vector at the next time predicted by the long short-term memory network; σ 2 Represents the variance of the prediction residual, used to standardize the abnormal score;

[0053] According to the abnormal score ξ and the preset abnormal threshold θ, it is determined whether to trigger the permission adjustment or warning mechanism; where the adjustment strategy is as follows:

[0054] If ξ ≤ θ, no adjustment is made and the current permission level is maintained

[0055] If ξ > θ, it is identified as an abnormal behavior, and the permission is tightened or the user is prompted to authenticate according to the severity of the abnormality.

[0056] Preferably, the acquisition of the user's emotional state includes audio and video feature acquisition, and then a fusion model based on the self-attention mechanism is used to capture the mutual influence between multi-modal features. The specific formula is expressed as follows:

[0057]

[0058] Among them, E represents the fused emotional feature vector, used to describe the user's current comprehensive emotional state; A represents the audio feature vector with dimension m; V represents the video feature vector with dimension n; W A and W V Represents the projection matrix, used to map the audio and video features to the same space, with dimensions m×d k and n×d k ; d k Represents the dimension of the key vector in the attention mechanism, used to scale the dot product operation; softmax represents the normalization function used to calculate the attention weight;

[0059] Using the fused emotional feature vector E, classify the emotional state through the emotion classifier model C(E) to generate an emotion label c, expressed as follows:

[0060]

[0061] Among them, c represents the emotional state label of the current user, expressed as a discrete category; w i represents the weight vector of the classifier model, the weight corresponding to the i-th category, with a dimension of d k ; b i represents the bias term of the classifier model, the bias corresponding to the i-th category;

[0062] According to the emotional state label c and the abnormal behavior response strategy, propose a mapping function of emotional state and permission response to dynamically adjust the permission level expressed as follows:

[0063]

[0064] Among them, represents the permission level dynamically adjusted according to the emotional state, considering the permission setting of the user's emotion; ρ(c) represents the permission adjustment increment function, which depends on the user's emotional state c and is expressed as an integer that can be positive or negative.

[0065] Preferably, the adaptive permission optimization model based on reinforcement learning combines the user's current permission level and emotional state, as well as the user's historical behavior pattern and system feedback, that is, the state vector Among them, H represents the historical behavior pattern feature vector of the user; F represents the system feedback feature vector; define the actions that the system can execute as a set of permission adjustment operations where each action a i represents the adjustment of the permission level; define the optimization objective function J(θ), and optimize the parameter θ of the policy π θ (S) to maximize the expected cumulative reward;

[0066] The formula of the optimization objective function J(θ) is as follows:

[0067]

[0068] Among them, J(θ) represents the optimization objective function of the policy parameter θ, representing the expected cumulative reward; S represents the current state vector, which is composed of the permission level, emotional state, historical behavior pattern and system feedback; represents the set of state distributions, representing the possible state space; γ represents the discount factor, and the value range is 0 ≤ γ ≤ 1, which is used to measure the importance of future rewards; r(S t ,a t)Indicates in state S t Execute action a t The immediate reward obtained when reflecting the effectiveness and reasonableness of the permission adjustment; a t Indicates the action executed at time t, determined by the policy π θ (S);

[0069] During the policy optimization process, the policy regularization term Ω(θ) is introduced to avoid the policy overfitting the user's short-term behavior pattern, and the policy parameter θ is updated to maximize the regularized optimization objective, which is expressed as follows:

[0070]

[0071] where θ t+1 Indicates the updated policy parameter; θ t Indicates the policy parameter at the current time t; α represents the learning rate, controlling the step size of parameter update; Indicates the policy gradient, representing the gradient of the optimization objective function with respect to the policy parameter; λ represents the regularization strength parameter, controlling the influence degree of the regularization term on the policy update; Ω(θ t )Indicates the policy regularization term, which is a sparsity regularization, encouraging the simplicity and generalization ability of the policy parameter, and is expressed as follows:

[0072]

[0073] where N represents the total number of policy parameters; θ i Indicates the i-th parameter in the policy parameters.

[0074] In the second aspect of the present invention, an intelligent speaker permission design and management system is provided, and the system includes:

[0075] A user data collection module, configured to collect user audio data, preprocess the audio data, and construct a Gaussian mixture model based on the preprocessed data to output an audio signature;

[0076] A configuration initialization module, configured to use the audio signature as an input representing the user's identity characteristics, collect real-time environmental data and perform preprocessing, construct a user-situation joint feature vector based on the user's identity characteristics and the preprocessed real-time environmental data, add a regularization term to the user-situation joint feature vector to adjust the recognition accuracy, and then perform situation awareness and dynamic permission configuration initialization to obtain an optimized permission level and permission score;

[0077] The permission initialization module is used to, according to the optimized permission levels and permission scores, design a fuzzy matching and dynamic adjustment model for real-time dynamic management of the smart speaker's permissions in the case of a complex voice environment with multiple languages and accents; among them, the fuzzy matching and dynamic adjustment model uses a feature weighted aggregation method to calculate the voice feature vector C new The credibility score γ of

[0078]

[0079] where γ represents the credibility score of the voice input, indicating the credibility of the current voice feature matching degree; M represents the total number of voice features; w j represents the weight of the j-th feature, indicating the importance of this feature when calculating the voice credibility; α j represents the fuzzy adjustment parameter of the j-th feature, controlling the sensitivity of feature matching; C new,j represents the component of the j-th voice feature vector; μ j represents the expected value of the j-th feature, indicating the typical value of this feature under normal circumstances;

[0080] Using the credibility score γ of the voice input and the optimized permission score According to the permission adjustment function Dynamically adjust the user's permission level It is expressed as follows:

[0081]

[0082] Among them, represents the dynamically adjusted permission level, indicating the permission setting adjusted according to the voice credibility and the original permission score; represents the initial permission level output by the second step; η represents the permission adjustment gain coefficient, controlling the permission adjustment range; γ represents the credibility score of the voice input; δ represents the intermediate threshold of the credibility score; represents the optimized permission score; τ represents the benchmark threshold of the permission score; According to the calculated new permission level The smart speaker system immediately adjusts the user's permission configuration and dynamically changes the accessibility of functions;

[0083] The permission management module is used to, according to the new permission levels optimized by combining the complex voice environment Combined with real-time user behavior data, construct a behavior monitoring and anomaly detection model to analyze the user's permissions in real time through feature engineering and time series modeling, perform anomaly detection and behavior monitoring on the user's behavior, and dynamically adjust the permission level according to the results of the anomaly detection and behavior monitoring According to the permission level Classify the emotional state through an emotional classifier model in combination with the user's emotional state, generate an emotional label, and propose a mapping function of emotional state and permission response to dynamically adjust the permission level according to the emotional state label and the abnormal behavior response strategy. The permission level calculated according to the mapping function of emotional state and permission response The smart speaker immediately applies the new permission configuration;

[0084] The system optimization module is used to Build an "adaptive permission optimization model based on reinforcement learning", combine the user's long-term behavior and usage patterns, and adaptively learn the permission policy of the smart speaker through reinforcement learning for long-term policy optimization.

[0085] The beneficial technical effects of the present invention are at least as follows:

[0086] First of all, the present invention adopts the user audio signature recognition technology, which can identify the user's identity in real time through audio features during the use of the speaker. This method can not only effectively distinguish different users in the home or office environment, but also automatically adjust the permission settings according to the usage habits and permission requirements of each user, realizing true personalized services and overcoming the lack of flexibility in traditional static permission settings.

[0087] Secondly, the present invention introduces the context-aware permission management technology. Through multi-sensor data fusion, such as microphones, cameras, light sensors, etc., combined with the user's behavior data and environmental information (such as time, location, schedule, etc.), the system can intelligently perceive the current context and dynamically adjust the permission settings. For example, in the evening home mode, the system will automatically enable the child protection mode to restrict access to inappropriate functions, thus greatly improving the intelligence of the system and the user experience, and overcoming the problem of lack of dynamic adjustment ability in the prior art.

[0088] The present invention designs an adaptive speech fuzzy matching algorithm, which can adaptively adjust the speech recognition accuracy under various accents and speech rates, ensure accurate understanding of the user's intention, and flexibly adjust the permission settings to ensure security. This adaptive mechanism solves the problem of low recognition accuracy in the prior art in multi-language and multi-accent usage scenarios, and at the same time reduces the risk of incorrect permission operations.

[0089] In addition, the present invention also integrates an abnormal permission behavior detection module and an emotion reasoning permission management module. The abnormal permission behavior detection module uses machine learning algorithms to monitor and analyze the user's permission usage behavior in real time, and can quickly identify abnormal operations and take corresponding security measures, enhancing the security and protection capabilities of the system. The emotion reasoning permission management module analyzes the user's emotional state through voice, and automatically adjusts permissions when the user is emotionally excited to prevent improper behavior caused by emotional fluctuations. These innovative points effectively solve the deficiencies in the prior art of lacking response to user emotions and abnormal behaviors, and significantly improve the security and user-friendly experience of the system.

[0090] Through these innovative points, the present invention provides a comprehensive solution that can effectively overcome various deficiencies of the prior art, realizing dynamic management of smart speaker permissions, personalized services, security protection, and multi-user support. This multi-level and all-round intelligent permission management system not only improves the user experience, but also significantly enhances the security and adaptability of the device, meeting the wide application requirements of smart speakers in diverse scenarios such as home and office. BRIEF DESCRIPTION OF THE DRAWINGS

[0091] The present invention will be further described with reference to the accompanying drawings. However, the embodiments in the drawings do not constitute any limitation to the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained according to the following drawings.

[0092] Figure 1 It is a flowchart of a method for designing and managing smart speaker permissions according to an embodiment of the present invention.

[0093] Figure 2 It is a framework diagram of a system for designing and managing smart speaker permissions according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0094] The embodiments of the present invention will be described in detail below. The examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention and should not be construed as a limitation to the present invention.

[0095] In one or more embodiments, as Figure 1 shown, a method for designing and managing smart speaker permissions of the present invention is disclosed. The method includes steps 1 to 6, including:

[0096] S1. Collect user audio data and preprocess the audio data, and construct a Gaussian mixture model based on the preprocessed data to output an audio signature.

[0097] Specifically, during the preprocessing, the collected audio data x(t) is processed to remove noise and normalize it.

[0098]

[0099] Among them, y(t) represents the denoised audio signal. x(t) represents the original audio signal, expressed as a function of time. represents the estimated noise signal, which is extracted from the environment using an adaptive filter.

[0100] In the signal normalization process, the normalized signal is expressed as:

[0101]

[0102] Among them, y norm (t) represents the normalized audio signal. μ(y) represents the mean of the denoised signal y(t), calculated as where N is the number of signal samples. σ(y) represents the standard deviation of the denoised signal y(t), calculated as

[0103] Further, a Gaussian mixture model is used to model the audio features of the user to form an audio signature. The probability density function of the GMM is expressed as:

[0104]

[0105] Among them, C represents the extracted Mel-frequency cepstral coefficients (MFCC) feature vectors, which are used to represent the voiceprint features of the audio. represents the parameter set of the GMM. π k represents the mixing weight of the k-th Gaussian component, satisfying and π k ≥ 0. μ k represents the mean vector of the k-th Gaussian component. Σ k represents the covariance matrix of the k-th Gaussian component. K represents the number of Gaussian components, indicating the number of Gaussian distributions in the model. represents the multivariate normal distribution, expressed as:

[0106]

[0107] Among them, d is the dimension of the feature vector C.

[0108] Further, for each new audio input, the user identity is identified by calculating the log-likelihood estimate of its feature vector under each user model:

[0109]

[0110] Among them, logL(C new |S i ) represents the log-likelihood value of the new feature vector under the i-th user model, which is used to measure the matching degree between the audio input and the user model. C new represents the feature vector of the new audio input. S i represents the audio signature of the i-th user, that is, the GMM parameter set Θ i . T represents the total number of frames of the new audio input, which is the number of frames after the audio signal is segmented. P(C new (t)|Θ i ) represents the probability of the feature vector C new (t) of the t-th frame under the user model Θ i .

[0111] S2. Using the audio signature as the input represents the identity characteristics of the user. Collect real-time environmental data and perform preprocessing. Combine the identity characteristics of the user with the preprocessed real-time environmental data to construct a user-situation joint feature vector, and add a regularization term to the user-situation joint feature vector to adjust the recognition accuracy. Then perform context awareness and dynamic permission configuration initialization to obtain the optimized permission level and permission score.

[0112] Among them, the smart speaker collects environmental data in real time through a variety of sensors (such as microphones, cameras, light sensors, temperature and humidity sensors, accelerometers, etc.) to form an environmental feature vector E = [e1, e2, …, e n . These environmental features include:

[0113] e1 represents the environmental noise level, with a range of 0 to 100 decibels.

[0114] e2 represents the light intensity, with a range of 0 to 10,000 lux.

[0115] e3 represents the distance between the user and the speaker, with a range of 0 to 5 meters.

[0116] e4 represents the current time, in 24-hour format, with a range of 0 to 23.99.

[0117] e5 represents the number of people detected, with an integer range of 0 to 10.

[0118] e6 represents the user's activity status, a categorical variable, with a value range of {stationary, walking, jumping, etc.}.

[0119] Other sensor data such as temperature e7, humidity e8, etc.

[0120] Furthermore, since the data ranges and units of different sensors are different, it is necessary to standardize the environmental feature vector E to unify the data scale:

[0121]

[0122] Among them, represents the standardized environmental feature. min(e i ) and max(e i ) represent the minimum and maximum values of the environmental feature e i , which are used for normalization.

[0123] Furthermore, by using the user audio signature S i obtained in Step 1 to fuse the user identity and environmental features, the present invention introduces a new fusion method to fuse the user identity features and environmental data to construct a context feature vector C e . The specific method is through a weighted multi-layer perceptron (MLP) model, which can dynamically adjust the weights of user features and environmental features during the learning process to reflect the user's behavior patterns in different situations.

[0124] C e = σ(W2·σ(W1·[E norm , S i + b1) + b2)

[0125] Among them, C e represents the generated context feature vector, which represents the user's behavior features in the current environment. [E norm , S i represents the input feature after merging the environmental feature vector and the user audio signature. W1 and W2 represent weight matrices with dimensions of m×(n + 1) and m×m respectively, which are learned through model training. b1 and b2 represent bias vectors with dimensions of m and m respectively, which are learned through model training. σ represents a non-linear activation function (such as ReLU or Sigmoid), which is used to capture non-linear relationships.

[0126] Furthermore, by combining the context feature vector C e and the user audio signature S i , a user-context joint feature vector F = [S i , C e is generated. This is a fused feature vector, which is used to represent the comprehensive state of a specific user in the current environment.

[0127] To better capture the complex relationship between user features and environmental features, a joint feature modeling method with an innovative regularization term is proposed. This regularization term achieves higher recognition accuracy through the weighted combination of user and environmental features. The model formula is as follows:

[0128] F opt = F + λ·(F⊙W reg )

[0129] Among them, F opt represents the optimized joint feature vector, which considers the weighted relationship between user and environmental features. λ represents the regularization parameter, which is used to balance the influence of the original features and the regularization term, and is usually determined by cross-validation. ⊙ represents the element-wise product operation (Hadamard product), which is used for element-wise weighting. W reg represents the weight matrix, which represents the mutual relationship between features and is obtained through optimized dynamic learning.

[0130] Furthermore, based on the optimized joint feature vector F opt , a permission scoring function with an innovative additional term is proposed to calculate the permission score of the user in the current context

[0131]

[0132] Among them, represents the permission score, which is used to reflect the permission level of the user in the current context. w represents the weight vector, which is obtained through model training and represents the importance of each feature. β represents the weight adjustment coefficient, which controls the overall weight of the joint features. ‖F opt ‖ 2 represents the second norm of the joint features, which is used as a measure of feature importance and adds a constraint on the balance between features. b represents the bias term, which is obtained through learning from the training data.

[0133] According to the calculated permission score and the preset permission level threshold, the permission is dynamically adjusted. The permission adjustment function is defined as:

[0134]

[0135] Among them, represents the dynamically adjusted permission level, and the specific permission configuration will be set according to this level. T1, T2 represent the permission thresholds, which are set according to user requirements and system configuration.

[0136] The finally determined permission level is applied to the smart speaker system to dynamically control the functions and data permissions that the user can access.

[0137] S3. According to the optimized permission level and permission score, if there is a complex speech environment with multiple languages and multiple accents, then design a fuzzy matching and dynamic adjustment model for real-time dynamic management of smart speaker permissions.

[0138] Specifically, when the user issues a voice command, the speaker collects the current voice input signal x(t), preprocesses the voice signal (such as denoising and normalization) using the method in the first step, and generates the processed voice feature vector C new .

[0139] Meanwhile, to adapt to multilingual and multi-accent environments, an adaptive voice fuzzy matching model is proposed. This model uses a feature weighted aggregation method to calculate the credibility score γ of the voice feature vector C new The formula is as follows:

[0140]

[0141] where γ represents the credibility score of the voice input, indicating the credibility of the current voice feature matching degree. M represents the total number of voice features. w j represents the weight of the j-th feature, indicating the importance of this feature in calculating the voice credibility. α j represents the fuzzy adjustment parameter of the j-th feature, controlling the sensitivity of feature matching. C new,j represents the component of the j-th voice feature vector. μ j represents the expected value of the j-th feature, indicating the typical value of this feature under normal circumstances.

[0142] Furthermore, a permission adjustment strategy based on credibility is defined. Using the credibility score γ of the voice input and the permission score in the second step A "permission adjustment function" is proposed to dynamically adjust the user's permission level

[0143]

[0144] where, represents the dynamically adjusted permission level, indicating the permission setting adjusted according to the voice credibility and the original permission score. represents the initial permission level output by the second step. η represents the permission adjustment gain coefficient, controlling the permission adjustment range. γ represents the credibility score of the voice input. δ represents the intermediate threshold of the credibility score, usually set to 0.5. represents the permission score calculated in the second step. τ represents the benchmark threshold of the permission score, usually set based on actual usage.

[0145] Furthermore, according to the calculated new permission level The smart speaker system immediately adjusts the user's permission configuration and dynamically changes the accessibility of functions.

[0146] To further improve the robustness and accuracy of the system, an "error correction feedback mechanism" is introduced. This mechanism dynamically adjusts the weights and parameters of the fuzzy matching model by detecting and analyzing the user feedback or usage behavior after permission adjustment. The core formula for error correction is as follows:

[0147] Δw = ∈·(γ target -γ)·C new

[0148] Where, Δw represents the weight adjustment vector, which is used to correct the weights of speech features. ∈ represents the learning rate, which is used to control the step size of weight adjustment. γ target represents the target confidence score, which is usually set to the desired confidence level of the system. γ represents the confidence score of the current speech input. C new represents the feature vector of the current speech input.

[0149] Update the parameters of the speech fuzzy matching model according to the error correction formula, so that the system can more accurately adjust the permission configuration when facing speech inputs from different users and contexts.

[0150] Furthermore, combining the user's instant feedback and long-term usage behavior data, continuously optimize the permission adjustment function and the fuzzy matching model. By real-time monitoring the user's operation habits and permission usage, dynamically adjust the system parameters to improve the adaptability and accuracy of permission adjustment.

[0151] Collect the user's usage data (such as the user's behavior under the adjusted permission level, satisfaction feedback, etc.), and regularly retrain the fuzzy matching model and the permission adjustment function to ensure that the model parameters are consistent with the dynamic changes of user needs.

[0152] S4. According to the new permission level optimized by combining with the complex speech environment and then combining with the real-time user behavior data, construct a behavior monitoring and anomaly detection model. Through feature engineering and time series modeling, analyze the user permissions in real time, detect anomalies and monitor user behavior, and dynamically adjust the permission level according to the results of anomaly detection and behavior monitoring

[0153] Specifically, extract the user behavior characteristics, including:

[0154] User behavior data collection: The smart speaker continuously collects the user's operation behavior data and generates a behavior feature vector B = [b1, b2, …, b k , and these features include:

[0155] b1 represents the type of speech command issued by the user (such as playing music, setting an alarm, etc.), which is represented as a categorical variable.

[0156] b2 represents the time interval (unit: second) when an instruction occurs, expressed as a continuous variable.

[0157] b3 represents the instruction frequency issued by the user within a specific time period (unit: times / minute), expressed as a continuous variable.

[0158] b4 represents the amplitude of the change in the permission level corresponding to each instruction, expressed as an integer.

[0159] b5 represents the recognized emotional state of the user, expressed as a categorical variable (such as calm, excited, etc.).

[0160] Standardization and weighting of behavioral features: Standardize the behavioral feature vector B to eliminate the scale differences between different features. Then, perform weighted combination according to the feature importance to generate the weighted behavioral feature vector B w :

[0161] B w = W B ·B norm

[0162] Among them, B w represents the weighted behavioral feature vector, expressed as the behavioral features weighted by different weights. W B represents the weight matrix, used for weighting different features. B norm represents the standardized behavioral feature vector, expressed as the standardized eigenvalue.

[0163] Furthermore, use the weighted behavioral feature vector B w , construct a behavior monitoring and anomaly detection model based on the long short-term memory network (LSTM) to capture the time dependence of user behavior. The goal of the model is to predict the behavioral features at the next moment, so as to identify the abnormal patterns of behavior. The update formula of the LSTM model is:

[0164] h t = f LSTM (B w,t , h t-1 )

[0165] Among them, h t represents the hidden state at the current moment t, capturing the time dependence of user behavior. B w,t represents the weighted behavioral feature vector at the current moment t. h t-1 represents the hidden state at the previous moment t-1. f LSTM represents the state update function of the LSTM, defined as the calculation formula of the conventional LSTM cell.

[0166] Anomaly behavior is judged by calculating the residual between the actual behavior characteristics and the predicted behavior characteristics. The anomaly score ξ is calculated by the following formula:

[0167]

[0168] where ξ represents the anomaly score, indicating the degree of difference between the current behavior and the predicted behavior. B w,t represents the actual weighted behavior feature vector at the current moment t. represents the weighted behavior feature vector at the next moment predicted by the LSTM model. σ 2 represents the variance of the prediction residual, which is used to standardize the anomaly score.

[0169] Furthermore, according to the anomaly score ξ and the preset anomaly threshold θ, it is decided whether to trigger the permission adjustment or warning mechanism. The adjustment strategy is as follows:

[0170] If ξ ≤ θ, no adjustment is made and the current permission level is maintained

[0171] If ξ > θ, it is identified as an anomaly behavior, and the permission is tightened or the user is prompted to authenticate according to the severity of the anomaly.

[0172] Furthermore, a permission dynamic adjustment function under anomaly behavior conditions is defined Dynamically adjust the permission level according to the anomaly score ξ

[0173]

[0174] where represents the dynamically adjusted permission level, which takes into account the permission settings after anomaly behavior. represents the permission level output by the third step. ζ represents the adjustment step parameter, which controls the amplitude of the permission adjustment. ξ represents the current anomaly score. θ represents the threshold for anomaly detection, which is usually set based on actual user behavior data.

[0175] Furthermore, combining the user's feedback and actual behavior data, continuously optimize the time-series behavior model and anomaly detection mechanism. Through online learning and regular model updates, ensure the adaptability of the model to changes in user behavior. Collect user behavior data (such as the number of anomaly behaviors, permission adjustment frequency, etc.), regularly update the parameters of the LSTM model, and adjust the anomaly detection threshold θ and the weight matrix W B , ensuring that the system can operate stably in a changing user behavior environment.

[0176] S5. According to the permission level Classify the emotional state through an emotional classifier model in combination with the user's emotional state, generate emotional labels, and propose a mapping function of emotional state and permission response to dynamically adjust the permission level according to the emotional state label and the abnormal behavior response strategy. The permission level calculated according to the mapping function of emotional state and permission response The smart speaker immediately applies the new permission configuration.

[0177] Furthermore, for the extraction of user emotional state features, the smart speaker is used to collect the user's multi-modal emotional data in real time in combination with the microphone and camera, including audio and video features. The audio feature vector A = [a1, a2, …, a m May include:

[0178] a1 represents the change in speech pitch frequency and is expressed as a continuous variable.

[0179] a2 represents the change in speech rate and is expressed as a continuous variable.

[0180] a3 represents the speech volume intensity and is expressed as a continuous variable.

[0181] Other audio features such as speech discontinuity and timbre.

[0182] The video feature vector V = [v1, v2, …, v n May include:

[0183] v1 represents facial expression features (such as the corners of the eyes rising and the corners of the mouth dropping), and is expressed as a categorical variable.

[0184] v2 represents eye movement features (such as the speed and direction of eye movement), and is expressed as a continuous variable.

[0185] v3 represents the head pose (such as the rotation angle and pitch angle), and is expressed as a continuous variable.

[0186] Furthermore, the audio feature vector A and the video feature vector V are fused to generate a fused emotional feature vector E. A fusion model based on the self-attention mechanism is used to capture the mutual influence between multi-modal features. The specific formula is as follows:

[0187]

[0188] Among them, E represents the fused emotional feature vector, which is used to describe the user's current comprehensive emotional state. A represents the audio feature vector, with a dimension of m. V represents the video feature vector, with a dimension of n. W A and W V represent projection matrices, which are used to map audio and video features into the same space, with dimensions of m×d k and n×d k . d kRepresents the dimension of the key vector in the attention mechanism, which is used to scale the dot product operation. Softmax represents the normalization function used to calculate the attention weights.

[0189] Further, using the fused sentiment feature vector E, through a sentiment classifier model to classify the sentiment state and generate a sentiment label c:

[0190]

[0191] Among them, c represents the sentiment state label of the current user, expressed as discrete categories (such as calm, angry, happy, etc.). w i represents the weight vector of the classifier model, the weight corresponding to the i-th class, with a dimension of d k . b i represents the bias term of the classifier model, the bias corresponding to the i-th class, which is a scalar.

[0192] Further, according to the sentiment state label c and the abnormal behavior response strategy, a mapping function of sentiment state and permission response is proposed to dynamically adjust the permission level

[0193]

[0194] Among them, represents the permission level dynamically adjusted according to the sentiment state, considering the permission setting of the user's sentiment. represents the permission level output in the fourth step. ρ(c) represents the permission adjustment increment function, which depends on the user's sentiment state c and is expressed as an integer that can be positive or negative.

[0195] Further, according to the mapping function of sentiment state and permission response The calculated permission level The smart speaker immediately applies the new permission configuration to ensure the safety and reasonableness of the user's operations in the current sentiment state.

[0196] Combined with the user's immediate feedback and long-term behavior pattern data, continuously optimize the sentiment classification model and the permission response strategy. By real-time monitoring of the user's behavior and sentiment changes, regularly update the parameters of the sentiment classifier and adjust the permission response function and the sentiment feature fusion model to ensure that the model can adapt to the dynamic changes of the user's sentiment state.

[0197] S6. According to the permission level Construct an adaptive permission optimization model based on reinforcement learning. Combining the user's long-term behavior and usage patterns, adaptively learn the permission policy of the smart speaker through reinforcement learning for long-term policy optimization.

[0198] Specifically, the state of the adaptive permission optimization model based on reinforcement learning is defined as the user's current permission level and emotional state, as well as the user's historical behavior patterns and system feedback, that is, the state vector Among them, represents the dynamically adjusted permission level output by the fifth step, and represents the current permission configuration optimized according to the user's emotional state. c represents the user's current emotional state label, which is expressed as discrete categories (such as calm, angry, happy, etc.). H represents the user's historical behavior pattern feature vector, which represents the statistical features of the user's behavior in the past period of time, such as operation frequency, preferred operation type, etc. F represents the system feedback feature vector, which represents the user feedback data of the system under different permission configurations, such as user satisfaction score, error operation frequency, etc.

[0199] Define the actions executable by the system as a set of permission adjustment operations Among them, each action a i represents the adjustment of the permission level (such as promotion, reduction, maintenance, etc.). The design of these actions needs to consider the security of permissions, user experience, and functional limitations of the smart speaker.

[0200] Furthermore, in order to enable the system to maintain the optimal permission configuration under changing user behaviors and emotional states, we adopt a policy optimization reinforcement learning method. Define the optimization objective function J(θ), and optimize the parameters θ of the policy π θ (S) to maximize the expected cumulative reward. The formula of the optimization objective function J(θ) is as follows:

[0201]

[0202] Among them, J(θ) represents the optimization objective function of the policy parameter θ, and represents the expected cumulative reward. S represents the current state vector, which is composed of permission level, emotional state, historical behavior pattern, and system feedback. represents the set of state distributions, and represents the possible state space. γ represents the discount factor, and its value range is 0 ≤ γ ≤ 1, which is used to measure the importance of future rewards. r(S t ,a t ) represents the immediate reward obtained when executing the action a t in the state S t , which reflects the effectiveness and rationality of the permission adjustment. a t represents the action executed at time t, which is determined by the policy π θ (S).

[0203] During the policy optimization process, to avoid the policy overfitting to the short-term behavior patterns of users, we introduce a novel policy regularization term Ω(θ) and update the policy parameter θ to maximize the regularized optimization objective:

[0204]

[0205] where θ t+1 represents the updated policy parameter. θ t represents the policy parameter at the current time step t. α represents the learning rate, which controls the step size of parameter update. represents the policy gradient, which is the gradient of the optimization objective function with respect to the policy parameter. λ represents the regularization strength parameter, which controls the influence degree of the regularization term on the policy update. Ω(θ t ) represents the policy regularization term, which is designed as a sparsity regularization to encourage the simplicity and generalization ability of the policy parameter. The formula is:

[0206]

[0207] where N represents the total number of policy parameters. θ i represents the i-th parameter in the policy parameters.

[0208] Furthermore, the design of the immediate reward function r(S t , a t ) combines the user behavior feedback, the change of emotional state, and the rationality of permission adjustment, aiming to balance the security of the system and the user experience. The formula of the reward function is as follows:

[0209]

[0210] where r(S t , a t ) represents the immediate reward, which measures the effect of executing the action a t in the state S t . U t represents the change of the user satisfaction score at the current time step t. The function f(U t ) represents the positive reward, which is positively correlated with the reasonable permission adjustment. E t represents the change of the user emotional state fluctuation. The function g(E t ) represents the negative reward, which is negatively correlated with the negative emotional reaction caused by the permission adjustment. represents the amplitude of the permission adjustment. The function represents the negative reward, which punishes the overly frequent permission adjustment operations. ω1, ω2, ω3 represent the weight coefficients, which control the influence degree of each factor on the immediate reward.

[0211] To further improve the adaptability and generalization ability of the model, we introduce an adaptive learning mechanism based on user behavior patterns. This mechanism dynamically adjusts the learning rate and regularization parameters of the policy model by continuously monitoring the changes in users' behavior characteristics and emotional states to meet the personalized needs of users. The dynamic adjustment formulas for the learning rate and regularization parameters are as follows:

[0212] α t+1 =α t ·(1 - η·|ΔH t |)

[0213] λ t+1 =λ t ·(1 + ζ·|ΔF t |)

[0214] Where α t+1 represents the learning rate at the next moment. α t represents the learning rate at the current moment. η represents the learning rate adjustment coefficient, which controls the adjustment range of the learning rate. |ΔH t | represents the change amplitude of the user behavior pattern feature vector, indicating the degree of change in users' behavior. λ t+1 represents the regularization strength parameter at the next moment. λ t represents the regularization strength parameter at the current moment. ζ represents the regularization parameter adjustment coefficient, which controls the adjustment range of the regularization strength. |ΔF t | represents the change amplitude of the system feedback feature vector, indicating the degree of change in users' feedback.

[0215] Furthermore, by combining the long-term behavior data and immediate feedback of users, a policy optimization and feedback loop is established. By continuously monitoring the performance of the system under different permission configurations, the parameters of the policy model are updated regularly to ensure the optimal performance of the system during long-term operation.

[0216] Collect long-term behavior data of users (such as permission adjustment frequency, changes in user satisfaction, emotional state fluctuations, etc.), and regularly update the parameters of the reinforcement learning model, including the policy parameter θ, the regularization parameter λ, and the learning rate α, to ensure that the model can adapt to the long-term changes in users' behavior and emotional states.

[0217] The embodiment of this application also provides an intelligent speaker permission design and management system, as Figure 2 shown. The system includes:

[0218] A user data collection module 101, which is used to collect user audio data, preprocess the audio data, and construct a Gaussian mixture model based on the preprocessed data to output an audio signature;

[0219] The configuration initialization module 102 is used to take the audio signature as an input to represent the user's identity characteristics, collect real-time environmental data and perform preprocessing, construct a user-context joint feature vector based on the user's identity characteristics combined with the preprocessed real-time environmental data, add a regularization term to the user-context joint feature vector to adjust the recognition accuracy, and then perform context awareness and dynamic permission configuration initialization to obtain an optimized permission level and permission score;

[0220] The permission initialization module 103 is used to, according to the optimized permission level and permission score, design a fuzzy matching and dynamic adjustment model for the real-time dynamic management of the intelligent speaker permissions if there is a complex voice environment with multiple languages and multiple accents; wherein, the fuzzy matching and dynamic adjustment model uses a feature weighted aggregation method to calculate the credibility score γ of the voice feature vector C new as follows:

[0221]

[0222] where γ represents the credibility score of the voice input, indicating the credibility of the current voice feature matching degree; M represents the total number of voice features; w j represents the weight of the j-th feature, indicating the importance of this feature when calculating the voice credibility; α j represents the fuzzy adjustment parameter of the j-th feature, controlling the sensitivity of feature matching; C new,j represents the component of the j-th voice feature vector; μ j represents the expected value of the j-th feature, indicating the typical value of this feature under normal circumstances;

[0223] Using the credibility score γ of the voice input and the optimized permission score According to the permission adjustment function Dynamically adjust the user's permission level as follows:

[0224]

[0225] where, represents the dynamically adjusted permission level, indicating the permission setting adjusted according to the voice credibility and the original permission score; represents the initial permission level output in the second step; η represents the permission adjustment gain coefficient, controlling the permission adjustment range; γ represents the credibility score of the voice input; δ represents the intermediate threshold of the credibility score; represents the optimized permission score; τ represents the benchmark threshold of the permission score; According to the calculated new permission level The intelligent speaker system immediately adjusts the user's permission configuration and dynamically changes the accessibility of functions;

[0226] The permission management module 104 is used to optimize the new permission level according to the combined complex voice environment Combined with the real-time user behavior data, a behavior monitoring and anomaly detection model is constructed to analyze the user permissions in real time through feature engineering and time series modeling, detect anomalies and monitor the user behavior, and dynamically adjust the permission level according to the results of anomaly detection and behavior monitoring According to the permission level Classify the emotional state of the user through the emotional classifier model, generate emotional labels, and propose a mapping function of emotional state and permission response to dynamically adjust the permission level according to the emotional state label and the abnormal behavior response strategy. According to the permission level calculated by the emotional state and permission response mapping function The intelligent speaker immediately applies the new permission configuration;

[0227] The system optimization module 105 is used to optimize the new permission level according to the combined complex voice environment Construct an "adaptive permission optimization model based on reinforcement learning", combine the long-term behavior and usage patterns of users, and adaptively learn the permission policy of the intelligent speaker through reinforcement learning for long-term policy optimization

[0228] The above-disclosed are only some preferred embodiments of the present invention. Of course, the scope of the rights of the present invention cannot be limited thereby. Those of ordinary skill in the art can understand all or part of the processes of implementing the above embodiments, and the equivalent changes made according to the claims of the present invention still fall within the scope covered by the present invention

Claims

1. A smart speaker permission design and management method, characterized in that: The method comprises: S1, collect user audio data and preprocess the audio data, build a Gaussian mixture model based on the preprocessed data to output the audio signature; S2, using the audio signature as input to represent the user's identity features, collecting real-time environment data and preprocessing it, constructing a user-context joint feature vector based on the user's identity features combined with the preprocessed real-time environment data, and adding a regularization term to the user-context joint feature vector to adjust the recognition accuracy, then performing context awareness and dynamic permission configuration initialization to obtain the initial permission level and permission score; S3. According to the initial permission level and permission score, if a complex voice environment with multiple languages ​​and multiple accents appears, a fuzzy matching and dynamic adjustment model is designed to dynamically manage the permissions of the smart speaker in real time; wherein the fuzzy matching and dynamic adjustment model uses a feature weighted aggregation method to calculate the voice feature vector C new The credibility score γ is expressed as follows: Among them, γ represents the credibility score of the speech input, which indicates the credibility of the current speech feature matching; M represents the total number of speech features; w j represents the weight of the jth feature, indicating the importance of this feature in calculating the credibility of speech; α j represents the fuzzy adjustment parameter of the jth feature, controlling the sensitivity of feature matching; C new,j represents the component of the jth speech feature vector; μ j represents the expected value of the jth feature, which represents the typical value of the feature under normal circumstances; Using the credibility score γ of voice input and the optimized authority score Adjusting functions based on permissions Dynamically adjust user permission levels It is expressed as follows: in, Indicates the dynamically adjusted permission level, indicating the permission setting adjusted according to the voice credibility and the original permission score; represents the initial authority level; η represents the authority adjustment gain coefficient, which controls the authority adjustment range; γ represents the credibility score of the voice input; δ represents the intermediate threshold of the credibility score; represents the optimized permission score; τ represents the baseline threshold of the permission score; based on the calculated new permission level The smart speaker system instantly adjusts the user's permission configuration and dynamically changes the accessibility of functions; S4. New permission level optimized based on complex voice environment Combined with real-time user behavior data, a behavior monitoring and anomaly detection model is built to analyze user permissions in real time through feature engineering and time series modeling, perform anomaly detection and behavior monitoring on user behavior, and dynamically adjust permission levels based on the results of anomaly detection and behavior monitoring. S5. According to the permission level Combined with the user's emotional state, the emotional state is classified through the emotion classifier model to generate emotional labels. According to the emotional labels and abnormal behavior response strategies, a mapping function between emotional state and permission response is proposed to dynamically adjust the permission level. The permission level calculated by the mapping function between emotional state and permission response is Smart speakers instantly apply new permission configurations; S6. According to the authority level Build an adaptive permission optimization model based on reinforcement learning, combine users' long-term behavior and usage patterns, and use reinforcement learning to adaptively learn the permission strategy of smart speakers for long-term strategy optimization.

2. A smart speaker permission design and management method according to claim 1, characterized in that: The preprocessing is to process the collected audio data x(t) to remove noise and perform standardization; Wherein, the probability density function of the Gaussian mixture model is expressed as: Wherein, C represents the extracted Mel frequency cepstral coefficient feature vector, which is used to represent the voiceprint feature of the audio; Represents the parameter set of the Gaussian mixture model; π k represents the mixing weight of the kth Gaussian component, satisfying And π k ≥0;μ k represents the mean vector of the kth Gaussian component; Σ k represents the covariance matrix of the kth Gaussian component; K represents the number of Gaussian components, which means the number of Gaussian distributions in the model; represents the multivariate normal distribution, expressed as: Where d represents the dimension of the feature vector C.

3. A smart speaker permission design management method according to claim 2, characterized in that: For each new audio input, the user identity is identified by calculating the log-likelihood estimate of the feature vector of the new audio under each user model, which is expressed as follows: Among them, logL(C new |S i ) represents the log-likelihood value of the feature vector of the new audio under the i-th user model, which is used to measure the matching degree between the audio input and the user model; C new The feature vector representing the new audio input; S i Represents the audio signature of the i-th user, i.e., the user model Θ i ; R represents the total number of frames of the new audio input, which represents the number of frames after the audio signal is segmented; P(C new (t)|Θ i ) represents the feature vector C of the tth frame new (t) In the user model Θ i The probability of the following.

4. The method for designing and managing permissions for a smart speaker according to claim 1, characterized in that: The real-time environmental data is collected through a variety of sensors, including microphones, cameras, light sensors, temperature and humidity sensors, and accelerometers. The real-time environmental feature vector is represented by E; The user-context joint feature vector is fused through a weighted multilayer perceptron model, which is expressed as follows: C e =σ(W2·σ(W1·[E norm ,S i ]+b1)+b2) Among them, C e represents the generated context feature vector, which represents the user's behavior characteristics in the current environment; [E norm ,S i ] represents the input feature after combining the environment feature vector and the user audio signature; W1, W2 represent weight matrices with dimensions of m×(n+1) and m×m, respectively, which are learned through model training; b1, b2 represent bias vectors with dimensions of m and m, respectively, which are learned through model training; σ represents a nonlinear activation function, which is used to capture nonlinear relationships; Combined with the context feature vector C e and user audio signature S i , generate user-context joint feature vector F = S i ,C e ]; In the user-context joint feature vector F = [S i ,C e ]Adding regularization terms can achieve higher recognition accuracy by weighted combination of user and environment features, as shown below: F opt =F+λ·(F⊙W reg ) Among them, F opt represents the optimized joint feature vector; λ represents the regularization parameter, which is used to balance the influence of the original features and the regularization term; ⊙ represents the element-wise product operation, which is used for element-by-element weighting; W reg Represents the weight matrix, which represents the relationship between features.

5. A smart speaker permission design management method according to claim 4, characterized in that: The configuration initialization includes: Design a permission scoring function based on the adjusted user-context joint feature vector to calculate the user's permission score in the current context It is expressed as follows: in, Indicates the permission score, which is used to reflect the user's permission level in the current context; represents the weight vector, indicating the importance of each feature; β represents the weight adjustment coefficient, controlling the overall weight of the joint features; The bi-norm of the joint feature vector is used as a measure of feature importance, which adds constraints on the balance between features. represents the bias term; Based on the calculated authority score and preset permission level threshold, and dynamically adjust permissions. It is expressed as follows: in, Indicates the initial permission level; Indicates the permission threshold, which is set according to user needs and system configuration.

6. A smart speaker permission design and management method according to claim 1, characterized in that: In S3, an error correction feedback mechanism is designed for the fuzzy matching and dynamic adjustment model to dynamically adjust the weights and parameters of the fuzzy matching model. The error correction mechanism is expressed as follows: Δw=∈·(γ target -c)·C new Among them, Δw represents the weight adjustment vector, which is used to correct the weight of speech features; ∈ represents the learning rate, which is used to control the step size of weight adjustment; γ target represents the target credibility score, which is usually set to the credibility level expected by the system; γ represents the credibility score of the current speech input; C new The feature vector representing the current speech input.

7. The method for designing and managing permissions for a smart speaker according to claim 1, characterized in that: According to the weighted behavior feature vector B w A monitoring and anomaly detection model based on long short-term memory network is constructed to predict the behavior characteristics of the next moment and identify abnormal behavior patterns. The update formula of the behavior monitoring and anomaly detection model is expressed as follows: h t =f LSTM (B w,t ,h t-1 ) Among them, h t represents the hidden state at the current time t, capturing the time dependency of user behavior; B w,t represents the weighted behavior feature vector at the current time t; h t-1 represents the hidden state at the previous moment t-1; f LSTM Represents the state update function of the long short-term memory network; By calculating the residual between the actual behavior characteristics and the predicted behavior characteristics, it is determined whether there is abnormal behavior; wherein the abnormal score ξ of the abnormal behavior is expressed as follows: Among them, ξ represents the anomaly score, which indicates the difference between the current behavior and the predicted behavior; B w,t Represents the actual weighted behavior feature vector at the current time t; Represents the weighted behavior feature vector of the next moment predicted by the long short-term memory network; σ 2 represents the variance of the prediction residuals, which is used to normalize the anomaly scores; According to the abnormal score ξ and the preset abnormal threshold θ, it is determined whether to trigger the permission adjustment or warning mechanism; the adjustment strategy is as follows: If ξ≤θ, no adjustment is made and the current permission level is maintained If ξ>θ, it is identified as abnormal behavior, and permissions are tightened or the user is prompted to authenticate based on the severity of the abnormality.

8. The smart speaker permission design and management method according to claim 1, characterized in that: The collection of the user's emotional state includes the collection of audio and video features, and then a fusion model based on the self-attention mechanism is used to capture the mutual influence between multimodal features. The specific formula is as follows: Where E represents the fused emotional feature vector, which is used to describe the user's current comprehensive emotional state; A represents the audio feature vector, with a dimension of m; V represents the video feature vector, with a dimension of n; W A and W V Represents the projection matrix, which is used to map audio and video features to the same space, with a dimension of m×d k and n×d k ;d k Represents the dimension of the key vector in the attention mechanism, which is used to scale the dot product operation; softmax represents the normalization function used to calculate the attention weight; Using the fused emotional feature vector E, the emotional state is classified through the emotional classifier model C(E) to generate the emotional label c, which is expressed as follows: Among them, c represents the emotion label of the current user, which is expressed as a discrete category; w i Represents the weight vector of the classifier model, the weight corresponding to the i-th class, with a dimension of d k ; b i Represents the bias term of the classifier model, the bias corresponding to the i-th class; According to the emotional label c and the abnormal behavior response strategy, a mapping function between emotional state and authority response is proposed Used to dynamically adjust permission levels It is expressed as follows: in, It represents the permission level dynamically adjusted according to the emotional state, and takes the user's emotions into consideration when setting the permission. ρ(c) represents the permission adjustment increment function, which depends on the user's emotional state c and is expressed as a positive or negative integer.

9. A smart speaker permission design management method according to claim 8, characterized in that: The adaptive permission optimization model based on reinforcement learning is based on the user's current permission level and emotional state, as well as the user's historical behavior pattern and system feedback, that is, the state vector Among them, H represents the user's historical behavior pattern feature vector; F represents the system feedback feature vector; the system's executable actions are defined as the permission adjustment operation set Each action a i Indicates the adjustment of the permission level; defines the optimization objective function J(θ), through the strategy π θ (S) is optimized to maximize the expected cumulative reward; The formula for optimizing the objective function J(θ) is as follows: Where J(θ) represents the optimization objective function of the policy parameter θ, which represents the expected cumulative reward; S represents the current state vector, which consists of the authority level, emotional state, historical behavior pattern, and system feedback; represents the state distribution set, which represents the possible state space; ω represents the discount factor, which ranges from 0≤ω≤1 and is used to measure the importance of future rewards; r(S t ,a t ) indicates that in state S t Execute action a t The immediate reward obtained when the authority is adjusted reflects the effectiveness and rationality of the authority adjustment; t represents the action performed at time t, given by the strategy π θ (S) confirm; In the process of policy optimization, the policy regularization term Ω(θ) is introduced to prevent the policy from overfitting the user's short-term behavior pattern, and the policy parameter θ is updated to maximize the regularized optimization objective, which is expressed as follows: Among them, θ t+1 represents the updated policy parameters; θ t represents the policy parameter at the current time t; α represents the learning rate, which controls the step size of parameter update; represents the policy gradient, which represents the gradient of the optimization objective function to the policy parameters; λ represents the regularization strength parameter, which controls the influence of the regularization term on the policy update; Ω(θ t ) represents the policy regularization term, which is a sparsity regularization that encourages simplicity and generalization of policy parameters, and is expressed as follows: Where N represents the total number of policy parameters; θ i Represents the i-th parameter in the strategy parameters.

10. A smart speaker authority design and management system, characterized in that: The system comprises: A user data collection module is used to collect user audio data and preprocess the audio data, and construct a Gaussian mixture model based on the preprocessed data to output an audio signature; Configuration initialization module, used to take audio signature as input to represent the identity characteristics of the user, collect real-time environment data and pre-process it, build a user-context joint feature vector based on the user's identity characteristics combined with the pre-processed real-time environment data, add regularization terms to the user-context joint feature vector to adjust the recognition accuracy, and then perform context awareness and dynamic permission configuration initialization to obtain the initial permission level and permission score; The permission initialization module is used to design a fuzzy matching and dynamic adjustment model to dynamically manage the permissions of the smart speaker in real time according to the initial permission level and permission score if a complex voice environment with multiple languages ​​and multiple accents appears; wherein the fuzzy matching and dynamic adjustment model uses a feature weighted aggregation method to calculate the voice feature vector C new The credibility score γ is expressed as follows: Among them, γ represents the credibility score of the speech input, which indicates the credibility of the current speech feature matching; M represents the total number of speech features; w j represents the weight of the jth feature, indicating the importance of this feature in calculating the credibility of speech; α j represents the fuzzy adjustment parameter of the jth feature, controlling the sensitivity of feature matching; C new,j represents the component of the jth speech feature vector; μ j represents the expected value of the jth feature, which represents the typical value of the feature under normal circumstances; Using the credibility score γ of voice input and the optimized authority score Adjusting functions based on permissions Dynamically adjust user permission levels It is expressed as follows: in, Indicates the dynamically adjusted permission level, indicating the permission setting adjusted according to the voice credibility and the original permission score; represents the initial authority level; η represents the authority adjustment gain coefficient, which controls the authority adjustment range; γ represents the credibility score of the voice input; δ represents the intermediate threshold of the credibility score; represents the optimized permission score; τ represents the baseline threshold of the permission score; based on the calculated new permission level The smart speaker system instantly adjusts the user's permission configuration and dynamically changes the accessibility of functions; The permission management module is used to optimize the new permission level based on the complex voice environment. Combined with real-time user behavior data, a behavior monitoring and anomaly detection model is built to analyze user permissions in real time through feature engineering and time series modeling, perform anomaly detection and behavior monitoring on user behavior, and dynamically adjust permission levels based on the results of anomaly detection and behavior monitoring. Based on permission level Combined with the user's emotional state, the emotional state is classified through the emotion classifier model to generate emotional labels. According to the emotional labels and abnormal behavior response strategies, a mapping function between emotional state and permission response is proposed to dynamically adjust the permission level. The permission level calculated by the mapping function between emotional state and permission response is Smart speakers instantly apply new permission configurations; System optimization module for Build an adaptive permission optimization model based on reinforcement learning, combine users' long-term behavior and usage patterns, and use reinforcement learning to adaptively learn the permission strategy of smart speakers for long-term strategy optimization.

Citation Information

Patent Citations

  • Systems and methods with integrated gaming engines and smart contracts

    CA3177530A1

  • Permission management method and device, mobile terminal and storage medium

    CN108763892A