Identity verification system and method based on Internet of Things and voiceprint recognition
By designing a voiceprint identification authentication system with multimodal convergence and distributed storage in the Internet of Things devices, the problem of insufficient comprehensiveness and reliability of identity authentication in the existing technology is solved, higher authentication accuracy and security is achieved, user portraits are dynamically updated, and security risks brought about by data centralization are reduced.
Patent Information
- Application Number
- CN202510105084.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-01-23
AI Technical Summary
When using voiceprint identification authentication in IoT devices in prior art, there are low comprehensiveness and reliability of verification, inability to dynamically update user profiles, poor system adaptability and generalization capabilities, low verification accuracy and fault tolerance capabilities, and security risks brought about by data centralization.
An identity authentication system based on the Internet of Things and voiceprint recognition is designed, including a collection preprocessing module, a multi-modal fusion module, a storage and transmission module, an identity screening module, an identity verification module, a dynamic risk assessment module, a model update module, a feedback correction module, an authorization record module and a monitoring self-optimization module. The system improves the accuracy and security of identity verification through technical means such as multimodal feature fusion, distributed storage and encrypted transmission, dynamic risk assessment and model collaborative optimization.
It achieves higher comprehensiveness and reliability of identity verification, dynamically updates user profiles, improves the system's adaptability and generalization capabilities, enhances the accuracy and fault tolerance of verification, and protects user privacy through distributed storage and encrypted transmission, and reduces the security risks brought about by data centralization.
Smart Images

Figure CN119989321A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of identity recognition, and in particular to an identity recognition system and method based on the Internet of Things and voiceprint recognition. Background Art
[0002] With the rapid development of Internet of Things (IoT) technology, a large number of smart devices and systems are interconnected, building a highly collaborative ecological environment. This ecology requires a reliable authentication mechanism to protect user privacy and system security. Although traditional password or biometric authentication methods are popular, they often have the risk of poor user experience and information theft. As a contactless and natural biometric, voiceprint recognition has become a research hotspot in the field of identity authentication because of its convenience and uniqueness. The distributed characteristics of IoT devices and the dynamic network environment make voiceprint recognition face challenges in large-scale identity authentication applications, such as differences in transmission quality between devices, computing resource limitations, and the impact of user behavior and environment on voiceprints. In addition, how to ensure data privacy in the case of multi-device collaboration while improving the accuracy and robustness of verification has also become an urgent problem to be solved.
[0003] After searching, China Publication No. CN114579947A discloses an identity authentication method and system based on the Internet of Things and voiceprint recognition. Although the invention only requires the person to be authenticated to speak according to the prompt, and the person to be authenticated can wear a protective user, it does not affect the identity authentication and reduces the possibility of disease transmission. However, the comprehensiveness and reliability of the verification are low, and it is impossible to dynamically update the user portrait, and it is impossible to build personalized verification rules for different users; and the existing identity authentication system and method cannot provide users with a smoother and more friendly identity authentication experience, the overall verification accuracy and fault tolerance of the system are reduced, and the computational pressure and time consumption of the subsequent deep comparison module are increased; in addition, the existing identity authentication system and method are prone to security risks caused by data centralization, the system's adaptability and generalization ability are poor, and the communication cost and resource consumption are increased. For this reason, we propose an identity authentication system and method based on the Internet of Things and voiceprint recognition. Summary of the invention
[0004] The purpose of the present invention is to solve the defects in the prior art and to propose an identity authentication system and method based on the Internet of Things and voiceprint recognition.
[0005] In order to achieve the above object, the present invention adopts the following technical solutions: The identity authentication system based on the Internet of Things and voiceprint recognition includes an acquisition preprocessing module, a multimodal fusion module, a storage and transmission module, an identity screening module, an identity authentication module, a dynamic risk assessment module, a model update module, a feedback correction module, an authorization record module, and a monitoring self-optimization module; The acquisition and preprocessing module is used to acquire the user's voice data in real time and preprocess the acquired voice data; The multimodal fusion module is used to extract voiceprint features from voice data and combine voice features with user behavior patterns to construct multimodal features; The storage and transmission module is used for distributed storage of user voiceprint features and constructed multimodal features, and for encrypted transmission of each set of feature data; The identity screening module is used to perform a preliminary comparison on the collected voiceprint features to screen out candidate identities; The identity verification module is used to construct a voiceprint recognition model, and to perform in-depth comparison and analysis on the candidate identities initially screened through the voiceprint recognition model to verify the user's identity; The dynamic risk assessment module is used to assess the risk level of identity authentication in real time and dynamically adjust the authentication strategy according to the risk; The model updating module is used to collaboratively optimize the voiceprint recognition model through distributed devices; The feedback correction module is used to provide instant feedback to users and correct potential errors; The authorization recording module is used to authorize the user's operation request after the verification is passed, and record the operation log at the same time; The monitoring self-optimization module is used to continuously monitor the system operation status and regularly optimize the parameter configuration in the identity authentication process.
[0006] As a further solution of the present invention, the specific steps of preprocessing the collected voice signal by the acquisition preprocessing module are as follows: real-time monitoring of the network transmission status between devices, and collecting network quality indicators such as bandwidth, delay, jitter and packet loss rate, calculating the transmission quality of each monitoring device network, and then dynamically adjusting the voice sampling rate according to the network quality threshold and the transmission quality of the monitoring device network, and collecting voice data in real time according to the adjusted voice sampling rate, performing short-time Fourier transform on the collected voice data to obtain the corresponding frequency domain signal, and then calculating the power spectrum and noise power spectrum of the voice signal by minimum mean square error to construct a corresponding Wiener filter, and then filtering the frequency domain signal by the Wiener filter to obtain a denoised voice spectrum. After filtering, convolving the voice signal with the room impulse response to construct a room reverberation signal model, and then using a known test signal to measure the impulse response of the room, and designing an inverse filter through frequency domain representation, and performing inverse filtering on the spectrum of the reverberated voice signal to obtain a dereverberated spectrum, and using an inverse short-time Fourier transform to restore the dereverberated spectrum to the original voice data.
[0007] As a further solution of the present invention, the multimodal fusion module constructs multimodal features in the following specific steps: S1.1: Extract the Mel frequency cepstrum coefficients, linear prediction coefficients, energy and pitch feature information from the preprocessed speech data, and record the extracted voiceprint features as , then extract the behavior features and state features from the user behavior pattern and device status respectively, and record them as and ; S1.2: Standardize each dimension of voiceprint features, behavioral features, and state features from different sources to eliminate feature scale differences, unify feature dimensions through zero padding, truncation, and dimensionality reduction operations, assign weights to each modality feature, and perform weighted fusion of features from different modalities based on the assigned weights; S1.3: Use the Min-Max normalization method to map the fused multimodal features to the interval [0, 1], and use the normalized multimodal features as The multimodal feature space is constructed according to the number of processed multimodal features, and a group of populations are initialized in the multimodal feature space, and the position of each group of individuals in the population is used as a multimodal feature of the multimodal feature space; S1.4: Based on the separation of feature representations between different users and the stability of feature representations input by the same user multiple times, the fitness value of each multimodal feature in the multimodal feature space is calculated, and the individuals in each group are sorted in descending order according to the fitness value. The top three groups of individuals with the best performance are selected from the population, and the remaining individuals are moved closer to the top three groups of individuals with the highest fitness values and the positions of the remaining individuals are updated; S1.5: After each position update, the individuals with the updated positions are randomly disturbed. If the multimodal feature fitness value corresponding to the disturbed position is better than the original value, the position update after the disturbance is accepted. After the position update is completed, the fitness values of each individual in the population are recalculated and sorted, and the top three groups of individuals with the best performance are reselected to update the position again; S1.6: Repeat the fitness ranking, best individual update and position update until the preset maximum number of iterations is reached or the change in the fitness value of the first-ranked individual is less than the set threshold. Then stop the iteration and take the first-ranked multimodal feature as the optimal feature representation and add it to the user feature library to update the user feature library.
[0008] As a further solution of the present invention, the specific steps of distributed storage of the storage transmission module are as follows: S2.1: Create a decentralized storage structure for voiceprint features in the blockchain network, and define each block in the blockchain to store voiceprint feature data, multimodal feature data and their hash values. Then, a proof-of-work consensus mechanism is adopted, and each node in the blockchain searches for a hash value that is less than the difficulty threshold set by the network through calculation; S2.2: Broadcast the block corresponding to the hash value that meets the conditions obtained through the consensus mechanism to the entire network and add it to the chain, making each storage block in the blockchain a state node , and the connection between nodes represents the transmission path of data in the blockchain network, where NodeID represents the unique identifier of the current storage node. Represents the cumulative reward value of the current node, Represents the number of times the current node has been visited; S2.3: Starting from the starting node of the blockchain as the root node, a search tree is constructed based on the subsequent child nodes and the root node in the blockchain network, the UCB value of each child node is calculated, and the child node with the highest UCB value is selected as the next storage node. If the current child node is not fully expanded, the child node is expanded and the new node is added to the search tree; S2.4: Starting from the newly added node, randomly simulate the storage path, calculate the storage cost and retrieval delay, update the cumulative reward value and access number of the search tree according to the simulation results, repeat the selection, expansion, simulation and backtracking steps until the preset maximum number of iterations is reached. After the iteration, select the path with the highest cumulative reward value as the storage path for voiceprint feature data and multimodal feature data, and dynamically adjust the data distribution plan in the blockchain.
[0009] As a further solution of the present invention, each block in the blockchain described in S2.1 is represented as follows: ;in, Representative blocks; Representative Voiceprint feature data or multimodal feature data of each block; Representative The hash value of the block; The hash value representing the voiceprint data of this block; Representative The random number in the proof of work of each block; The specific calculation formula of the UCB value described in S2.3 is as follows: ; In the formula, Represents the selection value of the current node s; Represents the cumulative reward value of the current node s; represents the number of times the current node s has been visited; c represents the parameter for balancing exploration and utilization, which is usually a constant; Represents the number of times the parent node of the current node has been visited; ln represents a logarithmic function, which is used to adjust the priority of unvisited nodes.
[0010] As a further solution of the present invention, the identity screening module performs preliminary comparison in the following specific steps: S3.1: Extract voiceprint features from the user's current voice data, select pre-registered voiceprint features that match the current user from the user feature library, and initialize the candidate identity pool based on the extracted voiceprint features. Measure the similarity between the features to be verified and the identity features in the candidate identity pool through dynamic weighted distance, and screen out the identity features in the candidate identity pool whose similarity does not meet the preset threshold; S3.2: Initialize the feature weight population and the candidate identity pool after the initial screening, and calculate the fitness value of each feature weight and candidate identity feature respectively. Then retain the feature weights whose fitness is higher than the preset threshold, randomly select two groups of individuals from the selected feature weights, perform linear combination to generate new feature weights, and then randomly adjust some weight components so that the new feature weights meet the preset constraints; S3.3: Select candidate identity features whose fitness is lower than the preset threshold, and replace them with new identity features that have not been selected in the user feature library, and re-update the feature weight population and the candidate identity pool based on the candidate identity features whose fitness is higher than the preset threshold and the matching rules, until the fitness value of the feature weight population changes less than the preset threshold in multiple rounds; S3.4: After stopping the update, the feature weight with the highest fitness is selected as the weight distribution of the final matching rule. Then, the matching distance between the feature to be verified and each identity feature in the candidate identity pool is calculated according to the optimized rule, and the maximum threshold of the matching distance is set. The candidate identities with a matching distance less than the maximum threshold are considered to have passed the screening.
[0011] As a further solution of the present invention, the specific calculation formula of the dynamic weighted distance in S3.1 is as follows: ; In the formula, Represents the voiceprint feature to be verified With The matching distance of candidate identities v, where is the index of candidate identity v in the candidate identity pool; d represents the voiceprint feature to be verified The number of dimensions; Representative The weight of the dimension feature is initially uniformly distributed; Represents the voiceprint feature to be verified No. Dimension value; Representative Candidate Status No. Dimension value; The specific calculation formulas for the feature weights and the fitness values of the candidate identity features described in S3.2 are as follows: , ; In the formula, Representative The fitness value of the feature weight w; represents the total number of eigenvectors; Represents the attenuation factor of the matching distance. The smaller the distance, the greater the corresponding contribution. Represents the voiceprint feature to be verified With The similarity score of the candidate identity v; Representative The fitness value of the candidate identity v is the minimum matching distance under the current optimal weight vector.
[0012] The identity authentication method based on the Internet of Things and voiceprint recognition has the following specific steps: Ⅰ. Collect user voice data and device environment information through IoT devices, and pre-process the voice data; Ⅱ. Extract voiceprint features from the processed voice data, combine the user's behavior pattern and device status information, form multimodal features and update the user feature library; III. Construct and optimize feature matching rules, make preliminary matching results based on feature matching rules, and screen out candidate identity sets from the user feature library; IV. Construct a voiceprint recognition model to conduct in-depth comparison and analysis on the initially screened candidate identity set to verify the user's identity; V. Real-time assessment of identity authentication risk level, dynamic adjustment of authentication strategy, and collaborative optimization of voiceprint recognition model; VI. Confirm the user's identity based on the verification result, feed back the result to the corresponding user device, authorize the user's operation request, and record the operation log; Ⅶ. Monitor the system operation status in real time and regularly optimize the parameter configurations in the identity authentication process.
[0013] As a further solution of the present invention, the specific steps of the depth comparison and analysis in step IV are as follows: S4.1: Design and construct a voiceprint recognition model based on the Bi-GRU architecture. The model includes an input layer, a Bi-GRU layer, a fully connected layer, a classification layer, and an output layer. Then, based on the public voiceprint dataset, construct a training set, a test set, and a validation set. S4.2: The training set is input into the voiceprint recognition model. The Bi-GRU layer of the voiceprint recognition model processes the input data in chronological order and reverse chronological order respectively, and outputs the concatenated vectors of the forward and reverse features. The fully connected layer aggregates the received concatenated vectors into a fixed-length voiceprint embedding vector. Then the classification layer outputs the final identity recognition result based on the candidate identity set. S4.3: The loss value between the final identification result of the model and the actual identification result is calculated through the cross entropy loss function. Then, based on the chain rule, the loss value is propagated from the output layer to the input layer of the voiceprint recognition model, and the gradient of each network layer of the model corresponding to the loss value is calculated. Then, the parameters of each network layer are optimized based on the Adam optimizer. The performance of the model after this round of training is evaluated through the verification machine. If the model performance does not reach the preset performance index, the model is retrained until the model loss value converges to the preset threshold, and the training is stopped. Then, the performance of the trained voiceprint recognition model on unknown data is tested through the test set, and the model is deployed on the actual voiceprint recognition system. S4.4: Use the current voiceprint feature as the current state Input to the voiceprint recognition model, based on the policy network , select the current action , that is, adjusting the parameters or hyperparameters of the voiceprint recognition model, recalculating the predicted output by adjusting the parameters or hyperparameters of the voiceprint recognition model, and based on the updated model performance, ; Calculate the reward value, where Represents the selection of the current action After the model performance evaluation reward, Represents the recognition accuracy of the model, Represents the similarity between the target domain features and the source domain embedding, Represent weight parameters respectively; S4.5: Feedback the reward value to the policy network for updating, and accumulate the reward value as the policy optimization target, with maximizing the cumulative reward as the target of the policy network, and use the gradient to update the policy network parameters. After that, the voiceprint recognition model is optimized in the target field according to the migrated task data, and combined with the contrast loss function, the model parameters are updated, and multiple rounds of reinforcement learning training are repeated until the model loss value converges to the preset range; S4.6: The voiceprint recognition model receives the current voiceprint features, and after forward propagation processing, generates an embedding vector of the voiceprint to be verified, calculates the similarity between the embedding vector and each candidate identity in the candidate identity set, and after comparing all candidate identities, screens out one or more matching identities with a similarity higher than a preset threshold, sorts the candidate identities in descending order according to the similarity score, and gives priority to the identity with the highest similarity score.
[0014] As a further solution of the present invention, the specific steps of collaboratively optimizing the voiceprint recognition model are as follows: S5.1: After the verification strategy is adjusted, each IoT device collects the voice data of its user interaction to form a local data set, and converts the original voice data into feature representation through voice preprocessing technology. Then, the processed local data set of each device is defined as ;in Representative Device D, Representative Voice features, Representative The user identity label corresponding to the voice feature, Representative The amount of data on device D; S5.2: Each device initializes a set of model architectures that are the same as the original voiceprint recognition model, and initializes the model parameters on each device to the unified weights distributed by the global server. At the same time, the local loss function, optimizer, and hyperparameters are set for each group of devices. After that, the local device uses the current model parameters to predict each sample in the data set locally, and evaluates the error value between the prediction result and the true label through the cross entropy loss function; S5.3: Based on the chain rule, the error value is propagated layer by layer, and the gradient of the parameters of each layer of the model is calculated according to the error value. Then, the model parameters are updated using the SGD optimizer, and the parameter update amount of the local model is calculated and stored. Each group of local IoT devices uploads the local model parameter update amount to the terminal server; S5.4: Based on the characteristics of each IoT device, the spatial range of local model parameters, and the local gradient trend information, a set of belief spaces are initialized for each IoT device. According to the voice characteristics of the device and the user behavior pattern, multiple groups of local behaviors are generated, and the fitness of each behavior in the local task is evaluated. The behavior with the highest fitness is selected as the optimal behavior of the IoT device. Each group of IoT devices generates a belief update vector based on its optimal behavior. The server summarizes the optimal behavior of each device and integrates the belief updates of the devices by averaging to generate a global belief vector. S5.5: Utilize global beliefs and local model updates to calculate the weight of each device. The terminal server then performs weighted aggregation on the model parameter updates of all devices based on the device weights to obtain the corresponding global model updates. The updated global model is distributed to all IoT devices, and the local model parameters of each IoT device are synchronized to the global model parameters, while updating and training are performed in real time.
[0015] Compared with the prior art, the present invention has the following beneficial effects: 1. The present invention constructs a multimodal feature space according to the number of processed multimodal features, and initializes a group of populations in the multimodal feature space, takes the position of each group of individuals in the population as a multimodal feature of the multimodal feature space, calculates the fitness value of each multimodal feature, and sorts each group of individuals in descending order according to the fitness value, selects the top three groups of individuals with the best performance from the population, and the remaining individuals approach the top three groups of individuals with the highest fitness values and update the positions of the remaining individuals, and after each position update, randomly perturbs the individual after the position update, and if the multimodal feature fitness value corresponding to the perturbed position is better than the original value, the position update after the perturbation is accepted, and after the position update is completed, The fitness value of each individual in the population is recalculated and sorted, and the top three groups of individuals with the best performance are reselected, and the position is updated again. The fitness arrangement, optimal individual update and position update are repeated until the preset maximum number of iterations is reached or the change in the fitness value of the first-ranked individual is less than the set threshold. The iteration is stopped, and the first-ranked multimodal feature is represented as the optimal feature and added to the user feature library to update the user feature library, improve the comprehensiveness and reliability of verification, and realize dynamic user portrait updates. To achieve mutual enhancement of multimodality, personalized verification rules can be built for different users, and the system's flexible response capabilities to specific user needs can be improved.
[0016] 2. The present invention extracts voiceprint features from the user's current voice data, selects pre-registered voiceprint features that match the current user from the user feature library, and initializes the candidate identity pool based on the extracted voiceprint features, measures the similarity between the features to be verified and the identity features in the candidate identity pool through a dynamic weighted distance, and screens out the identity features in the candidate identity pool whose similarity does not meet a preset threshold, initializes the feature weight population and the candidate identity pool after the initial screening, and calculates the fitness value of each feature weight and the candidate identity feature respectively, then retains the feature weights with a fitness higher than the preset threshold, randomly selects two groups of individuals from the selected feature weights, performs linear combination to generate new feature weights, and then randomly adjusts some weight components to make the new feature weights meet the preset constraints, selects the candidate identity features with a fitness lower than the preset threshold, and replaces them It is replaced with a new identity feature that has not been selected in the user feature library, and based on the candidate identity features and matching rules whose fitness is higher than the preset threshold, the feature weight population and the candidate identity pool are updated again until the fitness value of the feature weight population changes less than the preset threshold in multiple rounds. After stopping the update, the feature weight with the highest fitness is selected as the weight distribution of the final matching rule, and then the matching distance between the feature to be verified and each identity feature in the candidate identity pool is calculated according to the optimized rule, and the maximum threshold of the matching distance is set, and the candidate identities with a matching distance less than the maximum threshold are regarded as having passed the screening, which can provide users with a smoother and more friendly identity authentication experience, improve the overall verification accuracy and fault tolerance of the system, reduce the calculation pressure and time consumption of the subsequent deep comparison module, and can flexibly deal with environmental noise and data deviations.
[0017] 3. Each IoT device in the present invention collects voice data interacting with its user to form a local data set, and converts the original voice data into feature representation through voice preprocessing technology. Each device initializes a set of model architectures that are the same as the original voiceprint recognition model, and initializes the model parameters on each device to the unified weights distributed by the global server. At the same time, a local loss function, optimizer and hyperparameters are set for each group of devices. After that, the local device uses the current model parameters to predict each sample in the data set locally, and updates the model parameters, calculates and stores the parameter update amount of the local model, and each group of local IoT devices uploads the local model parameter update amount to the terminal server. According to the characteristics of each IoT device, the local model parameter space range and the local gradient trend information, a set of belief spaces are initialized for each IoT device, and multiple groups of local behavior patterns are generated according to the voice characteristics of the device and the user behavior pattern. And evaluate the fitness of each behavior in the local task, select the behavior with the highest fitness as the optimal behavior of the IoT device, each group of IoT devices generates a belief update vector according to its optimal behavior, the server summarizes the optimal behavior of each device, integrates the belief update of the device by averaging to generate a global belief vector, uses the global belief and local model update to calculate the weight of each device, and then the terminal server performs weighted aggregation on the model parameter update of all devices according to the device weight to obtain the corresponding global model update, distributes the updated global model to all IoT devices, and synchronizes the local model parameters of each IoT device to the global model parameters, and performs update training in real time, which can protect the privacy of users, avoid the security risks brought by data centralization, improve the adaptability and generalization ability of the system, reduce communication costs and resource consumption, solve the problem of uneven data distribution, and improve the overall performance of the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The accompanying drawings are used to provide further understanding of the present invention and constitute a part of the specification. They are used to explain the present invention together with the embodiments of the present invention and do not constitute a limitation of the present invention.
[0019] Figure 1 This is a system block diagram of the identity authentication system based on the Internet of Things and voiceprint recognition proposed by the present invention; Figure 2 This is a flowchart of the identity authentication method based on the Internet of Things and voiceprint recognition proposed by the present invention. DETAILED DESCRIPTION
[0020] The technical solutions in the embodiments of the present invention will be described clearly and completely below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments.
[0021] Example 1, reference Figure 1, an identity authentication system based on the Internet of Things and voiceprint recognition, including an acquisition preprocessing module, a multimodal fusion module, a storage and transmission module, an identity screening module, an identity authentication module, a dynamic risk assessment module, a model updating module, a feedback correction module, an authorization record module, and a monitoring self-optimization module; The collection and preprocessing module is used to collect the user's voice data in real time and preprocess the collected voice data.
[0022] It should be further explained that the specific steps of the acquisition preprocessing module preprocessing the collected voice signal are as follows: real-time monitoring of the network transmission status between devices, and collecting network quality indicators such as bandwidth, delay, jitter and packet loss rate, calculating the transmission quality of each monitoring device network, and then dynamically adjusting the voice sampling rate according to the network quality threshold and the transmission quality of the monitoring device network, and collecting voice data in real time according to the adjusted voice sampling rate, performing short-time Fourier transform on the collected voice data to obtain the corresponding frequency domain signal, and then calculating the power spectrum and noise power spectrum of the voice signal through the minimum mean square error to construct the corresponding Wiener filter, and then filtering the frequency domain signal through the Wiener filter to obtain the denoised voice spectrum. After filtering, the voice signal is convolved with the room impulse response to construct a room reverberation signal model, and then using a known test signal to measure the impulse response of the room, and through frequency domain representation, designing an inverse filter, and inverse filtering the spectrum of the reverberated voice signal to obtain the dereverberated spectrum, and using the inverse short-time Fourier transform to restore the dereverberated spectrum to the original voice data.
[0023] The multimodal fusion module is used to extract voiceprint features from speech data and combine speech features with user behavior patterns to construct multimodal features.
[0024] Specifically, the Mel frequency cepstrum coefficients, linear prediction coefficients, energy and pitch feature information are extracted from the preprocessed speech data, and the extracted voiceprint features are recorded as , then extract the behavior features and state features from the user behavior pattern and device status respectively, and record them as and , standardize each dimension of voiceprint features, behavioral features, and state features from different sources to eliminate feature scale differences, unify feature dimensions through zero padding, truncation, and dimensionality reduction operations, assign weights to each modal feature, and perform weighted fusion of features of different modalities based on the assigned weights. Use the Min-Max normalization method to map the fused multimodal features to the [0, 1] interval, and use the normalized multimodal features as The multimodal feature space is constructed according to the number of processed multimodal features, and a group of populations is initialized in the multimodal feature space. The position of each group of individuals in the population is used as a multimodal feature in the multimodal feature space. Based on the separation degree of feature representation between different users and the stability of feature representation of multiple inputs by the same user, the fitness value of each multimodal feature in the multimodal feature space is calculated, and each group of individuals is sorted in descending order according to the fitness value. The top three groups of individuals with the best performance are selected from the population, and the remaining individuals are respectively close to the top three groups of individuals with the highest fitness values and the positions of the remaining individuals are updated. After each position update, the position The updated individuals are randomly disturbed. If the multimodal feature fitness value corresponding to the disturbed position is better than the original value, the position update after the disturbance is accepted. After the position update is completed, the fitness values of each individual in the population are recalculated and sorted, and the top three groups of individuals with the best performance are reselected to update the position again. The fitness arrangement, optimal individual update and position update are repeated until the preset maximum number of iterations is reached or the fitness value change of the first-ranked individual is less than the set threshold. The iteration is stopped, and the first-ranked multimodal feature is represented as the optimal feature and added to the user feature library to update the user feature library.
[0025] The storage and transmission module is used for distributed storage of user voiceprint features and constructed multimodal features, and for encrypted transmission of each set of feature data.
[0026] Specifically, a decentralized storage structure of voiceprint features is created in the blockchain network, and each block in the blockchain is defined to store voiceprint feature data, multimodal feature data and their hash values. Then, the proof-of-work consensus mechanism is adopted. Each node in the blockchain calculates and finds a hash value that is less than the difficulty threshold set by the network. The block corresponding to the hash value that meets the conditions obtained through the consensus mechanism is broadcast to the entire network and added to the chain. Each storage block in the blockchain is used as a state node. , and the connection between nodes represents the transmission path of data in the blockchain network, where NodeID represents the unique identifier of the current storage node. Represents the cumulative reward value of the current node, Represents the number of times the current node has been visited. The starting node of the blockchain is used as the root node. A search tree is constructed based on the subsequent child nodes and the root node in the blockchain network. The UCB value of each child node is calculated, and the child node with the highest UCB value is selected as the next storage node. If the current child node is not fully expanded, the child node is expanded, and the new node is added to the search tree. Starting from the newly added node, the storage path is randomly simulated, the storage cost and retrieval delay are calculated, and the cumulative reward value and the number of visits of the search tree are updated according to the simulation results. The selection, expansion, simulation and backtracking steps are repeated until the preset maximum number of iterations is reached. After the iteration, the path with the highest cumulative reward value is selected as the storage path for the voiceprint feature data and the multimodal feature data, and the data distribution scheme in the blockchain is dynamically adjusted.
[0027] In this embodiment, each block in the blockchain is represented as follows: ;in, Representative blocks; Representative Voiceprint feature data or multimodal feature data of each block; Representative The hash value of the block; The hash value representing the voiceprint data of this block; Representative The random number in the proof of work of each block; The specific calculation formula of UCB value is as follows: ; In the formula, Represents the selection value of the current node s; Represents the cumulative reward value of the current node s; represents the number of times the current node s has been visited; c represents the parameter for balancing exploration and utilization, which is usually a constant; Represents the number of times the parent node of the current node has been visited; ln represents a logarithmic function, which is used to adjust the priority of unvisited nodes.
[0028] The identity screening module is used to perform a preliminary comparison of the collected voiceprint features and screen out candidate identities.
[0029] Specifically, voiceprint features are extracted from the user's current voice data, pre-registered voiceprint features matching the current user are selected from the user feature library, and the candidate identity pool is initialized based on the extracted voiceprint features. The similarity between the feature to be verified and the identity features in the candidate identity pool is measured by a dynamic weighted distance, and the identity features in the candidate identity pool whose similarity does not meet a preset threshold are screened out. The feature weight population and the candidate identity pool after the initial screening are initialized, and the fitness value of each feature weight and the candidate identity feature is calculated respectively. The feature weights whose fitness is higher than the preset threshold are then retained, and two groups of individuals are randomly selected from the selected feature weights, and a linear combination is performed to generate a new feature weight, and then some weight components are randomly adjusted. Make the new feature weights meet the preset constraints, select candidate identity features with fitness lower than the preset threshold, and replace them with new identity features that have not been selected in the user feature library, and re-update the feature weight population and the candidate identity pool based on the candidate identity features with fitness higher than the preset threshold and the matching rules, until the fitness value of the feature weight population changes less than the preset threshold in multiple rounds. After stopping the update, select the feature weight with the highest fitness as the weight distribution of the final matching rule, and then calculate the matching distance between the feature to be verified and each identity feature in the candidate identity pool according to the optimized rules, and set the maximum threshold of the matching distance, and consider the candidate identities with a matching distance less than the maximum threshold as having passed the screening.
[0030] It should be further explained that the specific calculation formula of dynamic weighted distance is as follows: ; In the formula, Represents the voiceprint feature to be verified With The matching distance of candidate identities v, where is the index of candidate identity v in the candidate identity pool; d represents the voiceprint feature to be verified The number of dimensions; Representative The weight of the dimension feature is initially uniformly distributed; Represents the voiceprint feature to be verified No. Dimension value; Representative Candidate Status No. Dimension value; The specific calculation formulas for feature weights and the fitness values of candidate identity features are as follows: , ; In the formula, Representative The fitness value of the feature weight w; represents the total number of eigenvectors; Represents the attenuation factor of the matching distance. The smaller the distance, the greater the corresponding contribution. Represents the voiceprint feature to be verified With The similarity score of the candidate identity v; Representative The fitness value of the candidate identity v is the minimum matching distance under the current optimal weight vector.
[0031] The identity authentication module is used to build a voiceprint recognition model, and use the voiceprint recognition model to conduct in-depth comparison and analysis of the candidate identities initially screened out to verify the user's identity; the dynamic risk assessment module is used to evaluate the risk level of identity authentication in real time, and dynamically adjust the verification strategy according to the risk; the model update module is used to collaboratively optimize the voiceprint recognition model through distributed devices; the feedback correction module is used to provide users with instant feedback and correct potential errors; the authorization recording module is used to authorize the user's operation request after the verification is passed, and record the operation log at the same time; the monitoring self-optimization module is used to continuously monitor the system operation status and regularly optimize the parameter configuration in the identity authentication process.
[0032] Example 2, reference Figure 2 , an identity verification method based on the Internet of Things and voiceprint recognition. The specific steps of this verification method are as follows: Collect user voice data and device environment information through IoT devices, and pre-process the voice data.
[0033] Voiceprint features are extracted from the processed voice data, combined with the user's behavior patterns and device status information to form multimodal features and update the user feature library.
[0034] Construct and optimize feature matching rules, make preliminary matching results based on feature matching rules, and filter out candidate identity sets from the user feature library.
[0035] Construct a voiceprint recognition model to conduct in-depth comparison and analysis on the initially screened candidate identity set to verify the user's identity.
[0036] Specifically, a set of voiceprint recognition models is designed and constructed based on the Bi-GRU architecture. The model includes an input layer, a Bi-GRU layer, a fully connected layer, a classification layer and an output layer. Then, according to the public voiceprint dataset, a training set, a test set and a validation set are constructed. The training set is input into the voiceprint recognition model. The Bi-GRU layer of the voiceprint recognition model processes the input data in chronological order and reverse chronological order respectively, and outputs the concatenated vectors of the forward and reverse features. The fully connected layer aggregates the received concatenated vectors into a fixed-length voiceprint embedding vector. After that, the classification layer outputs the final identity recognition result based on the candidate identity set, and the final identity of the model is calculated through the cross entropy loss function. The loss value of the recognition result and the actual identity recognition result is then propagated from the output layer to the input layer of the voiceprint recognition model based on the chain rule, and the gradient of each network layer of the model corresponding to the loss value is calculated. Then, the parameters of each network layer are optimized based on the Adam optimizer, and the performance of the model at the end of this round of training is evaluated through the verification machine. If the model performance does not reach the preset performance index, the model is retrained until the model loss value converges to the preset threshold, and the training is stopped. The performance of the trained voiceprint recognition model on unknown data is tested through the test set, and the model is deployed on the actual voiceprint recognition system, and the current voiceprint feature is used as the current state. Input to the voiceprint recognition model, based on the policy network , select the current action , that is, adjusting the parameters or hyperparameters of the voiceprint recognition model, recalculating the predicted output by adjusting the parameters or hyperparameters of the voiceprint recognition model, and based on the updated model performance, ; Calculate the reward value, where Represents the selection of the current action After the model performance evaluation reward, Represents the recognition accuracy of the model, Represents the similarity between the target domain features and the source domain embedding, They represent weight parameters respectively, feed back the reward value to the policy network for updating, and accumulate the reward value as the policy optimization goal, with maximizing the accumulated reward as the goal of the policy network, and use the gradient to update the policy network parameters. After that, the voiceprint recognition model optimizes the target field according to the migrated task data, and combines the contrast loss function to update the model parameters. Repeat multiple rounds of reinforcement learning training until the model loss value converges to the preset range. The voiceprint recognition model receives the current voiceprint feature, and after forward propagation processing, generates an embedding vector of the voiceprint to be verified, calculates the similarity between the embedding vector and each candidate identity in the candidate identity set, and after comparing all candidate identities, screens out one or more matching identities with a similarity higher than the preset threshold, sorts the candidate identities in descending order according to the similarity score, and gives priority to the identity with the highest similarity score.
[0037] Assess identity authentication risk levels in real time and dynamically adjust verification strategies while collaboratively optimizing voiceprint recognition models.
[0038] Specifically, after the verification strategy is adjusted, each IoT device collects the voice data of its user interaction to form a local data set, and converts the original voice data into feature representation through voice preprocessing technology. Then, the processed local data set of each device is defined as ;in Representative Device D, Representative Voice features, Representative The user identity label corresponding to the voice feature, Representative The data volume of each device D is large, each device initializes a set of model architectures that are the same as the original voiceprint recognition model, and initializes the model parameters on each device to the unified weights distributed by the global server. At the same time, local loss functions, optimizers, and hyperparameters are set for each group of devices. After that, the local device uses the current model parameters to predict each sample in the data set locally, and evaluates the error value between the prediction result and the true label through the cross entropy loss function. The error value is propagated layer by layer based on the chain rule, and the gradient of the parameters of each layer of the model is calculated according to the error value. Then, the SGD optimizer is used to update the model parameters, and the parameter update amount of the local model is calculated and stored. Each group of local IoT devices uploads the local model parameter update amount to the terminal server. According to the characteristics of each IoT device, the spatial range of the local model parameters, and the local gradient trend, each Item information, initialize a set of belief spaces for each IoT device respectively, generate multiple groups of local behaviors according to the device's voice features and user behavior patterns, evaluate the fitness of each behavior in the local task, and select the behavior with the highest fitness as the optimal behavior of the IoT device. Each group of IoT devices generates a belief update vector according to its optimal behavior. The server summarizes the optimal behavior of each device, integrates the belief updates of the devices by averaging to generate a global belief vector, calculates the weight of each device by using the global belief and local model update, and then the terminal server performs weighted aggregation on the model parameter update amounts of all devices according to the device weight to obtain the corresponding global model update, distributes the updated global model to all IoT devices, synchronizes the local model parameters of each IoT device with the global model parameters, and performs update training in real time.
[0039] The user's identity is confirmed based on the verification result, and the result is fed back to the corresponding user device. The user's operation request is authorized and the operation log is recorded.
[0040] Monitor the system operation status in real time and regularly optimize the parameter configurations in the identity authentication process.
Claims
1. The identity authentication system based on the Internet of Things and voiceprint recognition is characterized by: It includes acquisition preprocessing module, multimodal fusion module, storage transmission module, identity screening module, identity authentication module, dynamic risk assessment module, model update module, feedback correction module, authorization record module and monitoring self-optimization module; The acquisition and preprocessing module is used to acquire the user's voice data in real time and preprocess the acquired voice data; The multimodal fusion module is used to extract voiceprint features from voice data and combine voice features with user behavior patterns to construct multimodal features; The storage and transmission module is used for distributed storage of user voiceprint features and constructed multimodal features, and for encrypted transmission of each set of feature data; The identity screening module is used to perform a preliminary comparison on the collected voiceprint features to screen out candidate identities; The identity verification module is used to construct a voiceprint recognition model, and to perform in-depth comparison and analysis on the candidate identities initially screened through the voiceprint recognition model to verify the user's identity; The dynamic risk assessment module is used to assess the risk level of identity authentication in real time and dynamically adjust the authentication strategy according to the risk; The model updating module is used to collaboratively optimize the voiceprint recognition model through distributed devices; The feedback correction module is used to provide instant feedback to users and correct potential errors; The authorization recording module is used to authorize the user's operation request after the verification is passed, and record the operation log at the same time; The monitoring self-optimization module is used to continuously monitor the system operation status and regularly optimize the parameter configuration in the identity authentication process.
2. The identity authentication system based on the Internet of Things and voiceprint recognition according to claim 1 is characterized in that: The specific steps of constructing multimodal features by the multimodal fusion module are as follows: S1.1: Extract the Mel frequency cepstrum coefficients, linear prediction coefficients, energy and pitch feature information from the preprocessed speech data, and record the extracted voiceprint features as , then extract the behavior features and state features from the user behavior pattern and device status respectively, and record them as and ; S1.2: Standardize each dimension of voiceprint features, behavioral features, and state features from different sources to eliminate feature scale differences, unify feature dimensions through zero padding, truncation, and dimensionality reduction operations, assign weights to each modality feature, and perform weighted fusion of features from different modalities based on the assigned weights; S1.3: Use the Min-Max normalization method to map the fused multimodal features to the interval [0, 1], and use the normalized multimodal features as The multimodal feature space is constructed according to the number of processed multimodal features, and a group of populations are initialized in the multimodal feature space, and the position of each group of individuals in the population is used as a multimodal feature of the multimodal feature space; S1.4: Based on the separation of feature representations between different users and the stability of feature representations input by the same user multiple times, the fitness value of each multimodal feature in the multimodal feature space is calculated, and the individuals in each group are sorted in descending order according to the fitness value. The top three groups of individuals with the best performance are selected from the population, and the remaining individuals are moved closer to the top three groups of individuals with the highest fitness values and the positions of the remaining individuals are updated; S1.5: After each position update, the individuals with the updated positions are randomly disturbed. If the multimodal feature fitness value corresponding to the disturbed position is better than the original value, the position update after the disturbance is accepted. After the position update is completed, the fitness values of each individual in the population are recalculated and sorted, and the top three groups of individuals with the best performance are reselected to update the position again; S1.6: Repeat the fitness ranking, best individual update and position update until the preset maximum number of iterations is reached or the change in the fitness value of the first-ranked individual is less than the set threshold. Then stop the iteration and take the first-ranked multimodal feature as the optimal feature representation and add it to the user feature library to update the user feature library.
3. The identity authentication system based on Internet of Things and voiceprint recognition according to claim 2 is characterized in that: The specific steps of distributed storage of the storage transmission module are as follows: S2.1: Create a decentralized storage structure for voiceprint features in the blockchain network, and define each block in the blockchain to store voiceprint feature data, multimodal feature data and their hash values. Then, a proof-of-work consensus mechanism is adopted, and each node in the blockchain searches for a hash value that is less than the difficulty threshold set by the network through calculation; S2.2: Broadcast the block corresponding to the hash value that meets the conditions obtained through the consensus mechanism to the entire network and add it to the chain, making each storage block in the blockchain a state node , and the connection between nodes represents the transmission path of data in the blockchain network, where NodeID represents the unique identifier of the current storage node. Represents the cumulative reward value of the current node, Represents the number of times the current node has been visited; S2.3: Starting from the starting node of the blockchain as the root node, a search tree is constructed based on the subsequent child nodes and the root node in the blockchain network, the UCB value of each child node is calculated, and the child node with the highest UCB value is selected as the next storage node. If the current child node is not fully expanded, the child node is expanded and the new node is added to the search tree; S2.4: Starting from the newly added node, randomly simulate the storage path, calculate the storage cost and retrieval delay, update the cumulative reward value and access number of the search tree according to the simulation results, repeat the selection, expansion, simulation and backtracking steps until the preset maximum number of iterations is reached. After the iteration, select the path with the highest cumulative reward value as the storage path for voiceprint feature data and multimodal feature data, and dynamically adjust the data distribution plan in the blockchain.
4. The identity authentication system based on Internet of Things and voiceprint recognition according to claim 3 is characterized in that: Each block in the blockchain described in S2.1 is represented as follows: ;in, Representative blocks; Representative Voiceprint feature data or multimodal feature data of each block; Representative The hash value of the block; The hash value representing the voiceprint data of this block; Representative The random number in the proof of work of each block; The specific calculation formula of the UCB value described in S2.3 is as follows: ; In the formula, Represents the selection value of the current node s; Represents the cumulative reward value of the current node s; represents the number of times the current node s has been visited; c represents the parameter for balancing exploration and utilization, which is usually a constant; Represents the number of times the parent node of the current node has been visited; ln represents a logarithmic function, which is used to adjust the priority of unvisited nodes.
5. The identity authentication system based on Internet of Things and voiceprint recognition according to claim 3 is characterized in that: The specific steps of the initial comparison of the identity screening module are as follows: S3.1: Extract voiceprint features from the user's current voice data, select pre-registered voiceprint features that match the current user from the user feature library, and initialize the candidate identity pool based on the extracted voiceprint features. Measure the similarity between the features to be verified and the identity features in the candidate identity pool through dynamic weighted distance, and screen out the identity features in the candidate identity pool whose similarity does not meet the preset threshold; S3.2: Initialize the feature weight population and the candidate identity pool after the initial screening, and calculate the fitness value of each feature weight and candidate identity feature respectively. Then retain the feature weights whose fitness is higher than the preset threshold, randomly select two groups of individuals from the selected feature weights, perform linear combination to generate new feature weights, and then randomly adjust some weight components so that the new feature weights meet the preset constraints; S3.3: Select candidate identity features whose fitness is lower than the preset threshold, and replace them with new identity features that have not been selected in the user feature library, and re-update the feature weight population and the candidate identity pool based on the candidate identity features whose fitness is higher than the preset threshold and the matching rules, until the fitness value of the feature weight population changes less than the preset threshold in multiple rounds; S3.4: After stopping the update, the feature weight with the highest fitness is selected as the weight distribution of the final matching rule. Then, the matching distance between the feature to be verified and each identity feature in the candidate identity pool is calculated according to the optimized rule, and the maximum threshold of the matching distance is set. The candidate identities with a matching distance less than the maximum threshold are considered to have passed the screening.
6. The identity authentication system based on Internet of Things and voiceprint recognition according to claim 5, characterized in that: The specific calculation formula of the dynamic weighted distance described in S3.1 is as follows: ; In the formula, Represents the voiceprint feature to be verified With The matching distance of candidate identities v, where is the index of candidate identity v in the candidate identity pool; d represents the voiceprint feature to be verified The number of dimensions; Representative The weight of the dimension feature is initially uniformly distributed; Represents the voiceprint feature to be verified No. Dimension value; Representative Candidate Status No. Dimension value; The specific calculation formulas for the feature weights and the fitness values of the candidate identity features described in S3.2 are as follows: , ; In the formula, Representative The fitness value of the feature weight w; represents the total number of eigenvectors; Represents the attenuation factor of the matching distance. The smaller the distance, the greater the corresponding contribution. Represents the voiceprint feature to be verified With The similarity score of the candidate identity v; Representative The fitness value of the candidate identity v is the minimum matching distance under the current optimal weight vector.
7. An identity authentication method based on the Internet of Things and voiceprint recognition, used to implement the identity authentication system function based on the Internet of Things and voiceprint recognition as described in any one of claims 1 to 6, characterized in that: The specific steps of this verification method are as follows: Ⅰ. Collect user voice data and device environment information through IoT devices, and pre-process the voice data; Ⅱ. Extract voiceprint features from the processed voice data, combine the user's behavior pattern and device status information, form multimodal features and update the user feature library; III. Construct and optimize feature matching rules, make preliminary matching results based on feature matching rules, and screen out candidate identity sets from the user feature library; IV. Construct a voiceprint recognition model to conduct in-depth comparison and analysis on the initially screened candidate identity set to verify the user's identity; V. Real-time assessment of identity authentication risk level, dynamic adjustment of authentication strategy, and collaborative optimization of voiceprint recognition model; VI. Confirm the user's identity based on the verification result, feed back the result to the corresponding user device, authorize the user's operation request, and record the operation log; Ⅶ. Monitor the system operation status in real time and regularly optimize the parameter configurations in the identity authentication process.
8. The identity authentication method based on the Internet of Things and voiceprint recognition according to claim 7, characterized in that: The specific steps of the in-depth comparison and analysis in step IV are as follows: S4.1: Design and construct a voiceprint recognition model based on the Bi-GRU architecture. The model includes an input layer, a Bi-GRU layer, a fully connected layer, a classification layer, and an output layer. Then, based on the public voiceprint dataset, construct a training set, a test set, and a validation set. S4.2: The training set is input into the voiceprint recognition model. The Bi-GRU layer of the voiceprint recognition model processes the input data in chronological order and reverse chronological order respectively, and outputs the concatenated vectors of the forward and reverse features. The fully connected layer aggregates the received concatenated vectors into a fixed-length voiceprint embedding vector. Then the classification layer outputs the final identity recognition result based on the candidate identity set. S4.3: The loss value between the final identification result of the model and the actual identification result is calculated through the cross entropy loss function. Then, based on the chain rule, the loss value is propagated from the output layer to the input layer of the voiceprint recognition model, and the gradient of each network layer of the model corresponding to the loss value is calculated. Then, the parameters of each network layer are optimized based on the Adam optimizer. The performance of the model after this round of training is evaluated through the verification machine. If the model performance does not reach the preset performance index, the model is retrained until the model loss value converges to the preset threshold, and the training is stopped. Then, the performance of the trained voiceprint recognition model on unknown data is tested through the test set, and the model is deployed on the actual voiceprint recognition system. S4.4: Use the current voiceprint feature as the current state Input to the voiceprint recognition model, based on the policy network , select the current action , that is, adjusting the parameters or hyperparameters of the voiceprint recognition model, recalculating the predicted output by adjusting the parameters or hyperparameters of the voiceprint recognition model, and based on the updated model performance, ; Calculate the reward value, where Represents the selection of the current action After the model performance evaluation reward, Represents the recognition accuracy of the model, Represents the similarity between the target domain features and the source domain embedding, Represent weight parameters respectively; S4.5: Feedback the reward value to the policy network for updating, and accumulate the reward value as the policy optimization target, with maximizing the cumulative reward as the target of the policy network, and use the gradient to update the policy network parameters. After that, the voiceprint recognition model is optimized in the target field according to the migrated task data, and combined with the contrast loss function, the model parameters are updated, and multiple rounds of reinforcement learning training are repeated until the model loss value converges to the preset range; S4.6: The voiceprint recognition model receives the current voiceprint features, and after forward propagation processing, generates an embedding vector of the voiceprint to be verified, calculates the similarity between the embedding vector and each candidate identity in the candidate identity set, and after comparing all candidate identities, screens out one or more matching identities with a similarity higher than a preset threshold, sorts the candidate identities in descending order according to the similarity score, and gives priority to the identity with the highest similarity score.
9. The identity authentication method based on the Internet of Things and voiceprint recognition according to claim 7, characterized in that: The specific steps of collaborative optimization of the voiceprint recognition model are as follows: S5.1: After the verification strategy is adjusted, each IoT device collects the voice data of its user interaction to form a local data set, and converts the original voice data into feature representation through voice preprocessing technology. Then, the processed local data set of each device is defined as ;in Representative Device D, Representative Voice features, Representative The user identity label corresponding to the voice feature, Representative The amount of data on device D; S5.2: Each device initializes a set of model architectures that are the same as the original voiceprint recognition model, and initializes the model parameters on each device to the unified weights distributed by the global server. At the same time, the local loss function, optimizer, and hyperparameters are set for each group of devices. After that, the local device uses the current model parameters to predict each sample in the data set locally, and evaluates the error value between the prediction result and the true label through the cross entropy loss function; S5.3: Based on the chain rule, the error value is propagated layer by layer, and the gradient of the parameters of each layer of the model is calculated according to the error value. Then, the model parameters are updated using the SGD optimizer, and the parameter update amount of the local model is calculated and stored. Each group of local IoT devices uploads the local model parameter update amount to the terminal server; S5.4: Based on the characteristics of each IoT device, the spatial range of local model parameters, and the local gradient trend information, a set of belief spaces are initialized for each IoT device. According to the voice characteristics of the device and the user behavior pattern, multiple groups of local behaviors are generated, and the fitness of each behavior in the local task is evaluated. The behavior with the highest fitness is selected as the optimal behavior of the IoT device. Each group of IoT devices generates a belief update vector based on its optimal behavior. The server summarizes the optimal behavior of each device and integrates the belief updates of the devices by averaging to generate a global belief vector. S5.5: Utilize global beliefs and local model updates to calculate the weight of each device. The terminal server then performs weighted aggregation on the model parameter updates of all devices based on the device weights to obtain the corresponding global model updates. The updated global model is distributed to all IoT devices, and the local model parameters of each IoT device are synchronized to the global model parameters, while updating and training are performed in real time.
Citation Information
Patent Citations
Identity verification method and system based on Internet of Things and voiceprint recognition
CN114579947A
Multi-mode identity verification system and method
CN118245994A
Intelligent business management and scheduling decision-making method of heat supply system based on voice interaction
CN118607959A
Intelligent access control management method and system based on multi-mode identification and Internet of Things technology
CN118968665A
Real-time contextually aware artificial intelligence (AI) assistant system and a method for providing a contextualized response to a user using ai
US20240412720A1
Cited By
Organization member information acquisition management system
CN121435263A
An organization member information collection management system
CN121435263B
Industrial vehicle identity recognition and unlocking system based on image recognition and fingerprint comparison
CN121716643A