Identity verification system and method based on internet of things and voiceprint recognition

By using multimodal feature fusion and blockchain storage in the IoT identity verification system, the problems of low verification accuracy and poor data privacy and security in IoT voiceprint recognition are solved, thereby improving personalized user experience and system adaptability.

CN119989321BActive Publication Date: 2025-10-17STATE GRID SHANDONG ELECTRIC POWER CO
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510105084.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-10-17
Estimated Expiration
2045-01-23

AI Technical Summary

Technical Problem

Existing IoT voiceprint recognition authentication systems suffer from low verification accuracy, high computational resource consumption, poor data privacy and security, and insufficient adaptability in multi-device collaborative environments, failing to provide a smooth and personalized user experience.

Method used

An IoT-based identity verification system is adopted, including a data acquisition and preprocessing module, a multimodal fusion module, a storage and transmission module, an identity screening module, an identity verification module, a dynamic risk assessment module, a model update module, a feedback correction module, and a monitoring and self-optimization module. Through multimodal feature fusion, blockchain storage, dynamic risk assessment, and collaborative model optimization, the system achieves comprehensive, reliable, and personalized identity verification.

Benefits of technology

It improves the comprehensiveness and reliability of identity verification, provides personalized verification rules, reduces computing resource consumption and communication costs, enhances the system's adaptability and generalization capabilities, and protects user privacy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119989321B_ABST
    Figure CN119989321B_ABST
Patent Text Reader

Abstract

The application discloses an identity verification system and method based on an Internet of Things and voiceprint recognition, belongs to the technical field of identity recognition, and comprises a collection preprocessing module, a multi-modal fusion module, a storage transmission module and an identity screening module. The application improves the comprehensiveness and reliability of verification, realizes dynamic user portrait updating, realizes mutual enhancement of multi-modal, can construct personalized verification rules for different users, improves the flexible response ability of the system to specific user needs, can provide a more smooth and friendly identity verification experience for users, improves the overall verification accuracy and fault tolerance of the system, reduces the calculation pressure and time consumption of the subsequent deep comparison module, can flexibly cope with environmental noise and data deviation, can protect the privacy of users, avoids security risks caused by data centralization, improves the adaptability and generalization ability of the system, reduces communication costs and resource consumption, solves the problem of uneven data distribution, and improves the overall performance of the system.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of identity recognition, in particular to an identity verification system and method based on Internet of Things and voiceprint recognition. BACKGROUND

[0002] With the rapid development of Internet of Things (IoT) technology, a large number of intelligent devices and systems are interconnected, building a highly collaborative ecological environment. This ecology needs a reliable identity verification mechanism to protect user privacy and system security. Although traditional password or biometric authentication methods are popular, they often have poor user experience and the risk of information theft. As a non-contact and natural biometric feature, voiceprint recognition has become a research hotspot in the field of identity verification due to its convenience and uniqueness. The distributed nature of Internet of Things devices and dynamic network environment pose challenges to voiceprint recognition in large-scale identity verification applications, such as differences in transmission quality between devices, computational resource limitations, and the impact of user behavior and environment on voiceprints. In addition, how to ensure data privacy while improving the accuracy and robustness of verification in a multi-device collaboration scenario is also a problem to be solved.

[0003] After searching, CN114579947A discloses an identity verification method and system based on Internet of Things and voiceprint recognition. This invention only requires the verifier to speak as prompted, and the verifier can wear protective equipment without affecting identity verification, reducing the likelihood of disease transmission. However, the overall verification comprehensiveness and reliability are low, and dynamic user profile updating cannot be performed, and personalized verification rules cannot be constructed for different users. Existing identity verification systems and methods cannot provide a more smooth and friendly identity verification experience for users, reducing the overall verification accuracy and fault tolerance of the system, increasing the computational pressure and time consumption of the subsequent deep comparison module. In addition, existing identity verification systems and methods are prone to security risks caused by data centralization, have poor adaptability and generalization ability, and increase communication costs and resource consumption. Therefore, we propose an identity verification system and method based on Internet of Things and voiceprint recognition. SUMMARY

[0004] The purpose of the present application is to solve the defects in the prior art and to provide an identity verification system and method based on Internet of Things and voiceprint recognition.

[0005] To achieve the above purpose, the present application adopts the following technical solutions:

[0006] The identity verification system based on Internet of Things and voiceprint recognition comprises a collection and preprocessing module, a multi-modal fusion module, a storage and transmission module, an identity screening module, an identity verification module, a dynamic risk assessment module, a model updating module, a feedback correction module, an authorization record module, and a monitoring and self-optimization module.

[0007] The collection preprocessing module is configured to collect voice data of the user in real time and preprocess the collected voice data.

[0008] The multi-modal fusion module is configured to extract a voiceprint feature from the voice data and combine the voice feature with a user behavior pattern to construct a multi-modal feature.

[0009] The storage transmission module is configured to distribute storage of the user voiceprint feature and the constructed multi-modal feature, and to encrypt and transmit each group of feature data.

[0010] The identity screening module is configured to preliminarily compare the collected voiceprint feature and screen out a candidate identity.

[0011] The identity verification module is configured to construct a voiceprint recognition model and perform deep comparison and analysis on the preliminarily screened candidate identity through the voiceprint recognition model to verify the user identity.

[0012] The dynamic risk assessment module is configured to assess a risk level of the identity verification in real time and dynamically adjust a verification strategy according to the risk.

[0013] The model updating module is configured to collaboratively optimize the voiceprint recognition model through distributed devices.

[0014] The feedback correction module is configured to provide instant feedback for the user and correct potential errors.

[0015] The authorization record module is configured to authorize an operation request of the user after verification and record an operation log.

[0016] The monitoring self-optimization module is configured to continuously monitor a system running state and periodically optimize parameter configuration in the identity verification process.

[0017] As a further solution of the present invention, the specific steps of preprocessing the collected voice signals by the acquisition preprocessing module are as follows: real-time monitoring of the network transmission status between devices, and collecting network quality indicators such as bandwidth, delay, jitter and packet loss rate, calculating the transmission quality of each monitoring device network, and then dynamically adjusting the voice sampling rate according to the network quality threshold and the transmission quality of the monitoring device network, and collecting voice data in real time according to the adjusted voice sampling rate, performing short-time Fourier transform on the collected voice data to obtain the corresponding frequency domain signal, and then calculating the power spectrum and noise power spectrum of the voice signal by minimum mean square error to construct a corresponding Wiener filter, and then filtering the frequency domain signal by the Wiener filter to obtain a denoised voice spectrum, after filtering is completed, convolving the voice signal with the room impulse response to construct a room reverberation signal model, and then using a known test signal to measure the impulse response of the room, and designing an inverse filter through frequency domain representation, and performing inverse filtering on the spectrum of the reverberated voice signal to obtain a dereverberated spectrum, and using an inverse short-time Fourier transform to restore the dereverberated spectrum to the original voice data.

[0018] As a further solution of the present invention, the multimodal fusion module constructs multimodal features in the following specific steps:

[0019] S1.1: Extract the Mel frequency cepstral coefficients, linear prediction coefficients, energy and pitch feature information from the preprocessed speech data, and record the extracted voiceprint features as , and then extract behavioral features and state features from user behavior patterns and device states, and record them as and ;

[0020] S1.2: Standardize each dimension of voiceprint features, behavioral features, and state features from different sources to eliminate feature scale differences. Then, unify the feature dimensions through zero-padding, truncation, and dimensionality reduction operations. Assign weights to each modal feature and perform weighted fusion of features from different modalities based on the assigned weights.

[0021] S1.3: Use the Min-Max normalization method to map the fused multimodal features to the interval [0, 1], and use the normalized multimodal features to , a multimodal feature space is constructed based on the number of processed multimodal features, and a group of populations is initialized in the multimodal feature space, and the position of each group of individuals in the population is used as a multimodal feature of the multimodal feature space;

[0022] S1.4: Based on the separation degree of the feature representation between different users and the stability of the feature representation of multiple inputs of the same user, the fitness value of each multi-modal feature in the multi-modal feature space is calculated, and each group of individuals is ranked in descending order of fitness value, and the top three groups of individuals with the best performance are selected from the population, and the remaining individuals approach the top three groups of individuals with the highest fitness value and update the positions of the remaining individuals;

[0023] S1.5: After each position update, the individuals after position update are randomly disturbed, and if the fitness value of the multi-modal feature corresponding to the disturbed position is better than the original value, the position update after disturbance is accepted, and after position update, the fitness values of each individual in the population are recalculated and ranked, and the top three groups of individuals with the best performance are selected again, and position update is performed again;

[0024] S1.6: Repeat the fitness arrangement, optimal individual update and position update until the preset maximum iteration number is reached or the fitness value of the first ranked individual changes less than the set threshold, stop iteration, and take the first ranked multi-modal feature as the optimal feature representation, and add it to the user feature library to update the user feature library.

[0025] As a further scheme of the present application, the specific steps of the storage transmission module distributed storage are as follows:

[0026] S2.1: Create a decentralized storage structure for voiceprint features in the blockchain network, and define that each block in the blockchain stores voiceprint feature data, multi-modal feature data and its hash value, then use the proof of work consensus mechanism, and each node in the blockchain finds a hash value less than the difficulty threshold set by the network through calculation;

[0027] S2.2: Broadcast the block corresponding to the hash value meeting the condition obtained through the consensus mechanism to the entire network, and add it to the chain, and take each storage block in the blockchain as a state node , and the connection between nodes represents the transmission path of data in the blockchain network, wherein NodeID represents the unique identifier of the current storage node, represents the cumulative reward value of the current node, represents the number of times the current node is accessed;

[0028] S2.3: Take the starting node of the blockchain as the root node, build a search tree according to the subsequent child nodes in the blockchain network and the root node, calculate the UCB value of each child node, and select the child node with the highest UCB value as the next storage node, if the current child node is not fully expanded, expand the child node, and add the new node to the search tree;

[0029] S2.4: Starting from the newly added node, randomly simulate the storage path, calculate the storage cost and retrieval delay, update the cumulative reward value and access frequency of the search tree according to the simulation result, repeat the steps of selection, expansion, simulation and backtracking until the preset maximum iteration number is reached, after the iteration is completed, select the path with the highest cumulative reward value as the storage path of the voiceprint feature data and the multi-modal feature data, and dynamically adjust the distribution scheme of the data in the blockchain.

[0030] As a further scheme of the present application, the form of each block in the blockchain described in S2.1 is as follows:

[0031] ; wherein, represents the first block; represents the voiceprint feature data or multi-modal feature data of the first block; represents the hash value of the first block; represents the hash value of the voiceprint data of the current block; represents the random number in the proof of work of the first block;

[0032] The specific calculation formula of the UCB value described in S2.3 is as follows:

[0033] ; wherein, represents the selection value of the current node s; represents the cumulative reward value of the current node s; represents the number of times the current node s is accessed; c represents a parameter for balancing exploration and utilization, which is usually a constant; represents the number of times the parent node of the current node is accessed; ln represents the logarithmic function, which is used to adjust the priority of unvisited nodes.

[0034] As a further scheme of the present application, the specific steps of the identity screening module are as follows:

[0035] S3.1: Extract the voiceprint feature from the current voice data of the user, select the pre-registered voiceprint feature matching the current user from the user feature library, and initialize the candidate identity pool according to the extracted voiceprint feature. The similarity between the to-be-verified feature and each identity feature in the candidate identity pool is measured by dynamically weighting the distance, and the identity features in the candidate identity pool whose similarity does not satisfy the preset threshold are screened out;

[0036] S3.2: initialize the feature weight population and the candidate identity pool after preliminary screening, and calculate the fitness value of each feature weight and candidate identity feature respectively, then retain the feature weight with fitness value higher than the preset threshold, randomly select two groups of individuals from the selected feature weight, perform linear combination to generate a new feature weight, then randomly adjust part of the weight components to make the new feature weight meet the preset constraint;

[0037] S3.3: select the candidate identity feature with fitness value lower than the preset threshold, and replace it with a new identity feature in the user feature library which has not been selected, and based on the candidate identity feature with fitness value higher than the preset threshold and the matching rule, update the feature weight population and the candidate identity pool again until the fitness value of the feature weight population changes by less than the preset threshold in multiple rounds;

[0038] S3.4: after stopping updating, select the feature weight with the highest fitness value as the weight distribution of the final matching rule, then calculate the matching distance between the to-be-verified feature and each identity feature in the candidate identity pool according to the optimized rule, and set the maximum threshold of the matching distance, and the candidate identity with matching distance less than the maximum threshold is regarded as passing the screening.

[0039] As a further scheme of the application, the dynamic weighted distance calculation formula of S3.1 is as follows:

[0040] ; in the formula, represents the to-be-verified voiceprint feature and the matching distance of the first candidate identity v, wherein is the index of the candidate identity v in the candidate identity pool; d represents the dimension number of the to-be-verified voiceprint feature ; represents the weight of the first dimension feature, which is initially uniformly distributed; represents the value of the first dimension of the to-be-verified voiceprint feature ; represents the value of the first dimension of the candidate identity ;

[0041] The fitness value calculation formula of the feature weight and the candidate identity feature of S3.2 is as follows:

[0042] , ; in the formula, represents the fitness value of the first feature weight w; represents the total number of feature vectors; Represents the attenuation factor of the matching distance. The smaller the distance, the greater the corresponding contribution. Represents the voiceprint feature to be verified With the The similarity score of the candidate identity v; Representative The fitness value of the candidate identity v is the minimum matching distance under the current optimal weight vector.

[0043] The authentication method based on the Internet of Things and voiceprint recognition has the following specific steps:

[0044] Ⅰ. Collect user voice data and device environment information through IoT devices and pre-process the voice data;

[0045] II. Extract voiceprint features from the processed voice data, combine them with user behavior patterns and device status information to form multimodal features and update the user feature library;

[0046] Ⅲ. Construct and optimize feature matching rules, make preliminary matching results based on the feature matching rules, and filter out candidate identity sets from the user feature library;

[0047] IV. Build a voiceprint recognition model to conduct in-depth comparison and analysis on the initially screened candidate identity set to verify the user's identity;

[0048] V. Real-time assessment of identity verification risk levels, dynamic adjustment of verification strategies, and collaborative optimization of voiceprint recognition models;

[0049] VI. Confirm the user's identity based on the verification result, feed the result back to the corresponding user device, authorize the user's operation request, and record the operation log;

[0050] Ⅶ. Monitor the system operation status in real time and regularly optimize the parameter configurations in the identity authentication process.

[0051] As a further solution of the present invention, the specific steps of the in-depth comparison and analysis in step IV are as follows:

[0052] S4.1: Design and construct a voiceprint recognition model based on the Bi-GRU architecture. The model includes an input layer, a Bi-GRU layer, a fully connected layer, a classification layer, and an output layer. Then, based on the public voiceprint dataset, construct a training set, a test set, and a validation set.

[0053] S4.2: The training set is input into the voiceprint recognition model. The Bi-GRU layer of the voiceprint recognition model processes the input data in chronological order and reverse chronological order respectively, and outputs the concatenated vectors of the forward and reverse features. The fully connected layer aggregates the received concatenated vectors into a fixed-length voiceprint embedding vector. Then, the classification layer outputs the final identity recognition result based on the candidate identity set.

[0054] S4.3: Calculate the loss between the final identification result of the model and the actual identification result using the cross-entropy loss function. Then, based on the chain rule, propagate the loss value from the output layer to the input layer of the voiceprint recognition model, and calculate the gradient of each network layer of the model corresponding to the loss value. Then, optimize the parameters of each network layer using the Adam optimizer. Use the verification machine to evaluate the performance of the model after this round of training. If the model performance does not reach the preset performance index, retrain the model until the model loss value converges to the preset threshold. Then stop training, test the performance of the trained voiceprint recognition model on unknown data using the test set, and deploy the model on the actual voiceprint recognition system.

[0055] S4.4: Use the current voiceprint feature as the current state Input to the voiceprint recognition model, based on the policy network , select the current action , that is, adjusting the parameters or hyperparameters of the voiceprint recognition model, recalculating the predicted output by adjusting the parameters or hyperparameters of the voiceprint recognition model, and based on the updated model performance, ; Calculate the reward value, where Represents the selection of the current action After the model performance evaluation reward, Represents the recognition accuracy of the model, Represents the similarity between the target domain features and the source domain embedding, Represent weight parameters respectively;

[0056] S4.5: Feed the reward value back to the policy network for updating, and accumulate the reward value as the policy optimization target. Maximizing the cumulative reward is the goal of the policy network. The policy network parameters are updated using gradients. The voiceprint recognition model is then optimized for the target domain based on the migrated task data. The model parameters are updated in combination with the contrast loss function. Multiple rounds of reinforcement learning training are repeated until the model loss value converges to the preset range.

[0057] S4.6: The voiceprint recognition model receives the current voiceprint feature, and after processing by forward propagation, generates an embedding vector of the voiceprint to be verified, calculates the similarity between the embedding vector and each candidate identity in the candidate identity set, and after comparing all candidate identities, selects one or more matching identities with a similarity higher than a preset threshold, sorts the candidate identities in descending order according to the similarity scores, and preferentially selects the identity with the highest similarity score.

[0058] As a further scheme of the present application, the steps of cooperatively optimizing the voiceprint recognition model are as follows:

[0059] S5.1: After the verification strategy is adjusted, each Internet of Things device collects voice data interacted with its user, forms a local data set, and converts the original voice data into feature representation through voice preprocessing technology, and then defines the processed local data set of each device as ; wherein represents the th device D, represents the th voice feature, represents the th voice feature corresponding to the user identity label, represents the th device D data size;

[0060] S5.2: Each device initializes a group of model architectures identical to the original voiceprint recognition model, and initializes the model parameters on each device to the uniform weights distributed by the global server, and sets the local loss function, optimizer and hyperparameters for each group of devices, and then the local device uses the current model parameters to predict each sample in the local data set, and evaluates the error value between the prediction result and the true label through the cross-entropy loss function;

[0061] S5.3: Based on the chain rule, the error value is propagated layer by layer, and the gradient of each layer parameter of the model is calculated according to the error value, and then the model parameters are updated using the SGD optimizer, and the parameter update amount of the local model is calculated and stored, and each group of local Internet of Things devices uploads the local model parameter update amount to the terminal server;

[0062] S5.4: According to the characteristics of each Internet of Things device, the local model parameter space range and the local gradient trend information, a group of belief spaces are initialized for each Internet of Things device, and according to the voice features and user behavior patterns of the device, a plurality of local behaviors are generated, and the fitness of each behavior in the local task is evaluated, and the behavior with the highest fitness is selected as the optimal behavior of the Internet of Things device, and each group of Internet of Things devices generates a belief update vector according to its optimal behavior, and the server integrates the optimal behaviors of each device by averaging method to generate a global belief vector.

[0063] S5.5: Calculate the weight of each device by using global belief and local model update, then the terminal server weights the model parameter update amount of all devices according to the device weight to obtain the corresponding global model update, distributes the updated global model to all Internet of Things devices, synchronizes the local model parameters of each Internet of Things device to the global model parameters, and simultaneously performs real-time update training.

[0064] Compared with the prior art, the present application has the following beneficial effects:

[0065] 1、The application constructs a multi-modal feature space according to the processed multi-modal feature quantity, initializes a group of populations in the multi-modal feature space, takes the position of each group of individuals in the population as a multi-modal feature of the multi-modal feature space, calculates the fitness value of each multi-modal feature, and arranges the individuals in descending order of fitness value, selects the top three individuals with the best performance from the population, and updates the positions of the remaining individuals to approach the top three individuals with the highest fitness value, after each position update, the individuals with updated positions are randomly disturbed, if the fitness value of the multi-modal feature corresponding to the disturbed position is better than the original value, the position update after disturbance is accepted, after the position update is completed, the fitness values of the individuals in the population are recalculated and sorted, and the top three individuals with the best performance are selected again, and the position update is performed again, the fitness arrangement, optimal individual update and position update are repeated until the preset maximum iteration number is reached or the fitness value of the first-ranked individual changes less than the set threshold, the iteration is stopped, the first-ranked multi-modal feature is taken as the optimal feature representation, and it is added to the user feature library to update the user feature library, improve the comprehensiveness and reliability of verification, realize dynamic user portrait update, realize mutual enhancement of multi-modal, can construct personalized verification rules for different users, and improve the flexible response ability of the system to specific user demand.

[0066] 2、The voiceprint feature is extracted from the current voice data of the user, the pre-registered voiceprint feature matched with the current user is selected from the user feature library, the candidate identity pool is initialized according to the extracted voiceprint feature, the similarity of the to-be-verified feature and each identity feature in the candidate identity pool is measured by a dynamic weighted distance, the identity features in the candidate identity pool that do not satisfy the preset threshold of similarity are screened out, the feature weight population and the candidate identity pool after the initial screening are initialized, the fitness value of each feature weight and the candidate identity feature is calculated, then the feature weight with a fitness value higher than a preset threshold is reserved, two groups of individuals are randomly selected from the selected feature weight, a new feature weight is generated by linear combination, then part of the weight components are randomly adjusted to make the new feature weight satisfy the preset constraint, the candidate identity feature with a fitness value lower than the preset threshold is selected and replaced by a new identity feature in the user feature library that has not been selected, and the feature weight population and the candidate identity pool are updated based on the candidate identity feature with a fitness value higher than the preset threshold and the matching rule, until the fitness value of the feature weight population changes by less than a preset threshold in multiple rounds, the updating is stopped, the feature weight with the highest fitness value is selected as the weight distribution of the final matching rule, then the matching distance of the to-be-verified feature and each identity feature in the candidate identity pool is calculated according to the optimized rule, and the maximum threshold of the matching distance is set, the candidate identity with a matching distance less than the maximum threshold is regarded as passing the screening, which can provide a more smooth and friendly identity verification experience for the user, improve the overall verification accuracy and fault tolerance of the system, reduce the calculation pressure and time consumption of the subsequent deep comparison module, and be flexible in dealing with environmental noise and data deviation.

[0067] 3、The voice data collected by each Internet of Things device in the application in the interaction with its user forms a local data set, and the original voice data is converted into a feature representation through voice preprocessing technology, each device initializes a set of model architectures same as the original voiceprint recognition model, and initializes the model parameters on each device to the uniform weight distributed by the global server, at the same time, sets the local loss function, optimizer and hyperparameters for each group of devices, then the local device uses the current model parameters to predict each sample in the local data set, and updates the model parameters, calculates and stores the parameter update amount of the local model, each group of local Internet of Things devices uploads the local model parameter update amount to the terminal server, according to the characteristics of each Internet of Things device, the local model parameter space range and the local gradient trend information, initializes a set of belief space for each Internet of Things device, generates multiple local behaviors according to the voice features and user behavior patterns of the device, evaluates the fitness of each behavior in the local task, and selects the behavior with the highest fitness as the optimal behavior of the Internet of Things device, each group of Internet of Things devices generates a belief update vector according to the optimal behavior, the server aggregates the optimal behaviors of each device, integrates the belief update of the device by averaging method to generate a global belief vector, calculates the weight of each device by using the global belief and the local model update, and then the terminal server aggregates the model parameter update amount of all devices by weighting according to the device weight to obtain the corresponding global model update, distributes the updated global model to all Internet of Things devices, synchronizes the local model parameters of each Internet of Things device to the global model parameters, and at the same time, updates and trains in real time, which can protect the privacy of the user, avoid the security risks brought by data centralization, improve the adaptability and generalization ability of the system, reduce the communication cost and resource consumption, solve the problem of uneven data distribution, and improve the overall performance of the system. BRIEF DESCRIPTION OF DRAWINGS

[0068] The accompanying drawings are included to provide a further understanding of the application and are incorporated in and constitute a part of this specification, illustrate embodiments of the application and serve to explain the principles of the application, and do not constitute a limitation of the application.

[0069] Figure 1 The system block diagram of the identity verification system based on Internet of Things and voiceprint recognition proposed by the application;

[0070] Figure 2 The flowchart of the identity verification method based on Internet of Things and voiceprint recognition proposed by the application. DETAILED DESCRIPTION

[0071] The technical solutions in the embodiments of the application will be described clearly and completely below with reference to the drawings in the embodiments of the application. Obviously, the described embodiments are only a part of the embodiments of the application, not all the embodiments of the application.

[0072] Embodiment 1, refer toFigure 1 The identity authentication system based on the Internet of Things and voiceprint recognition comprises a collection preprocessing module, a multi-modal fusion module, a storage transmission module, an identity screening module, an identity authentication module, a dynamic risk assessment module, a model updating module, a feedback correction module, an authorization record module, and a monitoring self-optimization module.

[0073] The collection preprocessing module is configured to collect voice data of a user in real time and pre-process the collected voice data.

[0074] It should be further explained that the collection preprocessing module pre-processes the collected voice signal in the following steps: monitoring the network transmission state between devices in real time, collecting network quality indicators such as bandwidth, delay, jitter, and packet loss rate, calculating the transmission quality of the network of the monitoring device, then dynamically adjusting the voice sampling rate according to the network quality threshold and the transmission quality of the network of the monitoring device, and collecting voice data in real time according to the adjusted voice sampling rate, performing short-time Fourier transform on the collected voice data to obtain the corresponding frequency domain signal, then calculating the power spectrum and noise power spectrum of the voice signal by the least mean square error to construct a corresponding Wiener filter, filtering the frequency domain signal by the Wiener filter to obtain the noise-reduced voice spectrum, and after filtering, convolving the voice signal with the room impulse response to construct a room reverberation signal model, then measuring the impulse response of the room using a known test signal, designing an inverse filter through frequency domain representation, performing inverse filtering on the spectrum of the reverberation voice signal to obtain the dereverberated spectrum, and restoring the dereverberated spectrum to the original voice data using inverse short-time Fourier transform.

[0075] The multi-modal fusion module is configured to extract voiceprint features from the voice data and combine the voice features with user behavior patterns to construct multi-modal features.

[0076] Specifically, Mel frequency cepstral coefficients, linear prediction coefficients, energy, and pitch features are extracted from the pre-processed voice data, and the extracted voiceprint features are denoted as Then, behavior features and state features are extracted from user behavior patterns and device states, respectively, and are denoted as and Each dimension of the voiceprint features, behavior features, and state features from different sources is standardized to eliminate feature scale differences, and each feature dimension is unified through zero padding, truncation, and dimension reduction operations. Each modal feature is assigned a weight, and the features of different modalities are weighted and fused based on the assigned weights. The fused multi-modal features are mapped to the [0, 1] interval using the Min-Max normalization method, and the normalized multi-modal features are denoted as The processed multi-modal feature quantity is used to construct a multi-modal feature space, and a group of populations is initialized in the multi-modal feature space. The position of each individual in the group is regarded as a multi-modal feature of the multi-modal feature space. The fitness value of each multi-modal feature in the multi-modal feature space is calculated based on the separation degree of the feature representation between different users and the stability of the feature representation of multiple inputs of the same user. The individuals are sorted in descending order of the fitness value. The top three individuals with the best performance are selected from the population. The remaining individuals approach the top three individuals with the highest fitness value and update the positions of the remaining individuals. After each position update, the individuals with updated positions are randomly disturbed. If the fitness value of the multi-modal feature corresponding to the disturbed position is better than the original value, the position update after the disturbance is accepted. After the position update, the fitness values of the individuals in the population are recalculated and sorted, and the top three individuals with the best performance are selected again. The position update is performed again. The fitness arrangement, optimal individual update and position update are repeated until a preset maximum iteration number is reached or the fitness value of the first-ranked individual changes by less than a set threshold. The iteration is stopped, and the first-ranked multi-modal feature is regarded as the optimal feature representation. The first-ranked multi-modal feature is added to the user feature library to update the user feature library.

[0077] The storage transmission module is used to store and transmit the multi-modal feature and the user voiceprint feature in a distributed manner, and each group of feature data is encrypted and transmitted.

[0078] Specifically, a decentralized storage structure of the voiceprint feature is created in the blockchain network, and each block in the blockchain stores voiceprint feature data, multi-modal feature data and its hash value. Then, a proof of work consensus mechanism is used, and each node in the blockchain finds a hash value less than the difficulty threshold set by the network by calculation. The block corresponding to the hash value meeting the condition through the consensus mechanism is broadcast to the entire network and added to the chain. Each storage block in the blockchain is regarded as a state node , and the connection between the nodes represents the transmission path of the data in the blockchain network, wherein NodeID represents the unique identifier of the current storage node, represents the cumulative reward value of the current node, The number of times the current node is visited, from the starting node of the blockchain as the root node, according to the subsequent child nodes in the blockchain network and the root node to build a search tree, calculate the UCB value of each child node, and select the child node with the highest UCB value as the next storage node, if the current child node is not fully expanded, expand the child node, add the new node to the search tree, start from the newly added node, randomly simulate the storage path, calculate the storage cost and retrieval delay, update the cumulative reward value and the number of visits of the search tree according to the simulation result, repeat the selection, expansion, simulation and backtracking steps until the preset maximum iteration number is reached, after the iteration is completed, select the path with the highest cumulative reward value as the storage path of the voiceprint feature data and the multi-modal feature data, and dynamically adjust the distribution scheme of the data in the blockchain.

[0079] In the embodiment, the form of each block in the blockchain is as follows:

[0080] ; wherein, represents the first block; represents the voiceprint feature data or multi-modal feature data of the first block; represents the hash value of the first block; represents the hash value of the voiceprint data of the current block; represents the random number in the proof of work of the first block;

[0081] The specific calculation formula of the UCB value is as follows:

[0082] ; wherein, represents the selection value of the current node s; represents the cumulative reward value of the current node s; represents the number of times the current node s is visited; c represents a parameter for balancing exploration and utilization, which is usually a constant; represents the number of times the parent node of the current node is visited; ln represents a logarithmic function, which is used to adjust the priority of unvisited nodes.

[0083] The identity screening module is used to preliminarily compare the collected voiceprint features and screen out candidate identities.

[0084] Specifically, the voiceprint feature is extracted from the current voice data of the user, the pre-registered voiceprint feature matching the current user is selected from the user feature library, and the candidate identity pool is initialized according to the extracted voiceprint feature. The similarity between the to-be-verified feature and each identity feature in the candidate identity pool is measured by a dynamic weighted distance, and the identity features in the candidate identity pool that do not satisfy the preset threshold of similarity are filtered out. The feature weight population and the candidate identity pool after the preliminary screening are initialized, and the fitness value of each feature weight and the candidate identity feature is calculated. Then, the feature weights with fitness values higher than the preset threshold are retained, two groups of individuals are randomly selected from the selected feature weights, linear combination is performed to generate new feature weights, and then part of the weight components are randomly adjusted to make the new feature weights satisfy the preset constraint. The candidate identity features with fitness values lower than the preset threshold are selected and replaced by new identity features in the user feature library that have not been selected. Based on the candidate identity features with fitness values higher than the preset threshold and the matching rule, the feature weight population and the candidate identity pool are updated again until the fitness value of the feature weight population changes by less than the preset threshold in multiple rounds. After the update is stopped, the feature weight with the highest fitness value is selected as the weight distribution of the final matching rule. Then, the matching distance between the to-be-verified feature and each identity feature in the candidate identity pool is calculated according to the optimized rule, and the maximum threshold of the matching distance is set. The candidate identity with a matching distance less than the maximum threshold is considered to pass the screening.

[0085] It should be further explained that the specific calculation formula of the dynamic weighted distance is as follows:

[0086] In the formula, represents the to-be-verified voiceprint feature The matching distance between the to-be-verified voiceprint feature and the first candidate identity v, where is the index of the candidate identity v in the candidate identity pool; d represents the dimension number of the to-be-verified voiceprint feature ; w represents the weight of the first dimension feature, which is initially uniformly distributed; represents the first dimension value of the to-be-verified voiceprint feature ; and represents the first dimension value of the candidate identity .

[0087] The specific calculation formulas of the fitness values of the feature weights and the candidate identity features are as follows:

[0088] , In the formula, represents the fitness value of the first feature weight w; ​representative of the total number of feature vectors; decay factor representing the matching distance, the smaller the distance, the greater the corresponding contribution; representative of the voiceprint features to be verified similarity score with the first candidate identity v; representative of the fitness value of the first candidate identity v, taking the minimum matching distance under the current optimal weight vector.

[0089] The identity verification module is used to build a voiceprint recognition model and perform deep comparison and analysis on the preliminary screened candidate identities through the voiceprint recognition model to verify the user's identity. The dynamic risk assessment module is used to assess the risk level of identity verification in real time and dynamically adjust the verification strategy according to the risk. The model updating module is used to optimize the voiceprint recognition model through distributed devices. The feedback correction module is used to provide instant feedback for the user and correct potential errors. The authorization record module is used to authorize the user's operation request after verification and record the operation log. The monitoring and self-optimization module is used to continuously monitor the system running state and periodically optimize the parameter configuration in the identity verification process.

[0090] Embodiment 2, refer to Figure 2 , the identity verification method based on the Internet of Things and voiceprint recognition, the specific steps of the verification method are as follows:

[0091] Collect the user's voice data and device environment information through the Internet of Things device, and preprocess the voice data.

[0092] Extract voiceprint features from the processed voice data, and form multi-modal features by combining the user's behavior patterns and device state information, and update the user feature library.

[0093] Build and optimize the feature matching rules, perform preliminary matching based on the feature matching rules, and screen out a candidate identity set from the user feature library.

[0094] Build a voiceprint recognition model to perform deep comparison and analysis on the preliminary screened candidate identity set to verify the user's identity.

[0095] Specifically, a set of voiceprint recognition models is designed and constructed based on the Bi-GRU architecture. The model includes an input layer, a Bi-GRU layer, a fully connected layer, a classification layer and an output layer. Then, based on the public voiceprint dataset, a training set, a test set and a validation set are constructed. The training set is input into the voiceprint recognition model. The Bi-GRU layer of the voiceprint recognition model processes the input data in chronological order and reverse chronological order respectively, and outputs the concatenated vectors of the forward and reverse features. The fully connected layer summarizes the received concatenated vectors into a fixed-length voiceprint embedding vector. After that, the classification layer outputs the final identity recognition result based on the candidate identity set, and the final identity of the model is calculated through the cross-entropy loss function. The loss value of the recognition result and the actual identity recognition result is then propagated from the output layer to the input layer of the voiceprint recognition model based on the chain rule, and the gradient of each network layer of the model corresponding to the loss value is calculated. The parameters of each network layer are then optimized based on the Adam optimizer, and the performance of the model at the end of this round of training is evaluated by the verification machine. If the model performance does not reach the preset performance index, the model is retrained until the model loss value converges to the preset threshold, and the training is stopped. The performance of the trained voiceprint recognition model on unknown data is then tested through the test set, and the model is deployed on the actual voiceprint recognition system, with the current voiceprint feature as the current state. Input to the voiceprint recognition model, based on the policy network , select the current action , that is, adjusting the parameters or hyperparameters of the voiceprint recognition model, recalculating the predicted output by adjusting the parameters or hyperparameters of the voiceprint recognition model, and based on the updated model performance, ; Calculate the reward value, where Represents the selection of the current action After the model performance evaluation reward, Represents the recognition accuracy of the model, Represents the similarity between the target domain features and the source domain embedding, They represent weight parameters respectively, feed back the reward value to the policy network for updating, and accumulate the reward value as the policy optimization goal, with maximizing the cumulative reward as the goal of the policy network, and use the gradient to update the policy network parameters. After that, the voiceprint recognition model optimizes the target field according to the migrated task data, and combines the comparative loss function to update the model parameters. Repeatedly perform multiple rounds of reinforcement learning training until the model loss value converges to the preset range. The voiceprint recognition model receives the current voiceprint feature and generates an embedding vector of the voiceprint to be verified after forward propagation processing. The similarity between the embedding vector and each candidate identity in the candidate identity set is calculated. After comparing all candidate identities, one or more matching identities with a similarity higher than the preset threshold are screened out, and the candidate identities are sorted in descending order according to the similarity score, and the identity with the highest similarity score is given priority.

[0096] Real-time assess the risk level of identity verification, and dynamically adjust the verification strategy while collaboratively optimizing the voiceprint recognition model.

[0097] Specifically, after the verification strategy adjustment is completed, each Internet of Things device collects voice data interacted with its user, forms a local data set, and converts the original voice data into feature representation through voice preprocessing technology. Then the processed local data set of each device is defined as ; wherein represents the th device D, represents the th voice feature, represents the th voice feature corresponding to the user identity label, represents the th device D, each device initializes a group of model architectures same as the original voiceprint recognition model, and initializes the model parameters on each device to the uniform weight distributed by the global server. At the same time, set the local loss function, optimizer and hyperparameters for each group of devices. Then the local device uses the current model parameters to predict each sample in the local data set, and evaluates the error value between the prediction result and the true label through the cross-entropy loss function. Based on the chain rule, the error value is propagated layer by layer, and the gradient of each layer parameter of the model is calculated according to the error value. Then the model parameters are updated by using the SGD optimizer, and the parameter update amount of the local model is calculated and stored. Each group of local Internet of Things devices uploads the local model parameter update amount to the terminal server. According to the characteristics of each Internet of Things device, the local model parameter space range and the local gradient trend information, a group of belief spaces are initialized for each Internet of Things device. According to the voice features and user behavior patterns of the device, multiple local behaviors are generated, and the fitness of each behavior in the local task is evaluated. The behavior with the highest fitness is selected as the optimal behavior of the Internet of Things device. According to the optimal behavior of each group of Internet of Things devices, a belief update vector is generated. The server integrates the optimal behaviors of each device by taking the average method to integrate the belief update of the device to generate a global belief vector. The global belief and the local model update are used to calculate the weight of each device. Then the terminal server aggregates the model parameter update amount of all devices according to the device weight to obtain the corresponding global model update. The updated global model is distributed to all Internet of Things devices, and the local model parameters of each Internet of Things device are synchronized to the global model parameters, while the update training is performed in real time.

[0098] Based on the verification result, the user's identity is confirmed, and the result is fed back to the corresponding user device. The user's operation request is authorized, and the operation log is recorded.

[0099] Real-time monitoring system running state, and regularly optimize the identity authentication process in the configuration of the parameters.

Claims

1. The identity authentication system based on the Internet of Things and voiceprint recognition is characterized by: It includes acquisition preprocessing module, multimodal fusion module, storage and transmission module, identity screening module, identity authentication module, dynamic risk assessment module, model update module, feedback correction module, authorization record module and monitoring self-optimization module; The acquisition and preprocessing module is used to acquire the user's voice data in real time and preprocess the acquired voice data; The multimodal fusion module is used to extract voiceprint features from voice data and combine voice features with user behavior patterns to construct multimodal features; The storage and transmission module is used to distribute the user's voiceprint features and the constructed multimodal features, and encrypt and transmit each set of feature data; The identity screening module is used to perform a preliminary comparison of the collected voiceprint features to screen out candidate identities; The identity verification module is used to build a voiceprint recognition model and perform in-depth comparison and analysis on the candidate identities initially screened through the voiceprint recognition model to verify the user's identity; The dynamic risk assessment module is used to assess the risk level of identity authentication in real time and dynamically adjust the authentication strategy based on the risk; The model updating module is used to collaboratively optimize the voiceprint recognition model through distributed devices; The feedback correction module is used to provide immediate feedback to users and correct potential errors; The authorization recording module is used to authorize the user's operation request after verification and record the operation log; The monitoring self-optimization module is used to continuously monitor the system operation status and regularly optimize the parameter configuration in the identity authentication process; The specific steps of constructing multimodal features in the multimodal fusion module are as follows: S1.1: Extract the Mel frequency cepstral coefficients, linear prediction coefficients, energy and pitch feature information from the preprocessed speech data, and record the extracted voiceprint features as , and then extract behavioral features and state features from user behavior patterns and device states, and record them as and ; S1.2: Standardize each dimension of voiceprint features, behavioral features, and state features from different sources to eliminate feature scale differences. Then, unify the feature dimensions through zero-padding, truncation, and dimensionality reduction operations. Assign weights to each modal feature and perform weighted fusion of features from different modalities based on the assigned weights. S1.3: Use the Min-Max normalization method to map the fused multimodal features to the interval [0, 1], and use the normalized multimodal features to , a multimodal feature space is constructed based on the number of processed multimodal features, and a group of populations is initialized in the multimodal feature space, and the position of each group of individuals in the population is used as a multimodal feature of the multimodal feature space; S1.4: Based on the separation of feature representations between different users and the stability of feature representations input by the same user multiple times, calculate the fitness value of each multimodal feature in the multimodal feature space, and sort the individuals in each group in descending order of fitness value. Select the top three groups of individuals with the best performance from the population, and move the remaining individuals closer to the top three groups of individuals with the highest fitness values, and update the positions of the remaining individuals. S1.5: After each position update, randomly perturb the individuals whose positions have been updated. If the multimodal fitness value corresponding to the perturbed position is better than the original value, the position update after the perturbation is accepted. After the position update is completed, the fitness values ​​of each individual in the population are recalculated and sorted, and the top three groups of individuals with the best performance are reselected and the positions are updated again. S1.6: Repeat the fitness ranking, best individual update, and position update until the preset maximum number of iterations is reached or the fitness value change of the top-ranked individual is less than the set threshold. Then, the iteration stops and the top-ranked multimodal feature is represented as the best feature and added to the user feature library to update the user feature library. The specific steps of collaborative optimization of the voiceprint recognition model are as follows: S5.1: After the verification strategy is adjusted, each IoT device collects the voice data of its user interaction to form a local data set, and converts the raw voice data into feature representation through voice preprocessing technology. The processed local data set of each device is then defined as ,in Representative Device D, Representative voice features, Representative The user identity label corresponding to the voice feature, Representative The amount of data on device D; S5.2: Each device initializes a model architecture identical to the original voiceprint recognition model and initializes the model parameters on each device to the uniform weights distributed by the global server. Local loss functions, optimizers, and hyperparameters are set for each group of devices. The local device then uses the current model parameters to make predictions for each sample in the dataset locally, and uses the cross-entropy loss function to evaluate the error between the prediction result and the true label. S5.3: Based on the chain rule, the error value is propagated layer by layer, and the gradient of the model parameters at each layer is calculated based on the error value. The model parameters are then updated using the SGD optimizer. The parameter updates of the local model are calculated and stored. Each group of local IoT devices uploads the local model parameter updates to the terminal server. S5.4: Initialize a belief space for each IoT device based on its characteristics, the local model parameter space range, and local gradient trend information. Generate multiple sets of local behaviors based on the device's voice characteristics and user behavior patterns. Evaluate the fitness of each behavior in the local task and select the behavior with the highest fitness as the optimal behavior for the IoT device. Each group of IoT devices generates a belief update vector based on its optimal behavior. The server aggregates the optimal behaviors of each device and averages the belief updates of the devices to generate a global belief vector. S5.5: Utilize global beliefs and local model updates to calculate the weight of each device. The terminal server then performs weighted aggregation on the model parameter updates of all devices based on the device weights to obtain the corresponding global model update. The updated global model is distributed to all IoT devices, and the local model parameters of each IoT device are synchronized with the global model parameters, while updating and training are performed in real time.

2. The identity authentication system based on Internet of Things and voiceprint recognition according to claim 1, characterized in that: The specific steps of distributed storage in the storage transmission module are as follows: S2.1: Create a decentralized storage structure for voiceprint features in the blockchain network, and define each block in the blockchain to store voiceprint feature data, multimodal feature data, and their hash values. Then, using a proof-of-work consensus mechanism, each node in the blockchain calculates and searches for a hash value that is less than a difficulty threshold set by the network. S2.2: The block corresponding to the hash value that meets the conditions obtained through the consensus mechanism is broadcast to the entire network and added to the chain, and each storage block in the blockchain is used as a state node , and the connection between nodes represents the data transmission path in the blockchain network, where Represents the unique identifier of the current storage node, Represents the cumulative reward value of the current node, Represents the number of times the current node has been visited; S2.3: Starting from the starting node of the blockchain as the root node, a search tree is constructed based on the subsequent child nodes in the blockchain network and the root node. The UCB value of each child node is calculated, and the child node with the highest UCB value is selected as the next storage node. If the current child node is not fully expanded, the child node is expanded and the new node is added to the search tree. S2.4: Starting from the newly added node, randomly simulate the storage path, calculate the storage cost and retrieval delay, update the cumulative reward value and number of visits of the search tree according to the simulation results, repeat the selection, expansion, simulation and backtracking steps until the preset maximum number of iterations is reached. After the iteration is completed, select the path with the highest cumulative reward value as the storage path for voiceprint feature data and multimodal feature data, and dynamically adjust the data distribution plan in the blockchain.

3. The identity authentication system based on Internet of Things and voiceprint recognition according to claim 2, characterized in that: Each block in the blockchain described in S2.1 is represented as follows: in, Representative blocks; Representative Voiceprint feature data or multimodal feature data of each block; Represents the hash value of the block; The hash value representing the voiceprint data of this block; Representative The random number in the proof of work of each block; The specific calculation formula for the UCB value mentioned in S2.3 is as follows: Where, Represents the selection value of the current node; Represents the cumulative reward value of the current node s; represents the number of times the current node s has been visited; c represents the parameter for balancing exploration and utilization, which is usually a constant; Represents the number of times the parent node of the current node has been visited; ln represents a logarithmic function, which is used to adjust the priority of unvisited nodes.

4. The identity authentication system based on Internet of Things and voiceprint recognition according to claim 2, characterized in that: The specific steps of the identity screening module's preliminary comparison are as follows: S3.1: Extract voiceprint features from the user's current voice data, select pre-registered voiceprint features that match the current user from the user feature library, and initialize the candidate identity pool based on the extracted voiceprint features. Use dynamic weighted distance to measure the similarity between the features to be verified and the features of each identity in the candidate identity pool, and filter out identity features in the candidate identity pool whose similarity does not meet the preset threshold. S3.2: Initialize the feature weight population and the pool of candidate identities after the initial screening, and calculate the fitness value of each feature weight and candidate identity feature respectively. Then retain the feature weights whose fitness exceeds the preset threshold. Randomly select two groups of individuals from the selected feature weights, perform linear combination to generate new feature weights, and then randomly adjust some weight components so that the new feature weights meet the preset constraints. S3.3: Select candidate identity features with fitness below the preset threshold and replace them with new identity features that have not been selected in the user feature library. Based on the candidate identity features with fitness above the preset threshold and the matching rules, re-update the feature weight population and candidate identity pool until the fitness value of the feature weight population changes by less than the preset threshold over multiple rounds. S3.4: After stopping the update, the feature weight with the highest fitness is selected as the weight distribution of the final matching rule. Then, the matching distance between the feature to be verified and each identity feature in the candidate identity pool is calculated according to the optimized rule, and the maximum threshold of the matching distance is set. Candidate identities with a matching distance less than the maximum threshold are considered to have passed the screening.

5. The identity authentication system based on Internet of Things and voiceprint recognition according to claim 4, characterized in that: The specific calculation formula for the dynamic weighted distance described in S3.1 is as follows: Where, Represents the voiceprint feature to be verified With the The matching distance of candidate identities v, where is the index of candidate identity v in the candidate identity pool; d represents the voiceprint feature to be verified The number of dimensions; Representative The weight of the dimension feature is initially uniformly distributed; Represents the voiceprint feature to be verified No. Dimension value; Representative candidate status No. Dimension value; The specific calculation formulas for the feature weights and the fitness values ​​of candidate identity features described in S3.2 are as follows: Where, Representative The fitness value of the feature weight w; represents the total number of eigenvectors; Represents the attenuation factor of the matching distance. The smaller the distance, the greater the corresponding contribution. Represents the voiceprint feature to be verified With the The similarity score of the candidate identity v; Representative The fitness value of the candidate identity v is the minimum matching distance under the current optimal weight vector.

6. An identity authentication method based on the Internet of Things and voiceprint recognition, used to implement the identity authentication system function based on the Internet of Things and voiceprint recognition as described in any one of claims 1 to 5, characterized in that: The specific steps of this verification method are as follows: Ⅰ. Collect user voice data and device environment information through IoT devices and pre-process the voice data; II. Extract voiceprint features from the processed voice data, combine them with user behavior patterns and device status information to form multimodal features and update the user feature library; Ⅲ. Construct and optimize feature matching rules, make preliminary matching results based on the feature matching rules, and filter out candidate identity sets from the user feature library; IV. Build a voiceprint recognition model to conduct in-depth comparison and analysis on the initially screened candidate identity set to verify the user's identity; V. Real-time assessment of identity verification risk levels, dynamic adjustment of verification strategies, and collaborative optimization of voiceprint recognition models; VI. Confirm the user's identity based on the verification result, feed the result back to the corresponding user device, authorize the user's operation request, and record the operation log; Ⅶ. Monitor the system operation status in real time and regularly optimize the parameter configurations in the identity authentication process.

7. The identity authentication method based on Internet of Things and voiceprint recognition according to claim 6, characterized in that: The specific steps of the in-depth comparison and analysis in step IV are as follows: S4.1: Design and construct a voiceprint recognition model based on the Bi-GRU architecture. The model includes an input layer, a Bi-GRU layer, a fully connected layer, a classification layer, and an output layer. Then, based on the public voiceprint dataset, construct a training set, a test set, and a validation set. S4.2: The training set is input into the voiceprint recognition model. The Bi-GRU layer of the voiceprint recognition model processes the input data in chronological order and reverse chronological order respectively, and outputs the concatenated vectors of the forward and reverse features. The fully connected layer aggregates the received concatenated vectors into a fixed-length voiceprint embedding vector. Then, the classification layer outputs the final identity recognition result based on the candidate identity set. S4.3: Calculate the loss between the final identification result of the model and the actual identification result using the cross-entropy loss function. Then, based on the chain rule, propagate the loss value from the output layer to the input layer of the voiceprint recognition model, and calculate the gradient of each network layer of the model corresponding to the loss value. Then, optimize the parameters of each network layer using the Adam optimizer. Use the verification machine to evaluate the performance of the model after this round of training. If the model performance does not reach the preset performance index, retrain the model until the model loss value converges to the preset threshold. Then stop training, test the performance of the trained voiceprint recognition model on unknown data using the test set, and deploy the model on the actual voiceprint recognition system. S4.4: Use the current voiceprint feature as the current state Input to the voiceprint recognition model, based on the policy network , select the current action , that is, adjusting the parameters or hyperparameters of the voiceprint recognition model, recalculating the predicted output by adjusting the parameters or hyperparameters of the voiceprint recognition model, and based on the updated model performance, Calculate the reward value, where Represents the selection of the current action After the model performance evaluation reward, Represents the recognition accuracy of the model, Represents the similarity between the target domain features and the source domain embedding, and Represent weight parameters respectively; S4.5: Feed the reward value back to the policy network for updating, and accumulate the reward value as the policy optimization target. Maximizing the cumulative reward is the goal of the policy network. The policy network parameters are updated using gradients. The voiceprint recognition model is then optimized for the target domain based on the migrated task data. The model parameters are updated in combination with the contrast loss function. Multiple rounds of reinforcement learning training are repeated until the model loss value converges to the preset range. S4.6: The voiceprint recognition model receives the current voiceprint features and generates an embedding vector of the voiceprint to be verified after forward propagation processing. It calculates the similarity between the embedding vector and each candidate identity in the candidate identity set. After comparing all candidate identities, it selects one or more matching identities with a similarity higher than a preset threshold. The candidate identities are sorted in descending order according to the similarity score, and the identity with the highest similarity score is selected first.

Citation Information

Patent Citations

  • Identity verification method and system based on Internet of Things and voiceprint recognition

    CN114579947A

  • Multi-mode identity verification system and method

    CN118245994A

  • Intelligent business management and scheduling decision-making method of heat supply system based on voice interaction

    CN118607959A