Lightweight voiceprint and text dual-verification speaker identity recognition method, device and equipment
By adopting a lightweight voiceprint and text dual verification method in voiceprint recognition, combined with multi-teacher knowledge distillation and multi-level self-distillation technology, the problems of excessive model, high computational complexity and poor security in the existing technology are solved, and efficient and safe speaker identity recognition is achieved.
Patent Information
- Application Number
- CN202510364675.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-06-06
AI Technical Summary
The existing voiceprint recognition methods have too large models and high computational complexity in resource-constrained environments, and have poor security when facing forgery attacks such as recording and synthetic voice.
The lightweight voiceprint and text dual verification method is adopted, and through multi-teacher knowledge distillation and multi-level self-distillation technology, a lightweight speaker recognition network model and speech recognition model are built, combining voiceprint and text features for dual verification to achieve identity recognition.
While maintaining high accuracy, it significantly reduces the model size and computing complexity, improves real-time identification capabilities in resource-constrained environments, and enhances resistance to forged attacks and improves security.
Smart Images

Figure CN120108403A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of speaker identification, and in particular relates to a lightweight speaker identification method, device and equipment for dual verification of voiceprint and text. Background Art
[0002] Speaker identification aims to authenticate the speaker by identifying personal information in the speaker's speech. Currently, the commonly used speaker identification technology mainly uses voiceprint recognition. Voiceprint recognition is a biometric recognition technology that identifies which speaker the query sentence belongs to by comparing the voiceprint features of the query sentence with the registered speech in the speaker database. This technology has been applied to financial payment, telecommunications anti-fraud, intelligent security, criminal investigation and other fields. Current voiceprint recognition methods mainly focus on the construction and optimization of voiceprint recognition models, such as Thin-ResNet, ECAPA-TDNN, MFA-conformer and CNN-Conformer, and have made significant breakthroughs in recognition accuracy.
[0003] However, current voiceprint recognition methods still face many practical problems. On the one hand, the consumption of computing resources has become a major bottleneck in practical applications, especially in resource-limited environments such as mobile devices and embedded systems. Therefore, how to maintain high accuracy while significantly reducing the size and computational complexity of the model to achieve real-time and efficient speaker recognition is particularly important for large-scale, high-frequency biometric application scenarios.
[0004] On the other hand, as fraud methods become increasingly sophisticated, the security of authentication methods has become another major bottleneck in practical applications. Therefore, how to achieve accurate speaker identification and ensure high security of the method under attacks by various forgery methods such as recording and synthesized speech is particularly important for biometric application scenarios with high security requirements. Summary of the invention
[0005] The purpose of the present invention is to provide a lightweight speaker identification method, device and equipment with dual verification of voiceprint and text, which can solve the problem that the speaker identification model is too large and the calculation complexity is high in a resource-constrained environment, and the speaker identification method has poor security in the face of various forgery attacks such as recording and synthesized speech;
[0006] In order to achieve the above object, the technical solution adopted by the present invention is:
[0007] A lightweight speaker identification method for dual verification of voiceprint and text includes the following steps:
[0008] S1, obtaining the speech features of the speech data to be recognized;
[0009] S2, inputting the speech features of the speech to be recognized into a lightweight speaker recognition network model to obtain the voiceprint features of the speech to be recognized;
[0010] S3, inputting the speech features of the speech to be recognized into a lightweight speech recognition network model to obtain text features of the speech to be recognized;
[0011] S4, comparing the voiceprint features of the speech to be recognized with the voiceprint features in the pre-established speaker voiceprint library one by one, obtaining similarity scores between the voiceprint features of the speech to be recognized and the voiceprint features of all registered speakers, and storing them in a voiceprint result candidate set;
[0012] S5, comparing the text features of the speech to be recognized with the texts in the pre-established speaker text password library one by one, obtaining similarity scores between the text features of the speech to be recognized and the texts of all registered speakers, and storing them in a text result candidate set;
[0013] S6, using a threshold-based sorting and filtering method to discriminate the voiceprint result candidate set and the text result candidate set to obtain the final recognition result.
[0014] Optionally, the implementation method of step S2 is as follows:
[0015] S21, for each of the N sentences in a training batch, extract the speech data x i The phonetic features F i , and input them into the two pre-trained teacher models and the student model to be trained, respectively, to obtain the voiceprint features of teacher model 1 Voiceprint features of teacher model 2 Voiceprint characteristics of student model
[0016] S22, use the following formulas 1, 2 and 3 to extract the logits features of the two teacher models and the student model respectively and
[0017]
[0018] Where s is the scaling factor, W i is the learnable classification weight, cosθ i It is the voiceprint feature and W i The cosine similarity of , m is the marginal distance.
[0019] S23. For voice data x i , respectively construct two teacher models under x i Similarity vectors with all other utterances in the training batch and And select similar vectors respectively and The K most similar speech samples are then sorted according to similarity and finally selected to obtain x i The K nearest neighbor samples of .
[0020] The following formulas 4 and 5 are used to calculate the speech x i and discourse x j Similarity between (i≠j) and The larger the value, the more similar the two samples are;
[0021]
[0022] where <·,·> is the dot product operation, ||·|| 2 is L2 normalization.
[0023] S24. For discourse x i , use the following formulas 6 and 7 to calculate x under the two teacher models respectively i With another discourse x j The similarity difference of voiceprint features and Use the following formula 8 to calculate x under the student model i With another discourse x j The similarity difference of voiceprint features Get x i The similarity difference between the voiceprint features and the K nearest neighbor samples and
[0024]
[0025] where φ(·) represents channel average pooling.
[0026] S25. Use the following formulas 9 and 10 to calculate the domain relationship distillation loss of the voiceprint features under the two teacher models respectively;
[0027]
[0028] Where K is the number of neighbors and N is the number of sentences in a batch.
[0029] S26. For discourse x i , use the following formulas 11, 12 and 13 to calculate x under the two teacher models and the student model respectively i With another discourse x j The similarity difference of logits features Get x i The similarity difference between the logits features of the K nearest neighbor samples and
[0030]
[0031] where ρ ij ∈R 1×M , M is the number of categories, and softmax() represents the similarity of logits as a probability distribution.
[0032] S27. Use the following formulas 14 and 15 to calculate the domain relationship distillation loss of logits features under the two teacher models respectively;
[0033]
[0034] where JS(·,·) represents the Jensen-Shannon divergence;
[0035] S28. Calculate the classification loss of the student model using the following formula 16;
[0036]
[0037] where y i is the one-hot encoding vector of the target category
[0038] S29. Use the following formula 17 to obtain the loss L of multi-teacher knowledge distillation based on similarity relationship RMKD ;
[0039]
[0040] Among them, α 1 , α 2 , β 1 , β 2 and γ are the coefficients of each loss.
[0041] Optionally, the step S3, before inputting the speech features of the speech to be recognized into the lightweight speech recognition network model, further includes obtaining a lightweight speech recognition network model, and the specific implementation method is as follows:
[0042] S31, for N sentences in a training batch, extract each voice data x i The phonetic features F i , and input it into speech recognition models such as Efficient-conformer, and output the speech features obtained by the first stage, second stage and final stage of the conformer block respectively. and
[0043] S32, using the following formula 18 to calculate the shallow speech features and deep speech features The L2 loss between S ;
[0044]
[0045] S33, using the following formula 19 respectively in the shallow speech features and deep speech features Extract logits features based on and
[0046]
[0047] Among them, FC is the fully connected layer.
[0048] S34. Calculate the shallow logits feature using the following formula 20 And deep logits features The KL loss L between L ;
[0049]
[0050] S35, use the following formula 21 to respectively calculate the shallow logits features And deep logits features Extract classification layer features based on and
[0051]
[0052] S36, use the following formula 22 to calculate the shallow classification layer features and deep classification layer features With label Y i The CTC loss L between C ;
[0053]
[0054] S37, using the following formula 23 to obtain the loss L based on multi-level self-distillation MSKD ;
[0055] L MSKD =α×L S +β×L L +γL C (twenty three)
[0056] Among them, α, β and γ are the coefficients of speech feature loss, logits feature loss and classification layer feature loss respectively.
[0057] Optionally, before comparing the voiceprint features of the speech to be recognized with the voiceprint features in a pre-established speaker voiceprint library one by one, step S4 further includes:
[0058] S41, collecting a piece of voice data from each registered speaker;
[0059] S42, extracting speech features of speech data of all registered speakers;
[0060] S43, inputting the speech features of the speech data of all registered speakers into a lightweight speaker recognition network model to obtain the voiceprint features of all registered speakers;
[0061] S44, storing the voiceprint features of all registered speakers in a speaker voiceprint database.
[0062] Optionally, before comparing the text features of the speech to be recognized with the text in the pre-established speaker text password library one by one in step S5, the method further includes:
[0063] S51, collect a text password of each registered speaker, and note that the text is different from the text content of the voice data of the same registered speaker in S41;
[0064] S52, storing the text passwords of all registered speakers into a speaker text password database.
[0065] Optionally, the step S6, using a threshold-based sorting and filtering method to discriminate the voiceprint result candidate set and the text result candidate set, comprises the following steps:
[0066] S61, sorting the speaker IDs in the voiceprint result candidate set from high to low according to the similarity scores, and filtering out the results with similarity lower than a threshold, to obtain a second voiceprint result candidate set;
[0067] S62, sorting the speaker IDs in the text result candidate set from high to low according to the similarity scores, and filtering out the results with similarity lower than a threshold, to obtain a second text result candidate set;
[0068] S63, merging the second voiceprint result candidate set and the second text result candidate set according to ID, calculating the voiceprint similarity and text similarity scores at 50% each, and obtaining a comprehensive result candidate set;
[0069] S64, sorting the comprehensive result candidate set according to the similarity scores, and the speaker ID ranked first is the final recognition result.
[0070] According to another aspect of the present invention, there is also provided a lightweight speaker identification device for dual verification of voiceprint and text, comprising:
[0071] An acquisition module, used for acquiring voice data to be recognized;
[0072] An extraction module, used to extract voiceprint features and speech text information of the speaker's speech data respectively;
[0073] The recognition module is used to compare the voiceprint features and text information of the voice data to be recognized with the data in the database to obtain the double identity authentication result of the voiceprint and text information;
[0074] According to another aspect of the present invention, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned lightweight voiceprint and text dual identity authentication method are performed.
[0075] According to another aspect of the present invention, there is also provided a lightweight voiceprint and text dual identity authentication device, characterized in that it includes:
[0076] Memory, for storing software applications,
[0077] The processor is used to execute the software application, and each program of the software application correspondingly executes the steps in the above-mentioned lightweight voiceprint and text dual identity authentication method.
[0078] The present invention first proposes a lightweight voiceprint recognition model based on multi-teacher knowledge distillation of similarity relations and a lightweight speech recognition model based on multi-level self-distillation. Then the obtained lightweight voiceprint recognition model and lightweight speech recognition model are used for dual identity authentication. Specifically, the voiceprint features extracted by the lightweight voiceprint recognition model and the predicted text features generated by the lightweight speech recognition model are compared with the voiceprint features of registered users in the voiceprint database and the text in the registered text library for similarity, and the users in the database are sorted according to the similarity, and the top K users are retained respectively. The result sets are then merged, scored and the final score is calculated. The speaker with the highest score is the recognition result. This method takes into account both lightweight and high precision, and will provide a more competitive solution for practical applications, especially in resource-constrained hardware platforms such as smart mobile products and wearable devices. It has broad application prospects. BRIEF DESCRIPTION OF THE DRAWINGS
[0079] Figure 1 It is a schematic diagram of the process of the present invention. DETAILED DESCRIPTION
[0080] The present invention is further described below in conjunction with specific embodiments.
[0081] The lightweight speaker identification method for dual verification of voiceprint and text described in this embodiment includes the following steps:
[0082] S1, obtaining the speech feature F of the speech data to be recognized;
[0083] S2, input the speech feature F of the speech to be recognized into the lightweight speaker recognition network model Half-MFA-Conformer to obtain the voiceprint feature E of the speech to be recognized;
[0084] Among them, the step S2, before inputting the speech feature F of the speech to be recognized into the Half-MFA-Conformer in the lightweight speaker recognition network model, also includes obtaining the lightweight speaker recognition network model Half-MFA-Conformer, and the step of using the multi-teacher knowledge distillation lightweight speaker recognition method based on similarity relationship to obtain the lightweight speaker recognition network model includes the following steps S21 to S29:
[0085] S21, for each of the N sentences in a training batch, extract the speech data x i The phonetic features F i , and input them into two pre-trained teacher models and the student model to be trained, respectively, to obtain the voiceprint features of teacher model 1 such as CNN-Conformer Teacher model 2 such as voiceprint features of MFA-Conformer Voiceprint features of student models such as Half-MFA-Conformer
[0086] Specifically, this embodiment uses the VoxCeleb dataset as the training set, and the VoxCeleb dataset consists of two sub-datasets, VoxCeleb1 and VoxCeleb2. The speech data in the training set is divided into several batches of equal size, each batch contains N = 200 speech data, and for each speech in a batch of data, a speech segment of length 2s is randomly selected, and it is overlapped with a frame length of 25ms and a frame shift of 10ms. The Hamming window is added to extract 200 frames of 80-dimensional Fbank features, and after the speaker recognition model, a 192-dimensional voiceprint feature is obtained.
[0087] S22, use the following formulas 1, 2 and 3 to extract the logits features of the two teacher models and the student model respectively and
[0088]
[0089]
[0090] Where s is the scaling factor, W i is the learnable classification weight, cosθ i It is the voiceprint feature and W i The cosine similarity of , m is the marginal distance.
[0091] In this embodiment, s=30 and m=0.2 are set.
[0092] S23. For voice data x i , respectively construct two teacher models under x i Similarity vectors with all other utterances in the training batch and And select similar vectors respectively and The K most similar speech samples are then sorted according to similarity and finally selected to obtain x i The K nearest neighbor samples of .
[0093] The following formulas 4 and 5 are used to calculate the speech x i and discourse x j Similarity between (i≠j) and The larger the value, the more similar the two samples are;
[0094]
[0095] where <·,·> is the dot product operation, ||·|| 2 is L2 normalization.
[0096] In this embodiment, K=3 is set.
[0097] S24. For discourse x i , use the following formulas 6 and 7 to calculate x under the two teacher models respectively i With another discourse x j The similarity difference of voiceprint features and Use the following formula 8 to calculate x under the student model i With another discourse x j The similarity difference of voiceprint features Get x i The similarity difference between the voiceprint features and the K nearest neighbor samples and
[0098]
[0099]
[0100] where φ(·) represents channel average pooling.
[0101] S25. Use the following formulas 9 and 10 to calculate the domain relationship distillation loss of the voiceprint features under the two teacher models respectively;
[0102]
[0103] Where K is the number of neighbors and N is the number of sentences in a batch.
[0104] In this embodiment, K=3 and N=200 are set.
[0105] S26. For discourse x i , use the following formulas 11, 12 and 13 to calculate x under the two teacher models and the student model respectively i With another discourse x j The similarity difference of logits features and Get x i The similarity difference between the logits features of the K nearest neighbor samples and
[0106]
[0107] where ρ ij ∈R 1×M , M is the number of categories, and softmax() represents the similarity of logits as a probability distribution.
[0108] In this embodiment, M=7205 is set.
[0109] S27. Use the following formulas 14 and 15 to calculate the domain relationship distillation loss of logits features under the two teacher models respectively;
[0110]
[0111] where JS(·,·) represents the Jensen-Shannon divergence.
[0112] S28. Calculate the classification loss of the student model using the following formula 16;
[0113]
[0114] where y i is the one-hot encoding vector of the target category
[0115] S29. Use the following formula 17 to obtain the loss L of multi-teacher knowledge distillation based on similarity relationship RMKD ;
[0116]
[0117] Among them, α 1 , α 2 , β 1 , β 2 and γ are the coefficients of each loss.
[0118] In this embodiment, α is set 1 =α 2 =β 1 =β 2 =γ=1.
[0119] S3, input the speech feature F of the speech to be recognized into the lightweight speech recognition network model self-Efficient-conformer to obtain the text feature S of the speech to be recognized;
[0120] Among them, the step S3, before inputting the speech feature F of the speech to be recognized into the lightweight speech recognition network model self-Efficient-conformer, also includes obtaining a lightweight speech recognition network model self-Efficient-conformer, and the step of obtaining the lightweight speech recognition network model using a lightweight speech recognition method based on multi-level self-distillation includes the following steps S31 to S37:
[0121] S31, for N sentences in a training batch, extract each voice data x i The phonetic features F i , and input it into the Efficient-conformer speech recognition model, and output the speech features obtained by the first stage, second stage and final stage of the conformer block respectively. and
[0122] Specifically, this embodiment uses the LibriSpeech dataset as the training set. The speech data in the training set is divided into several batches of equal size, each batch contains N = 200 speech data, and for each speech in a batch of data, it is framed with an overlapping frame length of 25ms and a frame shift of 10ms, and a Hamming window is added to extract the 80-dimensional Fbank feature. After the speech recognition model, a 256-dimensional speech feature is obtained.
[0123] S32, using the following formula 18 to calculate the shallow speech features and deep speech features The L2 loss between S ;
[0124]
[0125] S33, using the following formula 19 respectively in the shallow speech features and deep speech features Extract logits features based on and
[0126]
[0127] Among them, FC is the fully connected layer.
[0128] S34. Calculate the shallow logits feature using the following formula 20 And deep logits features The KL loss L between L ;
[0129]
[0130] S35, use the following formula 21 to respectively logits features in the shallow layer And deep logits features Extract classification layer features based on and
[0131]
[0132] S36, use the following formula 22 to calculate the shallow classification layer features and deep classification layer features With label Y i The CTC loss L between C ;
[0133]
[0134] S37, using the following formula 23 to obtain the loss L based on multi-level self-distillation MSKD ;
[0135] L MSKD =α×L S +β×L L +γ×L C (twenty three)
[0136] Among them, α, β and γ are the coefficients of speech feature loss, logits feature loss and classification layer feature loss respectively.
[0137] In this embodiment, α=0.3, β=0.3, and γ=0.7 are set.
[0138] S4, compare the voiceprint feature E of the speech to be recognized with the voiceprint feature in the pre-established speaker voiceprint database Perform a one-to-one comparison to obtain the voiceprint feature E of the speech to be recognized and the voiceprint features E of all registered speakers. r The similarity score between them is stored in the voiceprint result candidate set
[0139] In step S4, the voiceprint feature E of the speech to be recognized is compared with the pre-established speaker voiceprint database. Before comparing the voiceprint features one by one, the following steps S41 to S44 of establishing a speaker voiceprint library are also included:
[0140] S41, collect a piece of voice data from each registered speaker
[0141] S42, extracting speech features of speech data of all M registered speakers
[0142] S43, voice features of all registered speaker voice data Input a lightweight speaker recognition network model to obtain the voiceprint features of all registered speakers
[0143] S44, storing the voiceprint features of all registered speakers in the speaker voiceprint database to obtain the speaker voiceprint database
[0144] S5, compare the text feature S of the speech to be recognized with the pre-established speaker text password library Compare the text in one by one to obtain the text features S of the speech to be recognized and the text S of all registered speakers r The similarity score between them is stored in the text result candidate set
[0145] In step S5, the text feature S of the speech to be recognized is compared with a pre-established speaker text password library. Before comparing the texts in the text one by one, the following steps S51 and S52 are also included to establish a speaker text password library:
[0146] S51, collect a text password from each registered speaker Note that this text is different from the text content of the voice data of the same registered speaker in S41;
[0147] S52, storing the text passwords of all registered speakers into the speaker text password library, obtaining the text password library
[0148] S6, using the threshold-based sorting and filtering method to sort the candidate set of voiceprint results and text result candidate set Make a judgment and get the final recognition result.
[0149] In step S6, the candidate set of voiceprint results is sorted and filtered using a threshold-based sorting method. and text result candidate set The determination also includes the following steps S61 to S64:
[0150] S61, the voiceprint result candidate set The speaker IDs (1 to M) in the above example are sorted from high to low according to the similarity scores, and the results with similarity below the threshold are filtered out to obtain the second voiceprint result candidate set.
[0151] S62, the text result candidate set The speaker IDs (1 to M) in the text are sorted from high to low according to the similarity scores, and the results with similarity below the threshold are filtered out to obtain the second text result candidate set.
[0152] S63: Merge the second voiceprint result candidate set and the second text result candidate set according to ID The voiceprint similarity and text similarity scores are calculated with 50% each to obtain the comprehensive result candidate set.
[0153] S64, sorting the comprehensive result candidate set according to the similarity scores, and the speaker ID ranked first is the final recognition result.
[0154] According to another aspect of the present invention, there is also provided a lightweight speaker identification device for dual verification of voiceprint and text, comprising:
[0155] An acquisition module, used for acquiring voice data to be recognized;
[0156] An extraction module, used to extract voiceprint features and speech text information of the speaker's speech data respectively;
[0157] The recognition module is used to compare the voiceprint features and text information of the voice data to be recognized with the data in the database to obtain the dual identity authentication results of the voiceprint and text information.
[0158] The implementation principle and technical effect of the speaker identification device provided in the embodiment of the present invention are similar to those of the above embodiment, and will not be described in detail here.
[0159] According to another aspect of the present invention, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned lightweight voiceprint and text dual identity authentication method are performed.
[0160] According to another aspect of the present invention, there is also provided a lightweight voiceprint and text dual identity authentication device, characterized in that it includes:
[0161] Memory, for storing software applications,
[0162] The processor is used to execute the software application, and each program of the software application correspondingly executes the steps in the above-mentioned lightweight voiceprint and text dual identity authentication method.
Claims
1. A lightweight speaker identification method with dual verification of voiceprint and text, characterized in that: The following steps are involved: S1, obtaining the speech features of the speech data to be recognized; S2, inputting the speech features of the speech to be recognized into a lightweight speaker recognition network model to obtain the voiceprint features of the speech to be recognized; S3, inputting the speech features of the speech to be recognized into a lightweight speech recognition network model to obtain text features of the speech to be recognized; S4, comparing the voiceprint features of the speech to be recognized with the voiceprint features in the pre-established speaker voiceprint library one by one, obtaining similarity scores between the voiceprint features of the speech to be recognized and the voiceprint features of all registered speakers, and storing them in a voiceprint result candidate set; S5, comparing the text features of the speech to be recognized with the texts in the pre-established speaker text password library one by one, obtaining similarity scores between the text features of the speech to be recognized and the texts of all registered speakers, and storing them in a text result candidate set; S6, using a threshold-based sorting and filtering method to discriminate the voiceprint result candidate set and the text result candidate set to obtain the final recognition result.
2. The lightweight speaker identification method with dual verification of voiceprint and text according to claim 1 is characterized in that: The implementation method of step S2 is as follows: S21, for each of the N sentences in a training batch, extract the speech data x i The phonetic features F i , and input them into the two pre-trained teacher models and the student model to be trained, respectively, to obtain the voiceprint features of teacher model 1 Voiceprint features of teacher model 2 Voiceprint characteristics of student model S22, use the following formulas 1, 2 and 3 to extract the logits features of the two teacher models and the student model respectively and Where s is the scaling factor, W i is the learnable classification weight, cosθ i It is the voiceprint feature and W i The cosine similarity of , m is the marginal distance; S23. For voice data x i , respectively construct two teacher models under x i Similarity vectors with all other utterances in the training batch and And select similar vectors respectively and The K most similar speech samples are then sorted according to similarity and finally selected to obtain x i The K nearest neighbor samples of ; The following formulas 4 and 5 are used to calculate the speech x i and discourse x j Similarity between (i≠j) and The larger the value, the more similar the two samples are; Where <·,·> is the dot product operation, ||·||2 is the L2 normalization; S24. For discourse x i , use the following formulas 6 and 7 to calculate x under the two teacher models respectively i With another discourse x j The similarity difference of voiceprint features and Use the following formula 8 to calculate x under the student model i With another discourse x j The similarity difference of voiceprint features Get x i The similarity difference between the voiceprint features and the K nearest neighbor samples and Where φ(·) represents channel average pooling; S25. Use the following formulas 9 and 10 to calculate the domain relationship distillation loss of the voiceprint features under the two teacher models respectively; Where K is the number of neighbors and N is the number of statements in a batch; S26. For discourse x i , use the following formulas 11, 12 and 13 to calculate x under the two teacher models and the student model respectively i With another discourse x j The similarity difference of logits features Get x i The similarity difference between the logits features of the K nearest neighbor samples and where ρ ij ∈R 1×M , M is the number of categories, and softmax() represents the logits similarity as a probability distribution; S27. Use the following formulas 14 and 15 to calculate the domain relationship distillation loss of logits features under the two teacher models respectively; where JS(·,·) represents the Jensen-Shannon divergence; S28. Calculate the classification loss of the student model using the following formula 16; where y i is the one-hot encoding vector of the target category S29. Use the following formula 17 to obtain the loss L of multi-teacher knowledge distillation based on similarity relationship RMKD ; Among them, α1, α2, β1, β2 and γ are the coefficients of each loss.
3. The lightweight speaker identification method with dual verification of voiceprint and text according to claim 1 is characterized in that: The step S3, before inputting the speech features of the speech to be recognized into the lightweight speech recognition network model, also includes obtaining a lightweight speech recognition network model, and the specific implementation method is as follows: S31, for N sentences in a training batch, extract each voice data x i The phonetic features F i , and input it into speech recognition models such as Efficient-conformer, and output the speech features obtained by the first stage, second stage and final stage of the conformer block respectively. and S32, using the following formula 18 to calculate the shallow speech features and deep speech features The L2 loss between s ; S33, using the following formula 19 respectively in the shallow speech features and deep speech features Extract logits features based on and Among them, FC is the fully connected layer; S34. Calculate the shallow logits feature using the following formula 20 And deep logits features The KL loss L between L ; S35, use the following formula 21 to respectively logits features in the shallow layer And deep logits features Extract classification layer features based on and S36, use the following formula 22 to calculate the shallow classification layer features and deep classification layer features With label Y i The CTC loss L between C ; S37, using the following formula 23 to obtain the loss L based on multi-level self-distillation MSKD ; L MSKD =α×L s +β×L L +γ×L C (23) Among them, α, β and γ are the coefficients of speech feature loss, logits feature loss and classification layer feature loss respectively.
4. The lightweight speaker identification method with dual verification of voiceprint and text according to claim 1 is characterized in that: Before the step S4, comparing the voiceprint features of the speech to be recognized with the voiceprint features in the pre-established speaker voiceprint library one by one, the method further includes: S41, collecting a piece of voice data from each registered speaker; S42, extracting speech features of speech data of all registered speakers; S43, inputting the speech features of the speech data of all registered speakers into a lightweight speaker recognition network model to obtain the voiceprint features of all registered speakers; S44, storing the voiceprint features of all registered speakers in a speaker voiceprint database.
5. The lightweight speaker identification method with dual verification of voiceprint and text according to claim 1 is characterized in that: Before the step S5, comparing the text features of the speech to be recognized with the text in the pre-established speaker text password library one by one, the method further includes: S51, collect a text password of each registered speaker, and note that the text is different from the text content of the voice data of the same registered speaker in S41; S52, storing the text passwords of all registered speakers into a speaker text password database.
6. The lightweight speaker identification method with dual verification of voiceprint and text according to claim 1 is characterized in that: The step S6, using a threshold-based sorting and filtering method to discriminate the voiceprint result candidate set and the text result candidate set, includes the following steps: S61, sorting the speaker IDs in the voiceprint result candidate set from high to low according to the similarity scores, and filtering out the results with similarity lower than a threshold, to obtain a second voiceprint result candidate set; S62, sorting the speaker IDs in the text result candidate set from high to low according to the similarity scores, and filtering out the results with similarity lower than a threshold, to obtain a second text result candidate set; S63, merging the second voiceprint result candidate set and the second text result candidate set according to ID, calculating the voiceprint similarity and text similarity scores at 50% each, and obtaining a comprehensive result candidate set; S64, sorting the comprehensive result candidate set according to the similarity scores, and the speaker ID ranked first is the final recognition result.
7. A lightweight speaker identification device with dual verification of voiceprint and text, comprising: An acquisition module, used for acquiring voice data to be recognized; An extraction module, used to extract voiceprint features and speech text information of the speaker's speech data respectively; The recognition module is used to compare the voiceprint features and text information of the voice data to be recognized with the data in the database to obtain the dual identity authentication results of the voiceprint and text information.
8. A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, performs the steps in the above-mentioned lightweight voiceprint and text dual identity authentication method.
9. A lightweight voiceprint and text dual identity authentication device, characterized in that: include: a memory, at least one memory for storing a software application, A processor, at least one processor, is used to execute the software application, and each program of the software application correspondingly executes the steps in the above-mentioned lightweight voiceprint and text dual identity authentication method.