Voiceprint recognition method and device based on large model, equipment and medium

By using the combination of Transformer architecture and language model, the clarity evaluation and weighted fusion of voice signals is achieved, solving the problem of low reliability of recognition results in existing voiceprint recognition systems, and improving the identity recognition accuracy of financial remote verification.

CN120340503APending Publication Date: 2025-07-18PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510694266.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-27
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing voiceprint recognition system does not consider the non-uniform mass distribution within the voice signal and lacks a clarity evaluation mechanism, which leads to low reliability of the recognition results, especially in financial remote verification scenarios, where there are problems such as limited recognition accuracy and errors in paragraph boundaries.

Method used

The automatic speech recognition model using Transformer architecture outputs a time-stamped character-level probability sequence, uses a pre-trained language model to perform semantic clauses, calculates the speech clarity score, and performs weighted fusion of the vocalprint feature vector based on the clarity score, and finally compares it with the vocalprint database for identity recognition.

Benefits of technology

It improves the clarity and reliability of voiceprint recognition results, especially in application scenarios such as finance with high security requirements, and improves the accuracy and robustness of identity recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120340503A_ABST
    Figure CN120340503A_ABST
Patent Text Reader

Abstract

The invention discloses a voiceprint recognition method, a voiceprint recognition device, voiceprint recognition equipment and a medium based on a large model. The voiceprint recognition method comprises the following steps of: processing an input audio by using an automatic voice recognition model of a Transform architecture, and outputting a character-level probability sequence with a timestamp; and converting the character-level probability sequence into an initial text, and inputting the initial text into a pre-trained language model for semantic phrasing to obtain a plurality of phrased texts and corresponding audio time intervals. And according to the character-level probability sequence, calculating a voice definition score of each clause through a preset definition scoring formula. And determining an audio interval of each clause according to the audio time interval, extracting voiceprint feature vectors from the audio intervals of the clauses, and performing weighted fusion on the voiceprint feature vectors according to the voice definition score to generate a final voiceprint feature. And finally comparing the final voiceprint feature with a voiceprint database to complete identity recognition. According to the method, the definition of the voiceprint recognition result is improved, and the reliability is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to a voiceprint recognition method, device, equipment and medium based on a large model. Background Art

[0002] Under the existing technology, in the financial remote verification scenario, since the traditional solutions generally adopted by the current voiceprint recognition system do not consider the non-uniform quality distribution inside the voice signal and lack a clarity evaluation mechanism for different segments in the voice signal. Moreover, the recognition accuracy of the traditional punctuation prediction model for professional terms is limited, resulting in incorrect paragraph boundaries in subsequent voiceprint feature extraction. At the same time, the traditional solution uses average pooling or a simple attention mechanism to fuse segment features, without considering the reliability differences of different segments, thus causing the problem of low reliability of voiceprint recognition results. Summary of the Invention

[0003] Embodiments of the present invention provide a voiceprint recognition method, device, equipment and medium based on a large model, aiming to solve the problem of low reliability of voiceprint recognition results of traditional solutions under the existing technology.

[0004] In a first aspect, an embodiment of the present invention provides a voiceprint recognition method based on a large model, including: processing an input audio through an automatic speech recognition model using a Transformer architecture to output a character-level probability sequence with timestamps; converting the character-level probability sequence into an initial text, inputting the initial text into a pre-trained language model for semantic sentence splitting to obtain texts of multiple sentences and corresponding audio time intervals; calculating the voice clarity score of each sentence according to the character-level probability sequence through a preset clarity scoring formula; determining the audio interval of each sentence according to the audio time interval, extracting a voiceprint feature vector in the audio interval of the sentence, and performing weighted fusion on the voiceprint feature vector according to the voice clarity score to generate a final voiceprint feature; comparing the final voiceprint feature with a voiceprint database to complete identity recognition.

[0005] In a second aspect, an embodiment of the present invention further provides a voiceprint recognition device based on a large model, which is used to execute the voiceprint recognition method based on a large model as described above.

[0006] In a third aspect, an embodiment of the present invention further provides a computer device, where the computer device includes a memory and a processor connected to the memory; the memory is used to store a computer program; the processor is used to run the computer program stored in the memory to execute the steps of the above voiceprint recognition method based on a large model.

[0007] Fourthly, an embodiment of the present invention further provides a computer-readable storage medium storing a computer program, the computer program including program instructions which, when executed by a processor, can implement the steps of the above-mentioned large model-based voiceprint recognition method.

[0008] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0009] In the technical solution of the present invention, an automatic speech recognition model using a Transformer architecture is used to process the input audio, and a character-level probability sequence with timestamps is output. Then, the character-level probability sequence is converted into an initial text, which is input into a pre-trained language model for semantic clause segmentation to obtain texts of multiple clauses and corresponding audio time intervals. Then, according to the character-level probability sequence, the speech clarity scores of each clause are calculated through a preset clarity scoring formula. Then, according to the audio time interval, the audio interval of each clause is determined, and voiceprint feature vectors are extracted from the audio interval of the clause, and the voiceprint feature vectors are weighted and fused according to the speech clarity scores to generate a final voiceprint feature. Finally, the final voiceprint feature is compared with the voiceprint database to complete identity recognition. This improves the clarity of the voiceprint recognition result and enhances the reliability. Description of the Drawings

[0010] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0011] Figure 1 It is a flowchart of the large model-based voiceprint recognition method provided by the present invention;

[0012] Figure 2 It is the first sub-flowchart of the large model-based voiceprint recognition method provided by the present invention;

[0013] Figure 3 It is a sub-flowchart of the second sub-flowchart of the large model-based voiceprint recognition method provided by the present invention;

[0014] Figure 4 It is the third sub-flowchart of the large model-based voiceprint recognition method provided by the present invention;

[0015] Figure 5 It is the fourth sub-flowchart of the large model-based voiceprint recognition method provided by the present invention;

[0016] Figure 6 It is the fifth sub-flowchart of the large model-based voiceprint recognition method provided by the present invention;

[0017] Figure 7 This is the sixth sub - flowchart of the voiceprint recognition method based on large models provided by the present invention;

[0018] Figure 8 This is a schematic block diagram of a unit of the voiceprint recognition device based on large models provided by the present invention;

[0019] Figure 9 This is a schematic block diagram of the computer device provided by the embodiments of the present invention. Detailed implementation manners

[0020] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0021] It should be understood that when used in this specification and the appended claims, the terms "comprising" and "including" indicate the presence of the described features, wholes, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.

[0022] It should also be understood that the terms used in this specification of the present invention are only for the purpose of describing the medical embodiments and are not intended to limit the present invention. As used in this specification of the present invention and the appended claims, unless clearly indicated otherwise by the context, the singular forms "a", "an", and "the" are intended to include the plural forms.

[0023] It should be further understood that the term " / and / or" used in this specification of the present invention and the appended claims refers to any combination and all possible combinations of one or more of the related listed items, and includes these combinations.

[0024] The present invention aims to solve the problem of low reliability of voiceprint recognition results in traditional solutions under the existing technology, and provides a voiceprint recognition method, device, equipment, and medium based on large models. The voiceprint recognition method based on large models includes the following steps.

[0025] S110. Process the input audio through an automatic speech recognition model using the Transformer architecture, and output a character - level probability sequence with timestamps;

[0026] S120. Convert the character-level probability sequence into an initial text, input it into a pre-trained language model for semantic clause segmentation, and obtain texts of multiple clauses and corresponding audio time intervals;

[0027] S130. Calculate the speech clarity scores of each clause according to the character-level probability sequence through a preset clarity scoring formula;

[0028] S140. Determine the audio interval of each clause according to the audio time interval, extract the voiceprint feature vector in the audio interval of the clause, and perform weighted fusion on the voiceprint feature vector according to the speech clarity score to generate a final voiceprint feature;

[0029] S150. Compare the final voiceprint feature with the voiceprint database to complete identity recognition.

[0030] First, an automatic speech recognition (ASR) model based on the Transformer architecture is used to process the input audio. The core feature of this model is to use a pure Transformer network for audio signal transcription. During this process, the automatic speech recognition model outputs word-by-word probabilities and records the timestamp for each character. The processing result of each audio frame not only contains the probability distribution of the character but also records the start and end timestamps corresponding to each character. After the model outputs the character probability distribution corresponding to each frame, a connectionist temporal classification algorithm is used to align and generate a character sequence with millisecond-level timestamps. For example, the two characters "confirm" in the audio may correspond to the timestamps 120 - 180ms and 181 - 240ms, and the character probability of each frame is also recorded. In this way, the automatic speech recognition model can accurately perform temporal alignment on the audio to ensure the accuracy of time information during the speech-to-text conversion process.

[0031] Then, splice the character sequence output by the automatic speech recognition model into an initial text, such as "I confirm the insured ID number 1234", and input it into a large language model such as GPT for semantic clause segmentation. The large model intelligently breaks the text according to the context, thereby dividing it into several clauses with semantic integrity. During this process, the model calculates a clarity score based on the semantic coherence and clarity of each clause. This clarity score is inferred based on the large model's semantic understanding of the text and combined with the character-level probability. Specifically, the large model inserts clause delimiters based on grammar rules and business scenarios to generate a text segmentation that conforms to the semantics. For example, if the grammar rules and business scenarios are financial jargon, the generated text segmentation that conforms to the semantics is "I confirm the insurance, ID number 1234". By tracing back the character timestamps, map the clause positions to the original audio to obtain the audio time intervals corresponding to each clause.

[0032] Afterwards, the speech intelligibility score of each sentence can be calculated based on the character-level probability information obtained from the automatic speech recognition model. The specific method is to take the geometric average of the probability values of all characters in each sentence to obtain the overall intelligibility score of the sentence. The intelligibility score will be used as the weighted basis for subsequent voiceprint feature extraction.

[0033] After obtaining the clarity score of each sentence, the next step is to extract the voiceprint features of the corresponding sentence in the audio. At this time, the system locates and cuts the audio segment according to the audio time interval of each sentence. The voiceprint features corresponding to each audio segment will be extracted and further weighted fused. The weighted fusion strategy is based on the speech clarity score of the sentence. Through this weighted fusion strategy, the contribution of sentences with high speech clarity and clear semantics to the final voiceprint features can be better enhanced, thereby improving the accuracy and robustness of voiceprint recognition.

[0034] Finally, the weighted fused voiceprint features will be compared with the pre-stored voiceprints in the voiceprint database. Through the comparison results, the user's identity can be identified. So far, this solution has effectively improved the accuracy of the identity recognition system, especially in application scenarios with high security requirements such as finance, and can provide strong technical support for remote identity authentication, telephone banking account opening, insurance remote insurance and other services.

[0035] In one embodiment, the step of S110 includes:

[0036] S111, performing frame division and Mel feature extraction on the input audio to obtain a feature sequence;

[0037] S112, inputting the feature sequence into a pre-trained automatic speech recognition model to obtain the frame-level character probability distribution output by the model, wherein the model is trained using a joint loss function of connection time series classification and cross entropy, and is optimized for background noise and time-frequency mask data;

[0038] S113, generating the start and end timestamps of each character by connecting the timing classification decoding, and forming the character-level probability sequence with the timestamp in combination with the frame-level character probability distribution.

[0039] In this embodiment, before processing the input audio using an automatic speech recognition model with a Transformer architecture, the input audio needs to be processed first. The input audio is usually a continuous speech signal, and this signal must be converted into features suitable for input to the automatic speech recognition (ASR) model through a series of steps. First, the audio signal is framed, and Mel features (Mel-spectrogram) are extracted using methods such as Mel Frequency Cepstral Coefficients (MFCC) or Mel filter banks to obtain a feature sequence representing the audio content. The feature vector of each frame captures the audio feature information at that moment and provides input for the subsequent speech recognition model.

[0040] Next, the extracted Mel feature sequence is input into a pre-trained automatic speech recognition model. The core of this model is the Transformer architecture, whose structure is based on the self-attention mechanism and can effectively capture the temporal features in the audio sequence. In this way, the model can generate a character-level probability distribution corresponding to each frame based on the input audio features. Specifically, the model outputs a character probability distribution for each frame where V is the size of the vocabulary and t represents the index of the time frame. In this way, the model can assign a probability distribution to each audio frame, indicating the probability that the frame belongs to each character.

[0041] To improve the recognition ability and noise resistance of the model, the automatic speech recognition model is trained using a combined loss function of Connectionist Temporal Classification and Cross-Entropy. The Connectionist Temporal Classification loss function allows the model to be trained without precise alignment labels, adapting to the alignment problem between audio and text. In addition, to improve the robustness of the model in the actual noise environment, data augmentation methods such as background noise and time-frequency masking, such as Babble Noise and SpecAugment, are introduced during the training process, which can effectively improve the model's adaptability to complex background noise.

[0042] After being processed by the automatic speech recognition model, what the system obtains is the character-level probability distribution for each frame. Next, the Connectionist Temporal Classification decoding algorithm is used to process these frame-level probability distributions to generate the start and end timestamps of each character. Specifically, the Connectionist Temporal Classification decoding will use the probability values of each frame to find the most likely character sequence and infer the start and end times of each character. In this way, a character-level probability sequence with timestamps can be formed. The generation of timestamps is crucial for subsequent voiceprint feature extraction because it can accurately assign a time interval to each character, making the subsequent cutting and analysis of the audio segments of each clause more accurate.

[0043] By introducing an automatic speech recognition architecture and a joint training method, this embodiment significantly improves the accuracy and robustness in the process of audio-to-text conversion. The character-level probability sequence with timestamps lays a solid foundation for subsequent clause recognition, speech clarity analysis, and voiceprint feature extraction, ensuring the efficiency and accuracy of the system.

[0044] In one embodiment, the steps of S120 include:

[0045] S121. Perform connectionist temporal classification beam search decoding on the character-level probability sequence to generate an initial text;

[0046] S122. Input the initial text into the pre-trained language model for semantic clause segmentation;

[0047] S123. Trace back the character-level timestamps according to the clause positions to determine the start and end times of the audio corresponding to each clause;

[0048] S124. Output the clause text and its associated set of audio time intervals.

[0049] First, perform connectionist temporal classification beam search decoding on the output character-level probability sequence with timestamps described above. Connectionist temporal classification beam search decoding is a decoding technique commonly used in speech recognition, aiming to select the most likely character sequence from a sequence of character probability distributions while maintaining its time alignment information. The core idea of connectionist temporal classification decoding is to use the beam search strategy to find the path with the highest probability among possible character sequences given the character-level probability distribution of each frame. The beam search strategy is to limit the search space and control the search range through a certain "beam width" to improve the decoding efficiency and accuracy. During the decoding process, the algorithm will retain several candidate paths with the highest probabilities and search for the optimal character sequence through dynamic programming methods, finally generating the initial text. At this time, although there may be some errors in grammar and semantics in the decoded text, the timestamps of its characters have been accurately aligned through connectionist temporal classification decoding.

[0050] Next, the initial text is input into a pre-trained large language model, such as GPT, for semantic clause segmentation. The large language model has strong context understanding ability and can perform intelligent sentence splitting according to the semantic structure and context of the text. For example, the model will split the original long text into multiple meaningful clauses according to the grammar and semantic rules of the sentences, ensuring that each clause is grammatically complete and semantically reasonable. In this process, the language model not only relies on lexical-level information but also uses the context for reasoning to perform reasonable clause segmentation on the text. Through this method, the original text generated by the ASR model can be intelligently adjusted and corrected, thereby improving the readability and semantic accuracy of the text. Specifically, the language model will insert clause markers at appropriate positions according to grammar rules and semantic understanding. In the financial scenario, the model will pay special attention to the demarcation of business-critical information, such as splitting "I confirm the insured ID number 1234" into "I confirm the insurance. ID number 1234". This process fully considers the differences between business terms and spoken expressions, ensuring that the clause segmentation results meet both grammar norms and business requirements.

[0051] The process of speech-to-text transcription not only involves the generation of text content but also must ensure the accurate alignment of text and audio. Therefore, after the large language model completes text clause segmentation, it is necessary to backtrack the character-level timestamps according to the position of each clause. That is, after clause segmentation is completed, according to the position of the clause marker, backtrack to the original character-level timestamp to determine the audio time interval corresponding to each clause. Specifically, each clause contains multiple characters in the initial text, and each character corresponds to a timestamp. In this process, by aligning the timestamp information corresponding to the characters in the clause with the text content, the start and end time intervals of the audio for each clause can be accurately determined. For example, by looking up the character timestamps before and after the punctuation mark ".", the start and end times of each clause can be accurately determined. The final output is structured data containing the clause text and its corresponding audio time interval, such as [("I confirm the insurance", 0.5s - 1.8s), ("ID number 1234", 2.0s - 3.5s)]. This process ensures the precise alignment of the speech content with the text clauses, providing an accurate time segment division basis for subsequent voiceprint feature extraction. The entire implementation process is completely based on the organic combination of acoustic features and semantic understanding, retaining both the temporal characteristics of the speech and integrating the deep semantic understanding ability of the language model.

[0052] Finally, after the processing of the above steps, the system will output the clause text and the set of their corresponding audio time intervals. Each clause not only contains its semantic information but also corresponds to its specific duration in the audio.

[0053] Furthermore, the steps of S122 include:

[0054] S1221. Identify the boundaries of semantically complete clauses in the initial text through the language model, and insert clause separators in the spoken text lacking punctuation marks.

[0055] S1222. Strengthen clause separation for the natural pauses after professional terms through the language model.

[0056] Since spoken text usually lacks punctuation marks and there may be many natural pauses in the sentences, the text directly generated by the ASR model often cannot accurately reflect the grammatical structure of the sentences. Therefore, it is necessary to input the initial text into a pre-trained large language model for processing. Through the understanding of the context, the large language model can automatically identify the semantically complete clauses in the text and determine reasonable sentence-breaking positions. In this process, the language model not only breaks sentences according to grammatical rules, but also comprehensively considers the semantic consistency to ensure natural logical connections between clauses. Specifically, through its internal multi-layer attention mechanism, the language model identifies the semantically complete units in the text and can accurately judge the clause boundaries even in colloquial expressions lacking punctuation marks.

[0057] In practical applications, spoken text often contains some professional terms or special nouns, and there are usually natural pauses after these terms. The pauses in pronunciation often mean the segmentation of sentences. By identifying these pause points and the logical segmentation after professional terms, the large language model can strengthen clause separation at these positions to ensure the clarity of the text in terms of semantics and logic. For example, the model will identify the business logic transition between "insure" and "ID number", and insert a clause separator here to generate a standardized text like "I confirm to insure. ID number: 1234". This process particularly considers the common punctuation-free continuous discourse scenarios in the financial field to ensure that the clause separation result not only conforms to grammatical norms but also retains the original semantics. This way of determining clause boundaries through the semantic understanding of the language model can significantly improve the readability and logic of the text. The resulting text not only conforms to grammatical rules but also can accurately reflect the pauses and key points in spoken expressions. Especially when dealing with colloquial text lacking punctuation marks, it can effectively restore the accurate clause structure and improve the performance of the voiceprint recognition system.

[0058] In one embodiment, the steps of S130 include:

[0059] S131. Extract the frame-level probability values of all characters in the interval according to the audio time interval corresponding to the clause.

[0060] S132. Calculate the geometric mean of the probability values as the clarity score:

[0061]

[0062] wherein, the S iis the speech clarity score for the i-th clause, P t is the probability of the character corresponding to the t-th frame, T is the total number of frames corresponding to the clause, and t is the frame number;

[0063] S133. Mark the clauses with clarity scores lower than the preset threshold as low-quality segments, and output the set of clarity scores for each clause.

[0064] For the audio time interval corresponding to each clause, extract the frame-level probability values of all characters within that period from the character-level probability sequence. When performing speaker recognition, it is necessary to extract the frame-level probability values of all characters within the interval according to the audio time interval of each clause output by the automatic speech recognition model. Specifically, the character-level probability sequence generated by the automatic speech recognition model provides the probability distribution of the character corresponding to each time frame, and these probability distributions reflect the possibility of each character appearing within that time frame. For example, if a clause contains n = 50 frames of audio, then extract the maximum character probability values {P1, P2, …, P 50}, where each P t represents the probability of the most likely character recognized in the t-th frame, with a value range of 0 to 1.

[0065] After obtaining all the frame-level probability values corresponding to the clause, next, process these probability values according to the preset clarity scoring formula. Specifically, the geometric mean is used to calculate the clarity score of the clause. The formula is as follows:

[0066]

[0067] where the S i is the speech clarity score for the i-th clause; Pt is the probability of the character corresponding to the t-th frame; T is the total number of frames corresponding to the clause; t is the frame number. The geometric mean is calculated by taking the n-th root of the product of the probability values of all frames, which can effectively balance the influence of different frames and avoid excessive influence of certain extreme values on the clarity score. The geometric mean is more sensitive than the arithmetic mean in reflecting the influence of low-probability frames on the overall quality. Therefore, it can comprehensively evaluate the speech clarity of the clause and ensure that the contribution of each time frame is considered evenly.

[0068] After calculating the clarity score for each clause, next, perform quality assessment. If the clarity score of a certain clause is lower than the preset threshold, the clause will be marked as a low-quality segment. This marking process is based on a configurable preset clarity threshold θ, for example, the clarity threshold θ = 0.6. If S i<θ, mark this clause as a low-quality segment, such as a passage with background noise or unclear pronunciation, but still retain its voiceprint features to participate in subsequent weighted fusion. This threshold can be set according to the actual application scenario. For example, in a financial remote identity authentication scenario with a relatively severe noise environment, a lower threshold may need to be set to avoid interference from low-quality voice segments. The purpose of marking low-quality segments is to ensure that the finally output voiceprint features only include clear and reliable audio segments, thereby improving the accuracy of voiceprint recognition. Finally, output the score set {S1, S2, …} of all clauses for voiceprint feature weighting.

[0069] In the solution of this embodiment, by quantitatively evaluating the clarity of audio clauses and calculating the clarity score in combination with the geometric mean, the impact of various factors such as noise interference, pronunciation problems, and channel impairments on the voice quality can be effectively handled. By marking low-quality segments and outputting the clarity score set, the system can dynamically adjust the weighted fusion of voiceprint features, enhance the contribution of high-quality segments to the recognition result, and thus significantly improve the performance and accuracy of the voiceprint recognition system in practical applications.

[0070] Furthermore, the steps of S140 include:

[0071] S141. Cut the corresponding voice segments from the input audio according to the audio time intervals of each clause;

[0072] S142. Extract frame-level voiceprint features through a preset separable convolutional layer of the ECAPA-TDNN model;

[0073] S143. Generate a clause-level voiceprint embedding vector h using an attention statistical pooling layer i ;

[0074] S144. Calculate the fusion weights according to the clarity score set:

[0075]

[0076] where τ is the temperature coefficient, K is the total number of clauses, and α i is the weight value of a certain sentence in the global context;

[0077] S145. Perform weighted summation on all the clause-level voiceprint embedding vectors h i :

[0078]

[0079] where e is the final voiceprint vector.

[0080] The system first precisely cuts out the corresponding speech segments from the original input audio according to the audio time intervals of each clause. When cutting, boundary alignment with a preset precision, such as 10 ms, is adopted to ensure that no valid speech frames are lost and at the same time avoid introducing interference from adjacent clauses.

[0081] Then, the preset 1D depthwise separable convolutional layer of the ECAPA-TDNN model is used to extract frame-level speaker features from the audio segments. The so-called ECAPA-TDNN model is a neural network architecture for speaker verification. It is based on the time-delay neural network and has been improved in many aspects to improve the system performance and accuracy. These frame-level features can provide in-depth information about each time frame in the speech signal and lay the foundation for the generation of further clause-level speaker embeddings.

[0082] Subsequently, an attention statistical pooling layer is used to aggregate the frame-level features: by calculating the weights of each frame through the attention mechanism, the mean and variance statistics are obtained by weighted averaging, and finally a 256-dimensional clause-level speaker embedding vector h is generated. i This pooling method can automatically focus on the stable phoneme segments in the speech, such as vowels, and suppress the influence of interference segments such as coughs. After obtaining the speaker embedding vector h of each clause, i the next step is to calculate the weight of each clause according to the speech clarity score S of the clause. i The weight coefficient α i is defined as the proportion of the clarity score of each clause to the sum of the clarity scores of all clauses. The specific formula is as follows:

[0083]

[0084] where τ is the temperature coefficient, which is used to control the steepness of the weight distribution, is an empirical value and is an adjustable constant; K is the total number of clauses; and α i is the weight value of a certain sentence in the global context.

[0085] After completing the weighted calculation of each clause, all clause-level speaker embedding vectors h i will be weighted and summed according to their corresponding weight coefficients α i to generate a global final speaker feature vector e. The formula is as follows:

[0086]

[0087] where e is the final speaker vector, which reflects the speaker information of all clauses in the entire audio and dynamically adjusts the feature fusion method according to the speech clarity of each clause.

[0088] The method steps of this embodiment effectively improve the contribution of high-quality speech segments through a dynamic weighting strategy, reduce the interference of low-quality segments on the recognition result, and thus improve the overall performance of the voiceprint recognition system.

[0089] In one embodiment, before S150, the steps further include:

[0090] S1501. Perform hash encryption processing on the final voiceprint feature to obtain an encrypted feature;

[0091] S1502. Transmit the encrypted feature to the voiceprint database through a secure transmission protocol.

[0092] During the identity recognition process, in order to protect the privacy of users, the final voiceprint feature vector e is first subjected to hash encryption processing before being compared with the voiceprint database. Hash encryption is a method of converting the original data into an irreversible fixed-length hash value, so as to ensure that even if the data is leaked, the original voiceprint feature cannot be restored. The process of hash encryption can be implemented through common encryption algorithms, such as the SHA-256 hash algorithm. Specifically, the operation of hash encryption is as follows:

[0093] h(e) = HashFunction(e)

[0094] Among them, h(e) is the encrypted hash value; e is the final voiceprint feature vector. The length of this hash value is fixed and determined by the selected hash algorithm.

[0095] The encrypted voiceprint feature h(e) will be sent to the voiceprint database through a secure transmission protocol. To ensure that the data is not stolen or tampered with during the transmission process, common security protocols, such as TLS or SSL, can be used to ensure the confidentiality and integrity of the data during network transmission.

[0096] The hashed feature h(e) encrypted and transmitted to the database is compared with the registered voiceprint features stored in the voiceprint database. During the voiceprint comparison process, the voiceprint features in the database also need to be hashed to ensure a secure match in the encrypted form. Due to the irreversibility of the hash function, the system cannot directly recover the original voiceprint feature from the hash value, ensuring the privacy of the data. If the hash values match, it means that the identity verification is successful, and the system can continue with corresponding operations, such as access permission authorization, transaction confirmation, etc.; if they do not match, it means that the identity verification fails, and the system will reject the request.

[0097] Figure 8 It is a schematic block diagram of a voiceprint recognition device 600 based on a large model provided by an embodiment of the present invention. As Figure 8As shown in the figure, corresponding to the above large model-based voiceprint recognition method, the present invention also provides a large model-based voiceprint recognition device 600. The large model-based voiceprint recognition device 600 includes units for executing the above large model-based voiceprint recognition method, and the device can be configured in terminals such as desktop computers, tablet computers, smart phones, etc.

[0098] Specifically, please refer to Figure 8 , the large model-based voiceprint recognition device 600 includes:

[0099] An automatic recognition unit 610, configured to process the input audio through an automatic speech recognition model using the Transformer architecture, and output a character-level probability sequence with timestamps;

[0100] A semantic clause segmentation unit 620, configured to convert the character-level probability sequence into an initial text, input it into a pre-trained language model for semantic clause segmentation, and obtain texts of multiple clauses and corresponding audio time intervals;

[0101] A clarity scoring unit 630, configured to calculate the speech clarity score of each clause according to the character-level probability sequence through a preset clarity scoring formula;

[0102] A weighted fusion unit 640, configured to determine the audio interval of each clause according to the audio time interval, extract the voiceprint feature vector in the audio interval of the clause, and perform weighted fusion on the voiceprint feature vector according to the speech clarity score to generate a final voiceprint feature;

[0103] An identity recognition unit 650, configured to compare the final voiceprint feature with the voiceprint database to complete identity recognition.

[0104] In one embodiment, the automatic recognition unit 610 includes:

[0105] A feature sequence acquisition unit, configured to perform frame splitting and Mel feature extraction on the input audio to obtain a feature sequence;

[0106] A frame-level character probability distribution acquisition model unit, configured to input the feature sequence into a pre-trained automatic speech recognition model, and obtain the frame-level character probability distribution output by the model. The model is trained using a joint loss function of connectionist temporal classification and cross entropy, and is optimized and trained through background noise and time-frequency mask data;

[0107] A character-level probability sequence generation unit, configured to generate the start and end timestamps of each character through connectionist temporal classification decoding, and form the character-level probability sequence with timestamps in combination with the frame-level character probability distribution.

[0108] In one embodiment, the semantic clause segmentation unit 620 includes:

[0109] An initial text generation unit for performing connectionist temporal classification beam search decoding on the character-level probability sequence to generate an initial text;

[0110] An initial text semantic clause segmentation unit for inputting the initial text into the pre-trained language model for semantic clause segmentation;

[0111] An audio start and end time determination unit for backtracking the character-level timestamps according to the clause positions to determine the audio start and end times corresponding to each clause;

[0112] A text and audio time interval set generation unit for outputting the clause text and its associated audio time interval set.

[0113] Further, the initial text semantic clause segmentation unit includes:

[0114] A clause separator insertion unit for identifying the boundaries of semantically complete sub-clauses in the initial text through the language model and inserting clause separators in the oral text lacking punctuation;

[0115] A reinforcement clause segmentation unit for strengthening clause segmentation of the natural pauses after professional terms through the language model.

[0116] In one embodiment, the clarity scoring unit 630 includes:

[0117] A frame-level probability value extraction unit for extracting the frame-level probability values of all characters in the interval according to the audio time interval corresponding to the clause;

[0118] A clarity score acquisition unit for calculating the geometric mean of the probability values as the clarity score:

[0119]

[0120] wherein, the S i is the speech clarity score of the i-th clause, P t is the probability of the character corresponding to the t-th frame, T is the total number of frames corresponding to the clause, and t is the frame number;

[0121] A low-quality segment marking unit for marking the clauses with clarity scores lower than the preset threshold as low-quality segments and outputting the clarity score set of each clause.

[0122] Further, the weighted fusion unit 640 includes:

[0123] A speech segment cutting unit for cutting the corresponding speech segments from the input audio according to the audio time intervals of each clause;

[0124] A frame-level voiceprint feature extraction unit, configured to extract frame-level voiceprint features through a preset separable convolutional layer of an ECAPA-TDNN model;

[0125] A voiceprint embedding vector generation unit, configured to generate a clause-level voiceprint embedding vector h by using an attention statistical pooling layer i ;

[0126] A fusion weight calculation unit, configured to calculate fusion weights according to the set of clarity scores:

[0127]

[0128] where τ is a temperature coefficient, K is the total number of clauses, and α i is the weight value of a certain sentence in the global context;

[0129] A final voiceprint vector acquisition unit, configured to perform weighted summation on all the clause-level voiceprint embedding vectors h i :

[0130]

[0131] where e is the final voiceprint vector.

[0132] In one embodiment, between the identity recognition unit 650 and the weighted fusion unit 640, there is further provided:

[0133] A hash encryption unit, configured to perform hash encryption processing on the final voiceprint features to obtain encrypted features;

[0134] A secure transmission unit, configured to transmit the encrypted features to the voiceprint database through a secure transmission protocol.

[0135] The above voiceprint recognition device 600 based on a large model can be implemented in the form of a computer program, and this computer program can run on a computer device as shown in Figure 9 .

[0136] Please refer to Figure 9 , Figure 9 which is a schematic block diagram of a computer device provided by an embodiment of the present application. The computer device 500 can be a terminal or a server. Among them, the terminal can be an electronic device with communication functions such as a desktop computer, a tablet computer, a smart phone, etc. The server can be an independent server or a server cluster composed of multiple servers.

[0137] Refer to Figure 9, the computer device 500 includes a processor 502, a memory, and a network interface 505 connected via a system bus 501. Among them, the memory may include a non-volatile storage medium 503 and an internal memory 504.

[0138] The non-volatile storage medium 503 can store an operating system 5031 and a computer program 5032. The computer program 5032 includes program instructions that, when executed, enable the processor 502 to execute a large model-based voiceprint recognition method.

[0139] The processor 502 is used to provide computing and control capabilities to support the operation of the entire computer device 500.

[0140] The internal memory 504 provides an environment for the operation of the computer program 5032 in the non-volatile storage medium 503. When the computer program 5032 is executed by the processor 502, it enables the processor 502 to execute a large model-based voiceprint recognition method.

[0141] The network interface 505 is used for network communication with other devices. Those skilled in the art can understand that Figure 9 the structure shown in is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device 500 to which the solution of this application is applied. The specific computer device 500 may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0142] Among them, the processor 502 is used to run the computer program 5032 stored in the memory to implement the steps of the above method.

[0143] It should be understood that in the embodiments of this application, the processor 502 may be a central processing unit (CPU), and the processor 502 may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), off-the-shelf programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0144] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above-described embodiments can be completed by instructing relevant hardware through a computer program. The computer program includes program instructions, and the computer program can be stored in a storage medium, which is a computer-readable storage medium. The program instructions are executed by at least one processor in the computer system to implement the flow steps of the embodiments of the above method.

[0145] Therefore, the present invention also provides a storage medium. The storage medium can be a computer-readable storage medium. The storage medium stores a computer program, where the computer program includes program instructions. When the program instructions are executed by a processor, the processor is caused to execute the steps of the above method.

[0146] The storage medium can be a variety of computer-readable storage media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a magnetic disk, or an optical disc, etc., that can store program codes.

[0147] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0148] In several embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of each unit is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed.

[0149] The steps in the method embodiments of the present invention can be adjusted, combined, and deleted according to actual needs. The units in the device embodiments of the present invention can be combined, divided, and deleted according to actual needs. In addition, the functional units in each embodiment of the present invention can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.

[0150] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a terminal, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention.

[0151] As described above, the above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.

Claims

1. A voiceprint recognition method based on a large model, characterized in that Including: Processing the input audio by using an automatic speech recognition model with a Transformer architecture to output a timestamped character-level probability sequence; Converting the character-level probability sequence into an initial text, inputting it into a pre-trained language model for semantic clause segmentation to obtain texts of multiple clauses and corresponding audio time intervals; Calculating the speech clarity score of each clause according to the character-level probability sequence through a preset clarity scoring formula; Determining the audio interval of each clause according to the audio time interval, extracting the voiceprint feature vector in the audio interval of the clause, and performing weighted fusion on the voiceprint feature vector according to the speech clarity score to generate the final voiceprint feature; Comparing the final voiceprint feature with the voiceprint database to complete identity recognition.

2. The voiceprint recognition method based on a large model according to claim 1, wherein The step of processing the input audio by using an automatic speech recognition model with a Transformer architecture to output a timestamped character-level probability sequence includes: Performing frame segmentation and Mel feature extraction on the input audio to obtain a feature sequence; Inputting the feature sequence into a pre-trained automatic speech recognition model to obtain the frame-level character probability distribution output by the model, and the model is trained with a joint loss function of connectionist temporal classification and cross-entropy and optimized and trained with background noise and time-frequency mask data; Generating the start and end timestamps of each character through connectionist temporal classification decoding, and combining the frame-level character probability distribution to form the timestamped character-level probability sequence.

3. The voiceprint recognition method based on a large model according to claim 2, wherein The step of converting the character-level probability sequence into an initial text, inputting it into a pre-trained language model for semantic clause segmentation to obtain multiple clauses and corresponding audio time intervals includes: Performing connectionist temporal classification beam search decoding on the character-level probability sequence to generate an initial text; Inputting the initial text into the pre-trained language model for semantic clause segmentation; Retrospecting the character-level timestamps according to the clause positions to determine the start and end audio times corresponding to each clause; Outputting the clause text and its associated set of audio time intervals.

4. The voiceprint recognition method based on a large model according to claim 3, wherein The step of inputting the initial text into the pre-trained language model for semantic clause segmentation includes: Identifying the boundaries of semantically complete sub-clauses in the initial text through the language model, and inserting clause delimiters in the spoken text lacking punctuation; Strengthening clause segmentation of the natural pauses after professional terms through the language model.

5. The voiceprint recognition method based on a large model according to claim 1, wherein The step of calculating the speech clarity score of each clause according to the character-level probability sequence through a preset clarity scoring formula includes: Extracting the frame-level probability values of all characters in the interval according to the audio time interval corresponding to the clause; Calculating the geometric mean of the probability values as the clarity score: Among them, the S i is the speech clarity score of the i-th clause, and P t is the probability of the character corresponding to the t-th frame, T is the total number of frames corresponding to this clause, and t is the frame number; Marking the clauses with clarity scores lower than a preset threshold as low-quality segments, and outputting the set of clarity scores of each clause.

6. The voiceprint recognition method based on a large model according to claim 5, wherein The step of determining the audio interval of each clause according to the audio time interval, extracting the voiceprint feature vector in the audio interval of the clause, and performing weighted fusion on the voiceprint feature vector according to the speech clarity score to generate the final voiceprint feature includes: Cutting the corresponding speech segments from the input audio according to the audio time interval of each clause; Extract frame-level voiceprint features through the preset separable convolutional layer of the ECAPA-TDNN model; Generate a clause-level voiceprint embedding vector h using the attention statistical pooling layer i ; Calculate the fusion weights according to the clarity score set: wherein, τ is the temperature coefficient, K is the total number of clauses, and α i is the weight value of a certain sentence in the whole context; For all the clause-level acoustic feature embedding vectors h i perform weighted summation: Wherein, the e is the final voiceprint vector.

7. The voiceprint recognition method based on a large model according to any one of claims 1 to 6, characterized in that Before comparing the final voiceprint feature with the voiceprint database to complete identity recognition, it further includes: Perform hash encryption processing on the final voiceprint feature to obtain an encrypted feature; Transmit the encrypted feature to the voiceprint database through a secure transmission protocol.

8. An apparatus for voiceprint recognition based on a large model, characterized in that, For executing the voiceprint recognition method based on a large model according to any one of claims 1 to 7.

9. A computer device, characterized in that, The computer device includes a memory and a processor connected to the memory; the memory is used to store a computer program; the processor is used to run the computer program stored in the memory to execute the steps of the method according to any one of claims 1 to 6.

10. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, and the computer program includes program instructions, and the program instructions can implement the steps of the method according to any one of claims 1 to 6 when executed by a processor.