A voiceprint identification and comparison recommendation method, device, electronic device and storage medium

New audio is obtained through speech recognition and recommendation algorithms, and the recommendation index is calculated in combination with text and phoneme matching, which solves the problem of insufficient manual screening in the prior art, and achieves efficient and accurate voiceprint identification, especially in identity and non-identity identification.

CN114648994BActive Publication Date: 2025-05-23XIAMEN KUAISHANGTONG TECH CORP LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210169791.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-23
Publication Date
2025-05-23
Estimated Expiration
2042-02-23

AI Technical Summary

Technical Problem

The existing voiceprint identification technology relies on manual screening and evaluation, which is time-consuming and easy to miss important information. The existing recommended solutions are insufficiently accurate and mainly rely on formant-related information calculations, and fail to effectively consider identity and non-identity identification.

Method used

New audio is obtained through speech recognition and recommendation algorithms, combined with self-built dictionary and the original dictionary to stutter participle, obtain the same phrases and single words in the text, and intercept the corresponding voice segments; at the same time, obtain the same triphones through phoneme matching, form a new voice segment, calculate the recommendation index of each voice segment, select the highest voice segment to form a new audio, and perform automatic voiceprint recognition to distinguish identity and non-identity recommendation.

Benefits of technology

It saves labor costs, improves identification accuracy, can effectively consider identity and non-identity identification, reduces redundancy and error, and reduces calculation costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114648994B_ABST
    Figure CN114648994B_ABST
Patent Text Reader

Abstract

The invention discloses a voiceprint identification comparison and recommendation method, comprising: obtaining a sample voice and a test material voice; identifying the text and corresponding phonemes of the sample voice, as well as the text and corresponding phonemes of the test material voice; performing stuttering word segmentation based on a dictionary according to the identified texts of the sample voice and the test material voice, obtaining the same phrases and single words in the texts of the sample voice and the test material voice, and forming a first voice segment; obtaining the same three phonemes of the sample voice and the test material voice according to the identified phonemes of the sample voice and the test material voice, and forming a second voice segment; for the first voice segment and the second voice segment, calculating the recommendation index of each voice segment through a recommendation algorithm; selecting the voice segment with the highest recommendation index to form a new audio; performing automatic voiceprint identification on the new audio to distinguish between identity recommendation and non-identity recommendation. The invention obtains a new audio through a recommendation algorithm, and then performs voiceprint identification, taking identity and non-identity identification into consideration, thereby saving labor costs and having high accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of voiceprint identification, and in particular to a voiceprint identification comparison recommendation method, device, electronic equipment and storage medium. Background Art

[0002] The application scenarios of voiceprint identification technology are increasing, and the requirements for related technical indicators are becoming higher and higher. At present, voiceprint identification technology generally involves manual screening of samples and then searching for identification evidence in the screened samples and samples. However, it takes a lot of time to manually select these identification evidence. With the improvement of speech recognition and automatic voiceprint recognition technology, these technologies are gradually applied to the screening and evaluation process of identification evidence.

[0003] Existing screening and assessment methods have certain shortcomings:

[0004] First, manual screening and evaluation rely on the experience of the identification personnel. When the inspection materials and sample voices are long, it will consume a lot of energy and time of the identification personnel.

[0005] Secondly, with the help of voiceprint recognition and automatic voiceprint identification methods, some identification evidence can be quickly selected. However, the existing identification evidence recommendation scheme mainly relies on the calculation of resonance peak related information, which can easily miss some important information.

[0006] Finally, in the recommended scheme, accuracy is evaluated and usually only identity identification is considered. Summary of the invention

[0007] The main purpose of the present invention is to overcome the above-mentioned defects in the prior art and propose a voiceprint identification and comparison recommendation method. First, new audio is obtained through speech recognition and recommendation algorithm, and then voiceprint identification is performed on the new audio. In addition, both identity identification and non-identity identification are considered in voiceprint identification, which saves labor costs and has high accuracy.

[0008] The present invention adopts the following technical solution:

[0009] A recommended method for voiceprint identification and comparison includes:

[0010] Obtain sample voice and inspection material voice;

[0011] Inputting the sample speech and the sample speech into the speech recognition system, the speech recognition system recognizes the text and corresponding phonemes of the sample speech, and the text and corresponding phonemes of the sample speech;

[0012] According to the recognized text of the sample speech and the test material speech, stuttering word segmentation is performed based on the self-built dictionary and the original dictionary, the same phrases and single words in the text of the sample speech and the test material speech are obtained, the corresponding starting and ending positions of the phrases and single words are recorded, and the corresponding speech segments are intercepted to form the first speech segment;

[0013] According to the recognized phonemes of the sample voice and the sample voice, three phonemes that are the same as the sample voice and the sample voice are obtained, and the corresponding starting and ending positions are recorded, and the corresponding voice segments are intercepted to form a second voice segment;

[0014] For the first speech segment and the second speech segment, a recommendation index of each speech segment is calculated by using a recommendation algorithm;

[0015] Select the first speech segment with the highest recommendation index and the second speech segment with the highest recommendation index to form a new audio;

[0016] Distinguish between identical recommendations and non-identical recommendations by automatically performing voiceprint recognition on new audio.

[0017] Specifically, the speech recognition system recognizes the text and corresponding phonemes of the sample speech, as well as the text and corresponding phonemes of the test material speech, specifically:

[0018] Acoustic features are obtained from speech; including but not limited to linear predictive coding and Mel-frequency cepstral coefficients;

[0019] The LSTM+CTC neural network acoustic model is used to convert acoustic features into phonemes;

[0020] Language models based on deep neural networks convert phonemes into phrases and words.

[0021] Specifically, for the first speech segment and the second speech segment, a recommendation index of each speech segment is calculated through a recommendation algorithm, and the recommendation index is based on recommendation indicators, which include: speech content consistency, spectrum clarity, resonance peak number index, context consistency, and speech speed consistency.

[0022] Specifically, the recommended indicators are:

[0023] Speech content consistency: Calculate the spectral cosine similarity of speech segments;

[0024] Spectral clarity: Calculate the standard deviation of the harmonic energy in the peak part of the broadband spectrogram of the speech segment;

[0025] Formant number index: count the number of formants in the speech segment and convert it into an index;

[0026] Context consistency: Calculate the context consistency of the phonemes in the speech segment, that is, whether the previous and following consonants / silences of the vowel phonemes are the same, and obtain the corresponding index;

[0027] Speech rate consistency: Calculates the duration index of speech segments.

[0028] Another embodiment of the present invention provides a voiceprint identification, comparison and recommendation device, including:

[0029] Acquisition unit: acquires sample voice and inspection material voice;

[0030] Speech recognition unit: the sample speech and the test material speech are input into the speech recognition system, and the speech recognition system recognizes the text and corresponding phonemes of the sample speech, as well as the text and corresponding phonemes of the test material speech;

[0031] The first speech segment acquisition unit: according to the recognized text of the sample speech and the test material speech, the self-built dictionary and the original dictionary are used to perform word segmentation, obtain the same phrases and single words in the text of the sample speech and the test material speech, record the corresponding starting and ending positions of the phrases and single words, and intercept the corresponding speech segments to form the first speech segment;

[0032] The second speech segment acquisition unit: according to the recognized phonemes of the sample speech and the test material speech, acquires the three phonemes that are the same in the sample speech and the test material speech, records the corresponding starting and ending positions, and intercepts the corresponding speech segment to form the second speech segment;

[0033] Recommendation index calculation unit: for the first speech segment and the second speech segment, calculates the recommendation index of each speech segment through a recommendation algorithm;

[0034] New audio acquisition unit: selects the first speech segment with the highest recommendation index and the second speech segment with the highest recommendation index to form a new audio;

[0035] Recommendation unit: Distinguish between identical recommendations and non-identical recommendations by automatically performing voiceprint recognition on new audio.

[0036] Specifically, the speech recognition system recognizes the text and corresponding phonemes of the sample speech, as well as the text and corresponding phonemes of the test material speech, specifically:

[0037] Acoustic features are obtained from speech, including but not limited to linear predictive coding and Mel-frequency cepstral coefficients;

[0038] The LSTM+CTC neural network acoustic model is used to convert acoustic features into phonemes;

[0039] Language models based on deep neural networks convert phonemes into phrases and words.

[0040] Specifically, in the recommendation index calculation unit, for the first speech segment and the second speech segment, the recommendation index of each speech segment is calculated through a recommendation algorithm, and the recommendation index is based on recommendation indicators, which include: speech content consistency, spectrum clarity, resonance peak number index, context consistency, and speech speed consistency.

[0041] Specifically, the recommended indicators are:

[0042] Speech content consistency: Calculate the spectral cosine similarity of speech segments;

[0043] Spectral clarity: Calculate the standard deviation of the harmonic energy in the peak part of the broadband spectrogram of the speech segment;

[0044] Formant number index: count the number of formants in the speech segment and convert it into an index;

[0045] Context consistency: Calculate the context consistency of the phonemes in the speech segment, that is, whether the previous and following consonants / silences of the vowel phonemes are the same, and obtain the corresponding index;

[0046] Speech rate consistency: Calculates the duration index of speech segments.

[0047] Yet another embodiment of the present invention provides an electronic device, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above-mentioned steps of a voiceprint identification and comparison recommendation method when executing the computer program.

[0048] Yet another embodiment of the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned voiceprint identification and comparison recommendation method are implemented.

[0049] It can be seen from the above description of the present invention that, compared with the prior art, the present invention has the following beneficial effects:

[0050] (1) The present invention provides a voiceprint identification comparison recommendation method, comprising: obtaining a sample voice and a test material voice; inputting the sample voice and the test material voice into a voice recognition system, the voice recognition system identifying the text and corresponding phonemes of the sample voice, as well as the text and corresponding phonemes of the test material voice; performing stuttering word segmentation based on the recognized text of the sample voice and the test material voice based on a self-built dictionary and an original dictionary, obtaining the same phrases and single words in the text of the sample voice and the test material voice, recording the corresponding starting and ending positions of the phrases and single words, and intercepting the corresponding voice segments to form a first voice segment; obtaining the corresponding phonemes of the sample voice and the test material voice based on the recognized sample voice and the test material voice The same three phonemes are identified, and the corresponding starting and ending positions are recorded, and the corresponding speech segments are intercepted to form the second speech segment; for the first speech segment and the second speech segment, the recommendation index of each speech segment is calculated through the recommendation algorithm; the speech segment with the highest recommendation index of the first speech segment and the speech segment with the highest recommendation index of the second speech segment are selected to form a new audio; the new audio is automatically distinguished by voiceprint recognition to distinguish the identity recommendation and the non-identity recommendation; the method provided by the present invention first obtains the new audio through speech recognition and the recommendation algorithm, and then performs voiceprint identification on the new audio, and the identity identification and the non-identity identification are considered in the voiceprint identification, which saves labor costs and has high accuracy.

[0051] (2) In the method provided by the present invention, the same phrases and words in the text of the sample speech and the test material speech are obtained to form a first speech segment, and the same three phonemes in the sample speech and the test material speech are obtained to form a second speech segment; after obtaining the key speech segments, identification is performed to reduce redundancy and errors, reduce computing costs and improve accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] Figure 1 A flowchart of a voiceprint identification and comparison recommendation method provided by an embodiment of the present invention;

[0053] Figure 2 Another flow chart of a voiceprint identification and comparison recommendation method provided by an embodiment of the present invention;

[0054] Figure 3 A system architecture diagram of a voiceprint identification, comparison and recommendation device provided by an embodiment of the present invention;

[0055] Figure 4 A schematic diagram of an electronic device provided by an embodiment of the present invention;

[0056] Figure 5 A schematic diagram of an embodiment of a computer-readable storage medium provided in an embodiment of the present invention.

[0057] The present invention is further described in detail below in conjunction with the accompanying drawings and specific embodiments. DETAILED DESCRIPTION

[0058] The present invention proposes a voiceprint identification comparison and recommendation method, which first obtains new audio through speech recognition and recommendation algorithm, and then performs voiceprint identification on the new audio. In addition, the voiceprint identification takes into account both identity identification and non-identity identification, which saves labor costs and has high accuracy.

[0059] It should be noted that, in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply that there is any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of more restrictions, the elements defined by the sentence "including one..." do not exclude the existence of other identical elements in the process, method, article or device including the elements. The above is only a specific implementation of the present application, so that those skilled in the art can understand or implement the present application. The various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features applied for herein.

[0060] The method provided in the present application can be applied to various scenarios that require identity authentication through voice data recognition, for example, it can be applied to criminal investigation scenarios, payment scenarios, security verification scenarios or password scenarios, etc. Taking the application in the security verification scenario as an example, when logging into the application, the user can use the voiceprint identification method provided in the present application to determine whether the user currently logging into the application is a user with login authority, wherein only when the voice data of the user currently logging into the application successfully matches the voice data of the user with login authority, can the login be successful; or when the user unlocks the door with voice, the voiceprint identification method provided in the present application is used to determine whether the user currently logging into the application is a user with unlocking authority, wherein the user with unlocking authority can include one user or multiple users. It should be understood that the above embodiments are only exemplary descriptions, and the specific application scope and application method of the present application can be flexibly adjusted according to user needs, and are not limited to those given in the above embodiments.

[0061] The following is a description of a voiceprint identification method provided by the embodiment of the present application in combination with a number of specific application examples. The execution subject of the method can be a terminal device, such as a mobile phone, a computer, a wearable device, etc., or a server, which is not limited here. Figure 1-2 , which is a flow chart of a voiceprint identification and comparison recommendation method provided by an embodiment of the present invention, specifically comprising:

[0062] S101: Acquire sample voice and test material voice;

[0063] For example, in some possible embodiments, the voice data to be identified can be input by the user through a terminal device, wherein the terminal device can be any smart terminal device with a voiceprint identification function, such as a mobile phone, a camera with a recording function, a wearable device, etc. For example, in a payment scenario, the voice data to be identified can be input by the user through a mobile phone. In a criminal investigation scenario, the voice data to be identified can be obtained by a computer after the video to be analyzed is separated from the video and uploaded through the computer. The specific method of obtaining the voice data to be identified can be flexibly adjusted according to user needs and is not limited to the above embodiments.

[0064] S102: Inputting the sample speech and the sample speech into a speech recognition system, and the speech recognition system recognizes the text and corresponding phonemes of the sample speech, and the text and corresponding phonemes of the sample speech;

[0065] Specifically, the speech recognition system recognizes the text and corresponding phonemes of the sample speech, as well as the text and corresponding phonemes of the test material speech, specifically:

[0066] Acoustic features are obtained from speech, including but not limited to linear predictive coding and Mel-frequency cepstral coefficients;

[0067] The LSTM+CTC neural network acoustic model is used to convert acoustic features into phonemes;

[0068] Language models based on deep neural networks convert phonemes into phrases and words.

[0069] S103: performing stuttering word segmentation based on the dictionary according to the recognized text of the sample speech and the sample speech, obtaining the same phrases and single words in the text of the sample speech and the sample speech, recording the corresponding starting and ending positions of the phrases and single words, and intercepting the corresponding speech segments to form a first speech segment;

[0070] The algorithms involved in Jieba Chinese word segmentation include efficient word graph scanning based on the Trie tree structure, generating a directed acyclic graph consisting of all possible word formation situations of Chinese characters in a sentence, using dynamic programming to find the maximum probability path, and finding the maximum segmentation combination based on word frequency. For unregistered words, the HMM model based on the word formation ability of Chinese characters is adopted, and the Viterbi algorithm is used. It is worth noting that some proper nouns will be separated due to word segmentation, resulting in recognition errors. Therefore, the dictionary here can be a self-built dictionary, and some proper nouns can be added to the dictionary. Although Jieba has the ability to recognize new words, adding new words by yourself can ensure a higher accuracy, especially for proper nouns.

[0071] S104: according to the recognized phonemes of the sample voice and the sample voice, obtain the three phonemes that are the same in the sample voice and the sample voice, record the corresponding starting and ending positions, and intercept the corresponding voice segment to form a second voice segment;

[0072] Triphone is a type of phoneme. Different from monophones (such as t, iy, n), triphone is expressed as t-iy+n, that is, it is composed of three monophones. It is similar to the monophone iy, but it takes into account the relationship of the context, that is, the previous context is t and the following context is n.

[0073] Both triphone and monophone are a hidden Markov model (HMM);

[0074] Triphones are used to consider contextual information (co-articulation); when extracting cepstrum features, the Hanning window contains redundant spectra to the left and right. Therefore, it is reasonable to replace monophones with triphones. After a monophone is copied to a triphone, the number of states increases exponentially, but the copied states. In order to solve the problem of data sparsity, the number of training required is huge, so the number of parameters needs to be reduced. Clustering is to reduce the number of all triphone parameters, that is, to reduce the number of triphone states. The role of the decision tree is to cluster the states of the triphone. After clustering, all triphone states are clustered into multiple clusters, and all triphone states in each cluster are bound (that is, multiple states share the parameters of one state).

[0075] S105: For the first speech segment and the second speech segment, calculating the recommendation index of each speech segment by using a recommendation algorithm;

[0076] Specifically, for the first speech segment and the second speech segment, a recommendation index of each speech segment is calculated through a recommendation algorithm, and the recommendation index is based on recommendation indicators, which include: speech content consistency, spectrum clarity, resonance peak number index, context consistency, and speech speed consistency.

[0077] Specifically, the recommended indicators are:

[0078] Speech content consistency: Calculate the spectral cosine similarity of speech segments;

[0079]

[0080] Among them, A is the n-dimensional spectrum vector of the sample speech segment in the first speech segment and the second speech segment, and B is the n-dimensional spectrum vector of the sample speech segment in the first speech segment and the second speech segment.

[0081] Spectral clarity: Calculate the standard deviation of the harmonic energy in the peak part of the broadband spectrogram of the speech segment;

[0082] Formant number index: Count the number of formants in the speech segment and convert it into an index; if the sample speech segment and the test material speech segment in the first speech segment and the second speech segment have two formants, the corresponding index is 0.8, if there are three formants, the corresponding index is 0.9, if there are four or more formants, the corresponding index is 1.0, if the index of the formant is 1 or there is no formant, the corresponding index is 0;

[0083] Context consistency: Calculate the context consistency of the phonemes in the speech segment, that is, whether the previous and following consonants / silences of the vowel phonemes are the same, and obtain the corresponding index;

[0084] Speech rate consistency: Calculates the duration index of speech segments.

[0085]

[0086] Among them, A duration is the speech duration of the sample speech segments in the first speech segment and the second speech segment, B duration It is the speech duration of the sample speech segment in the first speech segment and the second speech segment.

[0087] S106: Select the first speech segment with the highest recommendation index and the second speech segment with the highest recommendation index to form a new audio;

[0088] S107: Distinguish identical recommendations from non-identical recommendations by automatically performing voiceprint recognition on the new audio.

[0089] It is necessary to set in advance a preset range for judging whether the characteristic parameters are qualified. The preset range usually includes three units, namely, a preset frequency range, a preset energy range, and a preset peak sharpness difference range. The identity recommendation and the non-identity recommendation are distinguished based on whether the obtained characteristic parameters fall into the corresponding range.

[0090] like Figure 3 Another embodiment of the present invention provides a voiceprint identification, comparison and recommendation device, including:

[0091] Acquisition unit 301: Acquire sample speech and evidentiary speech;

[0092] Exemplarily, in some possible embodiments, the speech data to be identified may be input by a user through a terminal device. Among them, the terminal device may be any intelligent terminal device with a voiceprint identification function, such as a mobile phone, a camera with a recording function, a wearable device, etc. For example, in a payment scenario, the speech data to be identified may be input by the user through a mobile phone. In a criminal investigation scenario, the speech data to be identified may be separated from a video obtained by a computer and uploaded through the computer after analyzing the video. The specific acquisition method of the speech data to be identified can be flexibly adjusted according to the user's needs and is not limited to the above embodiments.

[0093] Speech recognition unit 302: Input the sample speech and evidentiary speech into a speech recognition system. The speech recognition system recognizes the text and corresponding phonemes of the sample speech, as well as the text and corresponding phonemes of the evidentiary speech;

[0094] Specifically, the speech recognition system recognizes the text and corresponding phonemes of the sample speech, as well as the text and corresponding phonemes of the evidentiary speech, specifically as follows:

[0095] Obtain acoustic features from the speech; including but not limited to, linear predictive coding and Mel frequency cepstral coefficients;

[0096] Use an LSTM+CTC neural network acoustic model to convert the acoustic features into phonemes;

[0097] Based on a deep neural network language model, convert the phonemes into phrases and single characters.

[0098] First speech segment acquisition unit 303: Based on the recognized text of the sample speech and evidentiary speech, perform jieba word segmentation based on a dictionary, obtain the same phrases and single characters in the text of the sample speech and evidentiary speech, record the corresponding start and end positions of the phrases and single characters, and intercept the corresponding speech segments to form the first speech segment;

[0099] The algorithms involved in jieba Chinese word segmentation include implementing efficient word graph scanning based on the Trie tree structure to generate a directed acyclic graph composed of all possible word formation situations of Chinese characters in a sentence, using dynamic programming to find the maximum probability path, finding the maximum segmentation combination based on word frequency, and for out-of-vocabulary words, using an HMM model based on the word formation ability of Chinese characters and the Viterbi algorithm; It should be noted that some proper nouns will be separated due to word segmentation, resulting in recognition errors. Therefore, the dictionary here can be a self-built dictionary, adding some proper nouns to the dictionary. Although jieba has the ability to recognize new words, adding new words by oneself can ensure higher accuracy, especially for proper nouns.

[0100] The second speech segment acquisition unit 304 acquires the three phonemes common to the sample speech and the test speech according to the recognized phonemes of the sample speech and the test speech, records the corresponding start and end positions, and intercepts the corresponding speech segment to form a second speech segment;

[0101] Triphone is a type of phoneme. Different from monophones (such as t, iy, n), triphone is expressed as t-iy+n, that is, it is composed of three monophones. It is similar to the monophone iy, but it takes into account the relationship of the context, that is, the previous context is t and the following context is n.

[0102] Both triphone and monophone are a hidden Markov model (HMM);

[0103] Triphones are used to consider contextual information (co-articulation); when extracting cepstrum features, the Hanning window contains redundant spectra to the left and right. Therefore, it is reasonable to replace monophones with triphones. After a monophone is copied to a triphone, the number of states increases exponentially, but the copied states. In order to solve the problem of data sparsity, the number of training required is huge, so the number of parameters needs to be reduced. Clustering is to reduce the number of all triphone parameters, that is, to reduce the number of triphone states. The role of the decision tree is to cluster the states of the triphone. After clustering, all triphone states are clustered into multiple clusters, and all triphone states in each cluster are bound (that is, multiple states share the parameters of one state).

[0104] Recommendation index calculation unit 305: for the first speech segment and the second speech segment, calculates the recommendation index of each speech segment by using a recommendation algorithm;

[0105] Specifically, for the first speech segment and the second speech segment, a recommendation index of each speech segment is calculated through a recommendation algorithm, and the recommendation index is based on recommendation indicators, which include: speech content consistency, spectrum clarity, resonance peak number index, context consistency, and speech speed consistency.

[0106] Specifically, the recommended indicators are:

[0107] Speech content consistency: Calculate the spectral cosine similarity of speech segments;

[0108]

[0109] Among them, A is the n-dimensional spectrum vector of the sample speech segment in the first speech segment and the second speech segment, and B is the n-dimensional spectrum vector of the sample speech segment in the first speech segment and the second speech segment.

[0110] Spectral clarity: Calculate the standard deviation of the harmonic energy in the peak part of the broadband spectrogram of the speech segment;

[0111] Formant number index: Count the number of formants in the speech segment and convert it into an index; if the sample speech segment and the test material speech segment in the first speech segment and the second speech segment have two formants, the corresponding index is 0.8; if there are three formants, the corresponding index is 0.9; if there are four or more formants, the corresponding index is 1.0; if the index of the formant is 1 or there is no formant, the corresponding index is 0;

[0112] Context consistency: Calculate the context consistency of the phonemes in the speech segment, that is, whether the previous and following consonants / silences of the vowel phonemes are the same, and obtain the corresponding index;

[0113] Speech rate consistency: Calculate the duration index of speech segments.

[0114]

[0115] Among them, A duration is the speech duration of the sample speech segments in the first speech segment and the second speech segment, B duration It is the speech duration of the sample speech segment in the first speech segment and the second speech segment.

[0116] New audio acquisition unit 306: selects the first speech segment with the highest recommendation index and the second speech segment with the highest recommendation index to form a new audio;

[0117] Recommendation unit 307: distinguishing between identical recommendations and non-identical recommendations by automatically performing voiceprint recognition on the new audio.

[0118] It is necessary to set in advance a preset range for judging whether the characteristic parameters are qualified. The preset range usually includes three units, namely, a preset frequency range, a preset energy range, and a preset peak sharpness difference range. The identity recommendation and the non-identity recommendation are distinguished based on whether the obtained characteristic parameters fall into the corresponding range.

[0119] Figure 4 As shown, an embodiment of the present invention provides an electronic device 400, including a memory 410, a processor 420, and a computer program 411 stored in the memory 420 and executable on the processor 420. When the processor 420 executes the computer program 411, a voiceprint identification and comparison recommendation method provided in an embodiment of the present invention is implemented.

[0120] Since the electronic device introduced in this embodiment is the device used to implement the embodiment of the present invention, based on the method introduced in the embodiment of the present invention, the technical personnel in this field can understand the specific implementation mode of the electronic device of this embodiment and its various variations. Therefore, how the electronic device implements the method in the embodiment of the present invention is not introduced in detail here. As long as the equipment used by the technical personnel in this field to implement the method in the embodiment of the present invention falls within the scope of protection of the present invention.

[0121] See also Figure 5 , Figure 5 A schematic diagram of an embodiment of a computer-readable storage medium provided in an embodiment of the present invention.

[0122] like Figure 5 As shown, this embodiment provides a computer-readable storage medium 500, on which a computer program 511 is stored. When the computer program 511 is executed by a processor, a voiceprint identification comparison recommendation method provided by an embodiment of the present invention is implemented;

[0123] It should be noted that in the above embodiments, the description of each embodiment has its own emphasis, and for parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0124] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0125] The present invention provides a voiceprint identification comparison recommendation method, comprising: obtaining a sample voice and a sample voice; inputting the sample voice and the sample voice into a voice recognition system, wherein the voice recognition system recognizes the text and corresponding phonemes of the sample voice, as well as the text and corresponding phonemes of the sample voice; performing stuttering word segmentation based on a self-built dictionary and an original dictionary according to the recognized text of the sample voice and the sample voice, obtaining the same phrases and single words in the text of the sample voice and the sample voice, recording the corresponding starting and ending positions of the phrases and single words, intercepting the corresponding voice segments, and forming a first voice segment; obtaining the same phrases and single words in the sample voice and the sample voice according to the recognized phonemes of the sample voice and the sample voice, Three phonemes are obtained, and the corresponding starting and ending positions are recorded, and the corresponding speech segments are intercepted to form a second speech segment; for the first speech segment and the second speech segment, the recommendation index of each speech segment is calculated through a recommendation algorithm; the speech segment with the highest recommendation index of the first speech segment and the speech segment with the highest recommendation index of the second speech segment are selected to form a new audio; the identity recommendation and the non-identity recommendation are distinguished by automatically performing voiceprint recognition on the new audio; the method provided by the present invention first obtains a new audio through speech recognition and a recommendation algorithm, and then performs voiceprint identification on the new audio, and the identity identification and the non-identity identification are considered at the same time in the voiceprint identification, which saves labor costs and has high accuracy.

[0126] In the method provided by the present invention, the same phrases and words in the text of the sample speech and the test material speech are obtained to form a first speech segment, and the same three phonemes in the sample speech and the test material speech are used to form a second speech segment; after the key speech segments are obtained, identification is performed to reduce redundancy and errors, reduce computing costs and improve accuracy.

[0127] The above is only a specific implementation of the present invention, but the design concept of the present invention is not limited thereto. Any non-substantial changes to the present invention using this concept shall be deemed as an infringement of the protection scope of the present invention.

Claims

1. A recommended method for voiceprint identification and comparison, It is characterized in that include: Obtain sample voice and inspection material voice; Inputting the sample speech and the sample speech into the speech recognition system, the speech recognition system recognizes the text and corresponding phonemes of the sample speech, and the text and corresponding phonemes of the sample speech; According to the recognized text of the sample speech and the test material speech, stuttering word segmentation is performed based on the dictionary, the same phrases and single words in the text of the sample speech and the test material speech are obtained, the corresponding starting and ending positions of the phrases and single words are recorded, and the corresponding speech segments are intercepted to form a first speech segment; According to the recognized phonemes of the sample voice and the sample voice, three phonemes that are the same as the sample voice and the sample voice are obtained, and the corresponding starting and ending positions are recorded, and the corresponding voice segments are intercepted to form a second voice segment; For the first speech segment and the second speech segment, the recommendation index of each speech segment is calculated by the recommendation algorithm. The recommendation index is based on the recommendation index, which includes: speech content consistency, spectrum clarity, formant number index, context consistency, and speech speed consistency. The recommendation index is specifically: Speech content consistency: Calculate the spectral cosine similarity of speech segments; Spectral clarity: Calculate the standard deviation of the harmonic energy in the peak part of the broadband spectrogram of the speech segment; Formant number index: count the number of formants in the speech segment and convert it into an index; Context consistency: Calculate the context consistency of the phonemes in the speech segment, that is, whether the previous and following consonants / silences of the vowel phonemes are the same, and obtain the corresponding index; Speech rate consistency: calculate the duration index of speech segments; Select the first speech segment with the highest recommendation index and the second speech segment with the highest recommendation index to form a new audio; Distinguish between identical recommendations and non-identical recommendations by automatically performing voiceprint recognition on new audio.

2. A voiceprint identification and comparison recommendation method according to claim 1, It is characterized in that The speech recognition system recognizes the text and corresponding phonemes of the sample speech, as well as the text and corresponding phonemes of the test material speech, specifically: Acoustic features are obtained from speech, including but not limited to linear predictive coding and Mel-frequency cepstral coefficients; The LSTM+CTC neural network acoustic model is used to convert acoustic features into phonemes; Language models based on deep neural networks convert phonemes into phrases and words.

3. A voiceprint identification and comparison recommendation device, It is characterized in that include: Acquisition unit: acquires sample voice and inspection material voice; Speech recognition unit: the sample speech and the test material speech are input into the speech recognition system, and the speech recognition system recognizes the text and corresponding phonemes of the sample speech, as well as the text and corresponding phonemes of the test material speech; A first speech segment acquisition unit: according to the recognized text of the sample speech and the test material speech, stuttering and word segmentation is performed based on the dictionary, the same phrases and single words in the text of the sample speech and the test material speech are acquired, the corresponding starting and ending positions of the phrases and single words are recorded, and the corresponding speech segments are intercepted to form the first speech segment; The second speech segment acquisition unit: according to the recognized phonemes of the sample speech and the test material speech, acquires the three phonemes that are the same in the sample speech and the test material speech, records the corresponding starting and ending positions, and intercepts the corresponding speech segment to form the second speech segment; The recommendation index calculation unit calculates the recommendation index of each speech segment through the recommendation algorithm for the first speech segment and the second speech segment. The recommendation index is based on the recommendation index, which includes: speech content consistency, spectrum clarity, formant number index, context consistency, and speech speed consistency. The recommendation index is specifically: Speech content consistency: Calculate the spectral cosine similarity of speech segments; Spectral clarity: Calculate the standard deviation of the harmonic energy in the peak part of the broadband spectrogram of the speech segment; Formant number index: count the number of formants in the speech segment and convert it into an index; Context consistency: Calculate the context consistency of the phonemes in the speech segment, that is, whether the previous and following consonants / silences of the vowel phonemes are the same, and obtain the corresponding index; Speech rate consistency: calculate the duration index of speech segments; New audio acquisition unit: selects the first speech segment with the highest recommendation index and the second speech segment with the highest recommendation index to form a new audio; Recommendation unit: Distinguish between identical recommendations and non-identical recommendations by automatically performing voiceprint recognition on new audio.

4. A voiceprint identification, comparison and recommendation device according to claim 3, It is characterized in that In the speech recognition unit, the speech recognition system recognizes the text and corresponding phonemes of the sample speech, as well as the text and corresponding phonemes of the sample speech, specifically: Acoustic features are obtained from speech; including but not limited to linear predictive coding and Mel-frequency cepstral coefficients; The LSTM+CTC neural network acoustic model is used to convert acoustic features into phonemes; Language models based on deep neural networks convert phonemes into phrases and words.

5. An electronic device, It is characterized in that include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the method steps of claim 1 or 2 are implemented when the processor executes the computer program.

6. A computer-readable storage medium, It is characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method steps of claim 1 or 2 are implemented.

Citation Information

Patent Citations

  • Voice fraud identification method and apparatus, terminal equipment and storage medium

    CN107680602A

  • Voice identity test method and device, electronic equipment and storage medium

    CN113921017A