Voiceprint recognition method and device, electronic equipment and storage medium

By performing final labeling and coverage calculation in the voiceprint recognition method, the problem of low voiceprint recognition efficiency in the prior art is solved, and faster and more accurate voiceprint recognition is achieved.

CN120375829APending Publication Date: 2025-07-25PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510511299.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The existing voiceprint recognition methods have low recognition efficiency due to the large number of types of phoneme characteristics in the speech.

Method used

By obtaining the voiceprint verification and registered voice of the target user, performing speech recognition, final labeling and coverage calculation are performed, the final distribution characteristics are extracted, the speech comparison is reduced, and the speech comparison is converted into final comparison to improve recognition efficiency.

Benefits of technology

It improves the speed and efficiency of voiceprint recognition, reduces the difficulty of extracting voiceprint features, and enhances the accuracy of recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120375829A_ABST
    Figure CN120375829A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a voiceprint recognition method and device, electronic equipment and a storage medium, belongs to the technical field of voiceprint recognition, and is suitable for the fields of financial science and technology and medical treatment. The method comprises the following steps: acquiring voiceprint verification voice and voiceprint registration voice of a target user; performing voice recognition on the voiceprint verification voice to obtain a voiceprint verification text; based on the voiceprint verification text, performing vowel labeling on the voiceprint verification voice to obtain a verification voice vowel label; performing vowel coverage rate calculation on the voiceprint verification voice based on the verification voice vowel label and a preset total number of basic vowels to obtain a verification voice vowel coverage rate; performing vowel distribution feature extraction on the voiceprint verification voice based on the verification voice vowel label to obtain verification voice vowel distribution; and performing voiceprint recognition on the target user based on the voiceprint registration voice, the total number of basic vowels, the verification voice vowel coverage rate and the verification voice vowel distribution characteristics. According to the embodiment of the invention, the voiceprint recognition efficiency can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of voiceprint recognition, and is applicable to the fields of fintech and healthcare, and particularly relates to a voiceprint recognition method, a device, an electronic device, and a storage medium. Background Art

[0002] Voiceprint recognition is a technology for identifying the identity of a speaker based on the voice signal in speech. For example, in the field of fintech, a financial system identifies the identity of a financial clerk or a customer through voiceprint recognition technology, so that the financial clerk or the customer can quickly log in to the medical system, improving the efficiency of the financial clerk or the customer accessing the financial system. In the field of healthcare, a medical system identifies the identity of medical staff through voiceprint recognition technology, so that the medical staff can quickly log in to the medical system, improving the efficiency of the medical staff accessing the medical system.

[0003] Currently, the method of voiceprint recognition usually analyzes the phoneme features in speech to identify the speaker's identity. Since phoneme features include various types such as initials, finals, and tones, the efficiency of voiceprint recognition is low. Therefore, how to improve the efficiency of voiceprint recognition has become an urgent technical problem to be solved. Summary of the Invention

[0004] The main purpose of the embodiments of the present application is to propose a voiceprint recognition method, a device, an electronic device, and a storage medium, aiming to improve the efficiency of voiceprint recognition.

[0005] To achieve the above object, in the first aspect of the embodiments of the present application, a voiceprint recognition method is proposed. Obtain the voiceprint verification speech and the voiceprint registration speech of a target user, where the voiceprint verification speech represents the speech data recorded at the current moment, and the voiceprint registration speech represents the speech data recorded at a historical moment, and the historical moment is before the current moment;

[0006] Perform speech recognition on the voiceprint verification speech to obtain a voiceprint verification text;

[0007] Based on the voiceprint verification text, perform finals annotation on the voiceprint verification speech to obtain a verification speech finals label;

[0008] Based on the verification speech finals label and the preset total number of basic finals, calculate the finals coverage rate of the voiceprint verification speech to obtain a verification speech finals coverage rate, where the verification speech finals coverage rate is used to represent the proportion of the number of types of finals labels in the total number of basic finals;

[0009] Based on the verified voice vowel labels, extract the vowel distribution features of the voiceprint verification voice to obtain the verified voice vowel distribution features, where the verified voice vowel distribution features are used to characterize the frequency of occurrence of each of the verified voice vowel labels in the voiceprint verification voice;

[0010] Based on the voiceprint registration voice, the total number of basic vowels, the vowel coverage rate of the verification voice, and the verified voice vowel distribution features, perform voiceprint recognition on the target user.

[0011] In some embodiments, the performing voiceprint recognition on the target user based on the voiceprint registration voice, the total number of basic vowels, the vowel coverage rate of the verification voice, and the verified voice vowel distribution features includes:

[0012] Perform vowel annotation on the voiceprint registration voice to obtain the registered voice vowel labels;

[0013] Based on the registered voice vowel labels, extract features from the voiceprint registration voice to obtain the vowel coverage rate of the registered voice and the vowel distribution features of the registered voice;

[0014] Based on the vowel coverage rate of the verification voice and the vowel coverage rate of the registered voice, perform similarity evaluation on the target user to obtain the coverage rate evaluation data;

[0015] Based on the verified voice vowel distribution features and the vowel distribution features of the registered voice, perform similarity evaluation on the target user to obtain the distribution feature evaluation data;

[0016] Based on the verified voice vowel labels, the registered voice vowel labels, and the total number of basic vowels, perform co-coverage evaluation on the voiceprint verification voice and the voiceprint registration voice to obtain the vowel co-coverage evaluation data;

[0017] Generate voiceprint recognition information based on the coverage rate evaluation data, the distribution feature evaluation data, and the vowel co-coverage evaluation data.

[0018] In some embodiments, the performing co-coverage evaluation on the voiceprint verification voice and the voiceprint registration voice based on the verified voice vowel labels, the registered voice vowel labels, and the total number of basic vowels to obtain the vowel co-coverage evaluation data includes:

[0019] Search for the same labels between the verified voice vowel labels and the registered voice vowel labels to obtain the same vowel label data;

[0020] Based on the same vowel label data and the total number of basic vowels, calculate the coverage rate of the voiceprint verification speech and the voiceprint registration speech to obtain the vowel common coverage rate evaluation data.

[0021] In some embodiments, the calculation of the vowel coverage rate of the voiceprint verification speech based on the vowel labels of the verification speech and the preset total number of basic vowels to obtain the vowel coverage rate of the verification speech includes:

[0022] Based on the vowel labels of the verification speech, identify the vowel types of the voiceprint verification speech to obtain the vowel type data of the verification speech;

[0023] Based on the total number of basic vowels, calculate the proportion of the vowel type data of the verification speech to obtain the vowel coverage rate of the verification speech.

[0024] In some embodiments, the extraction of the vowel distribution characteristics of the voiceprint verification speech based on the vowel labels of the verification speech to obtain the vowel distribution characteristics of the verification speech includes:

[0025] Based on the vowel labels of the verification speech, calculate the vowel occurrence frequency of the voiceprint verification speech to obtain the vowel occurrence frequency of the verification speech;

[0026] Perform information entropy calculation and processing on the vowel occurrence frequency of the verification speech to obtain the vowel distribution characteristics of the verification speech.

[0027] In some embodiments, the annotation of the vowel labels of the voiceprint verification speech based on the voiceprint verification text to obtain the vowel labels of the verification speech includes:

[0028] Based on the voiceprint verification speech, identify the single-character speech duration of the voiceprint verification text to obtain the single-character timestamp data;

[0029] Identify the basic vowel phonemes of the voiceprint verification text;

[0030] Based on the single-character timestamp data, segment the voiceprint verification speech to obtain single-character speech data;

[0031] Based on the basic vowel phonemes, label the vowels of the single-character speech data to obtain the vowel labels of the verification speech.

[0032] In some embodiments, the voice recognition of the voiceprint verification speech to obtain the voiceprint verification text includes:

[0033] Extract the features of the voiceprint verification speech to obtain the verification speech features;

[0034] Decode the verified voice feature to obtain the voiceprint verification text.

[0035] To achieve the above object, a second aspect of the embodiments of the present application provides a voiceprint recognition device, which includes:

[0036] A data acquisition module, configured to acquire the voiceprint verification voice and the voiceprint registration voice of the target user, where the voiceprint verification voice represents the voice data recorded at the current moment, and the voiceprint registration voice represents the voice data recorded at a historical moment, and the historical moment is before the current moment;

[0037] A voice conversion module, configured to perform voice recognition on the voiceprint verification voice to obtain a voiceprint verification text;

[0038] A vowel annotation module, configured to perform vowel annotation on the voiceprint verification voice based on the voiceprint verification text to obtain a verified voice vowel label;

[0039] A coverage rate calculation module, configured to calculate the vowel coverage rate of the voiceprint verification voice based on the verified voice vowel label and the preset total number of basic vowels to obtain the verified voice vowel coverage rate, where the verified voice vowel coverage rate is used to represent the proportion of the number of types of vowel labels in the total number of basic vowels;

[0040] A distribution feature calculation module, configured to extract the vowel distribution feature of the voiceprint verification voice based on the verified voice vowel label to obtain the verified voice vowel distribution feature, where the verified voice vowel distribution feature is used to represent the frequency of occurrence of each verified voice vowel label in the voiceprint verification voice;

[0041] A voiceprint recognition module, configured to perform voiceprint recognition on the target user based on the voiceprint registration voice, the total number of basic vowels, the verified voice vowel coverage rate, and the verified voice vowel distribution feature.

[0042] To achieve the above object, a third aspect of the embodiments of the present application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the method described in the first aspect above is implemented.

[0043] To achieve the above object, a fourth aspect of the embodiments of the present application provides a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, the method described in the first aspect above is implemented.

[0044] The voiceprint recognition method and device, electronic device, and storage medium proposed in this application obtain the voiceprint verification speech and voiceprint registration speech of the target user, perform speech recognition on the voiceprint verification speech to obtain the voiceprint verification text, and then perform vowel annotation on the voiceprint verification speech according to the voiceprint verification text to obtain the verification speech vowel label. Further, according to the verification speech vowel label and the preset total number of basic vowels, calculate the vowel coverage rate of the voiceprint verification speech to obtain the verification speech vowel coverage rate, and extract the vowel distribution feature of the voiceprint verification speech according to the verification speech vowel label to obtain the verification speech vowel distribution feature, which reduces the difficulty of extracting voiceprint features, improves the extraction efficiency of voiceprint features, and thus speeds up the voiceprint recognition speed. Finally, perform voiceprint recognition on the target user according to the voiceprint registration speech, the total number of basic vowels, the verification speech vowel coverage rate, and the verification speech vowel distribution feature, convert the comparison of voice phonemes in the voiceprint recognition process into the comparison of voice vowels, reduce the difficulty of voice comparison, and thus improve the voice comparison efficiency, and also improve the efficiency of voiceprint recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1 is a flowchart of the voiceprint recognition method provided by an embodiment of the present application;

[0046] Figure 2 is Figure 1 a flowchart of step S102 in

[0047] Figure 3 is Figure 1 a flowchart of step S103 in

[0048] Figure 4 is Figure 1 a flowchart of step S104 in

[0049] Figure 5 is Figure 1 a flowchart of step S105 in

[0050] Figure 6 is Figure 1 a flowchart of step S106 in

[0051] Figure 7 is Figure 6 a flowchart of step S605 in

[0052] Figure 8 is a schematic structural diagram of the voiceprint recognition device provided by an embodiment of the present application;

[0053] Figure 9 is a schematic hardware structure diagram of the electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0054] In order to make the objectives, technical solutions and advantages of this application clearer and more understandable, the following further elaborates on this application in combination with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not used to limit this application.

[0055] It should be noted that although functional modules are divided in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different module division in the device or a different order in the flowchart. Terms such as "first" and "second" in the description, claims and the above-mentioned drawings are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence.

[0056] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.

[0057] First, several terms involved in this application are analyzed:

[0058] Voiceprint recognition system: A voiceprint recognition system is an identity recognition technology based on the voice characteristics of a speaker. The voiceprint recognition system analyzes the voice signal of the speaker, extracts unique acoustic features, constructs and compares the voiceprint model of the speaker, so as to achieve fast and accurate identity verification. The voiceprint recognition system is widely used in fields such as finance, security, and medical care, providing a safe and convenient identity recognition solution for users.

[0059] Voiceprint verification voice: Voiceprint verification voice refers to the voice data recorded by a user when identity verification is required, which is used to compare with the pre-registered voiceprint data. Voiceprint verification voice usually contains specific phrases or random voice content, and the voiceprint recognition system confirms the identity of the speaker by analyzing its acoustic features. Voiceprint verification voice is a key input in the voiceprint recognition system, used to ensure that only authorized personnel can access sensitive information or systems.

[0060] Voiceprint registration voice: Voiceprint registration voice refers to a piece of voice pre-recorded by a user in a voiceprint recognition system, which is used to create and train an individual's voiceprint model. Voiceprint registration voice usually contains a series of specific phrases or random voice content, and the voiceprint recognition system constructs the voiceprint feature template of this user by analyzing the acoustic features of this voice. The quality and content of the voiceprint registration voice directly affect the accuracy and reliability of the voiceprint model, and are the basis for the voiceprint recognition system to achieve accurate identity verification.

[0061] Vowel coverage rate: The vowel coverage rate refers to the proportion of the number of basic vowel types contained in a piece of speech or text to the total number of basic vowels. The vowel coverage rate is used to measure the coverage breadth of speech data on basic vowels, reflecting the diversity and representativeness of speech data. A higher vowel coverage rate indicates that the speech data covers more basic vowels, which helps to improve the accuracy of voiceprint recognition and the generalization ability of the model.

[0062] Vowel distribution characteristics: The vowel distribution characteristics refer to the frequency of occurrence of each basic vowel and its overall distribution pattern in a piece of speech or text. The vowel distribution characteristics reflect the degree of uniformity of the vowel distribution in speech data by counting the number of occurrences and frequencies of each vowel. Entropy is usually used to quantify the vowel distribution characteristics, and the higher the entropy value, the more uniform the vowel distribution. The vowel distribution characteristics can evaluate the balance and diversity of speech data, providing an important basis for the optimization of the voiceprint recognition model.

[0063] Voiceprint recognition is a technology for identifying the identity of a speaker based on the sound signal in speech. For example, in the medical field, the medical system uses voiceprint recognition technology to identify the identity of medical staff, enabling medical staff to quickly log in to the medical system and improving the efficiency of medical staff accessing the medical system.

[0064] Currently, the method of voiceprint recognition usually analyzes the phoneme characteristics in speech to identify the speaker's identity. Since phoneme characteristics include various types such as initials, finals, and tones, the efficiency of voiceprint recognition is relatively low. Therefore, how to improve the efficiency of voiceprint recognition has become an urgent technical problem to be solved.

[0065] Based on this, the embodiments of the present application provide a voiceprint recognition method, device, electronic device, and storage medium, aiming to improve the efficiency of voiceprint recognition.

[0066] The voiceprint recognition method, device, electronic device, and storage medium provided by the embodiments of the present application are specifically described through the following embodiments. First, the voiceprint recognition method in the embodiments of the present application is described.

[0067] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Among them, artificial intelligence (AI) is a theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results.

[0068] The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technologies, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, robotics, biometric technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0069] The voiceprint recognition method provided by the embodiments of this application relates to the field of voiceprint recognition technology and is applicable to the fields of fintech and healthcare. The voiceprint recognition method provided by the embodiments of this application can be applied to a terminal, or to a server, or can be software running on a terminal or a server. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, etc.; the server can be configured as an independent physical server, or can be configured as a server cluster or a distributed system composed of multiple physical servers, or can be configured as a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application that implements the voiceprint recognition method, etc., but is not limited to the above forms.

[0070] This application can be used in many general or specific computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This application can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0071] It should be noted that in each specific embodiment of the present application, when it comes to relevant processing based on data related to the user's identity or characteristics, such as user information, user behavior data, user historical data, and user location information, the user's permission or consent will be obtained first. Moreover, the collection, use, and processing of these data will comply with relevant laws, regulations, and standards. In addition, when the embodiments of the present application need to obtain the user's sensitive personal information, the user's separate permission or separate consent will be obtained through methods such as pop-up windows or jumping to a confirmation page. After clearly obtaining the user's separate permission or separate consent, the necessary user-related data for the normal operation of the embodiments of the present application will be obtained.

[0072] Figure 1 FIG. 4 is an optional flowchart of a voiceprint recognition method provided by an embodiment of the present application. This method can be used in a voiceprint recognition system. Figure 1 The method in FIG. 4 may include, but is not limited to, steps S101 to S106.

[0073] Step S101: Obtain the voiceprint verification voice and voiceprint registration voice of the target user. Among them, the voiceprint verification voice represents the voice data recorded at the current moment, and the voiceprint registration voice represents the voice data recorded at a historical moment, where the historical moment is before the current moment.

[0074] Step S102: Perform speech recognition on the voiceprint verification voice to obtain a voiceprint verification text.

[0075] Step S103: Based on the voiceprint verification text, perform vowel annotation on the voiceprint verification voice to obtain a verification voice vowel label.

[0076] Step S104: Based on the verification voice vowel label and the preset total number of basic vowels, calculate the vowel coverage rate of the voiceprint verification voice to obtain a verification voice vowel coverage rate. Among them, the verification voice vowel coverage rate is used to represent the proportion of the number of types of vowel labels in the total number of basic vowels.

[0077] Step S105: Based on the verification voice vowel label, extract the vowel distribution feature of the voiceprint verification voice to obtain a verification voice vowel distribution feature. Among them, the verification voice vowel distribution feature is used to represent the frequency of occurrence of each verification voice vowel label in the voiceprint verification voice.

[0078] Step S106: Perform voiceprint recognition on the target user based on the voiceprint registration voice, the total number of basic vowels, the verification voice vowel coverage rate, and the verification voice vowel distribution feature.

[0079] Steps S101 to S106 shown in the embodiments of the present application obtain the voiceprint verification voice and voiceprint registration voice of the target user, perform speech recognition on the voiceprint verification voice to obtain the voiceprint verification text, then, according to the voiceprint verification text, perform vowel annotation on the voiceprint verification voice to obtain the verification voice vowel label. Further, according to the verification voice vowel label and the preset total number of basic vowels, calculate the vowel coverage rate of the voiceprint verification voice to obtain the verification voice vowel coverage rate. Extract the vowel distribution characteristics of the voiceprint verification voice according to the verification voice vowel label to obtain the verification voice vowel distribution characteristics, which reduces the difficulty of extracting voiceprint features and improves the extraction efficiency of voiceprint features, thereby accelerating the speed of voiceprint recognition. Finally, perform voiceprint recognition on the target user according to the voiceprint registration voice, the total number of basic vowels, the verification voice vowel coverage rate, and the verification voice vowel distribution characteristics, convert the comparison of speech phonemes in the voiceprint recognition process into the comparison of speech vowels, reduce the difficulty of speech comparison, thereby improving the speech comparison efficiency, and also improving the efficiency of voiceprint recognition.

[0080] In step S101 of some embodiments, the voiceprint verification voice refers to the voice data recorded by the target user at the current moment for verifying identity information, where the current moment refers to the time point when the target user performs voiceprint recognition. For example, in the electronic medical record system of a hospital, the voiceprint verification voice can be the voice data entered by a doctor when logging in to the electronic medical record system through voiceprint recognition. The voiceprint registration voice refers to the voice data pre-recorded and stored in the voiceprint recognition system. The recording time of the voiceprint registration voice can be the time point when the target user enables the voiceprint recognition function. For example, in the voice recognition system of a hospital, the voiceprint registration voice can be the voice data recorded by a doctor when enabling the voiceprint recognition login function in the electronic medical record system.

[0081] In the embodiments of the present application, the voiceprint recognition system receives the voiceprint recognition request sent by the target user and records the voice data of the target user in real time to obtain the voiceprint verification voice. Further, the voiceprint registration voice of the target user can be obtained by extracting the voice data stored in the system itself. It should be noted that when the target user enters the voiceprint registration voice, the voiceprint recognition system can store the voiceprint registration voice in its own storage space for voiceprint recognition of the target user.

[0082] In step S102 of some embodiments, the voiceprint verification text refers to the text content converted from the above voiceprint verification voice.

[0083] In the embodiments of the present application, the voiceprint verification text can be obtained by extracting the voice features of the voiceprint verification voice and then performing text conversion on the extracted voice features.

[0084] Specifically, please refer toFigure 2 In some embodiments, step S102 may include but is not limited to steps S201 to S202:

[0085] Step S201, extracting features of the voiceprint verification speech to obtain verification speech features;

[0086] Step S202: Decode the verification voice features to obtain voiceprint verification text.

[0087] In step S201 of some embodiments, the verification speech feature refers to the characteristic parameters of the voiceprint verification speech, such as Mel-frequency cepstral coefficients, linear prediction coefficients, etc.

[0088] The embodiments of the present application can obtain the Mel-frequency cepstrum coefficients of the voiceprint verification speech by performing Fourier transform, Mel filtering, logarithmic energy calculation, discrete cosine transform and other processing on the voiceprint verification speech, and can also obtain the linear prediction coefficients of the voiceprint verification speech by performing autocorrelation calculation, linear prediction analysis, LPC coefficient calculation and other processing on the voiceprint verification speech.

[0089] In step S202 of some embodiments, a trained acoustic model can be used to estimate the probability of each possible phoneme or syllable at each time point based on the verified speech features, and then a decoder is used to find the most likely text sequence from the probability output by the acoustic model, wherein the acoustic model can be a hidden Markov model, a deep neural network, a convolutional neural network, a recurrent neural network and their variants.

[0090] In steps S201 to S202 shown in the embodiment of the present application, by extracting features of the voiceprint verification speech to obtain verification speech features, and then decoding the verification speech features to obtain voiceprint verification text, the accuracy of speech-to-text conversion can be improved, thereby improving the accuracy of the voiceprint verification text.

[0091] In step S103 of some embodiments, the verification phonetic final label refers to the final of each Chinese character in the voiceprint verification text. For example, in the voiceprint verification text "Hello world", the verification phonetic final label of "你" is "i", the verification phonetic final label of "好" is "ao", the verification phonetic final label of "世" is "i", and the verification phonetic final label of "界" is "ie".

[0092] In an embodiment of the present application, by identifying the duration data occupied by each Chinese character in the voiceprint verification speech, single-word timestamp data can be obtained, and the voiceprint verification speech is segmented according to the single-word timestamp data to obtain single-word speech data, and by marking the finals of each single-word speech data, the final label of the verification speech can be obtained.

[0093] For details, seeFigure 3 In some embodiments, step S103 may include but is not limited to steps S301 to S304:

[0094] Step S301, based on the voiceprint verification speech, single-word voice duration recognition is performed on the voiceprint verification text to obtain single-word timestamp data;

[0095] Step S302, performing vowel recognition on the voiceprint verification text to obtain basic vowel phonemes;

[0096] Step S303, segmenting the voiceprint verification speech based on the single-word timestamp data to obtain single-word speech data;

[0097] Step S304, based on the basic final phonemes, the single-word voice data is marked with finals to obtain the final labels of the verified voice.

[0098] In step S301 of some embodiments, the single-word timestamp data refers to the time duration data occupied by a single Chinese character in the voiceprint verification speech. For example, the total duration of the voiceprint verification speech "Hello world" is 2 seconds, of which the single-word timestamp data of "你" can be 0.4s, the single-word timestamp data of "好" can be 0.6s, the single-word timestamp data of "世" can be 0.4s, and the single-word timestamp data of "界" can be 0.6s.

[0099] In an embodiment of the present application, the active part in the voiceprint verification speech can be detected to remove the silent segment of the voiceprint verification speech. Furthermore, the active part is divided into multiple short time frames, and the energy of each short time frame is calculated to determine the starting point and end point of the effective speech in the voiceprint verification speech. Finally, the start and end time of each word in the effective speech part is recorded to obtain the single word timestamp data of each word.

[0100] In step S302 of some embodiments, the basic vowel phoneme refers to the vowel corresponding to each Chinese character in the voiceprint verification text.

[0101] In the embodiment of the present application, by converting the voiceprint verification text into pinyin and then extracting the finals from the pinyin, the basic final phonemes of each Chinese character in the voiceprint verification text can be obtained.

[0102] In step S303 of some embodiments, the single-word voice data refers to the voice data of each Chinese character in the voiceprint verification voice.

[0103] In an embodiment of the present application, by matching single-word timestamp data with the voiceprint verification speech, the speech frame occupied by each Chinese character is obtained. Furthermore, according to the speech frame occupied by each Chinese character, the voiceprint verification speech is segmented to obtain the speech data of each character, that is, the single-word speech data.

[0104] In step S304 of some embodiments, each single-character voice data is matched with the corresponding basic vowel phoneme, and according to the matching relationship, a corresponding vowel label, i.e., a verified voice vowel label, is generated for each single-character voice data.

[0105] In steps S301 to S304 illustrated in the embodiments of the present application, based on the voiceprint-verified voice, the single-character voice duration of the voiceprint-verified text is recognized to obtain single-character timestamp data, the vowel of the voiceprint-verified text is recognized to obtain basic vowel phonemes, then based on the single-character timestamp data, the voiceprint-verified voice is segmented to obtain single-character voice data, and finally, based on the basic vowel phonemes, the single-character voice data is labeled with vowels to obtain a verified voice vowel label, so that there is a clear interval between the verified voice vowel labels to avoid confusion of vowel labels and affect voiceprint recognition.

[0106] In step S104 of some embodiments, the total number of basic vowels refers to the total number of vowels in Chinese pinyin. Generally, the total number of basic vowels is represented by T, and the value of the total number of basic vowels T is 25. The coverage rate of verified voice vowels refers to the ratio of the number of types of verified voice vowel labels included in the voiceprint-verified voice to the total number of basic vowels.

[0107] In the embodiments of the present application, by identifying the number of types of verified voice vowel labels, the number of types of basic vowels in the voiceprint-verified voice is determined, and then based on the determined number of types of basic vowels in the voiceprint-verified voice and the total number of basic vowels, the coverage rate of vowels in the voiceprint-verified voice can be calculated.

[0108] Specifically, please refer to Figure 4 , in some embodiments, step S104 may include but is not limited to steps S401 to S402:

[0109] Step S401, based on the verified voice vowel label, perform vowel type recognition on the voiceprint-verified voice to obtain verified voice vowel type data;

[0110] Step S402, based on the total number of basic vowels, perform a ratio calculation on the verified voice vowel type data to obtain the coverage rate of verified voice vowels.

[0111] In step S401 of some embodiments, the verified voice vowel type data refers to the number of types of verified voice vowel labels. For example, the verified voice vowel type data in the voiceprint-verified voice "Hello world" is 3.

[0112] In the embodiments of the present application, by identifying the number of different verified voice vowel labels in the voiceprint-verified voice, the verified voice vowel type data can be obtained.

[0113] In step S402 of some embodiments, after obtaining the verified speech vowel type data, based on the pre-known total number of basic vowels and the verified speech vowel type data, the vowel coverage rate of the verified speech vowel label can be calculated.

[0114] Specifically, the following formula can be used to calculate the vowel coverage rate of the voiceprint verification speech:

[0115]

[0116] Where V represents the vowel coverage rate of the verified speech, N represents the verified speech vowel type data, and T represents the total number of basic vowels.

[0117] In the embodiments of the present application, when the voiceprint verification speech is "Hello world", the verified speech vowel type data of the voiceprint verification speech is 3, and the total number of basic vowels is 25. At this time, the vowel coverage rate of the verified speech can be 3 / 25 = 0.12.

[0118] In steps S401 to S402 illustrated in the embodiments of the present application, based on the verified speech vowel label, the vowel type of the voiceprint verification speech is recognized to obtain the verified speech vowel type data, and then based on the total number of basic vowels, the proportion of the verified speech vowel type data is calculated to obtain the vowel coverage rate of the verified speech, which can clarify whether the voiceprint verification speech has vowel diversity. Furthermore, based on the vowel diversity, the voice quality of the voiceprint verification speech is evaluated. Specifically, the higher the vowel diversity, the higher the voice quality of the voiceprint verification speech, and thus the higher the accuracy of voiceprint recognition. In addition, by calculating the coverage rate of the vowels in the voiceprint verification data, the processing difficulty of the voiceprint features is reduced, thereby improving the efficiency of voiceprint recognition.

[0119] In step S105 of some embodiments, the verified speech vowel distribution feature refers to the frequency of various verified speech vowel labels appearing in the voiceprint verification speech, which can represent the distribution uniformity of the verified speech vowel labels in the voiceprint verification speech.

[0120] In the embodiments of the present application, based on the verified speech vowel label, calculate the proportion of the number of each verified speech vowel label to the total number of all verified speech vowel labels, that is, the verified speech vowel appearance frequency, and then perform information entropy calculation processing on the verified speech vowel appearance frequency, and the distribution feature of each verified speech vowel label can be obtained.

[0121] Specifically, please refer to Figure 5 , in some embodiments, step S105 may include but is not limited to steps S501 to S502:

[0122] Step S501: Calculate the occurrence frequency of the vowels in the voiceprint verification voice based on the verified voice vowel tags to obtain the occurrence frequency of the verified voice vowels.

[0123] Step S502: Perform information entropy calculation and processing on the occurrence frequency of the verified voice vowels to obtain the distribution characteristics of the verified voice vowels.

[0124] In step S501 of some embodiments, by identifying the types of the verified voice vowel tags, the quantity of each type of verified voice vowel tag can be obtained. Further, by calculating the ratio of the quantity of each type of verified voice vowel tag to the quantity of all the verified voice vowel tags in the voiceprint verification voice, the occurrence frequency of the verified voice vowels of each type of verified voice vowel tag can be obtained.

[0125] In step S502 of some embodiments, by performing logarithmic processing on the occurrence frequency of the verified voice vowels of each type of verified voice vowel tag, the logarithm of the occurrence frequency of the verified voice vowels of each type of verified voice vowel tag can be obtained. Further, by performing multiplication calculation on the logarithm of the occurrence frequency of the verified voice vowels of each type of verified voice vowel tag and the occurrence frequency of the verified voice vowels, the intermediate representation of the occurrence frequency of the verified voice vowels of each type of verified voice vowel tag can be obtained. Then, by merging the intermediate representations of the occurrence frequency of the verified voice vowels of each type of verified voice vowel tag and taking the opposite number, the distribution characteristics of the verified voice vowels can be obtained.

[0126] Specifically, the following formula can be used to calculate the distribution characteristics of the verified voice vowels of the voiceprint verification voice:

[0127]

[0128] where S represents the distribution characteristics of the verified voice vowels, i represents the i-th type of verified voice vowel tag, n represents the number of types of verified voice vowel tags, and p i represents the occurrence frequency of the verified voice vowels of the i-th type of verified voice vowel tag.

[0129] In steps S501 to S502 shown in the embodiments of the present application, according to the verified voice vowel labels, the vowel occurrence frequency of the voiceprint verification voice is calculated to obtain the vowel occurrence frequency of the verification voice, and then the information entropy calculation process is performed on the vowel occurrence frequency of the verification voice to obtain the vowel distribution characteristics of the verification voice, which is convenient to understand the occurrence frequency of various verification voice vowel labels in the voiceprint verification voice and the overall distribution pattern of the verification voice vowel labels, so as to realize the evaluation of the quality of the voiceprint verification voice from the dimension of the distribution model. Specifically, the more uniform the distribution of various verification voice vowel labels in the vowel distribution characteristics of the verification voice indicates, the better the quality of the voiceprint verification voice; the more uneven the distribution of various verification voice vowel labels in the vowel distribution characteristics of the verification voice indicates, the worse the quality of the voiceprint verification voice. Further, when the quality of the voiceprint verification voice is poor, by replacing the voiceprint verification voice, the accuracy of voiceprint recognition can be improved. In addition, by calculating the distribution characteristics of vowels in the voiceprint verification data, the processing difficulty of voiceprint features is reduced, thereby improving the efficiency of voiceprint recognition.

[0130] In step S106 of some embodiments, after obtaining the vowel coverage rate of the verification voice and the vowel distribution characteristics of the verification voice of the voiceprint verification voice, by calculating the vowel characteristics of the voiceprint registration voice, the vowel coverage rate of the registration voice and the vowel distribution characteristics of the registration voice of the voiceprint registration voice can be obtained. Further, by comparing the vowel coverage rate of the verification voice and the vowel coverage rate of the registration voice, and comparing the vowel distribution characteristics of the verification voice and the vowel distribution characteristics of the registration voice, it can be used to preliminarily judge whether the voiceprint verification voice data and the voiceprint registration voice data are the coverage rate evaluation data and the distribution characteristic evaluation data of the same person. Secondly, by calculating the common coverage rate of the basic vowels in the voiceprint verification voice and the voiceprint registration voice data, the common coverage rate evaluation data of the vowels is obtained. Finally, according to the above-mentioned coverage rate evaluation data, distribution characteristic evaluation data and common coverage rate, it is jointly identified whether the inputters of the voiceprint verification voice and the voiceprint registration voice are the same user.

[0131] Specifically, please refer to Figure 6 , in some embodiments, step S106 may include but is not limited to steps S601 to S606:

[0132] Step S601, perform vowel annotation on the voiceprint registration voice to obtain the registration voice vowel label;

[0133] Step S602, based on the registration voice vowel label, perform feature extraction on the voiceprint registration voice to obtain the vowel coverage rate of the registration voice and the vowel distribution characteristics of the registration voice;

[0134] Step S603: Based on the verified voice vowel coverage rate and the registered voice vowel coverage rate, perform a similarity evaluation on the target user to obtain coverage rate evaluation data;

[0135] Step S604: Based on the verified voice vowel distribution characteristics and the registered voice vowel distribution characteristics, perform a similarity evaluation on the target user to obtain distribution characteristic evaluation data;

[0136] Step S605: Based on the verified voice vowel label, the registered voice vowel label, and the total number of basic vowels, perform a common coverage rate evaluation on the voiceprint verification voice and the voiceprint registration voice to obtain vowel common coverage rate evaluation data;

[0137] Step S606: Generate voiceprint recognition information based on the coverage rate evaluation data, the distribution characteristic evaluation data, and the vowel common coverage rate evaluation data.

[0138] In steps S601 and S602 of some embodiments, the registered voice vowel label refers to the basic vowels in the voiceprint registration voice. The registered voice vowel coverage rate refers to the proportion of the number of types of registered voice vowel labels included in the voiceprint registration voice to the total number of basic vowels. The registered voice vowel distribution characteristic refers to the frequency of occurrence of various registered voice vowel labels in the voiceprint registration voice, which can represent the distribution uniformity of the registered voice vowel labels in the voiceprint registration voice.

[0139] In the embodiments of the present application, the annotation method of the registered voice vowel label is similar to that of the verified voice vowel label. Specifically, perform voice recognition on the voiceprint registration voice to obtain the voiceprint registration text. Then, perform single-character voice duration recognition on the voiceprint registration text to obtain the single-character timestamp data of the voiceprint registration voice, and perform vowel recognition on the voiceprint registration text to obtain the basic vowel phonemes. Secondly, based on the single-character timestamp data of the voiceprint registration voice, segment the voiceprint registration voice to obtain the single-character voice data of the voiceprint registration voice. Finally, based on the basic vowel phonemes, perform vowel tagging on the single-character voice data of the voiceprint registration voice to obtain the registered voice vowel label.

[0140] In another embodiment of the present application, the calculation method of the registered voice vowel coverage rate is similar to that of the verified voice vowel coverage rate. Specifically, according to the registered voice vowel label, perform vowel type recognition on the voiceprint registration voice to obtain the registered voice vowel type data, and then calculate the proportion according to the total number of basic vowels to obtain the registered voice vowel coverage rate. In addition, the calculation method of the registered voice vowel distribution characteristic is also similar to that of the verified voice vowel distribution characteristic. Specifically, based on the registered voice vowel label, calculate the vowel occurrence frequency of the voiceprint registration voice to obtain the registered voice vowel occurrence frequency, and then perform information entropy calculation processing on the registered voice vowel occurrence frequency to obtain the registered voice vowel distribution characteristic.

[0141] In steps S603 and S604 of some embodiments, after obtaining the coverage rate of the registered phonetic finals and the distribution characteristics of the registered phonetic finals of the voiceprint registration voice, comparing the coverage rate of the verification phonetic finals with the coverage rate of the registered phonetic finals can obtain the similarity of the coverage rate of the voiceprint verification voice and the voiceprint registration voice, that is, the coverage rate evaluation data. It should be noted that the larger the coverage rate evaluation data, the more similar the coverage rate of the verification phonetic finals and the coverage rate of the registered phonetic finals, that is, the greater the probability that the voiceprint verification voice and the voiceprint registration voice are voice data of the same user. The smaller the coverage rate evaluation data, the greater the difference between the coverage rate of the verification phonetic finals and the coverage rate of the registered phonetic finals, that is, the smaller the probability that the voiceprint verification voice and the voiceprint registration voice are voice data of the same user.

[0142] Similarly, comparing the distribution characteristics of the verification phonetic finals with the distribution characteristics of the registered phonetic finals can obtain the similarity of the distribution characteristics of the voiceprint verification voice and the voiceprint registration voice, that is, the distribution characteristic evaluation data, which is the evaluation data of the distribution characteristics. It should be noted that the larger the evaluation data, the more similar the distribution characteristics of the verification phonetic finals and the distribution characteristics of the registered phonetic finals, that is, the greater the probability that the voiceprint verification voice and the voiceprint registration voice are voice data of the same user. The smaller the distribution characteristic evaluation data, the greater the difference between the distribution characteristics of the verification phonetic finals and the distribution characteristics of the registered phonetic finals, that is, the smaller the probability that the voiceprint verification voice and the voiceprint registration voice are voice data of the same user.

[0143] In step S605 of some embodiments, the evaluation data of the common coverage rate of phonetic finals refers to the proportion of the number of the same basic phonetic finals in the voiceprint verification voice and the voiceprint registration voice to the total number of basic phonetic finals.

[0144] In the embodiments of the present application, by querying the phonetic final labels of the verification voice and the phonetic final labels of the registered voice, the phonetic final labels of the verification voice and the phonetic final labels of the registered voice with the same type can be found, that is, the same phonetic final label data. Further, calculating the proportion of the above same phonetic final label data in the total number of basic phonetic finals can obtain the evaluation data of the common coverage rate of phonetic finals.

[0145] Specifically, please refer to Figure 7 , in some embodiments, step S605 may include but is not limited to steps S701 to S702:

[0146] Step S701, perform the same label search on the phonetic final labels of the verification voice and the phonetic final labels of the registered voice to obtain the same phonetic final label data;

[0147] Step S702, based on the same phonetic final label data and the total number of basic phonetic finals, calculate the coverage rate of the voiceprint verification voice and the voiceprint registration voice to obtain the evaluation data of the common coverage rate of phonetic finals.

[0148] In steps S701 and S702 of some embodiments, the following formula can be used to calculate the evaluation data of the common coverage rate of the finals of the voiceprint verification voice and the voiceprint registration voice:

[0149]

[0150] Where V ′ represents the evaluation data of the common coverage rate of the finals, N ′ represents the data of the same finals label, and T represents the total number of basic finals.

[0151] In the embodiments of the present application, when the voiceprint verification voice is "Hello world" and the voiceprint registration voice is "Hello world", the data of the same finals label of the voiceprint verification voice and the voiceprint registration voice is 3, and the total number of basic finals is 25. At this time, the coverage rate of the verification voice finals can be 3 / 25 = 0.12.

[0152] In steps S701 to S702 shown in the embodiments of the present application, by searching for the same labels of the finals labels of the verification voice and the registration voice, the data of the same finals label is obtained, and then based on the data of the same finals label and the total number of basic finals, the coverage rate of the voiceprint verification voice and the voiceprint registration voice is calculated to obtain the evaluation data of the common coverage rate of the finals. Compared with comparing the phoneme features of the voiceprint verification voice and the voiceprint registration voice, calculating the evaluation data of the common coverage rate of the finals is simpler and more convenient, thus improving the efficiency of voiceprint recognition.

[0153] In step S606 of some embodiments, after obtaining the coverage rate evaluation data, the distribution feature evaluation data, and the evaluation data of the common coverage rate of the finals, the total evaluation data can be calculated according to the proportion of each evaluation data preset in voiceprint recognition. Further, based on the total evaluation data, it is determined whether the voiceprint verification voice and the voiceprint registration voice are voice data of the same user.

[0154] In steps S601 to S606 shown in the embodiments of the present application, by extracting the finals features of the voiceprint registration voice, the coverage rate of the registration voice finals and the distribution features of the registration voice finals are obtained. Then, the similarity between the coverage rate of the verification voice finals and the coverage rate of the registration voice finals is evaluated to obtain the coverage rate evaluation data, the similarity between the distribution features of the verification voice finals and the distribution features of the registration voice finals is evaluated to obtain the distribution feature evaluation data, and the evaluation data of the common coverage rate of the voiceprint verification voice and the voiceprint registration voice is calculated. Finally, based on the coverage rate evaluation data, the distribution feature evaluation data, and the evaluation data of the common coverage rate of the finals, voiceprint recognition information is generated, reducing the difficulty of comparing voiceprint features in the voiceprint recognition process, thus improving the efficiency of voiceprint recognition.

[0155] This application obtains the voiceprint verification voice and voiceprint registration voice of the target user, performs speech recognition on the voiceprint verification voice to obtain the voiceprint verification text, then based on the voiceprint verification text, performs vowel annotation on the voiceprint verification voice to obtain the verification voice vowel label. Further, based on the verification voice vowel label and the preset total number of basic vowels, calculates the vowel coverage rate of the voiceprint verification voice to obtain the verification voice vowel coverage rate, and extracts the vowel distribution feature of the voiceprint verification voice based on the verification voice vowel label to obtain the verification voice vowel distribution feature, reducing the difficulty of extracting voiceprint features and improving the extraction efficiency of voiceprint features, thereby accelerating the speed of voiceprint recognition. Finally, based on the voiceprint registration voice, the total number of basic vowels, the verification voice vowel coverage rate, and the verification voice vowel distribution feature, performs voiceprint recognition on the target user, converting the comparison of speech phonemes into the comparison of speech vowels during the voiceprint recognition process, reducing the difficulty of speech comparison, thereby improving the speech comparison efficiency and also enhancing the efficiency of voiceprint recognition.

[0156] Please refer to Figure 8 , this embodiment of the application also provides a voiceprint recognition device that can implement the above voiceprint recognition method. The device includes:

[0157] A data acquisition module 801, configured to acquire the voiceprint verification voice and voiceprint registration voice of the target user, where the voiceprint verification voice represents the voice data recorded at the current moment, and the voiceprint registration voice represents the voice data recorded at a historical moment, and the historical moment is before the current moment;

[0158] A speech conversion module 802, configured to perform speech recognition on the voiceprint verification voice to obtain the voiceprint verification text;

[0159] A vowel annotation module 803, configured to perform vowel annotation on the voiceprint verification voice based on the voiceprint verification text to obtain the verification voice vowel label;

[0160] A coverage rate calculation module 804, configured to calculate the vowel coverage rate of the voiceprint verification voice based on the verification voice vowel label and the preset total number of basic vowels to obtain the verification voice vowel coverage rate, where the verification voice vowel coverage rate is used to represent the proportion of the number of types of vowel labels in the total number of basic vowels;

[0161] A distribution feature calculation module 805, configured to extract the vowel distribution feature of the voiceprint verification voice based on the verification voice vowel label to obtain the verification voice vowel distribution feature, where the verification voice vowel distribution feature is used to represent the frequency of occurrence of each verification voice vowel label in the voiceprint verification voice;

[0162] The voiceprint recognition module 806 is configured to perform voiceprint recognition on a target user based on voiceprint-registered speech, the total number of basic vowels, the coverage rate of verification speech vowels, and the distribution characteristics of verification speech vowels.

[0163] The specific implementation manner of this voiceprint recognition device is basically the same as the specific embodiments of the above voiceprint recognition method, and will not be elaborated here.

[0164] An embodiment of this application also provides an electronic device. The electronic device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the above voiceprint recognition method is implemented. This electronic device can be any intelligent terminal including a tablet computer, an in-vehicle computer, etc.

[0165] Please refer to Figure 9 , Figure 9 which schematically shows the hardware structure of an electronic device in another embodiment. The electronic device includes:

[0166] A processor 901, which can be implemented in ways such as a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is configured to execute relevant programs to implement the technical solutions provided by the embodiments of this application;

[0167] A memory 902, which can be implemented in forms such as a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 902 can store an operating system and other application programs. When implementing the technical solutions provided by the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 902, and are called by the processor 901 to execute the voiceprint recognition method of the embodiments of this application;

[0168] An input / output interface 903, which is configured to implement information input and output;

[0169] A communication interface 904, which is configured to implement communication interaction between this device and other devices, and can implement communication through a wired manner (such as USB, network cable, etc.) or through a wireless manner (such as mobile network, WIFI, Bluetooth, etc.);

[0170] A bus 905, which transmits information between various components of the device (such as the processor 901, the memory 902, the input / output interface 903, and the communication interface 904);

[0171] Among them, the processor 901, the memory 902, the input / output interface 903, and the communication interface 904 are communicatively connected to each other inside the device through the bus 905.

[0172] An embodiment of the present application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the above-mentioned voiceprint recognition method is implemented.

[0173] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory may optionally include a memory remotely disposed relative to the processor, and these remote memories may be connected to the processor through a network. Examples of the above-mentioned network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0174] The voiceprint recognition method, voiceprint recognition device, electronic device, and storage medium provided by the embodiments of the present application obtain a voiceprint verification voice and a voiceprint registration voice of a target user, perform voice recognition on the voiceprint verification voice to obtain a voiceprint verification text. Further, based on the voiceprint verification text, perform vowel annotation on the voiceprint verification voice to obtain a verification voice vowel label. Secondly, based on the verification voice vowel label and the preset total number of basic vowels, calculate the vowel coverage rate of the voiceprint verification voice to obtain a verification voice vowel coverage rate, and based on the verification voice vowel label, extract the vowel distribution feature of the voiceprint verification voice to obtain a verification voice vowel distribution feature. Finally, perform voiceprint recognition on the target user based on the voiceprint registration voice, the total number of basic vowels, the verification voice vowel coverage rate, and the verification voice vowel distribution feature.

[0175] The embodiments described in the embodiments of the present application are for more clearly illustrating the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art can know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.

[0176] Those skilled in the art can understand that the technical solutions shown in the figure do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than shown in the figure, or combine certain steps, or different steps.

[0177] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0178] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, and appropriate combinations thereof.

[0179] It should be understood that in this application, the terms "first", "second", "third", "fourth", etc. (if any) in the specification of this application and the above drawings are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of this application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0180] It should be understood that in this application, "at least one (item)" means one or more, and "multiple" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that there can be three relationships. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or similar expressions refer to any combination of these items, including any combination of single item (one) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0181] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the above-mentioned units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. The displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of devices or units can be in electrical, mechanical or other forms.

[0182] The units described above as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0183] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0184] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in each embodiment of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs that can store programs.

[0185] The preferred embodiments of the embodiments of the present application have been described above with reference to the drawings, and thus do not limit the scope of the rights of the embodiments of the present application. Any modifications, equivalent replacements, and improvements made by those skilled in the art without departing from the scope and essence of the embodiments of the present application shall be within the scope of the rights of the embodiments of the present application.

Claims

1. A voiceprint recognition method, characterized in that, The method includes: Obtaining the voiceprint verification voice and the voiceprint registration voice of the target user, where the voiceprint verification voice represents the voice data recorded at the current moment, the voiceprint registration voice represents the voice data recorded at a historical moment before the current moment; Performing speech recognition on the voiceprint verification voice to obtain a voiceprint verification text; Based on the voiceprint verification text, performing vowel annotation on the voiceprint verification voice to obtain a verification voice vowel label; Based on the verification voice vowel label and the preset total number of basic vowels, calculating the vowel coverage rate of the voiceprint verification voice to obtain a verification voice vowel coverage rate, where the verification voice vowel coverage rate is used to represent the proportion of the number of types of vowel labels in the total number of basic vowels; Based on the verification voice vowel label, extracting the vowel distribution feature of the voiceprint verification voice to obtain a verification voice vowel distribution feature, where the verification voice vowel distribution feature is used to represent the frequency of occurrence of each verification voice vowel label in the voiceprint verification voice; Performing voiceprint recognition on the target user based on the voiceprint registration voice, the total number of basic vowels, the verification voice vowel coverage rate, and the verification voice vowel distribution feature.

2. The method according to claim 1, wherein The performing voiceprint recognition on the target user based on the voiceprint registration voice, the total number of basic vowels, the verification voice vowel coverage rate, and the verification voice vowel distribution feature includes: Performing vowel annotation on the voiceprint registration voice to obtain a registration voice vowel label; Based on the registration voice vowel label, extracting features from the voiceprint registration voice to obtain a registration voice vowel coverage rate and the registration voice vowel distribution feature; Based on the verification voice vowel coverage rate and the registration voice vowel coverage rate, performing similarity evaluation on the target user to obtain coverage rate evaluation data; Based on the verification voice vowel distribution feature and the registration voice vowel distribution feature, performing similarity evaluation on the target user to obtain distribution feature evaluation data; Based on the verification voice vowel label, the registration voice vowel label, and the total number of basic vowels, performing common coverage rate evaluation on the voiceprint verification voice and the voiceprint registration voice to obtain vowel common coverage rate evaluation data; Generating voiceprint recognition information based on the coverage rate evaluation data, the distribution feature evaluation data, and the vowel common coverage rate evaluation data.

3. The method according to claim 2, wherein The performing common coverage rate evaluation on the voiceprint verification voice and the voiceprint registration voice based on the verification voice vowel label, the registration voice vowel label, and the total number of basic vowels to obtain vowel common coverage rate evaluation data includes: Searching for the same labels in the verification voice vowel label and the registration voice vowel label to obtain the same vowel label data; Based on the same vowel label data and the total number of basic vowels, calculating the coverage rate of the voiceprint verification voice and the voiceprint registration voice to obtain the vowel common coverage rate evaluation data.

4. The method according to claim 1, wherein Calculating the vowel coverage rate of the voiceprint verification voice based on the verified voice vowel label and the preset total number of basic vowels to obtain the verified voice vowel coverage rate, including: Based on the verified voice vowel label, identifying the vowel type of the voiceprint verification voice to obtain the verified voice vowel type data; Based on the total number of basic vowels, calculating the proportion of the verified voice vowel type data to obtain the verified voice vowel coverage rate.

5. The method according to claim 1, wherein Extracting the vowel distribution feature of the voiceprint verification voice based on the verified voice vowel label to obtain the verified voice vowel distribution feature, including: Calculating the vowel occurrence frequency of the voiceprint verification voice based on the verified voice vowel label to obtain the verified voice vowel occurrence frequency; Performing information entropy calculation and processing on the verified voice vowel occurrence frequency to obtain the verified voice vowel distribution feature.

6. The method according to any one of claims 1 to 5, characterized in that, Labeling the vowel of the voiceprint verification voice based on the voiceprint verification text to obtain the verified voice vowel label, including: Based on the voiceprint verification voice, identifying the single-character voice duration of the voiceprint verification text to obtain the single-character timestamp data; Identifying the vowel of the voiceprint verification text to obtain the basic vowel phoneme; Based on the single-character timestamp data, splitting the voiceprint verification voice to obtain the single-character voice data; Based on the basic vowel phoneme, labeling the vowel of the single-character voice data to obtain the verified voice vowel label.

7. The method according to any one of claims 1-5, characterized in that, Performing voice recognition on the voiceprint verification voice to obtain the voiceprint verification text, including: Extracting the feature of the voiceprint verification voice to obtain the verified voice feature; Performing decoding processing on the verified voice feature to obtain the voiceprint verification text.

8. A voiceprint recognition device, characterized in that, The device includes: A data acquisition module, configured to acquire the voiceprint verification voice and the voiceprint registration voice of the target user, where the voiceprint verification voice represents the voice data recorded at the current moment, and the voiceprint registration voice represents the voice data recorded at a historical moment, and the historical moment is before the current moment; A voice conversion module, configured to perform voice recognition on the voiceprint verification voice to obtain the voiceprint verification text; A vowel labeling module, configured to label the vowel of the voiceprint verification voice based on the voiceprint verification text to obtain the verified voice vowel label; A coverage rate calculation module, configured to calculate the vowel coverage rate of the voiceprint verification voice based on the verified voice vowel label and the preset total number of basic vowels to obtain the verified voice vowel coverage rate, where the verified voice vowel coverage rate is used to represent the proportion of the number of types of vowel labels in the total number of basic vowels; A distribution feature calculation module, configured to extract the vowel distribution feature of the voiceprint verification voice based on the verified voice vowel label to obtain the verified voice vowel distribution feature, where the verified voice vowel distribution feature is used to represent the occurrence frequency of each verified voice vowel label in the voiceprint verification voice; A voiceprint recognition module, configured to perform voiceprint recognition on the target user based on the voiceprint-registered speech, the total number of basic finals, the coverage rate of the finals in the verification speech, and the distribution characteristics of the finals in the verification speech.

9. An electronic device, characterized in that, The electronic device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the voiceprint recognition method according to any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, the voiceprint recognition method according to any one of claims 1 to 7 is implemented.