Voiceprint recognition method, device and equipment and storage medium
By integrating voiceprint features with user attributes, state, and speech rate features into voiceprint recognition technology, the problem of insufficient accuracy in voiceprint recognition in security and control scenarios is solved, the recognition accuracy is improved, and security is ensured.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SOUNDAI TECH CO LTD
- Filing Date
- 2022-12-26
- Publication Date
- 2026-04-17
AI Technical Summary
In security and control scenarios, the accuracy of existing voiceprint recognition technology is insufficient, which can easily lead to economic losses or public security problems.
By acquiring audio data and extracting voiceprint features, and combining user attribute analysis, user status analysis, and user speech rate analysis, a fused voiceprint feature is obtained. The voiceprint database is used for recognition, and auxiliary information such as gender, age, and speech rate are considered to help judge the voiceprint recognition result.
It significantly improves the accuracy of voiceprint recognition, ensures the safe operation of security and prevention scenarios, and avoids economic losses or public security problems.
Smart Images

Figure CN116246635B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence, and in particular to a voiceprint recognition method, apparatus, device, and storage medium. Background Technology
[0002] In security-related scenarios such as smart elevators and banking transactions, voiceprint recognition technology is often used to verify user identity. A voiceprint is a sound wave spectrum carrying language information, and voiceprint recognition, also known as speaker identification, is a biometric technology that identifies a speaker through their voice.
[0003] In particular, security and control scenarios typically require high accuracy in voiceprint recognition. This is because, in such scenarios, errors in voiceprint recognition can lead to significant economic losses or serious public security problems. Therefore, improving the accuracy of voiceprint recognition is essential for these situations. Summary of the Invention
[0004] This application provides a voiceprint recognition method, apparatus, device, and storage medium, which can improve the accuracy of voiceprint recognition. The technical solution is as follows:
[0005] On the one hand, a voiceprint recognition method is provided, the method comprising:
[0006] Acquire audio data collected in the current scene;
[0007] Based on a voiceprint model that matches the current scene, voiceprint features are extracted from the audio data to obtain the original voiceprint features of the audio data.
[0008] User attribute analysis, user status analysis, and user speech rate analysis are performed on the audio data respectively to obtain the user attribute features, user status features, and user speech rate features of the audio data.
[0009] The original voiceprint features, user attribute features, user state features, and user speech rate features of the audio data are fused to obtain the fused voiceprint features of the audio data.
[0010] Based on the voiceprint database, the fused voiceprint features of the audio data are identified to obtain the voiceprint recognition result; wherein, the voiceprint database is used to store the fused voiceprint features of registered users.
[0011] In one possible implementation, the step of performing user attribute analysis, user state analysis, and user speech rate analysis on the audio data to obtain user attribute features, user state features, and user speech rate features of the audio data includes:
[0012] The audio data is analyzed to determine the user age characteristics of the audio data.
[0013] Perform user gender analysis on the audio data to obtain the user gender characteristics of the audio data;
[0014] User sentiment analysis is performed on the audio data to obtain the user sentiment characteristics of the audio data;
[0015] User fatigue level analysis is performed on the audio data to obtain the user fatigue level characteristics of the audio data;
[0016] User speech rate analysis is performed on the audio data to obtain the user speech rate characteristics of the audio data.
[0017] In one possible implementation, the method further includes:
[0018] Acquire image data captured in the current scene;
[0019] Filter the facial image of the current user from the collected image data;
[0020] The facial image is subjected to eye state recognition to obtain an image recognition result; wherein, the image recognition result is used to indicate the eye opening and closing state of the current user.
[0021] Based on the image recognition results, the user fatigue level characteristics are corrected to obtain the corrected user fatigue level characteristics.
[0022] In one possible implementation, the extraction of voiceprint features from the audio data based on a voiceprint model matching the current acquisition scene includes:
[0023] In response to the current scene being a quiet scene, the audio data is subjected to voiceprint feature extraction based on a first voiceprint model that matches the quiet scene;
[0024] In response to the current scene being a noisy scene, the audio data is subjected to voiceprint feature extraction based on a second voiceprint model that matches the noisy scene.
[0025] In one possible implementation, the method further includes:
[0026] For any audio frame in the audio data, obtain the energy of the audio frame;
[0027] The ratio of the energy of the audio frame to the reference energy of the noise is used as the signal-to-noise ratio of the audio frame.
[0028] In response to the signal-to-noise ratio of the audio frame being greater than a first threshold, the audio frame is determined to be a speech frame;
[0029] In response to the signal-to-noise ratio of the audio frame being less than the first threshold, the audio frame is determined to be a noise frame;
[0030] If the number of audio frames is greater than the number of noise frames, the current scene is determined to be the quiet scene.
[0031] If the number of audio frames is less than the number of noise frames, the current scene is determined to be the noisy scene.
[0032] In one possible implementation, the feature fusion of the original voiceprint features, user attribute features, user state features, and user speech rate features of the audio data to obtain the fused voiceprint features of the audio data includes:
[0033] The original voiceprint features, user attribute features, user status features, and user speech rate features of the audio data are concatenated to obtain the fused voiceprint features of the audio data.
[0034] In one possible implementation, the step of identifying the fused voiceprint features of the audio data based on a voiceprint database to obtain a voiceprint recognition result includes:
[0035] Obtain the similarity score between the fused voiceprint features of the audio data and the fused voiceprint features stored in the voiceprint database;
[0036] In response to determining, based on the obtained similarity score, that the voiceprint database stores fused voiceprint features with a similarity score greater than a second threshold, it is determined that the current user is identified through voiceprint recognition.
[0037] If the similarity score between the fused voiceprint feature of the audio data and the fused voiceprint feature of the target registered user is the highest, the current user is identified as the target registered user.
[0038] On the other hand, a voiceprint recognition device is provided, the device comprising:
[0039] The acquisition module is configured to acquire audio data collected in the current scene;
[0040] The first extraction module is configured to extract voiceprint features from the audio data based on a voiceprint model that matches the current scene, thereby obtaining the original voiceprint features of the audio data.
[0041] The second extraction module is configured to perform user attribute analysis, user status analysis and user speech rate analysis on the audio data respectively, to obtain the user attribute features, user status features and user speech rate features of the audio data.
[0042] The fusion module is configured to perform feature fusion on the original voiceprint features, user attribute features, user state features and user speech rate features of the audio data to obtain the fused voiceprint features of the audio data.
[0043] The recognition module is configured to recognize the fused voiceprint features of the audio data based on a voiceprint library to obtain a voiceprint recognition result; wherein, the voiceprint library is used to store the fused voiceprint features of registered users.
[0044] In one possible implementation, the second extraction module is configured as follows:
[0045] The audio data is analyzed to determine the user age characteristics of the audio data.
[0046] Perform user gender analysis on the audio data to obtain the user gender characteristics of the audio data;
[0047] User sentiment analysis is performed on the audio data to obtain the user sentiment characteristics of the audio data;
[0048] User fatigue level analysis is performed on the audio data to obtain the user fatigue level characteristics of the audio data;
[0049] User speech rate analysis is performed on the audio data to obtain the user speech rate characteristics of the audio data.
[0050] In one possible implementation, the second extraction module is further configured as follows:
[0051] Acquire image data captured in the current scene;
[0052] Filter the facial image of the current user from the collected image data;
[0053] The facial image is subjected to eye state recognition to obtain an image recognition result; wherein, the image recognition result is used to indicate the eye opening and closing state of the current user.
[0054] Based on the image recognition results, the user fatigue level characteristics are corrected to obtain the corrected user fatigue level characteristics.
[0055] In one possible implementation, the first extraction module is configured as follows:
[0056] In response to the current scene being a quiet scene, the audio data is subjected to voiceprint feature extraction based on a first voiceprint model that matches the quiet scene;
[0057] In response to the current scene being a noisy scene, the audio data is subjected to voiceprint feature extraction based on a second voiceprint model that matches the noisy scene.
[0058] In one possible implementation, the first extraction module is further configured as follows:
[0059] For any audio frame in the audio data, obtain the energy of the audio frame;
[0060] The ratio of the energy of the audio frame to the reference energy of the noise is used as the signal-to-noise ratio of the audio frame.
[0061] In response to the signal-to-noise ratio of the audio frame being greater than a first threshold, the audio frame is determined to be a speech frame;
[0062] In response to the signal-to-noise ratio of the audio frame being less than the first threshold, the audio frame is determined to be a noise frame;
[0063] If the number of audio frames is greater than the number of noise frames, the current scene is determined to be the quiet scene.
[0064] If the number of audio frames is less than the number of noise frames, the current scene is determined to be the noisy scene.
[0065] In one possible implementation, the fusion module is configured as follows:
[0066] The original voiceprint features, user attribute features, user status features, and user speech rate features of the audio data are concatenated to obtain the fused voiceprint features of the audio data.
[0067] In one possible implementation, the identification module is configured as follows:
[0068] Obtain the similarity score between the fused voiceprint features of the audio data and the fused voiceprint features stored in the voiceprint database;
[0069] In response to determining, based on the obtained similarity score, that the voiceprint database stores fused voiceprint features with a similarity score greater than a second threshold, it is determined that the current user is identified through voiceprint recognition.
[0070] If the similarity score between the fused voiceprint feature of the audio data and the fused voiceprint feature of the target registered user is the highest, the current user is identified as the target registered user.
[0071] On the other hand, a computer device is provided, the device including a processor and a memory, the memory storing at least one piece of program code, the at least one piece of program code being loaded and executed by the processor to implement the above-described voiceprint recognition method.
[0072] On the other hand, a computer-readable storage medium is provided, wherein at least one piece of program code is stored in the storage medium, the at least one piece of program code being loaded and executed by a processor to implement the above-described voiceprint recognition method.
[0073] On the other hand, a computer program product or computer program is provided, which includes computer program code stored in a computer-readable storage medium. A processor of a computer device reads the computer program code from the computer-readable storage medium and executes the computer program code, causing the computer device to perform the aforementioned voiceprint recognition method.
[0074] The voiceprint recognition method provided in this application can improve the accuracy of voiceprint recognition.
[0075] In detail, the process first acquires audio data collected in the current scene. Then, based on a voiceprint model matching the current scene, voiceprint features are extracted from the audio data to obtain the original voiceprint features. Further, in addition to voiceprint features, this embodiment also performs user attribute analysis, user state analysis, and user speech rate analysis on the audio data to obtain user attribute features, user state features, and user speech rate features. Next, the original voiceprint features, user attribute features, user state features, and user speech rate features of the audio data are fused to obtain the fused voiceprint features. Finally, based on a voiceprint database, the fused voiceprint features of the audio data are identified to obtain the voiceprint recognition result. Because this voiceprint recognition scheme not only considers voiceprint features but also additional auxiliary information such as user attribute features, user state features, and user speech rate features to assist in judging the voiceprint recognition result, it significantly improves the accuracy of voiceprint recognition and is suitable for scenarios with high requirements for voiceprint recognition accuracy, such as security and control. In particular, for such scenarios, the high accuracy of this voiceprint recognition solution can avoid serious economic losses or security problems, ensuring the safe operation of security and prevention scenarios. Attached Figure Description
[0076] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0077] Figure 1 This is a schematic diagram of the implementation environment of a voiceprint recognition method provided in an embodiment of this application;
[0078] Figure 2This is a schematic diagram of the overall execution flow of a voiceprint recognition method provided in an embodiment of this application;
[0079] Figure 3 This is a flowchart of a voiceprint recognition method provided in an embodiment of this application;
[0080] Figure 4 This is a schematic diagram of the structure of a voiceprint recognition device provided in an embodiment of this application;
[0081] Figure 5 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation
[0082] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.
[0083] In this application, the terms "first," "second," etc., are used to distinguish identical or similar items that have essentially the same function. It should be understood that there is no logical or temporal dependency between "first," "second," and "nth," nor does it limit the quantity or execution order. It should also be understood that although the following description uses the terms "first," "second," etc., to describe various elements, these elements should not be limited by the terms.
[0084] These terms are simply used to distinguish one element from another. For example, without departing from the scope of various examples, the first element can be referred to as the second element, and similarly, the second element can be referred to as the first element. Both the first and second elements can be elements, and in some cases, they can be separate and distinct elements.
[0085] "At least one" refers to one or more elements. For example, at least one element can be one element, two elements, three elements, or any integer number of elements greater than or equal to one. "Multiple" refers to two or more elements. For example, multiple elements can be two elements, three elements, or any integer number of elements greater than or equal to two.
[0086] In this article, "and / or" indicates that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. The character " / " generally indicates that the preceding and following objects are in an "or" relationship.
[0087] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, data stored, data displayed, etc.) and signals involved in this application are all authorized by the user or fully authorized by all parties, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.
[0088] The voiceprint recognition method provided in this application is applied to a voiceprint recognition device.
[0089] See Figure 1 The voiceprint recognition device 101 is a computer device with machine learning capabilities, such as a tablet computer or a smartphone; this application does not limit the scope of the application to this type of device. For example, this voiceprint recognition method can be applied in security-related scenarios such as smart elevators and bank transactions.
[0090] Among the metrics used to determine voiceprint recognition are the False Rejection Rate (FRR) and the False Acceptance Rate (FAR). FRR is the ratio of the number of samples that are incorrectly rejected to the number of samples that should be accepted, while FAR is the ratio of the number of samples that are incorrectly accepted to the number of samples that should be rejected.
[0091] In detail, if two samples are of the same type (the same person), but are mistakenly identified as different types (not the same person), they are considered falsely rejected samples (samples that should not be rejected but are rejected). For example, if you cannot authenticate your phone using your fingerprint, this is called a false rejection. Conversely, if two samples are of different types, but are mistakenly identified as the same type, they are considered falsely accepted samples (samples that should not be accepted but are accepted). For example, if a stranger finds your phone and successfully unlocks it using their fingerprint, this is called a false acceptance.
[0092] It should be noted that the performance of a voiceprint model is typically measured by a combined metric of FRR and FAR, known as the "equi-error rate." A lower equi-error rate indicates better performance. For example, after the collected audio data is fed into the voiceprint model to extract voiceprint features, the model proceeds to the pattern recognition step. During pattern recognition, the similarity between the extracted voiceprint features and those stored in the voiceprint database can be calculated, and the voiceprint recognition result is determined based on the obtained similarity score.
[0093] In particular, security and control scenarios typically require high accuracy in voiceprint recognition. This is because, in such scenarios, errors in voiceprint recognition can lead to significant economic losses or serious public security problems. Therefore, improving the accuracy of voiceprint recognition is crucial for these scenarios. For example, in speaker verification, higher similarity scores for voiceprint features are desirable in these scenarios, thus favoring models with lower FAR (Failure to Recognize and Identify) scores.
[0094] In view of the above problems, embodiments of this application provide a new voiceprint recognition scheme, such as... Figure 2 As shown, this voiceprint recognition scheme extracts additional features such as gender, age, and speech rate on top of the extracted voiceprint features. These features are then fused to obtain fused voiceprint features. Furthermore, the features stored in the voiceprint database are also fused voiceprint features; therefore, the voiceprint database is also called a fused voiceprint feature database. Correspondingly, speaker confirmation or identification is based on the fused voiceprint features during pattern recognition.
[0095] In summary, the voiceprint recognition scheme provided in this application applies multiple models such as gender classification model, age discrimination model, speech rate discrimination model, and voiceprint model to the voiceprint recognition system. After the collected audio data is sent into the above models, the voiceprint features, gender features, age features, and speech rate features corresponding to the audio data can be obtained. By fusing the above features, a comprehensive feature containing information such as gender, age, speech rate, and voiceprint can be obtained (referred to as fused voiceprint feature in this document). The fused voiceprint feature is used to establish a voiceprint database and to perform speaker confirmation or speaker identification.
[0096] Furthermore, since the fusion of voiceprint features not only includes voiceprint characteristics but also incorporates information such as gender, age, and speech rate, the similarity score of voiceprint features from individuals of the same gender and similar age will be higher than that of voiceprint features from individuals of different genders and significantly different ages. This is because the pronunciation characteristics, pitch, frequency, and other features of men and women, as well as different age groups, are not the same.
[0097] Since the voiceprint recognition scheme provided in this application not only considers the similarity of voiceprints, but also uses auxiliary information such as age, gender, and speech rate to assist in judging the voiceprint recognition result, it significantly improves the accuracy of voiceprint recognition and ensures the safe operation of scenarios such as security and prevention.
[0098] Figure 3 This is a flowchart illustrating a voiceprint recognition method provided in an embodiment of this application. The method is executed by a computer device, such as a voiceprint recognition device. See also... Figure 3 The method process includes:
[0099] 301. Obtain the audio data collected in the current scene.
[0100] For example, the aforementioned computer equipment may be a smart terminal used by the user, such as a smartphone or tablet computer; or, the aforementioned computer equipment may also be a voiceprint recognition device built into a smart elevator; or, the aforementioned computer equipment may also be a business processing device with voiceprint recognition function placed in a bank lobby, and this application does not limit it in this way.
[0101] In this embodiment, audio data can be collected using a built-in microphone (such as a microphone) in the aforementioned computer device. Furthermore, to ensure that usable voiceprint features can be extracted, the duration of the collected audio data should not be less than a preset duration, such as 3 seconds or 5 seconds; however, this application does not impose any limitation on this.
[0102] 302. Based on the voiceprint model that matches the current scene, extract the voiceprint features from the audio data to obtain the original voiceprint features of the audio data.
[0103] For example, embodiments of this application divide the current scene into a quiet scene or a noisy scene.
[0104] One possible implementation includes, but is not limited to, determining whether the current scene is a quiet scene or a noisy scene using the following methods:
[0105] For any audio frame in the audio data, obtain the energy of the audio frame; use the ratio of the energy of the audio frame to the reference energy of the noise as the signal-to-noise ratio (SNR) of the audio frame; if the SNR of the audio frame is greater than a first threshold, determine that the audio frame is a speech frame; if the SNR of the audio frame is less than the first threshold, determine that the audio frame is a noise frame; if the number of speech frames is greater than the number of noise frames, determine the current scene as a quiet scene; if the number of speech frames is less than the number of noise frames, determine the current scene as a noisy scene.
[0106] It should be noted that, for ease of distinction, this paper refers to the signal-to-noise ratio threshold here as the first threshold, and the similarity threshold appearing later as the second threshold.
[0107] For example, each audio frame corresponds to an energy value, such as the root mean square energy of the audio signal, representing the average energy of the audio signal waveform over a short period of time. Alternatively, a noise estimation algorithm can be used to estimate the energy of the noise (referred to as the reference energy in this paper); for example, this noise estimation algorithm could be a minimum value tracking algorithm.
[0108] Accordingly, based on the voiceprint model matching the current acquisition scenario, voiceprint features are extracted from the audio data, including but not limited to the following methods:
[0109] In response to the current scene being a quiet scene, voiceprint features are extracted from the audio data based on a first voiceprint model that matches the quiet scene; in response to the current scene being a noisy scene, voiceprint features are extracted from the audio data based on a second voiceprint model that matches the noisy scene.
[0110] In another possible implementation, quiet scenes and noisy scenes can be further subdivided into multiple levels, resulting in multiple scene types, which is not limited in this application.
[0111] In another possible implementation, the training process for the first voiceprint model matching a quiet scene can be as follows: acquiring a first sample audio set collected in a quiet scene; preprocessing each sample audio in the first sample audio set; and training the model using a pre-training fine-tuning method based on the preprocessed first sample audio set to obtain the first voiceprint model. Similarly, the training process for the second voiceprint model matching a noisy scene can be as follows: acquiring a second sample audio set collected in a noisy scene; preprocessing each sample audio in the second sample audio set; and training the model using a pre-training fine-tuning method based on the preprocessed second sample audio set to obtain the second voiceprint model.
[0112] For example, the above preprocessing methods include, but are not limited to: frame segmentation, windowing, noise reduction, and removal of silent segments in each sample audio, which are not limited in this application.
[0113] 303. Perform user attribute analysis, user status analysis, and user speech rate analysis on the audio data respectively to obtain the user attribute characteristics, user status characteristics, and user speech rate characteristics of the audio data.
[0114] For example, the aforementioned user attributes include, but are not limited to, gender, age, etc., and the aforementioned user status includes, but is not limited to, emotions, fatigue level, etc., which are not limited in this application.
[0115] In this embodiment of the application, a large number of sample audios with different user attributes, different user states, and different user speech rates are collected to train models such as convolutional neural networks, support vector machines (SVM), and random forest trees (RFT) to extract user attribute features, user state features, and user speech rate features of the audio data to be identified.
[0116] Accordingly, user attribute analysis, user status analysis, and user speech rate analysis are performed on the audio data respectively to obtain the user attribute characteristics, user status characteristics, and user speech rate characteristics of the audio data, including but not limited to the following methods:
[0117] Based on the age discrimination model, user age analysis was performed on the audio data to obtain the user age characteristics of the audio data.
[0118] Based on the gender classification model, user gender analysis was performed on the audio data to obtain the user gender characteristics of the audio data.
[0119] Based on the emotion discrimination model, user emotion analysis is performed on the audio data to obtain the user emotion characteristics of the audio data.
[0120] Based on the fatigue level discrimination model, user fatigue level analysis was performed on the audio data to obtain the user fatigue level characteristics of the audio data.
[0121] Based on the speech rate discrimination model, user speech rate analysis is performed on the audio data to obtain the user speech rate characteristics of the audio data.
[0122] In one possible implementation, taking the gender classification model as an example, the training process of the gender classification model will be illustrated below. The training process of other models is similar and will not be repeated.
[0123] Obtain a sample audio set for training the gender classification model. This sample audio set includes multiple male and female voice sample audios. Input the sample audio set into a convolutional neural network and obtain the predicted classification result of the convolutional neural network output for the sample audio set. Determine whether the labeled classification result of the sample audio set is consistent with the predicted classification result. When the labeled classification result is inconsistent with the predicted classification result, iteratively update the weight values of the convolutional neural network repeatedly until the labeled classification result is consistent with the predicted classification result, thus obtaining the gender classification model.
[0124] For example, based on a gender classification model, user gender analysis is performed on the audio data to obtain the user gender characteristics of the audio data, including: inputting the audio data into the gender classification model for feature extraction, and using the output of the penultimate layer of the gender classification model as the user gender characteristics of the audio data.
[0125] In another possible implementation, the embodiments of this application also support the correction of user fatigue levels based on speech recognition, that is, the embodiments of this application further include:
[0126] Image data of the current scene is acquired through an image acquisition device; facial images of the current user are filtered from the acquired image data; eye state recognition is performed on the facial images of the current user to obtain image recognition results; the image recognition results are used to indicate the opening and closing state of the current user's eyes; based on the image recognition results, the above-mentioned user fatigue characteristics are corrected to obtain corrected user fatigue characteristics.
[0127] Exemplary, the aforementioned image acquisition device can be independent of the voiceprint recognition device or integrated with it; this application does not limit this. In the embodiments of this application, if the current user's eye opening and closing status reflects that the user's eye opening time is less than the eye closing time, then it can be determined that the user is in a state of fatigue. Exemplary, the user's degree of fatigue can be determined based on the ratio or difference between the eye opening time and the eye closing time over a certain period of time; this application does not limit this.
[0128] In another possible implementation, the above correction process can be to weight the user fatigue level features obtained from speech and the user fatigue level features obtained from images to obtain the corrected user fatigue level features. For example, the weights corresponding to the user fatigue level features obtained from speech are less than the weights corresponding to the user fatigue level features obtained from images; this application does not limit this.
[0129] 304. Perform feature fusion on the original voiceprint features, user attribute features, user status features, and user speech rate features of the audio data to obtain the fused voiceprint features of the audio data.
[0130] In this embodiment of the application, the original voiceprint features, user attribute features, user state features, and user speech rate features of the audio data are fused to obtain the fused voiceprint features of the audio data, including but not limited to the following methods:
[0131] The original voiceprint features, user attribute features, user status features, and user speech rate features of the audio data are concatenated to obtain the fused voiceprint features of the audio data.
[0132] 305. Based on the voiceprint database, the fused voiceprint features of the audio data are identified to obtain the voiceprint recognition result; wherein, the voiceprint database is used to store the fused voiceprint features of registered users.
[0133] In this embodiment of the application, based on a voiceprint database, the fused voiceprint features of audio data are identified to obtain voiceprint recognition results, including but not limited to the following methods:
[0134] The similarity score between the fused voiceprint features of the audio data and the fused voiceprint features stored in the voiceprint database is obtained; in response to determining that there are fused voiceprint features in the voiceprint database with a similarity score greater than a second threshold based on the obtained similarity score, the current voice user is determined to be identified by voiceprint; in response to the highest similarity score between the fused voiceprint features of the audio data and the fused voiceprint features of the target registered user, the current voice user is determined to be the target registered user.
[0135] The aforementioned similarity can be the cosine similarity between features, and this application does not limit it to this.
[0136] The voiceprint recognition method provided in this application can improve the accuracy of voiceprint recognition. Specifically, it first acquires audio data collected in the current scene, then extracts voiceprint features from the audio data based on a voiceprint model matching the current scene, obtaining the original voiceprint features. Further, in addition to voiceprint features, this application also performs user attribute analysis, user state analysis, and user speech rate analysis on the audio data, thereby obtaining the user attribute features, user state features, and user speech rate features of the audio data. Next, the original voiceprint features, user attribute features, user state features, and user speech rate features of the audio data are fused to obtain the fused voiceprint features of the audio data. Finally, based on a voiceprint database, the fused voiceprint features of the audio data are recognized to obtain the voiceprint recognition result. Because this voiceprint recognition scheme not only considers voiceprint features but also additional auxiliary information such as user attribute features, user state features, and user speech rate features to assist in judging the voiceprint recognition result, it significantly improves the accuracy of voiceprint recognition and is suitable for scenarios with high requirements for voiceprint recognition accuracy, such as security and control. In particular, for such scenarios, the high accuracy of this voiceprint recognition solution can avoid serious economic losses or security problems, ensuring the safe operation of security and prevention scenarios.
[0137] Figure 4 This is a schematic diagram of the structure of a voiceprint recognition device provided in an embodiment of this application. See also... Figure 4 The device includes:
[0138] The acquisition module 401 is configured to acquire audio data collected in the current scene;
[0139] The first extraction module 402 is configured to extract voiceprint features from the audio data based on a voiceprint model that matches the current scene, thereby obtaining the original voiceprint features of the audio data.
[0140] The second extraction module 403 is configured to perform user attribute analysis, user status analysis and user speech rate analysis on the audio data respectively, and obtain the user attribute features, user status features and user speech rate features of the audio data.
[0141] The fusion module 404 is configured to perform feature fusion on the original voiceprint features, user attribute features, user status features and user speech rate features of the audio data to obtain the fused voiceprint features of the audio data.
[0142] The recognition module 405 is configured to recognize the fused voiceprint features of the audio data based on the voiceprint library to obtain the voiceprint recognition result; wherein, the voiceprint library is used to store the fused voiceprint features of registered users.
[0143] The voiceprint recognition method provided in this application can improve the accuracy of voiceprint recognition. Specifically, it first acquires audio data collected in the current scene, then extracts voiceprint features from the audio data based on a voiceprint model matching the current scene, obtaining the original voiceprint features. Further, in addition to voiceprint features, this application also performs user attribute analysis, user state analysis, and user speech rate analysis on the audio data, thereby obtaining the user attribute features, user state features, and user speech rate features of the audio data. Next, the original voiceprint features, user attribute features, user state features, and user speech rate features of the audio data are fused to obtain the fused voiceprint features of the audio data. Finally, based on a voiceprint database, the fused voiceprint features of the audio data are recognized to obtain the voiceprint recognition result. Because this voiceprint recognition scheme not only considers voiceprint features but also additional auxiliary information such as user attribute features, user state features, and user speech rate features to assist in judging the voiceprint recognition result, it significantly improves the accuracy of voiceprint recognition and is suitable for scenarios with high requirements for voiceprint recognition accuracy, such as security and control. In particular, for such scenarios, the high accuracy of this voiceprint recognition solution can avoid serious economic losses or security problems, ensuring the safe operation of security and prevention scenarios.
[0144] In one possible implementation, the second extraction module 403 is configured as follows:
[0145] The audio data is analyzed to determine the user age characteristics of the audio data.
[0146] Perform user gender analysis on the audio data to obtain the user gender characteristics of the audio data;
[0147] User sentiment analysis is performed on the audio data to obtain the user sentiment characteristics of the audio data;
[0148] User fatigue level analysis is performed on the audio data to obtain the user fatigue level characteristics of the audio data;
[0149] User speech rate analysis is performed on the audio data to obtain the user speech rate characteristics of the audio data.
[0150] In one possible implementation, the second extraction module 403 is further configured as follows:
[0151] Acquire image data captured in the current scene;
[0152] Filter the facial image of the current user from the collected image data;
[0153] The facial image is subjected to eye state recognition to obtain an image recognition result; wherein, the image recognition result is used to indicate the eye opening and closing state of the current user.
[0154] Based on the image recognition results, the user fatigue level characteristics are corrected to obtain the corrected user fatigue level characteristics.
[0155] In one possible implementation, the first extraction module 402 is configured as follows:
[0156] In response to the current scene being a quiet scene, the audio data is subjected to voiceprint feature extraction based on a first voiceprint model that matches the quiet scene;
[0157] In response to the current scene being a noisy scene, the audio data is subjected to voiceprint feature extraction based on a second voiceprint model that matches the noisy scene.
[0158] In one possible implementation, the first extraction module 402 is further configured as follows:
[0159] For any audio frame in the audio data, obtain the energy of the audio frame;
[0160] The ratio of the energy of the audio frame to the reference energy of the noise is used as the signal-to-noise ratio of the audio frame.
[0161] In response to the signal-to-noise ratio of the audio frame being greater than a first threshold, the audio frame is determined to be a speech frame;
[0162] In response to the signal-to-noise ratio of the audio frame being less than the first threshold, the audio frame is determined to be a noise frame;
[0163] If the number of audio frames is greater than the number of noise frames, the current scene is determined to be the quiet scene.
[0164] If the number of audio frames is less than the number of noise frames, the current scene is determined to be the noisy scene.
[0165] In one possible implementation, the fusion module 404 is configured as follows:
[0166] The original voiceprint features, user attribute features, user status features, and user speech rate features of the audio data are concatenated to obtain the fused voiceprint features of the audio data.
[0167] In one possible implementation, the identification module 405 is configured as follows:
[0168] Obtain the similarity score between the fused voiceprint features of the audio data and the fused voiceprint features stored in the voiceprint database;
[0169] In response to determining, based on the obtained similarity score, that the voiceprint database stores fused voiceprint features with a similarity score greater than a second threshold, it is determined that the current user is identified through voiceprint recognition.
[0170] If the similarity score between the fused voiceprint feature of the audio data and the fused voiceprint feature of the target registered user is the highest, the current user is identified as the target registered user.
[0171] All of the above-mentioned optional technical solutions can be combined in any way to form optional embodiments of this disclosure, and will not be described in detail here.
[0172] It should be noted that the voiceprint recognition device provided in the above embodiments is only illustrated by the division of the above functional modules. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the voiceprint recognition device and the voiceprint recognition method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.
[0173] Figure 5 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Typically, the computer device 500 includes a processor 501 and a memory 502.
[0174] Processor 501 may include one or more processing cores, such as a quad-core processor or an octa-core processor. Processor 501 may be implemented using at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), or PLA (Programmable Logic Array). Processor 501 may also include a main processor and a coprocessor. The main processor, also known as a CPU (Central Processing Unit), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In one possible implementation, processor 501 may integrate a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the screen. In another possible implementation, processor 501 may also include an AI (Artificial Intelligence) processor, which handles computational operations related to machine learning.
[0175] The memory 502 may include one or more computer-readable storage media, which may be non-transitory. The memory 502 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In one possible implementation, the non-transitory computer-readable storage media in the memory 502 is used to store at least one program code, which is executed by the processor 501 to implement the voiceprint recognition method provided in the method embodiments of this application.
[0176] In one possible implementation, the computer device 500 may also optionally include: a peripheral device interface 503 and at least one peripheral device. The processor 501, memory 502, and peripheral device interface 503 can be connected via a bus or signal line. Each peripheral device can be connected to the peripheral device interface 503 via a bus, signal line, or circuit board. Specifically, the peripheral device includes at least one of: a radio frequency circuit 504, a display screen 505, a camera assembly 506, an audio circuit 507, a positioning assembly 508, and a power supply 509.
[0177] Peripheral device interface 503 can be used to connect at least one I / O (Input / Output) related peripheral device to processor 501 and memory 502. In one possible implementation, processor 501, memory 502, and peripheral device interface 503 are integrated on the same chip or circuit board; in some other embodiments, any one or two of processor 501, memory 502, and peripheral device interface 503 can be implemented on separate chips or circuit boards, which is not limited in this embodiment.
[0178] The radio frequency (RF) circuit 504 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The RF circuit 504 communicates with communication networks and other communication devices via electromagnetic signals. The RF circuit 504 converts electrical signals into electromagnetic signals for transmission, or converts received electromagnetic signals back into electrical signals. Optionally, the RF circuit 504 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, etc. The RF circuit 504 can communicate with other terminals through at least one wireless communication protocol. This wireless communication protocol includes, but is not limited to: the World Wide Web, metropolitan area networks, intranets, various generations of mobile communication networks (2G, 3G, 4G, and 5G), wireless local area networks, and / or WiFi (Wireless Fidelity) networks. In one possible implementation, the RF circuit 504 may also include circuitry related to NFC (Near Field Communication), which is not limited in this application.
[0179] Display screen 505 is used to display a UI (User Interface). This UI can include graphics, text, icons, videos, and any combination thereof. When display screen 505 is a touch display, it also has the ability to collect touch signals on or above its surface. These touch signals can be input as control signals to processor 501 for processing. In this case, display screen 505 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In one possible implementation, there can be one display screen 505, located on the front panel of computer device 500; in another possible implementation, there can be at least two display screens, respectively located on different surfaces of computer device 500 or in a folded design; in yet another possible implementation, display screen 505 can be a flexible display screen, located on a curved or folded surface of computer device 500. Furthermore, display screen 505 can also be configured as a non-rectangular, irregular shape, i.e., a non-rectangular screen. Display screen 505 can be made of materials such as LCD (Liquid Crystal Display) or OLED (Organic Light-Emitting Diode).
[0180] Camera assembly 506 is used to acquire images or videos. Optionally, camera assembly 506 includes a front-facing camera and a rear-facing camera. Typically, the front-facing camera is located on the front panel of the terminal, and the rear-facing camera is located on the back of the terminal. In one possible implementation, there are at least two rear-facing cameras, which are any one of a main camera, a depth-sensing camera, a wide-angle camera, and a telephoto camera, to achieve background blurring by fusion of the main camera and the depth-sensing camera, panoramic shooting by fusion of the main camera and the wide-angle camera, VR (Virtual Reality) shooting, or other fusion shooting functions. In one possible implementation, camera assembly 506 may also include a flash. The flash can be a single-color temperature flash or a dual-color temperature flash. A dual-color temperature flash refers to a combination of a warm-light flash and a cool-light flash, which can be used for light compensation at different color temperatures.
[0181] The audio circuit 507 may include a microphone and a speaker. The microphone is used to collect sound waves from the user and the environment, converting them into electrical signals that are input to the processor 501 for processing, or to the radio frequency circuit 504 for voice communication. For stereo sound acquisition or noise reduction purposes, multiple microphones may be used, positioned at different locations within the computer device 500. The microphone may also be an array microphone or an omnidirectional microphone. The speaker is used to convert electrical signals from the processor 501 or the radio frequency circuit 504 into sound waves. The speaker may be a traditional film speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can convert electrical signals not only into audible sound waves but also into inaudible sound waves for purposes such as distance measurement. In one possible implementation, the audio circuit 507 may also include a headphone jack.
[0182] The positioning component 508 is used to locate the current geographical location of the computer device 500 in order to enable navigation or LBS (Location Based Service). The positioning component 508 can be a positioning component based on the US GPS (Global Positioning System), China's BeiDou system, Russia's Granas system, or the European Union's Galileo system.
[0183] Power supply 509 is used to supply power to the various components in computer device 500. Power supply 509 can be AC power, DC power, a disposable battery, or a rechargeable battery. When power supply 509 includes a rechargeable battery, the rechargeable battery can be a wired rechargeable battery or a wireless rechargeable battery. A wired rechargeable battery is a battery that is charged via a wired line, and a wireless rechargeable battery is a battery that is charged via a wireless coil. The rechargeable battery can also be used to support fast charging technology.
[0184] In one possible implementation, the computer device 500 further includes one or more sensors 510. The one or more sensors 510 include, but are not limited to: an accelerometer 511, a gyroscope 512, a pressure sensor 513, a fingerprint sensor 514, an optical sensor 515, and a proximity sensor 516.
[0185] Accelerometer 511 can detect the magnitude of acceleration along the three coordinate axes of a coordinate system established by computer device 500. For example, accelerometer 511 can be used to detect the components of gravitational acceleration along the three coordinate axes. Processor 501 can control display screen 505 to display the user interface in either a landscape or portrait view based on the gravitational acceleration signal acquired by accelerometer 511. Accelerometer 511 can also be used for games or for acquiring user motion data.
[0186] The gyroscope sensor 512 can detect the orientation and rotation angle of the computer device 500. The gyroscope sensor 512, in conjunction with the accelerometer sensor 511, can collect 3D motion data from the user on the computer device 500. Based on the data collected by the gyroscope sensor 512, the processor 501 can perform the following functions: motion sensing (e.g., changing the UI based on the user's tilt), image stabilization during shooting, game control, and inertial navigation.
[0187] The pressure sensor 513 can be disposed on the side bezel of the computer device 500 and / or on the lower layer of the display screen 505. When the pressure sensor 513 is disposed on the side bezel of the computer device 500, it can detect the user's grip signal on the computer device 500, and the processor 501 can perform left / right hand recognition or quick operation based on the grip signal collected by the pressure sensor 513. When the pressure sensor 513 is disposed on the lower layer of the display screen 505, the processor 501 can control the operable controls on the UI interface based on the user's pressure operation on the display screen 505. The operable controls include at least one of button controls, scroll bar controls, icon controls, and menu controls.
[0188] The fingerprint sensor 514 is used to collect the user's fingerprint. The processor 501 identifies the user's identity based on the fingerprint collected by the fingerprint sensor 514, or the fingerprint sensor 514 identifies the user's identity based on the collected fingerprint. When the user's identity is identified as trusted, the processor 501 authorizes the user to perform relevant sensitive operations, including unlocking the screen, viewing encrypted information, downloading software, making payments, and changing settings. The fingerprint sensor 514 can be located on the front, back, or side of the computer device 500. When the computer device 500 has physical buttons or a manufacturer's logo, the fingerprint sensor 514 can be integrated with the physical buttons or manufacturer's logo.
[0189] An optical sensor 515 is used to collect ambient light intensity. In one embodiment, the processor 501 can control the display brightness of the display screen 505 based on the ambient light intensity collected by the optical sensor 515. Specifically, when the ambient light intensity is high, the display brightness of the display screen 505 is increased; when the ambient light intensity is low, the display brightness of the display screen 505 is decreased. In another embodiment, the processor 501 can also dynamically adjust the shooting parameters of the camera assembly 506 based on the ambient light intensity collected by the optical sensor 515.
[0190] The proximity sensor 516, also known as a distance sensor, is typically located on the front panel of the computer device 500. The proximity sensor 516 is used to detect the distance between the user and the front of the computer device 500. In one embodiment, when the proximity sensor 516 detects that the distance between the user and the front of the computer device 500 is gradually decreasing, the processor 501 controls the display screen 505 to switch from a screen-on state to a screen-off state; when the proximity sensor 516 detects that the distance between the user and the front of the computer device 500 is gradually increasing, the processor 501 controls the display screen 505 to switch from a screen-off state to a screen-on state.
[0191] Those skilled in the art will understand that Figure 5 The structure shown does not constitute a limitation on the computer device 500, and may include more or fewer components than shown, or combine certain components, or use different component arrangements.
[0192] In an exemplary embodiment, a computer-readable storage medium is also provided, such as a memory including program code that can be executed by a processor in a computer device to perform the voiceprint recognition method in the above embodiments. For example, the computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CD-ROM), magnetic tape, floppy disk, and optical data storage device, etc.
[0193] In an exemplary embodiment, a computer program product or computer program is also provided, which includes computer program code stored in a computer-readable storage medium. The processor of a computer device reads the computer program code from the computer-readable storage medium and executes the computer program code, causing the computer device to perform the above-described voiceprint recognition method.
[0194] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.
[0195] The above description is merely an optional embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. A voiceprint recognition method, characterized in that, The method includes: Acquire audio data collected in the current scene; In response to the current scene being a quiet scene, based on a first voiceprint model that matches the quiet scene, voiceprint features are extracted from the audio data to obtain the original voiceprint features of the audio data; In response to the current scene being a noisy scene, based on the second voiceprint model matching the noisy scene, voiceprint features are extracted from the audio data to obtain the original voiceprint features of the audio data; User attribute analysis, user status analysis, and user speech rate analysis are performed on the audio data respectively to obtain the user attribute features, user status features, and user speech rate features of the audio data. The original voiceprint features, user attribute features, user state features, and user speech rate features of the audio data are fused to obtain the fused voiceprint features of the audio data. Based on the voiceprint database, the fused voiceprint features of the audio data are identified to obtain the voiceprint recognition result; wherein, the voiceprint database is used to store the fused voiceprint features of registered users; The step of performing user attribute analysis, user state analysis, and user speech rate analysis on the audio data to obtain user attribute features, user state features, and user speech rate features of the audio data includes: The audio data is analyzed to determine the user age characteristics of the audio data. Perform user gender analysis on the audio data to obtain the user gender characteristics of the audio data; User sentiment analysis is performed on the audio data to obtain the user sentiment characteristics of the audio data; The audio data is analyzed to determine user fatigue level characteristics; image data collected in the current scene is acquired; facial images of the current speaker are filtered from the collected image data; eye state recognition is performed on the facial images to obtain image recognition results; wherein, the image recognition results are used to indicate the eye opening and closing state of the current speaker; based on the image recognition results, the user fatigue level characteristics are corrected to obtain corrected user fatigue level characteristics. Perform user speech rate analysis on the audio data to obtain the user speech rate characteristics of the audio data; The process of determining the quiet scene and the noisy scene includes: For any audio frame in the audio data, obtain the energy of the audio frame; The ratio of the energy of the audio frame to the reference energy of the noise is used as the signal-to-noise ratio of the audio frame. In response to the signal-to-noise ratio of the audio frame being greater than a first threshold, the audio frame is determined to be a speech frame; In response to the signal-to-noise ratio of the audio frame being less than the first threshold, the audio frame is determined to be a noise frame; If the number of audio frames is greater than the number of noise frames, the current scene is determined to be the quiet scene. If the number of audio frames is less than the number of noise frames, the current scene is determined to be the noisy scene.
2. The method of claim 1, wherein, The process of fusing the original voiceprint features, user attribute features, user state features, and user speech rate features of the audio data to obtain the fused voiceprint features of the audio data includes: The original voiceprint features, user attribute features, user status features, and user speech rate features of the audio data are concatenated to obtain the fused voiceprint features of the audio data.
3. The method of claim 1, wherein, The process of identifying the fused voiceprint features of the audio data based on a voiceprint database to obtain voiceprint recognition results includes: Obtain the similarity score between the fused voiceprint features of the audio data and the fused voiceprint features stored in the voiceprint database; In response to determining, based on the obtained similarity score, that the voiceprint database stores fused voiceprint features with a similarity score greater than a second threshold, it is determined that the current user is identified through voiceprint recognition. If the similarity score between the fused voiceprint feature of the audio data and the fused voiceprint feature of the target registered user is the highest, the current user is identified as the target registered user.
4. A voiceprint recognition apparatus, characterized by comprising: The device includes: The acquisition module is configured to acquire audio data collected in the current scene; The first extraction module is configured to extract voiceprint features from the audio data based on a voiceprint model that matches the current scene, thereby obtaining the original voiceprint features of the audio data. The second extraction module is configured to perform user attribute analysis, user status analysis and user speech rate analysis on the audio data respectively, to obtain the user attribute features, user status features and user speech rate features of the audio data. The fusion module is configured to perform feature fusion on the original voiceprint features, user attribute features, user state features and user speech rate features of the audio data to obtain the fused voiceprint features of the audio data. The recognition module is configured to recognize the fused voiceprint features of the audio data based on a voiceprint database to obtain a voiceprint recognition result; wherein, the voiceprint database is used to store the fused voiceprint features of registered users; The second extraction module is configured to perform user age analysis on the audio data to obtain user age characteristics; perform user gender analysis on the audio data to obtain user gender characteristics; perform user emotion analysis on the audio data to obtain user emotion characteristics; perform user fatigue level analysis on the audio data to obtain user fatigue level characteristics; and perform user speech rate analysis on the audio data to obtain user speech rate characteristics. The second extraction module is further configured to acquire image data collected in the current scene; filter facial images of the current user in the acquired image data; perform eye state recognition on the facial images to obtain image recognition results; wherein the image recognition results are used to indicate the opening and closing state of the current user's eyes; and based on the image recognition results, correct the user fatigue level characteristics to obtain corrected user fatigue level characteristics. The first extraction module is configured to extract voiceprint features from the audio data based on a first voiceprint model that matches the quiet scene when the current scene is a quiet scene; and to extract voiceprint features from the audio data based on a second voiceprint model that matches the noisy scene when the current scene is a noisy scene. The first extraction module is further configured to: for any audio frame in the audio data, acquire the energy of the audio frame; use the ratio of the energy of the audio frame to the reference energy of the noise as the signal-to-noise ratio (SNR) of the audio frame; determine the audio frame as a speech frame in response to the SNR of the audio frame being greater than a first threshold; determine the audio frame as a noise frame in response to the SNR of the audio frame being less than the first threshold; determine the current scene as the quiet scene when the number of speech frames is greater than the number of noise frames; and determine the current scene as the noisy scene when the number of speech frames is less than the number of noise frames.
5. A computer device, comprising: The device includes a processor and a memory, the memory storing at least one piece of program code, the at least one piece of program code being loaded and executed by the processor to implement the voiceprint recognition method as claimed in any one of claims 1 to 3.
6. A computer-readable storage medium, characterized in that, The storage medium stores at least one piece of program code, which is loaded and executed by a processor to implement the voiceprint recognition method as described in any one of claims 1 to 3.
Citation Information
Patent Citations
Voice processing method and device, storage medium and electronic equipment
CN108922525A
Fatigue level recognition method and device, computer equipment, and storage medium
CN109119095A
Voiceprint recognition method and device
CN114495948A
Voice control method and device
CN115132212A