A voice-based identity verification method, device and equipment

By judging the quality of the voice data to be verified and generating voiceprint fusion features, the impact of audio interference on identity verification is resolved, thereby improving the accuracy and credibility of voice identity verification.

CN114547568BActive Publication Date: 2026-03-27ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-09
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing voice-based authentication methods fail to effectively consider the impact of interference information contained in audio files on authentication results, resulting in insufficient accuracy and reliability.

Method used

By acquiring the voice data to be verified, it is determined whether it meets the preset voice data quality conditions. If it does, a voiceprint fusion feature to be verified is generated and compared with the pre-stored benchmark voiceprint fusion feature of the target user to generate an identity verification result.

Benefits of technology

It improves the accuracy and reliability of identity verification results, avoids erroneous verification caused by interfering information, and enhances the reliability of identity verification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114547568B_ABST
    Figure CN114547568B_ABST
Patent Text Reader

Abstract

Embodiments of the present specification disclose a voice-based identity verification method, device and equipment. The scheme can include: after obtaining an identity verification request for a target user, judging whether the voice data to be verified carried in the identity verification request meets the preset voice data quality condition, if so, generating a voiceprint fusion feature to be verified according to the voice data to be verified; and then comparing the voiceprint fusion feature to be verified with the pre-stored reference voiceprint fusion feature of the target user to generate an identity verification result for the target user.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of identity verification, and in particular to a voice-based identity verification method, device and equipment. BACKGROUND

[0002] With the advent of the information age, more and more scenarios need to verify the identity of a user to ensure that the identity claimed by the user is real rather than fictitious, thereby protecting the rights and interests of the user and ensuring the smooth operation of the business. There are various existing identity verification methods, such as an identity verification method based on an identification card, an identity verification method based on facial recognition technology, and an identity verification method based on voice recognition. However, when verifying the identity of a user based on voice, the voiceprint features in the audio file provided by the user are usually used to generate an identity verification result, without considering that the interference information contained in the audio file will affect the identity verification result.

[0003] Therefore, how to improve the accuracy of the identity verification result generated based on voice has become a technical problem to be solved. SUMMARY

[0004] The voice-based identity verification method, device and equipment provided by the embodiments of the present application can improve the accuracy of the identity verification result generated based on voice.

[0005] To solve the above technical problem, the embodiments of the present application are implemented as follows:

[0006] The voice-based identity verification method provided by the embodiments of the present application comprises the following steps.

[0007] An identity verification request for a target user is obtained, and the identity verification request carries to-be-verified voice data.

[0008] It is determined whether the to-be-verified voice data meets a preset voice data quality condition, and a first determination result is obtained.

[0009] If the first determination result indicates that the to-be-verified voice data meets the preset voice data quality condition, a to-be-verified voiceprint fusion feature is generated according to the to-be-verified voice data.

[0010] The to-be-verified voiceprint fusion feature is compared with a pre-stored reference voiceprint fusion feature of the target user, and an identity verification result for the target user is obtained.

[0011] The voice-based identity verification device provided by the embodiments of the present application comprises the following steps.

[0012] The first obtaining module is configured to obtain an identity authentication request for a target user, the identity authentication request carrying to-be-verified voice data;

[0013] The determining module is configured to determine whether the to-be-verified voice data meets a preset voice data quality condition, to obtain a first determination result.

[0014] The first generating module is configured to generate to-be-verified voiceprint fusion features according to the to-be-verified voice data, if the first determination result indicates that the to-be-verified voice data meets the preset voice data quality condition.

[0015] The comparing module is configured to compare the to-be-verified voiceprint fusion features with pre-stored reference voiceprint fusion features of the target user, to obtain an identity authentication result for the target user.

[0016] An identity authentication device based on voice provided by an embodiment of the present specification comprises:

[0017] at least one processor; and

[0018] a memory in communication connection with the at least one processor; wherein

[0019] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to:

[0020] obtain an identity authentication request for a target user, the identity authentication request carrying to-be-verified voice data;

[0021] determine whether the to-be-verified voice data meets a preset voice data quality condition, to obtain a first determination result.

[0022] generate to-be-verified voiceprint fusion features according to the to-be-verified voice data, if the first determination result indicates that the to-be-verified voice data meets the preset voice data quality condition.

[0023] compare the to-be-verified voiceprint fusion features with pre-stored reference voiceprint fusion features of the target user, to obtain an identity authentication result for the target user.

[0024] At least one embodiment provided in the present specification can achieve the following beneficial effects:

[0025] After obtaining the identity authentication request for the target user, it is judged whether the to-be-verified voice data carried in the identity authentication request meets the preset voice data quality condition. If yes, it means that the accuracy of the identity authentication result generated based on the to-be-verified voice data is better, so that the to-be-verified voiceprint fusion feature with better accuracy is generated based on the to-be-verified voice data. Then, the to-be-verified voiceprint fusion feature with better accuracy is compared with the pre-stored reference voiceprint fusion feature of the target user, so as to improve the accuracy of the obtained user identity authentication result. BRIEF DESCRIPTION OF DRAWINGS

[0026] In order to more clearly illustrate the technical solutions in the embodiments of the present specification or the prior art, the drawings needed to be used in the embodiments or prior art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments described in the present specification, and for those skilled in the art, other drawings can also be obtained without creative labor.

[0027] Figure 1 A scene schematic diagram of a voice-based identity authentication method provided by an embodiment of the present specification;

[0028] Figure 2 A flowchart of a voice-based identity authentication method provided by an embodiment of the present specification;

[0029] Figure 3 A swim lane flowchart of a voice-based identity authentication method provided by an embodiment of the present specification;

[0030] Figure 4 A structural schematic diagram of a voice-based identity authentication device corresponding to Figure 2 provided by an embodiment of the present specification;

[0031] Figure 5 A structural schematic diagram of a voice-based identity authentication device corresponding to Figure 2 provided by an embodiment of the present specification. DETAILED DESCRIPTION

[0032] In order to make the purpose, technical solutions and advantages of one or more embodiments of the present specification clearer, the technical solutions of one or more embodiments of the present specification will be described clearly and completely in combination with the specific embodiments of the present specification and corresponding drawings. Obviously, the described embodiments are only some embodiments of the present specification, not all embodiments. Based on the embodiments in the present specification, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of one or more embodiments of the present specification.

[0033] The technical solutions provided by the embodiments of the present specification are described in detail below with reference to the drawings.

[0034] In the service industry operating based on a network order placement platform, such as express delivery, take-out, chauffeur, online car-hailing, and housekeeping, users of the network order placement platform (such as express delivery personnel, take-out personnel, etc.) will be exposed to a large amount of third-party personal information. Therefore, it is necessary to verify the identity of these users to ensure the security of the third-party personal information.

[0035] In the prior art, the identity of the users such as express delivery personnel and take-out personnel is usually verified by performing face recognition on them several times randomly. However, this identity verification method often disturbs the users and affects their work efficiency. Therefore, there is an urgent need for a non-intrusive and real-time all-weather identity verification method.

[0036] Since the express delivery personnel and take-out personnel often need to communicate with customers through voice during their work, voice-based identity verification methods are increasingly popular. However, in the existing voice-based identity verification schemes, the influence of interference information contained in the audio file on the identity verification result is often not considered, so that the accuracy and reliability of the generated identity verification result cannot be guaranteed.

[0037] To solve the defects in the prior art, the present scheme provides the following embodiments:

[0038] Figure 1 A scene diagram of a voice-based identity verification method provided in an embodiment of the present specification is shown in FIG. 1. Figure 1 As shown in FIG. 1, a terminal device 101 bound to a target user can be communicatively connected to an identity verification server 102 and a terminal device 103 of a third party. The target user and the third party can use the terminal device 101 and the terminal device 103 to make a phone call or send voice chat messages to communicate with each other. The audio file formed during the communication between the target user and the third party can be used as verification voice data and used to verify the identity of the target user.

[0039] Specifically, the terminal device 101 bound to the target user can send an identity verification request carrying the verification voice data to the identity verification server 102 through a client or a server (not shown) of a specified application. Figure 1 The identity verification server 102 determines whether the verification voice data meets a preset voice data quality condition. If the verification voice data meets the preset voice data quality condition, it means that a verification voiceprint fusion feature with high accuracy can be generated based on the verification voice data. Then, the identity verification server 102 can compare the verification voiceprint fusion feature with a pre-stored reference voiceprint fusion feature of the target user to generate an identity verification result with high accuracy and reliability for the target user.

[0040] Next, a voice-based identity verification method provided by the present specification embodiment will be specifically described in combination with the accompanying drawings:

[0041] Figure 2 A flowchart of a voice-based identity verification method provided by the present specification embodiment. From a program perspective, the execution subject of the flowchart can be a terminal device bound to a target user or an identity verification server, or an application program loaded at the terminal device bound to the target user or the identity verification server. As shown in the figure, the flowchart can include the following steps: Figure 2

[0042] Step 202: Obtain an identity verification request for a target user, wherein the identity verification request carries to-be-verified voice data.

[0043] In the present specification embodiment, when identity verification is needed for a target user, an identity verification request can be generated based on to-be-verified voice data collected by a terminal device bound to the target user within a preset time period. The terminal device bound to the target user can refer to a terminal device logged in a registered account of the target user, or a terminal device with a target unique device identifier designated by the target user in advance. The preset time period can be set according to actual needs, for example, the last half hour, the last 5 minutes, etc., which is not limited specifically.

[0044] In actual application, the identity verification request can include not only to-be-verified voice data, but also identity information of the target user, so as to verify whether the holder of the terminal device bound to the target user is the target user through the to-be-verified voice data.

[0045] In actual application, the to-be-verified voice data can be data containing only the voice of the holder of the terminal device, which is extracted from voice call recording or voice chat message of the terminal device bound to the target user. Specifically, for voice call recording, the uplink and downlink channels can be separated to obtain the to-be-verified voice data by extracting the voice in the specified channel; or the single-channel voice containing the voice data of the user and the third party can be processed by a channel separation algorithm to obtain the to-be-verified voice data. For voice chat message, the to-be-verified voice data can be obtained by extracting the audio data corresponding to the voice chat message sent by the terminal device bound to the target user.

[0046] Step 204: Determine whether the to-be-verified voice data meets a preset voice data quality condition, and obtain a first determination result.

[0047] ​In the embodiments of the present specification, since the partial voice data may contain more interference information, so as to be unable to generate an identity verification result with better accuracy based on such voice data, thus, the voice data is not suitable for identity verification. Based on this, the preset voice data quality condition can be used to check whether the voice data to be verified is suitable for the identity recognition process. Specifically, the preset voice data quality condition can check and evaluate the voice data to be verified from the aspects of voice duration, sampling rate, sampling accuracy, voice intelligibility, whether it is a live voice, etc.

[0048] Step 206: If the first judgment result indicates that the voice data to be verified meets the preset voice data quality condition, generating a voiceprint fusion feature to be verified according to the voice data to be verified.

[0049] In the embodiments of the present specification, if the first judgment result indicates that the voice data to be verified meets the preset voice data quality condition, it usually indicates that an identity verification result with better accuracy can be generated based on the voice data to be verified, so as to allow the identity verification result to be generated according to the voice data to be verified.

[0050] In order to facilitate the understanding of the present application, first, the voiceprint recognition technology is simply introduced. Since the human vocal organs (such as tongue, teeth, larynx, lungs, nasal cavity, etc.) are very different in size and shape when people speak, and the difference in pronunciation habits of each person is also very large, therefore, different individuals can be distinguished through voiceprint recognition technology. Voiceprint recognition refers to comparing voiceprint fusion features with pre-stored reference voiceprint fusion features to confirm whether they belong to the same user, thereby realizing the identity verification function.

[0051] The voiceprint fusion feature can refer to a feature vector that can represent the structure of a specific organ or the habit of a speaker. Specifically, the Mel-frequency cepstral coefficients (MFCC) and Mel-scale Filter Bank (FBank) of the voice data to be verified can be extracted to generate the voiceprint fusion feature to be verified.

[0052] The pre-stored reference voiceprint fusion feature can refer to the voiceprint fusion feature of the target user pre-stored as a comparison reference. The generation principle of the pre-stored reference voiceprint fusion feature and the voiceprint fusion feature to be verified can be the same.

[0053] In the embodiments of the present specification, if the first judgment result indicates that the to-be-verified voice data does not meet the preset voice data quality condition, it generally indicates that an identity verification result with better accuracy cannot be generated based on the to-be-verified voice data, so that the identity verification process can be terminated to ensure the accuracy of the identity verification result.

[0054] Step 208: Comparing the to-be-verified voiceprint fusion feature with the pre-stored reference voiceprint fusion feature of the target user to obtain an identity verification result for the target user.

[0055] The identity verification result can be used to reflect whether the to-be-verified voice data carried in the identity verification request in step 202 belongs to the target user, and further can reflect whether the terminal device holder bound to the target user is the target user.

[0056] In actual application, generally, if the identity verification result indicates that the to-be-verified voice data carried in the identity verification request is data generated based on the voice of the target user, it can be determined that the identity verification is passed; otherwise, the identity verification fails.

[0057] Figure 2 In the method in the first aspect of the present specification, after obtaining the identity verification request for the target user, it is judged whether the to-be-verified voice data carried in the identity verification request meets the preset voice data quality condition. If it meets, it indicates that the accuracy of the identity verification result generated based on the to-be-verified voice data is better, so that the to-be-verified voiceprint fusion feature with better accuracy is generated according to the to-be-verified voice data. Then, by comparing the to-be-verified voiceprint fusion feature with better accuracy with the pre-stored reference voiceprint fusion feature of the target user, the accuracy of the obtained user identity verification result is improved.

[0058] Based on the method in the first aspect of the present specification, the embodiments of the present specification further provide some specific implementation schemes of the method, which are described as follows. Figure 2

[0059] In actual application, the to-be-verified voice data can only contain silence data, and does not contain valid voice data that can be used for identity verification. In this case, it is impossible to generate an identity verification result with better accuracy based on the to-be-verified voice data, so that the silence data needs to be pre-screened out.

[0060] Therefore, before the step 204 of judging whether the to-be-verified voice data meets the preset voice data quality condition, the method can further include:

[0061] Performing silence detection on the to-be-verified voice data to obtain a silence detection result.

[0062] ​Correspondingly, step 204: determining whether the voice data to be verified meets the preset voice data quality conditions, may specifically include:

[0063] If the silence detection result indicates that the voice data to be verified is not silent data, then it is determined whether the voice data to be verified meets the preset voice data quality conditions.

[0064] In the embodiments of this specification, the silence detection can be used to detect whether the speech data to be verified is silent data. Specifically, silence detection can be implemented using methods such as Gaussian mixture model, threshold discrimination algorithm, model matching algorithm, and higher-order statistical methods.

[0065] In this embodiment of the specification, the voice data to be verified is subjected to silence detection. After obtaining the silence detection result, if it is determined that the voice data to be verified is silent data, that is, it does not contain voice data, it cannot be used for identity verification under normal circumstances, and the process can be skipped to the end step to terminate the identity verification process; otherwise, if it is determined that the voice data to be verified is not silent data, it can be further determined whether the voice data to be verified meets the preset voice data quality conditions.

[0066] In the embodiments of this specification, after performing silence detection on the voice data to be verified, the authentication is terminated by terminating the voice data to be verified that is silent. This not only helps to ensure the accuracy of the authentication result, but also avoids unnecessary waste of device resources by avoiding the execution of subsequent authentication steps.

[0067] In the embodiments of this specification, a silence detection model based on a Gaussian classification model is further provided.

[0068] Specifically, the step of performing silence detection on the speech data to be verified and obtaining a silence detection result may include:

[0069] Spectral features are extracted from the speech data to be verified to obtain the target spectral feature data of the speech data to be verified.

[0070] The target spectral feature data is input into the silence detection model to obtain the silence detection result output by the silence detection model. The silence detection model is obtained by training a Gaussian classification model in advance using spectral feature data samples carrying a first classification label. The spectral feature data samples are feature data obtained by extracting spectral features from the first speech data sample. The first classification label can be used to indicate whether the first speech data sample is silent data.

[0071] In the embodiments of the present specification, the first voice data sample can include user voice data pre-labeled with the first classification label. The first classification label can include a label indicating that the first voice data sample is silent data, and a label indicating that the first voice data sample is non-silent data.

[0072] The user voice data as the first voice data sample can be user voice data collected by different devices in different scenarios. Specifically, it can include but is not limited to the following scenarios: user voice data collected by a Bluetooth headset during voice communication while riding or walking outdoors, or user voice data collected by a mobile phone during communication in an indoor environment, or user chat voice data sent by a user in an instant messaging application, and the like. Training a silent detection model using user voice data collected in multiple scenarios and by different devices can help improve the accuracy of the trained silent detection model.

[0073] The training process of the silent detection model can include inputting the first voice data sample carrying the first classification label into a Gaussian classification model to obtain a predicted classification result output by the Gaussian classification model for whether the first voice data sample belongs to silent data, and optimizing parameters of the Gaussian classification model to minimize the difference between the predicted classification result and the first classification label. When the prediction accuracy of the predicted classification result output by the Gaussian classification model is greater than a preset value, it can be indicated that the prediction result of the trained silent detection model is relatively accurate, so that the model training can be stopped and put into use.

[0074] In actual applications, after generating a silent detection result for the to-be-verified voice data using the silent detection model, although the silent detection result indicates that the to-be-verified voice data does not belong to silent data, the to-be-verified voice data usually still contains silent segments more or less. In order to facilitate subsequent generation of an identity verification result with good accuracy based on the to-be-verified voice data, each silent segment contained in the to-be-verified voice data also needs to be removed.

[0075] Therefore, before determining whether the to-be-verified voice data meets the preset voice data quality condition, the method can further include:

[0076] Performing time-frequency domain feature extraction on the to-be-verified voice data to obtain time domain feature data and frequency domain feature data of the to-be-verified voice data.

[0077] Extracting user voice data from the to-be-verified voice data according to the time domain feature data and the frequency domain feature data to obtain a user voice data set.

[0078] Correspondingly, the determining whether the to-be-verified voice data satisfies the preset voice data quality condition to obtain a first determination result can specifically include:

[0079] The determining whether the user voice data set satisfies the preset voice data quality condition to obtain a second determination result.

[0080] In the embodiments of the present specification, the time-frequency domain feature extraction can be a process of time domain and frequency domain analysis and processing on the to-be-verified voice data. The time domain feature data can include one or more of a short-time zero-crossing rate, a short-time energy, a short-time average amplitude difference function, and a short-time autocorrelation function. The frequency domain feature data can include one or more of a cepstrum distance, a frequency variance, and a spectral entropy.

[0081] In the embodiments of the present specification, since there is a large difference between the time-frequency domain features of the silence segment and the voice segment, when extracting the user voice data, the starting point and the ending point of the voice segment can be identified from the to-be-verified voice data according to the time domain feature data and the frequency domain feature data of the to-be-verified voice data, and each voice segment in the to-be-verified voice data is extracted as the user voice data to obtain the user voice data set. The segments in the to-be-verified voice data that are not extracted belong to the silence segment, so that each silence segment contained in the to-be-verified voice data can be eliminated.

[0082] In actual application, the user voice data set can include several pieces of extracted user voice data. The extraction of the user voice data can be implemented by a threshold discrimination algorithm, a model matching algorithm, a Gaussian mixture model method, a high-order statistical method, and the like.

[0083] In the embodiments of the present specification, when determining whether the user voice data set satisfies the preset voice data quality condition, it can be determined according to whether each piece of user voice data included in the user voice data set satisfies the preset voice data quality condition, to combine the determination results corresponding to each piece of user voice data to generate a final determination result representing whether the user voice data set satisfies the preset voice data quality condition. Alternatively, each piece of user voice data in the user voice data set can be spliced to obtain a comprehensive user voice data, so that it is determined according to whether the comprehensive user voice data satisfies the preset voice data quality condition to generate a determination result representing whether the user voice data set satisfies the preset voice data quality condition, which is not specifically limited.

[0084] In the embodiments of the present specification, by extracting the user voice data from the to-be-verified voice data, each silence segment contained in the to-be-verified voice data can be eliminated, and then the accuracy of the generated identity verification result can be improved based on the user voice data set that does not contain the silence segment.

[0085] In actual application, due to the short duration of the user voice, the generated to-be-verified voiceprint fusion feature can be unstable, thereby affecting the accuracy of the identity verification result generated based on the to-be-verified voiceprint fusion feature.

[0086] Therefore, the preset voice data quality condition can include that the user voice duration is greater than or equal to a first threshold.

[0087] Correspondingly, the determining whether the user voice data set meets the preset voice data quality condition can specifically include:

[0088] Determining whether the total voice duration of the user voice data set is greater than or equal to the first threshold.

[0089] In the embodiments of the present specification, the user voice duration can refer to the sum of the durations of each piece of user voice data included in the user voice data set, that is, the duration of the effective voice used to generate the to-be-verified voiceprint fusion feature. The selection of the first threshold is related to the algorithm and accuracy requirement of voiceprint recognition, and can be set according to actual conditions, which is not specifically limited here.

[0090] In the embodiments of the present specification, if the total voice duration of the user voice data set is less than the first threshold, it usually indicates that the effective voice duration of the user is too short. When the to-be-verified voiceprint fusion feature generated based on the user voice data set is used to generate a user identity verification result, the accuracy of the user identity verification result is usually poor. Therefore, the user voice data set cannot meet the requirements of identity verification, so the process can jump to the end step, thereby terminating the identity verification process. On the contrary, if the total voice duration of the user voice data set is greater than or equal to the first threshold, the to-be-verified voiceprint fusion feature with better accuracy can be generated based on the user voice data set, thereby facilitating the improvement of the accuracy of the user identity verification result generated based on the to-be-verified voiceprint fusion feature.

[0091] In actual application, the user can talk to a third party in a noisy environment, or the sound collection capability of the terminal device used by the user can be poor, thereby possibly causing the user voice data extracted from the to-be-verified voice data to have large noise and poor quality.

[0092] Therefore, the preset voice data quality condition can include that the user voice quality score is greater than or equal to a second threshold.

[0093] Correspondingly, the determining whether the user voice data set meets the preset voice data quality condition can specifically include:

[0094] perform voice quality analysis on the user voice data set by a voice quality analysis model to obtain a voice quality score of the user voice data set; the voice quality analysis model is obtained by pre-training a deep learning model using second voice data samples carrying voice quality score labels, and the voice quality score labels can be used to represent a preset voice quality score of the second voice data samples.

[0095] determine whether the voice quality score of the user voice data set is greater than or equal to the second threshold value.

[0096] In the embodiments of the present specification, the second voice data samples can include user voice data pre-labeled with voice quality score labels. The voice quality score labels can be generated according to one or more of the following parameters: signal-to-noise ratio, segmented signal-to-noise ratio, PESQ (Perceptual evaluation of speech quality), log-likelihood ratio measure, log-spectral distance, short-time objective intelligibility, weighted spectral tilt measure, perceptual objective speech quality evaluation, sampling rate and sampling accuracy, etc.

[0097] The user voice data as the second voice data samples can be user voice data collected by different devices in different scenarios, specifically, can include but not limited to the following scenarios: voice communication through Bluetooth headset while riding or driving a car outdoors, voice communication through mobile phone in indoor environment, etc. The second voice data samples can also be voice data without silent segments.

[0098] The training process of the voice quality analysis model can be that the second voice data samples carrying voice quality score labels are used as the input of the deep learning model, the output of the deep learning model can be the voice quality prediction score of the second voice data samples, and the deep learning model is constantly optimized based on the loss function. When the error between the output voice quality prediction score of the second voice data samples and the voice quality score label carried by the second voice data samples is less than a preset value, it can be indicated that the prediction result of the trained voice quality analysis model is relatively accurate, so that the model training can be stopped and put into use.

[0099] In the embodiments of the present specification, when the voice quality analysis model is used to analyze the voice quality of the set of user voice data, the voice quality of each piece of user voice data in the set of user voice data can be analyzed respectively by using the voice quality analysis model to obtain the voice quality score of each piece of user voice data, and then the voice quality score of the set of user voice data can be obtained by calculating (for example, average value, weighted average value, etc.) the voice quality scores of each piece of user voice data. Alternatively, each piece of user voice data in the set of user voice data can be spliced to obtain a comprehensive user voice data, and the voice quality score of the comprehensive user voice data can be generated by using the voice quality analysis model to obtain the voice quality score of the set of user voice data, which is not limited specifically.

[0100] In the embodiments of the present specification, the voice quality score of the set of user voice data is generated by using the trained voice quality analysis model. If the voice quality score is less than the second threshold value, even if the to-be-verified voiceprint fusion feature is generated according to the set of user voice data, the accuracy of the generated to-be-verified voiceprint fusion feature is poor, and thus the requirement of voiceprint recognition cannot be met. Therefore, the identity verification process can be terminated directly by jumping to the end step; otherwise, if the voice quality score is greater than or equal to the second threshold value, the to-be-verified voiceprint fusion feature with good accuracy can be further generated according to the set of user voice data, thereby improving the accuracy of the user identity verification result generated based on the to-be-verified voiceprint fusion feature.

[0101] In the embodiments of the present specification, the voice quality score of the set of user voice data is generated by using the trained voice quality analysis model, which is beneficial to improving the accuracy of the obtained voice quality score and further ensuring the accuracy of the generated identity verification result.

[0102] In actual applications, the user of the terminal device bound to the target user may have cheating behaviors such as playing a recording, waveform splicing, and voice synthesis when performing identity verification. Therefore, in order to ensure the accuracy of the generated identity verification result, the set of user voice data extracted from the to-be-verified voice data also needs to be subjected to live voice detection.

[0103] Therefore, the preset voice data quality condition can include that the user voice is a live voice.

[0104] The method can further include:

[0105] The user voice data set is subjected to live voice detection through a live voice detection model to obtain a live voice detection result of the user voice data set; the live voice detection model is obtained by pre-training a deep learning model using third voice data samples carrying second classification labels, and the second classification labels can be used to indicate whether the third voice data samples belong to live voice.

[0106] It is determined whether the live voice detection result indicates that the user voice data set belongs to live voice.

[0107] In the embodiments of the present specification, the third voice data samples can include user voice data pre-labeled with second classification labels. The second classification labels can be used to indicate whether the third voice data samples belong to live voice, and for non-live voice, the specific form of false voice attack can be further labeled.

[0108] The user voice data as the third voice data samples can include user voice data (i.e., live voice data samples) collected by different devices in different scenarios, and can also include voice data (i.e., non-live voice data samples) formed after processing the directly collected user voice data, such as recording and playing back, waveform splicing, voice synthesis, voice imitation, etc.

[0109] The training process of the live voice detection model is similar to the training process of the silence detection model, which will not be repeated here.

[0110] In the embodiments of the present specification, the trained live voice detection model is used to detect live voice, and then it is determined whether the user voice contained in the to-be-verified voice data belongs to live voice. If it does not belong to live voice, it indicates that there is cheating or attack behavior, so the identity verification process can be terminated directly to effectively resist false voice attacks such as recording and playing back, waveform splicing, voice synthesis, etc. Otherwise, if it belongs to live voice, an identity verification result with better accuracy can be generated according to the user voice data set extracted from the to-be-verified voice data.

[0111] As mentioned above, the user voice data set extracted from the to-be-verified voice data usually no longer contains silence segments, but only contains voice segments. However, in addition to the voice data of the holder of the terminal device bound to the target user, the user voice data set can also contain voice data of other people who are making sounds in the environment of the holder. Therefore, the pronunciation data of each person in the user voice data set needs to be distinguished to ensure the accuracy of the identity verification result generated subsequently.

[0112] Based on this, after obtaining the second judgment result by judging whether the user voice data set meets the preset voice data quality condition, the method can further include:

[0113] If the second judgment result indicates that the user voice data set meets the preset voice data quality condition, voiceprint feature extraction is performed on each piece of user voice data included in the user voice data set to obtain voiceprint feature data of the each piece of user voice data.

[0114] According to the voiceprint feature data of the each piece of user voice data, the each piece of user voice data is divided to obtain at least one user voice data sub-set; the user voice data included in each user voice data sub-set belongs to the same user, and the user voice data included in different user voice data sub-sets belongs to different users.

[0115] A third judgment result is obtained by judging whether the voice data quantity of the target voice data sub-set is greater than a third threshold value; the target voice data sub-set is the user voice data sub-set with the largest voice data quantity.

[0116] Correspondingly, the generating of the to-be-verified voiceprint fusion feature according to the to-be-verified voice data can specifically include:

[0117] If the third judgment result indicates that the voice data quantity of the target voice data sub-set is greater than the third threshold value, a to-be-verified voiceprint fusion feature is generated according to the voiceprint feature data of the user voice data included in the target voice data sub-set.

[0118] In the embodiments of the present specification, the voiceprint feature can include Mel-frequency cepstral coefficients (MFCC), Mel-scale FilterBank, and the like.

[0119] The division of the each piece of user voice data can refer to dividing the user voice data belonging to the same user into the same user voice data sub-set, and dividing the user voice data of different users into different user voice data sub-sets, so that each user voice data sub-set corresponds to one user. In actual application, the division of the each piece of user voice data can be realized by speaker clustering technology, or can be realized according to the consistency comparison of the voiceprint feature data of the each piece of user voice data, and no specific limitation is made thereto.

[0120] Generally, the data quantity of the voice data of the holder of the terminal device bound to the target user in the user voice data set is the largest, and the data quantity of the voice data of other mixed personnel is generally smaller. Therefore, the user voice data sub-set with the largest voice data quantity can be considered as the set of the voice of the holder, so that the user voice data sub-set with the largest voice data quantity can be used as the target voice data sub-set to generate the to-be-verified voiceprint fusion feature. In actual application, the largest voice data quantity can refer to the longest voice duration or the largest number of voice clips, and no specific limitation is made.

[0121] Similarly, to avoid the problem of poor accuracy of the identity verification result caused by too small effective voice data quantity, the to-be-verified voiceprint fusion feature can also be generated according to the voiceprint feature data of the user voice data included in the target voice data sub-set after it is determined that the voice data quantity of the target voice data sub-set is greater than the third threshold. If the voice data quantity of the target voice data sub-set is less than the third threshold, the identity verification process can be terminated.

[0122] The generation process of the to-be-verified voiceprint fusion feature can be fusion calculation of the voiceprint feature data of the user voice data included in the target voice data sub-set to generate the to-be-verified voiceprint fusion feature, and the fusion calculation can include averaging, weighted averaging, etc.

[0123] In the embodiments of the present specification, the user voice data is divided into user voice data sub-sets corresponding to the speakers according to the voiceprint feature data of each piece of user voice data; and the target voice data sub-set corresponding to the user to be verified (i.e., the holder of the terminal device bound to the target user) is selected from the multiple user voice data sub-sets according to the voice data quantity of the user voice data sub-set; and then the to-be-verified voiceprint fusion feature is generated according to the target voice data sub-set, thereby facilitating the improvement of the accuracy of the identity verification result.

[0124] In the embodiments of the present specification, the data format of the to-be-verified voice data carried in the identity verification request can not be a preset format, or the to-be-verified voice data can also include the voice data of a third party in communication with the user, so that the to-be-verified voice data can need to be preprocessed before identity verification.

[0125] Based on this, the identity verification request carries a to-be-processed audio file containing the to-be-verified voice data; the to-be-processed audio file is an audio file obtained by audio acquisition of the terminal device bound to the target user in the process of user communication.

[0126] Correspondingly, before the determining whether the to-be-verified voice data satisfies the preset voice data quality condition, the method can further include:

[0127] determining whether the format type of the to-be-processed audio file belongs to a preset format type, to obtain a fourth determination result.

[0128] If the fourth determination result indicates that the format type of the to-be-processed audio file does not belong to the preset format type, performing audio decoding processing on the to-be-processed audio file to obtain an audio file of the preset format type.

[0129] performing channel separation processing on the audio file of the preset format type to obtain the to-be-verified voice data.

[0130] In the embodiments of the present specification, the audio decoding can be used to convert the to-be-processed audio file that does not belong to the preset format type into an audio file that belongs to the preset format type.

[0131] The channel separation can be used to process a single-channel / multi-channel audio file containing voice data of a user to be authenticated and a third party to obtain voice data of only the user to be authenticated as the to-be-verified voice data.

[0132] In actual applications, if it is determined that the format type of the to-be-processed audio file belongs to the preset format type, the process of audio decoding can be directly skipped, and the channel separation processing can be performed. Of course, if the to-be-processed audio file does not contain voice data of a third party in a conversation with the user, the channel separation processing can also be skipped, and details are not described herein.

[0133] In the embodiments of the present specification, since the voice-based identity verification needs to be implemented based on the pre-stored reference voiceprint fusion feature of the target user, before the method in Figure 2 is executed, the user needs to be registered to generate and save the reference voiceprint fusion feature of the target user. It is worth noting that the generation principle and processing process of the pre-stored reference voiceprint fusion feature of the target user and the to-be-verified voiceprint fusion feature can be the same, so as to guarantee the accuracy of the pre-stored reference voiceprint fusion feature of the target user, and further improve the accuracy of the voice-based identity verification result.

[0134] Specifically, before the comparing the to-be-verified voiceprint fusion feature with the pre-stored reference voiceprint fusion feature of the target user, the method can further include:

[0135] obtaining a user voice data sample set of the target user that satisfies the preset voice data quality condition.

[0136] extracting a voiceprint feature from each piece of user voice data sample in the user voice data sample set to obtain a voiceprint feature data sample of the user voice data sample.

[0137] dividing the user voice data samples according to the voiceprint feature data samples of the user voice data samples to obtain at least one user voice data sample subset; the user voice data samples in each user voice data sample subset belong to the same user, and the user voice data samples in different user voice data sample subsets belong to different users.

[0138] judging whether the voice data quantity of the target voice data sample subset is greater than a fourth threshold value to obtain a fourth judgment result; the target voice data sample subset is the user voice data sample subset with the largest voice data quantity.

[0139] if the fourth judgment result indicates that the voice data quantity of the target voice data sample subset is greater than the fourth threshold value, generating the reference voiceprint fusion feature of the target user according to the voiceprint feature data samples of the user voice data samples in the target voice data sample subset.

[0140] storing the reference voiceprint fusion feature of the target user to obtain a pre-stored reference voiceprint fusion feature of the target user.

[0141] In the embodiments of the present specification, the acquisition process of the user voice data sample set meeting the requirement is substantially the same as the acquisition process of the user voice data set mentioned above. Specifically, obtaining the to-be-registered voice data submitted in the target user registration process, and performing audio decoding, channel separation, silence detection and the like on the to-be-registered voice data. If the to-be-registered voice data does not belong to silence data, the to-be-registered voice data can be subjected to time-frequency domain feature extraction, and user voice data samples are extracted according to the extracted time domain feature data and frequency domain feature data, and the silence segments are removed to obtain a user voice data sample set. Continue to judge whether the user voice data sample set meets the preset voice data quality condition (for example, the user voice duration is greater than or equal to a first threshold value, the user voice quality score is greater than or equal to a second threshold value, and the user voice belongs to a living body voice), until a user voice data sample set meeting the preset voice data quality condition is determined.

[0142] Similarly, for the user voice data sample set meeting the preset voice data quality condition, each piece of user voice sample in the user voice data sample set is divided into a user voice data sample subset corresponding to a speaker according to the speaker. And the user voice data sample subset with the largest voice data quantity is selected as the target voice data sample subset corresponding to the target user.

[0143] Then it is judged whether the voice data quantity of the target voice data sample subset is greater than a fourth threshold value. If less than or equal to the fourth threshold value, jump to the end step, thereby terminating the identity registration process; and if greater than the fourth threshold value, only then generate the reference voiceprint fusion feature of the target user according to the target voice data sample subset.

[0144] In the embodiments of the present specification, although the target voice data sample subset meets the preset voice data quality condition, the quality of each piece of user voice data sample in the target voice data sample subset still has differences, and therefore the reference voiceprint fusion feature of the target user can be generated based on the number of user voice data samples with better quality in the target voice data sample subset, so as to improve the accuracy of the pre-stored reference voiceprint fusion feature of the target user.

[0145] Based on this, the reference voiceprint fusion feature of the target user is generated according to the voiceprint feature data samples of the user voice data samples in the target voice data sample subset, and specifically can include:

[0146] Obtain the voice quality score of the user voice data sample in the target voice data sample subset.

[0147] Sort the user voice data samples in the target voice data sample subset in descending order of the voice quality score to obtain a user voice data sample sequence.

[0148] Generate the reference voiceprint fusion feature of the target user according to the voiceprint feature data samples of the first N user voice data samples in the user voice data sample sequence.

[0149] In the embodiments of the present specification, the voice quality score of the user voice data sample in the target voice data sample subset can be obtained by performing voice quality analysis on the user voice data sample through a voice quality analysis model.

[0150] In the embodiments of the present specification, the generation of the reference voiceprint fusion feature can refer to fusion calculation of the voiceprint feature data samples of the first N user voice data samples in the user voice data sample sequence to generate the reference voiceprint fusion feature of the target user, and the fusion calculation can include average value calculation, weighted average value calculation, etc. Wherein, N can be limited according to actual requirements, and no specific limitation is made here.

[0151] In the embodiments of the present specification, the voiceprint feature data samples of the first N user voice data samples with higher voice quality scores are selected to generate the reference voiceprint fusion feature of the target user, which is beneficial to guarantee the accuracy of the reference voiceprint fusion feature of the target user, and further beneficial to improve the accuracy of the obtained user identity verification result.

[0152] Figure 3 A swim lane flow diagram of a voice-based identity verification method corresponding to Figure 2 provided in the embodiments of the present specification. Figure 3 In the embodiments of the present specification, the scenario in which the voice data to be verified is uploaded to an identity verification server for identity verification is taken as an example for explanation and description, Figure 3 The execution subject of the flow shown in the embodiments of the present specification can include: a terminal device bound to the target user, an identity verification server, and the like.

[0153] As shown in Figure 3 , in the preprocessing stage, the terminal device bound to the target user can extract the telephone call recording or voice chat message generated within a preset time period to obtain a to-be-processed audio file containing voice data to be verified, and generate and send an identity verification request for the target user to the identity verification server, wherein the identity verification request can contain the to-be-processed audio file.

[0154] After the identity verification server obtains the identity verification request for the target user, it can determine whether the format type of the to-be-processed audio file belongs to a preset format type. If the format type of the to-be-processed audio file does not belong to the preset format type, the to-be-processed audio file can be subjected to audio decoding processing to obtain an audio file of the preset format type. If the format type of the to-be-processed audio file belongs to the preset format type, the audio decoding processing can be directly skipped. The identity verification server can also perform channel separation processing on the audio file of the preset format type to obtain the voice data to be verified. Of course, if the audio file of the preset format type does not contain third-party voice data, the channel separation processing step can be skipped, and the voice data to be verified can be directly obtained.

[0155] Spectrum feature extraction is performed on the to-be-verified voice data to obtain target spectrum feature data of the to-be-verified voice data; the target spectrum feature data is input into a silence detection model to perform silence detection on the to-be-verified voice data. It is judged whether the to-be-verified voice data is silence data. If it is silence data, the identity verification process can be terminated. If it is not silence data, time-frequency domain feature extraction can be performed on the to-be-verified voice data to obtain time domain feature data and frequency domain feature data of the to-be-verified voice data; and user voice data is extracted from the to-be-verified voice data according to the time domain feature data and the frequency domain feature data to obtain a user voice data set to eliminate silence segments.

[0156] Thereafter, it is continuously judged whether the user voice data set meets a preset voice data quality condition; the preset voice data quality condition can include a user voice duration condition, a user voice quality score condition, and a live voice condition. The user voice duration condition can specifically refer to that the user voice duration is greater than or equal to a first threshold value; the user voice quality score condition can specifically refer to that the user voice quality score is greater than or equal to a second threshold value; and the live voice condition can specifically refer to that the user voice is a live voice. If the user voice data set does not meet the preset voice data quality condition, the identity verification process is terminated. If the preset voice data quality condition is met, the operation of the voiceprint feature extraction stage can be performed.

[0157] In the voiceprint feature extraction stage, a same-person judgment needs to be performed. Specifically, the identity verification server can extract voiceprint feature data of each piece of user voice data included in the user voice data set; divide the user voice data according to the voiceprint feature data of the user voice data to obtain at least one user voice data sub-set; and take the user voice data sub-set with the largest voice data amount as a target voice data sub-set, so that the voice data of a third party can be screened out.

[0158] It is judged whether the voice data amount of the target voice data sub-set is greater than a fourth threshold value. If it is less than or equal to the fourth threshold value, the identity verification process is terminated. If it is greater than the fourth threshold value, a to-be-verified voiceprint fusion feature is generated according to the voiceprint feature data of the user voice data included in the target voice data sub-set.

[0159] In the feature comparison stage, the identity verification server can compare the to-be-verified voiceprint fusion feature with a pre-stored reference voiceprint fusion feature of the target user to obtain an identity verification result for the target user.

[0160] Figure 3The scheme in the method can further include, before identity verification based on voice, a stage of identity registration by the target user to generate a pre-stored reference voiceprint fusion feature of the target user. The pre-stored reference voiceprint fusion feature is compared with Figure 3 The generation principle of the voiceprint fusion feature to be verified in the method can be the same, but in the generation process of the pre-stored reference voiceprint fusion feature of the target user, in addition to the audio decoding, channel separation, silence detection, voice data quality detection, same person judgment, and voiceprint feature fusion processing shown in the method, the method can further include the following steps. Figure 3 In addition to the audio decoding, channel separation, silence detection, voice data quality detection, same person judgment, and voiceprint feature fusion processing shown in the method, the method can further include the following steps.

[0161] Based on the same idea, the embodiments of the present specification also provide a device corresponding to the above method. Figure 4 The device corresponding to the voice-based identity verification device of the method provided by the embodiments of the present specification is shown in Figure 2 The device can include the following components. Figure 4

[0162] The first acquisition module 402 can be configured to acquire an identity verification request for a target user, the identity verification request carrying voice data to be verified.

[0163] The judgment module 404 can be configured to judge whether the voice data to be verified meets a preset voice data quality condition, to obtain a first judgment result.

[0164] The first generation module 406 can be configured to, if the first judgment result indicates that the voice data to be verified meets the preset voice data quality condition, generate a voiceprint fusion feature to be verified according to the voice data to be verified.

[0165] The comparison module 408 can be configured to compare the voiceprint fusion feature to be verified with a pre-stored reference voiceprint fusion feature of the target user, to obtain an identity verification result for the target user.

[0166] Based on the device shown in Figure 4 The embodiments of the present specification also provide some specific implementation solutions of the device, which are described as follows.

[0167] Optionally, the device shown in Figure 4 The device shown in can further include the following components.

[0168] ​The silence detection module can be configured to perform silence detection on the to-be-verified voice data, and obtain a silence detection result.

[0169] Correspondingly, the determining module 404 can be configured to:

[0170] If the silence detection result indicates that the to-be-verified voice data is not silence data, the determining module 404 can be configured to determine whether the to-be-verified voice data satisfies a preset voice data quality condition.

[0171] Optionally, the silence detection module can include:

[0172] The spectrum feature extraction unit can be configured to perform spectrum feature extraction on the to-be-verified voice data, and obtain target spectrum feature data of the to-be-verified voice data.

[0173] The silence detection result generation unit can be configured to input the target spectrum feature data into a silence detection model, and obtain a silence detection result output by the silence detection model. The silence detection model is obtained by training a Gaussian classification model using spectrum feature data samples carrying first classification labels. The spectrum feature data samples are feature data obtained by performing spectrum feature extraction on first voice data samples. The first classification labels can be used to indicate whether the first voice data samples are silence data.

[0174] Optionally, Figure 4 The apparatus shown can further include:

[0175] The time-frequency domain feature extraction module can be configured to perform time-frequency domain feature extraction on the to-be-verified voice data, and obtain time domain feature data and frequency domain feature data of the to-be-verified voice data.

[0176] The user voice data extraction module can be configured to extract user voice data from the to-be-verified voice data according to the time domain feature data and the frequency domain feature data, and obtain a user voice data set.

[0177] Correspondingly, the determining module 404 can include:

[0178] The determining unit can be configured to determine whether the user voice data set satisfies a preset voice data quality condition, and obtain a second determination result.

[0179] Optionally, the preset voice data quality condition can include that a user voice duration is greater than or equal to a first threshold.

[0180] Correspondingly, the determining unit can be configured to:

[0181] Determine whether a total voice duration of the user voice data set is greater than or equal to the first threshold.

[0182] Optionally, the preset voice data quality condition can include: a user voice quality score being greater than or equal to a second threshold value.

[0183] Correspondingly, the determining module 404 can further include:

[0184] The voice quality analysis unit can be configured to perform voice quality analysis on the user voice data set by using a voice quality analysis model to obtain a voice quality score of the user voice data set. The voice quality analysis model is obtained by pre-training a deep learning model using second voice data samples carrying voice quality score labels. The voice quality score labels can be used to represent preset voice quality scores of the second voice data samples.

[0185] Correspondingly, the determining unit can be specifically configured to:

[0186] determine whether the voice quality score of the user voice data set is greater than or equal to the second threshold value.

[0187] Optionally, the preset voice data quality condition can include: user voice being a live voice.

[0188] Correspondingly, the determining module 404 can further include:

[0189] The live voice detection unit can be configured to perform live voice detection on the user voice data set by using a live voice detection model to obtain a live voice detection result of the user voice data set. The live voice detection model is obtained by pre-training a deep learning model using third voice data samples carrying second classification labels. The second classification labels can be used to represent whether the third voice data samples belong to a live voice.

[0190] Correspondingly, the determining unit can be specifically configured to:

[0191] determine whether the live voice detection result indicates that the user voice data set belongs to a live voice.

[0192] Optionally, Figure 4 The apparatus shown can further include:

[0193] The first voiceprint feature extraction module can be configured to perform voiceprint feature extraction on each piece of user voice data included in the user voice data set to obtain voiceprint feature data of the each piece of user voice data, if the second determination result indicates that the user voice data set meets the preset voice data quality condition.

[0194] The first dividing module can be used for dividing the user voice data according to the voiceprint feature data of the user voice data, to obtain at least one user voice data sub-set; the user voice data in each user voice data sub-set belongs to the same user, and the user voice data in different user voice data sub-sets belongs to different users.

[0195] The first voice data amount judgment module can be used for judging whether the voice data amount of the target voice data sub-set is greater than a third threshold value, to obtain a third judgment result; the target voice data sub-set is the user voice data sub-set with the largest voice data amount.

[0196] Correspondingly, the first generating module 406 can be specifically used for:

[0197] If the third judgment result indicates that the voice data amount of the target voice data sub-set is greater than the third threshold value, the voiceprint feature data of the user voice data contained in the target voice data sub-set is used to generate a to-be-verified voiceprint fusion feature.

[0198] Optionally, the identity verification request carries a to-be-processed audio file containing the to-be-verified voice data; the to-be-processed audio file is an audio file obtained by audio acquisition of a terminal device bound to the target user in a user call process; the apparatus can further include:

[0199] The format type judgment module can be used for judging whether the format type of the to-be-processed audio file belongs to a preset format type, to obtain a fourth judgment result.

[0200] The audio decoding module can be used for performing audio decoding processing on the to-be-processed audio file to obtain an audio file of the preset format type, if the fourth judgment result indicates that the format type of the to-be-processed audio file does not belong to the preset format type.

[0201] The channel separation module can be used for performing channel separation processing on the audio file of the preset format type, to obtain the to-be-verified voice data.

[0202] Optionally, Figure 4 The apparatus shown can further include:

[0203] The second obtaining module can be used for obtaining a user voice data sample set of the target user satisfying the preset voice data quality condition.

[0204] The second voiceprint feature extraction module can be used for extracting voiceprint features of each user voice data sample in the user voice data sample set, to obtain voiceprint feature data samples of the user voice data samples.

[0205] The second dividing module can be configured to divide the user voice data samples according to the voiceprint feature data samples of the user voice data samples, to obtain at least one user voice data sample subset; the user voice data samples in each user voice data sample subset belong to the same user, and the user voice data samples in different user voice data sample subsets belong to different users.

[0206] The second voice data amount judging module can be configured to judge whether the voice data amount of a target voice data sample subset is greater than a fourth threshold value, to obtain a fourth judgment result; the target voice data sample subset is the user voice data sample subset with the largest voice data amount.

[0207] The second generating module can be configured to, if the fourth judgment result indicates that the voice data amount of the target voice data sample subset is greater than the fourth threshold value, generate the reference voiceprint fusion feature of the target user according to the voiceprint feature data samples of the user voice data samples in the target voice data sample subset.

[0208] The storage module can be configured to store the reference voiceprint fusion feature of the target user, to obtain a pre-stored reference voiceprint fusion feature of the target user.

[0209] Optionally, the second generating module can be specifically configured to:

[0210] Obtain the voice quality scores of the user voice data samples in the target voice data sample subset.

[0211] Sort the user voice data samples in the target voice data sample subset in descending order of the voice quality scores, to obtain a user voice data sample sequence.

[0212] Generate the reference voiceprint fusion feature of the target user according to the voiceprint feature data samples of the first N user voice data samples in the user voice data sample sequence.

[0213] Based on the same idea, the embodiments of the present specification also provide a device corresponding to the above method.

[0214] Figure 5 A structure schematic diagram of a voice-based identity verification device corresponding to the above method is provided for the embodiments of the present specification. As shown in Figure 2 , the device 500 can include: Figure 5

[0215] ​at least one processor 510; and a memory 530 connected with the at least one processor in communication; wherein the memory 530 stores instructions 520 executable by the at least one processor 510, the instructions being executed by the at least one processor 510 to enable the at least one processor 510 to:

[0216] obtain an identity authentication request for a target user, the identity authentication request carrying to-be-verified voice data.

[0217] determine whether the to-be-verified voice data meets a preset voice data quality condition, to obtain a first determination result.

[0218] if the first determination result indicates that the to-be-verified voice data meets the preset voice data quality condition, generate to-be-verified voiceprint fusion features according to the to-be-verified voice data.

[0219] compare the to-be-verified voiceprint fusion features with pre-stored reference voiceprint fusion features of the target user, to obtain an identity authentication result for the target user.

[0220] Each of the embodiments in the specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other. Each of the embodiments focuses on the difference from other embodiments. In particular, for the device shown in the specification, since it is basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the part of the method embodiments. Figure 5 The device shown in the specification is basically similar to the method embodiments, and the description is relatively simple. The relevant parts can be referred to the part of the method embodiments.

[0221] In the 1990s, it was possible to distinguish whether an improvement in a technology was a hardware improvement (e.g., an improvement in the circuit structure of a diode, transistor, switch, etc.) or a software improvement (an improvement in a method flow). However, as technology has advanced, many improvements in method flows today can be considered as direct improvements in hardware circuit structures. Designers almost always obtain a corresponding hardware circuit structure by programming an improved method flow into a hardware circuit. Therefore, it cannot be said that an improvement in a method flow cannot be implemented using a hardware entity module. For example, a programmable logic device (PLD) (e.g., a field programmable gate array (FPGA)) is an integrated circuit whose logic function is determined by user programming of the device. A designer programs a digital system "integrated" on a PLD by himself / herself, without having to ask a chip manufacturer to design and manufacture a special integrated circuit chip. Moreover, instead of manually manufacturing an integrated circuit chip, this programming is now mostly implemented using "logic compiler" software, which is similar to a software compiler used when developing a program, and the original code before compilation must also be written in a specific programming language, which is called a hardware description language (HDL), and there are many types of HDL, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc., and the most commonly used are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also be aware that it is only necessary to logically program a method flow using the above-mentioned hardware description languages and program it into an integrated circuit to easily obtain a hardware circuit that implements the logical method flow.

[0222] The controller can be implemented in any suitable way, for example, the controller can take the form of a microprocessor or processor and a computer readable medium storing computer readable program code, such as software or firmware, executable by the (micro)processor, logic gates, switches, an application specific integrated circuit (ASIC), a programmable logic controller and an embedded microcontroller, examples of which can include but are not limited to the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20 and Silicone Labs C8051F320, the memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also know that, in addition to implementing the controller in pure computer readable program code, it is also possible to implement the controller in the form of logic gates, switches, application specific integrated circuits, programmable logic controllers and embedded microcontrollers, etc. to perform the same functions by means of logical programming of the method steps. Such a controller can therefore be considered as a hardware component, while the means that can be included therein for implementing various functions can also be considered as structures within the hardware component. Alternatively, even the means that can be used to implement various functions can be considered as both a software module implementing a method and a structure within a hardware component.

[0223] The systems, apparatuses, modules or units illustrated by the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.

[0224] For the sake of brevity, the above apparatuses are described in functional form in various units. Of course, the functions of each unit can be implemented in the same or multiple software and / or hardware.

[0225] Those skilled in the art will appreciate that embodiments of the present application can be provided as a method, a system, or a computer program product. Accordingly, the present application can take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present application can take the form of a computer program product on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROMs, optical storage devices, etc.) embodying computer readable program code.

[0226] The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks. Figure 1 one or more flow or blocks Figure 1 means for functionally implementing the steps in the flowchart block or blocks

[0227] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instructions which implement the Figure 1 one or more flow or blocks Figure 1 means for functionally implementing the steps in the flowchart block or blocks

[0228] The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks. Figure 1 one or more flow or blocks Figure 1 means for functionally implementing the steps in the flowchart block or blocks

[0229] In one typical configuration, a computing device can include one or more processors (CPUs), input / output interfaces, network interfaces, and memory.

[0230] The memory can include a computer-readable medium, non-transitory memory, random access memory (RAM), and / or non-volatile memory such as read-only memory (ROM) or flash memory. The memory is an example of computer-readable media.

[0231] Computer-readable media can include permanent and non-permanent, movable and non-movable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media can include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette, magnetic tape disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information that can be accessed by a computing device. According to the definition herein, computer-readable media can not include transitory media, such as modulated data signals and carriers.

[0232] It should also be noted that the terms "comprises", "comprising", or any other variations thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements, but can include other elements not expressly listed or inherent to such process, method, article, or apparatus. Without more limitations, an element defined by the statement "comprises a" does not exclude the presence of additional identical elements in the process, method, article, or apparatus that can include the element.

[0233] Those skilled in the art will appreciate that embodiments of the present application can be provided as a method, system or computer program product. Accordingly, the present application can take the form of an entirely hardware embodiment, an entirely software embodiment or an embodiment combining software and hardware aspects. Furthermore, the present application can take the form of a computer program product embodied in one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROMs, optical storage devices, etc.) having computer usable program code embodied thereon.

[0234] The present application can be described in the general context of computer-executable instructions, such as program modules, being executed by a computer. Generally, program modules can include routines, programs, objects, components, data structures, etc., that perform particular tasks or implement particular abstract data types. The present application can also be practiced in distributed computing environments where tasks are performed by remote processing devices that are linked through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including memory storage devices.

[0235] The above merely provides an example of the present application, and should not be used to limit the present application. Any modification, equivalent replacement, improvement, etc. within the spirit and principle of the present application should be included in the scope of claims of the present application.

Claims

1. A voice-based authentication method, comprising: Obtain an authentication request for the target user, wherein the authentication request carries voice data to be verified; Determine whether the voice data to be verified meets the preset voice data quality conditions to obtain a first determination result; If the first judgment result indicates that the voice data to be verified meets the preset voice data quality conditions, then voiceprint feature extraction is performed on each piece of user voice data contained in the user voice data set to obtain the voiceprint feature data of each piece of user voice data; the user voice data set is extracted from the voice data to be verified. Based on the voiceprint feature data of each user voice data, the user voice data is divided to obtain at least one user voice data subset; the user voice data contained in each user voice data subset belongs to the same user; Based on the amount of voice data in the user voice data subset, a target voice data subset is selected from the at least one user voice data subset; Based on the target speech data subset, generate the voiceprint fusion features to be verified; The voiceprint fusion feature to be verified is compared with the pre-stored benchmark voiceprint fusion feature of the target user to obtain the identity verification result for the target user.

2. The method as described in claim 1, further comprising, before determining whether the voice data to be verified meets the preset voice data quality conditions: The speech data to be verified is subjected to silence detection to obtain silence detection results; The determination of whether the voice data to be verified meets the preset voice data quality conditions specifically includes: If the silence detection result indicates that the voice data to be verified is not silent data, then it is determined whether the voice data to be verified meets the preset voice data quality conditions.

3. The method as described in claim 2, wherein performing silence detection on the speech data to be verified to obtain a silence detection result specifically includes: Spectral feature extraction is performed on the speech data to be verified to obtain the target spectral feature data of the speech data to be verified. The target spectral feature data is input into the silence detection model to obtain the silence detection result output by the silence detection model. The silence detection model is obtained by training a Gaussian mixture classification model in advance using spectral feature data samples carrying a first classification label. The spectral feature data samples are feature data obtained by extracting spectral features from the first speech data sample. The first classification label is used to indicate whether the first speech data sample is silent data.

4. The method according to any one of claims 1-3, wherein before determining whether the voice data to be verified meets the preset voice data quality conditions, it further includes: Time-frequency domain features are extracted from the speech data to be verified to obtain the time-domain feature data and frequency-domain feature data of the speech data to be verified. Based on the time-domain feature data and the frequency-domain feature data, user voice data is extracted from the voice data to be verified to obtain a user voice data set. The step of determining whether the voice data to be verified meets the preset voice data quality conditions to obtain a first determination result specifically includes: Determine whether the user voice data set meets the preset voice data quality conditions to obtain a second determination result.

5. The method as described in claim 4, wherein the preset voice data quality conditions include: The duration of the user's voice message is greater than or equal to the first threshold; The step of determining whether the user voice data set meets the preset voice data quality conditions specifically includes: Determine whether the total duration of the user voice data set is greater than or equal to the first threshold.

6. The method as described in claim 4, wherein the preset voice data quality conditions include: The user's voice quality score is greater than or equal to the second threshold; The step of determining whether the user voice data set meets the preset voice data quality conditions specifically includes: The user speech dataset is analyzed using a speech quality analysis model to obtain a speech quality score for the user speech dataset. The speech quality analysis model is trained on a deep learning model using a second speech data sample carrying a speech quality score label. The speech quality score label is used to represent the preset speech quality score of the second speech data sample. Determine whether the voice quality score of the user voice data set is greater than or equal to the second threshold.

7. The method as described in claim 4, wherein the preset voice data quality conditions include: The user's voice is considered live voice. The step of determining whether the user voice data set meets the preset voice data quality conditions further includes: Live speech detection is performed on the user speech dataset using a live speech detection model to obtain the live speech detection result of the user speech dataset; The live speech detection model is obtained by training a deep learning model in advance using a third speech data sample carrying a second classification label. The second classification label is used to indicate whether the third speech data sample belongs to live speech. Determine whether the liveness detection result indicates that the user voice data set belongs to liveness.

8. The method as described in claim 4, after determining whether the user voice data set meets the preset voice data quality conditions and obtaining the second determination result, further comprising: If the second judgment result indicates that the user voice data set meets the preset voice data quality conditions, then the voiceprint features of each user voice data set contained in the user voice data set are extracted to obtain the voiceprint feature data of each user voice data set. Based on the voiceprint feature data of each user voice data, the user voice data is divided to obtain at least one user voice data subset; the user voice data contained in each user voice data subset belongs to the same user, and the user voice data contained in different user voice data subsets belong to different users. Determine whether the amount of voice data in the target voice data subset is greater than a third threshold to obtain a third determination result; the target voice data subset is the user voice data subset with the largest amount of voice data. The step of generating the voiceprint fusion feature to be verified based on the voice data to be verified specifically includes: If the third judgment result indicates that the amount of voice data in the target voice data subset is greater than the third threshold, then a voiceprint fusion feature to be verified is generated based on the voiceprint feature data of the user voice data contained in the target voice data subset.

9. The method as described in claim 1, wherein the authentication request carries an audio file to be processed containing the voice data to be verified; the audio file to be processed is a file obtained by the terminal device bound to the target user during the user's call. Before determining whether the voice data to be verified meets the preset voice data quality conditions, the method further includes: Determine whether the format type of the audio file to be processed belongs to a preset format type to obtain a fourth determination result; If the fourth determination result indicates that the format type of the audio file to be processed does not belong to the preset format type, then the audio file to be processed is subjected to audio decoding processing to obtain the audio file of the preset format type; The audio file of the preset format type is subjected to channel separation processing to obtain the voice data to be verified.

10. The method of claim 1, further comprising, before comparing the voiceprint fusion feature to be verified with the pre-stored benchmark voiceprint fusion feature of the target user: Obtain a set of user voice data samples that meet the preset voice data quality conditions for the target user; Voiceprint features are extracted from each user voice data sample in the user voice data sample set to obtain voiceprint feature data samples of each user voice data sample. Based on the voiceprint feature data samples of each user voice data sample, the user voice data samples are divided to obtain at least one subset of user voice data samples; the user voice data samples contained in each subset of user voice data samples belong to the same user, and the user voice data samples contained in different subsets of user voice data samples belong to different users. Determine whether the amount of voice data in the target voice data sample subset is greater than a fourth threshold to obtain a fourth determination result; the target voice data sample subset is the user voice data sample subset with the largest amount of voice data. If the fourth judgment result indicates that the amount of voice data in the target voice data sample subset is greater than the fourth threshold, then the baseline voiceprint fusion feature of the target user is generated based on the voiceprint feature data sample of the user voice data sample in the target voice data sample subset. The baseline voiceprint fusion features of the target user are stored to obtain the pre-stored baseline voiceprint fusion features of the target user.

11. The method of claim 10, wherein generating the baseline voiceprint fusion feature of the target user based on the voiceprint feature data samples of the user voice data samples in the target voice data sample subset specifically includes: Obtain the speech quality score of the user speech data sample in the target speech data sample subset; The user voice data samples in the target voice data sample subset are sorted in descending order of the voice quality scores to obtain a user voice data sample sequence. Based on the voiceprint feature data samples of the first N user voice data samples in the user voice data sample sequence, the baseline voiceprint fusion feature of the target user is generated.

12. A voice-based authentication device, comprising: The first acquisition module is used to acquire an authentication request for a target user, wherein the authentication request carries voice data to be verified. The judgment module is used to determine whether the voice data to be verified meets the preset voice data quality conditions and obtain a first judgment result; The first generation module is configured to extract voiceprint features from each piece of user voice data contained in the user voice data set if the first judgment result indicates that the voice data to be verified meets the preset voice data quality conditions, thereby obtaining voiceprint feature data of each piece of user voice data; the user voice data set is extracted from the voice data to be verified. Based on the voiceprint feature data of each user voice data, the user voice data is divided to obtain at least one user voice data subset; the user voice data contained in each user voice data subset belongs to the same user; based on the amount of voice data in each user voice data subset, a target voice data subset is selected from the at least one user voice data subset. Based on the target speech data subset, generate the voiceprint fusion features to be verified; The comparison module is used to compare the voiceprint fusion feature to be verified with the pre-stored benchmark voiceprint fusion feature of the target user to obtain the identity verification result for the target user.

13. The apparatus of claim 12, further comprising: A silence detection module is used to perform silence detection on the voice data to be verified and obtain a silence detection result; The judgment module is specifically used for: If the silence detection result indicates that the voice data to be verified is not silent data, then it is determined whether the voice data to be verified meets the preset voice data quality conditions.

14. The apparatus of claim 13, wherein the silent detection module comprises: A spectrum feature extraction unit is used to extract spectrum features from the speech data to be verified to obtain target spectrum feature data of the speech data to be verified. A silence detection result generation unit is used to input the target spectral feature data into a silence detection model to obtain the silence detection result output by the silence detection model. The silence detection model is obtained by training a Gaussian mixture classification model in advance using spectral feature data samples carrying a first classification label. The spectral feature data samples are feature data obtained by extracting spectral features from a first speech data sample. The first classification label is used to indicate whether the first speech data sample is silent data.

15. The apparatus according to any one of claims 12-14, further comprising: The time-frequency domain feature extraction module is used to extract time-frequency domain features from the speech data to be verified, and obtain the time-domain feature data and frequency-domain feature data of the speech data to be verified. The user voice data extraction module is used to extract user voice data from the voice data to be verified based on the time-domain feature data and the frequency-domain feature data, so as to obtain a user voice data set. The judgment module includes: The judgment unit is used to determine whether the user voice data set meets the preset voice data quality conditions and obtain a second judgment result.

16. The apparatus of claim 15, wherein the preset voice data quality conditions include: The duration of the user's voice message is greater than or equal to the first threshold; The judgment unit is specifically used for: Determine whether the total duration of the user voice data set is greater than or equal to the first threshold.

17. The apparatus of claim 15, wherein the preset voice data quality conditions include: The user's voice quality score is greater than or equal to the second threshold; The judgment module further includes: The speech quality analysis unit is used to perform speech quality analysis on the user speech data set through a speech quality analysis model to obtain a speech quality score of the user speech data set; the speech quality analysis model is obtained by training a deep learning model in advance using a second speech data sample carrying a speech quality score label, and the speech quality score label is used to represent the preset speech quality score of the second speech data sample; The judgment unit is specifically used for: Determine whether the voice quality score of the user voice data set is greater than or equal to the second threshold.

18. The apparatus of claim 15, wherein the preset voice data quality conditions include: The user's voice is considered live voice. The judgment module further includes: The live speech detection unit is used to perform live speech detection on the user speech data set using a live speech detection model to obtain the live speech detection result of the user speech data set; the live speech detection model is obtained by training a deep learning model in advance using a third speech data sample carrying a second classification label, and the second classification label is used to indicate whether the third speech data sample belongs to live speech. The judgment unit is specifically used for: Determine whether the liveness detection result indicates that the user voice data set belongs to liveness.

19. The apparatus of claim 15, further comprising: The first voiceprint feature extraction module is used to extract voiceprint features from each piece of user voice data contained in the user voice data set if the second judgment result indicates that the user voice data set meets the preset voice data quality conditions, so as to obtain the voiceprint feature data of each piece of user voice data. The first segmentation module is used to segment each piece of user voice data according to the voiceprint feature data of each piece of user voice data to obtain at least one subset of user voice data; the user voice data contained in each subset of user voice data belongs to the same user, and the user voice data contained in different subsets of user voice data belongs to different users. The first voice data volume judgment module is used to determine whether the voice data volume of the target voice data subset is greater than the third threshold, and to obtain the third judgment result; the target voice data subset is the user voice data subset with the largest voice data volume; The first generation module is specifically used for: If the third judgment result indicates that the amount of voice data in the target voice data subset is greater than the third threshold, then a voiceprint fusion feature to be verified is generated based on the voiceprint feature data of the user voice data contained in the target voice data subset.

20. The apparatus of claim 12, wherein the authentication request carries an audio file to be processed containing the voice data to be verified; the audio file to be processed is a file obtained by a terminal device bound to the target user during audio acquisition during a user's call; The device further includes: The format type determination module is used to determine whether the format type of the audio file to be processed belongs to a preset format type, and to obtain a fourth determination result; An audio decoding module is used to perform audio decoding on the audio file to be processed to obtain an audio file of the preset format type if the fourth judgment result indicates that the format type of the audio file to be processed does not belong to the preset format type. The channel separation module is used to perform channel separation processing on the audio file of the preset format type to obtain the voice data to be verified.

21. The apparatus of claim 12, further comprising: The second acquisition module is used to acquire a set of user voice data samples of the target user that meet the preset voice data quality conditions; The second voiceprint feature extraction module is used to extract voiceprint features from each user voice data sample contained in the user voice data sample set, and obtain the voiceprint feature data sample of each user voice data sample. The second segmentation module is used to segment the user voice data samples according to the voiceprint feature data samples of each user voice data sample to obtain at least one subset of user voice data samples; the user voice data samples contained in each subset of user voice data samples belong to the same user, and the user voice data samples contained in different subsets of user voice data samples belong to different users. The second voice data volume judgment module is used to judge whether the voice data volume of the target voice data sample subset is greater than the fourth threshold, and to obtain the fourth judgment result; the target voice data sample subset is the user voice data sample subset with the largest voice data volume; The second generation module is used to generate the baseline voiceprint fusion feature of the target user based on the voiceprint feature data sample of the user voice data sample in the target voice data sample subset if the fourth judgment result indicates that the amount of voice data in the target voice data sample subset is greater than the fourth threshold. The storage module is used to store the baseline voiceprint fusion features of the target user, thereby obtaining the pre-stored baseline voiceprint fusion features of the target user.

22. The apparatus of claim 21, wherein the second generating module is specifically configured to: Obtain the speech quality score of the user speech data sample in the target speech data sample subset; The user voice data samples in the target voice data sample subset are sorted in descending order of the voice quality scores to obtain a user voice data sample sequence. Based on the voiceprint feature data samples of the first N user voice data samples in the user voice data sample sequence, the baseline voiceprint fusion feature of the target user is generated.

23. A voice-based authentication device, comprising: At least one processor; as well as, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, enable the at least one processor to: Obtain an authentication request for the target user, wherein the authentication request carries voice data to be verified; Determine whether the voice data to be verified meets the preset voice data quality conditions to obtain a first determination result; If the first judgment result indicates that the voice data to be verified meets the preset voice data quality conditions, then voiceprint feature extraction is performed on each piece of user voice data contained in the user voice data set to obtain the voiceprint feature data of each piece of user voice data; the user voice data set is extracted from the voice data to be verified. Based on the voiceprint feature data of each user voice data, the user voice data is divided to obtain at least one user voice data subset; the user voice data contained in each user voice data subset belongs to the same user; Based on the amount of voice data in the user voice data subset, a target voice data subset is selected from the at least one user voice data subset; Based on the target speech data subset, generate the voiceprint fusion features to be verified; The voiceprint fusion feature to be verified is compared with the pre-stored benchmark voiceprint fusion feature of the target user to obtain the identity verification result for the target user.

Citation Information

Patent Citations

  • Living identity authentication method and device, computer device and readable storage medium

    CN109346089A

  • Speech recognition and authentication method and system

    CN110473552A

  • Telephone channel voiceprint recognition method and device

    CN112509586A