Voiceprint processing method, storage medium, program product, electronic device and vehicle

By acquiring and fusing voiceprint features and calculating deviation values, the problem of the accuracy of voiceprint recognition deteriorates in different states is solved, and more accurate voiceprint recognition is achieved.

CN120452471APending Publication Date: 2025-08-08BYD CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510499971.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

The existing voiceprint recognition technology has a large difference in the voice of the speaker at different times or emotional states, resulting in a decrease in recognition accuracy and it is difficult to accurately distinguish different voices when the voices of different speakers are approximate.

Method used

By obtaining multiple audio data of the reference speaker, standard voiceprint characteristics and voiceprint deviation values are determined, multiple voiceprint characteristics are fused and deviation values are calculated to perform voiceprint matching and improve recognition accuracy.

Benefits of technology

More accurate voiceprint recognition in different states is achieved, recognition errors caused by state changes or sound approximation are avoided, and the accuracy and distinction ability of voiceprint recognition are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120452471A_ABST
    Figure CN120452471A_ABST
Patent Text Reader

Abstract

The invention relates to a voiceprint processing method, a storage medium, a program product, electronic equipment and a vehicle. The voiceprint processing method comprises the steps of determining a standard voiceprint feature and a voiceprint deviation value of a reference speaker according to multiple pieces of audio data of the reference speaker; wherein the standard voiceprint features comprise a plurality of voiceprint features, and the voiceprint deviation value is used for indicating the deviation among the plurality of voiceprint features; the standard voiceprint feature and the voiceprint deviation value are used for voiceprint matching. According to the voiceprint recognition method and device, the voiceprint recognition accuracy is improved by obtaining the voiceprint features and the voiceprint deviation values of the speakers, and different speakers can be distinguished more accurately.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of vehicle technology, and in particular to a voiceprint processing method, storage medium, program product, electronic device and vehicle. Background Art

[0002] Voiceprint recognition is a voice biometric technology that verifies a speaker's identity by analyzing their voice characteristics. Because each person's voice has unique characteristics, voiceprints can serve as a relatively reliable form of identity verification. Ensuring the accuracy of voiceprint recognition is a technical challenge that the industry is currently working to address. Summary of the Invention

[0003] The embodiments of the present application provide a voiceprint processing method, storage medium, program product, electronic device, and vehicle, which improve the accuracy of voiceprint recognition and help to more accurately distinguish different speakers, thereby at least partially solving the above-mentioned technical problems.

[0004] In order to achieve the above-mentioned purpose, according to the first aspect of the present application, a voiceprint processing method is provided, comprising: determining the standard voiceprint features and voiceprint deviation values of a reference speaker based on multiple audio data of the reference speaker; wherein the standard voiceprint features include multiple voiceprint features, and the voiceprint deviation values are used to indicate the deviation between the multiple voiceprint features; the standard voiceprint features and the voiceprint deviation values are used for voiceprint matching.

[0005] Optionally, determining the standard voiceprint features of the reference speaker includes: extracting multiple original voiceprint features of the reference speaker from multiple audio data of the reference speaker; wherein the standard voiceprint features include multiple original voiceprint features.

[0006] Optionally, the method further comprises: fusing a plurality of the original voiceprint features of the reference speaker to obtain a fused voiceprint feature of the reference speaker; wherein the standard voiceprint feature includes the fused voiceprint feature.

[0007] Optionally, the fusing of the multiple original voiceprint features of the reference speaker includes: performing weighted sum processing on the multiple original voiceprint features of the reference speaker to obtain the fused voiceprint features of the reference speaker; or performing averaging processing on the multiple original voiceprint features of the reference speaker to obtain the fused voiceprint features of the reference speaker.

[0008] Optionally, determining the voiceprint deviation value of the reference speaker includes: intercepting multiple audio sub-data from multiple audio data of the reference speaker; extracting multiple voiceprint sub-features from the multiple audio sub-data; and determining the voiceprint deviation value of the reference speaker based on the multiple voiceprint sub-features.

[0009] Optionally, the determining the voiceprint deviation value of the reference speaker includes: determining the voiceprint deviation value of the reference speaker according to a plurality of the voiceprint sub-features and a plurality of the original voiceprint features.

[0010] Optionally, determining the voiceprint deviation value of the reference speaker includes: performing error statistical processing on similarities between a plurality of the voiceprint sub-features and a plurality of the original voiceprint features to obtain the voiceprint deviation value of the reference speaker.

[0011] Optionally, the error statistical processing includes at least one of the following: root mean square error statistical processing, mean absolute percentage error statistical processing.

[0012] Optionally, the error statistical processing of the similarities between the multiple voiceprint sub-features and the multiple original voiceprint features includes: performing the root mean square error statistical processing on the similarities between the multiple voiceprint sub-features and the multiple original voiceprint features to obtain a first deviation value; performing the mean absolute percentage error statistical processing on the similarities between the multiple voiceprint sub-features and the multiple original voiceprint features to obtain a second deviation value; and determining the voiceprint deviation value based on the first deviation value and the second deviation value.

[0013] Optionally, the method further comprises: performing voiceprint matching on the speaker to be identified and the reference speaker based on the voiceprint features of the speaker to be identified and the standard voiceprint features of the reference speaker.

[0014] Optionally, the voiceprint matching of the speaker to be identified and the reference speaker includes: determining a voiceprint matching score between the speaker to be identified and the reference speaker based on the voiceprint features of the speaker to be identified and the standard voiceprint features of the reference speaker; and performing voiceprint matching of the speaker to be identified and the reference speaker based on the voiceprint matching score between the speaker to be identified and the reference speaker.

[0015] Optionally, determining the voiceprint matching score between the speaker to be identified and the reference speaker includes: determining the voiceprint similarity between the speaker to be identified and the reference speaker based on the feature similarity between the voiceprint features of the speaker to be identified and the standard voiceprint features of the reference speaker; and determining the voiceprint matching score between the speaker to be identified and the reference speaker based on the voiceprint similarity between the speaker to be identified and the reference speaker.

[0016] Optionally, determining the voiceprint similarity between the speaker to be identified and the reference speaker includes: determining the voiceprint similarity between the speaker to be identified and the reference speaker based on a first similarity threshold and the feature similarity between the voiceprint features of the speaker to be identified and the standard voiceprint features of the reference speaker.

[0017] Optionally, determining the voiceprint similarity between the speaker to be identified and the reference speaker includes: screening the feature similarities between the voiceprint features of the speaker to be identified and multiple voiceprint features of the reference speaker according to a first similarity threshold; and determining the voiceprint similarity between the speaker to be identified and the reference speaker based on the feature similarities retained after screening.

[0018] Optionally, the screening of the feature similarities between the voiceprint features of the speaker to be identified and the multiple voiceprint features of the reference speaker includes: determining a statistical number of the feature similarities having values greater than a first similarity threshold based on the feature similarities between the voiceprint features of the speaker to be identified and the multiple voiceprint features of the reference speaker; and screening the feature similarities between the voiceprint features of the speaker to be identified and the multiple voiceprint features of the reference speaker based on the statistical number.

[0019] Optionally, screening the feature similarities between the voiceprint feature of the speaker to be identified and the multiple voiceprint features of the reference speaker based on the statistical quantity includes: if the statistical quantity is greater than a first quantity threshold, excluding the feature similarity with the smallest value from the feature similarities between the voiceprint feature of the speaker to be identified and the multiple voiceprint features of the reference speaker;

[0020] Optionally, the feature similarities between the voiceprint features of the speaker to be identified and the multiple voiceprint features of the reference speaker are screened based on the statistical quantity, including: if the statistical quantity is less than a first quantity threshold, then excluding the feature similarity with the largest value from the feature similarities between the voiceprint features of the speaker to be identified and the multiple voiceprint features of the reference speaker.

[0021] Optionally, the determining the voiceprint similarity between the speaker to be identified and the reference speaker includes: averaging the feature similarities retained after screening to obtain the voiceprint similarity between the speaker to be identified and the reference speaker.

[0022] Optionally, determining the voiceprint matching score between the speaker to be identified and the reference speaker includes: determining the voiceprint matching score between the speaker to be identified and the reference speaker based on a second similarity threshold and the voiceprint similarity between the speaker to be identified and the reference speaker.

[0023] Optionally, determining the voiceprint matching score between the speaker to be identified and the reference speaker includes: if the voiceprint similarity between the speaker to be identified and the reference speaker is greater than a second similarity threshold, determining the voiceprint matching score between the speaker to be identified and the reference speaker based on the voiceprint deviation value of the reference speaker and the voiceprint similarity between the speaker to be identified and the reference speaker.

[0024] Optionally, determining the voiceprint matching score between the speaker to be identified and the reference speaker includes: summing the product of the voiceprint deviation value of the reference speaker and the voiceprint similarity between the speaker to be identified and the reference speaker, and the voiceprint similarity between the speaker to be identified and the reference speaker, to obtain the voiceprint matching score between the speaker to be identified and the reference speaker.

[0025] Optionally, determining the voiceprint matching score between the speaker to be identified and the reference speaker includes: if the voiceprint similarity between the speaker to be identified and the reference speaker is less than the second similarity threshold, then using the voiceprint similarity between the speaker to be identified and the reference speaker as the voiceprint matching score between the speaker to be identified and the reference speaker.

[0026] Optionally, there are multiple reference speakers; the voiceprint matching of the speaker to be identified and the reference speakers includes: taking the reference speaker corresponding to the largest voiceprint matching score among the voiceprint matching scores between the speaker to be identified and the multiple reference speakers as the reference speaker matched with the speaker to be identified.

[0027] Optionally, the method further includes: extracting voiceprint features of the speaker to be identified from audio data of the speaker to be identified.

[0028] Optionally, the method further includes: updating the standard voiceprint feature of the reference speaker that matches the speaker to be identified based on the voiceprint feature of the speaker to be identified.

[0029] Optionally, the updating of the standard voiceprint features of the reference speaker that matches the speaker to be identified includes: determining the number of target voiceprint features in the standard voiceprint features of the reference speaker that matches the speaker to be identified; wherein the target voiceprint features refer to the voiceprint features added during the updating process; and updating the standard voiceprint features of the reference speaker that matches the speaker to be identified based on the number of target voiceprint features and the voiceprint features of the speaker to be identified.

[0030] Optionally, the updating of the standard voiceprint features of the reference speaker that matches the speaker to be identified includes: if the number of the target voiceprint features is less than a second quantity threshold, adding the voiceprint features of the speaker to be identified to the standard voiceprint features of the reference speaker that matches the speaker to be identified.

[0031] Optionally, the updating of the standard voiceprint features of the reference speaker that matches the speaker to be identified includes: if the number of the target voiceprint features is greater than a second number threshold, replacing any one of the target voiceprint features of the reference speaker that matches the speaker to be identified with the voiceprint features of the speaker to be identified.

[0032] According to a second aspect of the present application, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the voiceprint processing method described above is implemented.

[0033] According to a third aspect of the present application, a computer program product is provided, comprising a computer program, wherein the computer program implements the above-mentioned voiceprint processing method when executed by a processor.

[0034] According to a fourth aspect of the present application, an electronic device is provided, comprising: a memory storing a computer program; and a processor configured to execute the computer program in the memory to implement the above-mentioned voiceprint processing method.

[0035] According to a fifth aspect of the present application, a vehicle is provided, comprising the above-mentioned electronic device.

[0036] The voiceprint processing method provided in the embodiment of the present application determines the standard voiceprint features and voiceprint deviation values of the reference speaker based on multiple audio data of the reference speaker. Among them, the standard voiceprint features of the reference speaker include multiple voiceprint features, and the voiceprint deviation value is used to indicate the deviation between the multiple voiceprint features. The standard voiceprint features and voiceprint deviation values can be used for voiceprint matching to achieve voiceprint recognition. For each speaker, the embodiment of the present application obtains multiple voiceprint features and voiceprint deviation values of the speaker based on the multiple audio data of the speaker. On the one hand, it can cover the various states of the speaker as much as possible, avoiding voiceprint recognition errors due to changes in the speaker's state; on the other hand, it can more accurately distinguish different speakers, avoiding voiceprint recognition errors due to similar voices of speakers. The embodiment of the present application improves the accuracy of voiceprint recognition by obtaining multiple voiceprint features and voiceprint deviation values of the speaker, and helps to achieve more accurate distinction between different speakers.

[0037] Other features and advantages of the present application will be described in detail in the subsequent detailed description. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] To more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present application. Those skilled in the art can also derive other drawings based on these drawings without inventive effort.

[0039] In order to more completely understand the present application and its beneficial effects, the following description will be given in conjunction with the accompanying drawings, wherein the same drawing numbers represent the same parts in the following description.

[0040] Figure 1 This is a flow chart of a voiceprint processing method provided in an embodiment of the present application;

[0041] Figure 2 This is a flowchart of another voiceprint processing method provided by an embodiment of the present application;

[0042] Figure 3 This is a schematic diagram of a voiceprint processing process provided by an embodiment of the present application;

[0043] Figure 4 This is a schematic diagram of a voiceprint registration process provided by an embodiment of the present application;

[0044] Figure 5 This is a schematic diagram of a voiceprint matching process provided by an embodiment of the present application;

[0045] Figure 6 This is a schematic diagram of a voiceprint database update process provided by an embodiment of the present application;

[0046] Figure 7 It is a schematic diagram of a vehicle provided in an embodiment of the present application. DETAILED DESCRIPTION

[0047] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the embodiments described are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present application.

[0048] Voiceprint recognition is a voice biometric technology that verifies a speaker's identity by analyzing their voice characteristics. These include, but are not limited to, pitch, frequency, intensity, and articulation. Because each person's voice has unique characteristics, voiceprints can serve as a relatively reliable form of identity verification.

[0049] In recent years, with the development of machine learning, especially deep learning technology, the accuracy of voiceprint recognition has been significantly improved. Voiceprint recognition models can better capture complex human voice features and maintain good performance in noisy environments. In related technologies, voiceprint recognition models based on deep learning can usually be divided into two parts. The first part is voiceprint registration. The speaker reads the specified text as required, and the generated audio data is input into the model through front-end processing. The model extracts features from the audio, generates an embedding vector, and saves it. The second part is voiceprint matching. The speaker speaks normally, and the generated audio data is input into the model through front-end processing and generates an embedding vector. The similarity between the embedding vector and the saved embedding vector is calculated, and the similarity score is compared with the set threshold to determine the identity of the speaker.

[0050] However, the voice of the same speaker may vary at different times or in different emotional states. For example, the voice of a person with a cold or fatigue is different from that of a normal person. This results in significant differences between voice samples of the same speaker. In addition, the voices of different speakers may be very similar. These factors can all lead to poor voiceprint recognition results.

[0051] In view of this, embodiments of the present application provide a voiceprint processing method, a storage medium, a program product, an electronic device, and a vehicle to at least partially solve the above technical problems.

[0052] According to the first aspect of the present application, an embodiment of the present application provides a voiceprint processing method.

[0053] See also Figure 1 , Figure 1This is a flow chart of a voiceprint processing method provided by an embodiment of the present application. Figure 1 As shown, the voiceprint processing method may include the following steps:

[0054] Step S100: determining a standard voiceprint feature and a voiceprint deviation value of a reference speaker based on multiple audio data of the reference speaker.

[0055] The reference speaker can be any speaker. When performing voiceprint registration, multiple audio data of the reference speaker can be obtained. For example, the speaker can dictate several texts according to the instructions of the device, and the device can obtain multiple audio data during the speaker's dictation; or, the device can randomly obtain multiple audio data of the speaker while the speaker is using the device. In some embodiments, the speaker's multiple audio data can be pre-processed with noise reduction, signal processing, interception, etc., so as to facilitate the subsequent rapid and accurate extraction of voiceprint features. The multiple audio data can cover the widest possible pronunciation of the reference speaker, including but not limited to the reference speaker's different emotional states, volume levels, and background noise environment.

[0056] The embodiment of the present application processes multiple audio data of a reference speaker to determine the standard voiceprint features and voiceprint deviation values of the reference speaker. The standard voiceprint features include multiple voiceprint features, and the voiceprint deviation values are used to indicate the deviations between the multiple voiceprint features. The standard voiceprint features and voiceprint deviation values are used for voiceprint matching. For each reference speaker, the embodiment of the present application determines multiple voiceprint features and the voiceprint deviation values between the multiple voiceprint features based on the multiple audio data of the reference speaker. The voiceprint features can be embedded vectors that can convert the original audio data into a fixed-length numerical representation, which can be used to capture speech features and support subsequent calculations. To facilitate subsequent voiceprint recognition, the embodiment of the present application can store the standard voiceprint features and voiceprint deviation values of the reference speaker. For example, the standard voiceprint features and voiceprint deviation values of the reference speaker are stored in a voiceprint database, so that the voiceprint database can include the standard voiceprint features and voiceprint deviation values of one or more reference speakers.

[0057] This embodiment of the present application obtains multiple voiceprint features and voiceprint deviation values for each speaker based on the speaker's multiple audio data. This not only covers the speaker's various states as much as possible, avoiding voiceprint recognition errors due to changes in the speaker's state; it also allows for more accurate distinction between different speakers, avoiding voiceprint recognition errors due to similar voices. By obtaining multiple voiceprint features and voiceprint deviation values for each speaker, this embodiment of the present application improves the accuracy of voiceprint recognition and helps to achieve more accurate distinction between different speakers.

[0058] In some embodiments, the above-mentioned step S100 may include step S110: extracting multiple original voiceprint features of the reference speaker from multiple audio data of the reference speaker; wherein the standard voiceprint features include multiple original voiceprint features. By extracting multiple original voiceprint features from multiple audio data, the standard voiceprint features can more accurately reflect the essential characteristics of the reference speaker. The embodiment of the present application does not limit the method of extracting the original voiceprint features. In actual application, it can be flexibly selected according to needs. For example, the original voiceprint features can be extracted through models, spectrum conversion, linear prediction, etc. In some embodiments, the above-mentioned step S110 may include: inputting multiple audio data of the reference speaker into the voiceprint recognition model, so that the voiceprint recognition model outputs multiple original voiceprint features of the reference speaker. The voiceprint recognition model may be an artificial intelligence model, for example, a voiceprint recognition model based on deep learning, such as ResNet34 (Residual Network 34), ResNet221 (Residual Network 221), ECAPA (Efficient CAPacity Architecture for Speaker Verification with Time-Delay Neural Networks), etc. In the embodiment of the present application, a pre-trained or trained voiceprint recognition model may be obtained, and multiple audio data of a reference speaker may be processed using the voiceprint recognition model to obtain the original voiceprint features corresponding to each audio data.

[0059] To more fully represent the voice characteristics of the reference speaker, in some embodiments, step S100 may include step S120: fusing multiple original voiceprint features of the reference speaker to obtain a fused voiceprint feature of the reference speaker; wherein the standard voiceprint feature includes the fused voiceprint feature. Thus, the fused voiceprint feature incorporates the voice characteristics of multiple audio data sets of the reference speaker. Based on this, in embodiments of the present application, the standard voiceprint feature of each reference speaker may include multiple original voiceprint features and one fused voiceprint feature. The embodiments of the present application do not limit the specific implementation of the fusion process. In some embodiments, step S120 may include: performing a weighted summation process on the multiple original voiceprint features of the reference speaker to obtain the fused voiceprint feature of the reference speaker; or performing an averaging process on the multiple original voiceprint features to obtain the fused voiceprint feature of the reference speaker. By performing a weighted summation or averaging process on the multiple original voiceprint features, the multiple original voiceprint features can be integrated, allowing the fused voiceprint feature to more accurately and fully represent the voice characteristics of the reference speaker.

[0060] In some embodiments, to improve the accuracy of the voiceprint deviation value, the above step S100 may include steps S131 to S133:

[0061] Step S131: extracting a plurality of audio sub-data from a plurality of audio data of a reference speaker;

[0062] Step S132: extracting multiple voiceprint sub-features from the multiple audio sub-data;

[0063] Step S133: Determine the voiceprint deviation value of the reference speaker based on the multiple voiceprint sub-features.

[0064] In step S131, multiple audio sub-data may be intercepted from the multiple audio data. The duration of the audio sub-data may be less than or equal to the duration of the audio data in which the audio sub-data is located. In some embodiments, one or more audio sub-data may be intercepted from each audio data, and the audio sub-data may be intercepted randomly or according to a preset rule. For example, one second of audio sub-data may be intercepted from the beginning, end, or middle of the audio data.

[0065] In step S132, for each intercepted audio sub-data, voiceprint sub-features are extracted from the audio sub-data. The embodiment of the present application does not limit the method for extracting voiceprint sub-features. In actual application, it can be flexibly selected according to the needs. For example, voiceprint sub-features can be extracted through models, spectrum conversion, linear prediction, etc. The method for extracting voiceprint sub-features can be the same as the method for extracting the original voiceprint features mentioned above. In some embodiments, the above step S132 may include: inputting the audio sub-data into the above-mentioned voiceprint recognition model, so that the voiceprint recognition model outputs the voiceprint sub-features corresponding to the audio sub-data. The voiceprint sub-features can be embedded vectors.

[0066] In step S133, a voiceprint deviation value for the reference speaker is determined based on the multiple voiceprint sub-features of the reference speaker. For example, similarity calculations may be performed on multiple voiceprint sub-features, multiple voiceprint sub-features and multiple original voiceprint features, and / or multiple voiceprint sub-features and a fused voiceprint feature, and then the voiceprint deviation value is determined based on the calculated similarities. In some embodiments, step S133 includes determining the voiceprint deviation value for the reference speaker based on the multiple voiceprint sub-features and multiple original voiceprint features. Determining the voiceprint deviation value using multiple voiceprint sub-features and multiple original voiceprint features can simplify and accurately calculate the voiceprint deviation value. For example, similarities between multiple voiceprint sub-features and multiple original voiceprint features may be calculated, and then the voiceprint deviation value is determined based on the calculated similarities. In some embodiments, determining the voiceprint deviation value for the reference speaker based on the multiple voiceprint sub-features and multiple original voiceprint features includes performing error statistics on the similarities between the multiple voiceprint sub-features and the multiple original voiceprint features to obtain the voiceprint deviation value for the reference speaker. The similarity calculation may be based on cosine similarity, Euclidean distance, etc. The embodiment of the present application does not limit the specific method of similarity calculation.

[0067] In some embodiments, the similarity between multiple voiceprint sub-features and multiple original voiceprint features includes: the similarity between each voiceprint sub-feature and the remaining original voiceprint features, where the remaining original voiceprint features refer to original voiceprint features other than the original voiceprint features corresponding to the voiceprint sub-feature. Exemplarily, for a reference speaker, the audio data A, B and C of the reference speaker are obtained; a segment of audio sub-data is respectively intercepted from the audio data A, B and C to obtain audio sub-data A1, B1 and C1; the audio data A, B and C are input into the voiceprint recognition model to obtain original voiceprint features A, B and C; the audio sub-data A1, B1 and C1 are input into the voiceprint recognition model to obtain voiceprint sub-features A1, B1 and C1; thus, the similarities between multiple voiceprint sub-features and multiple original voiceprint features include: the similarity between voiceprint sub-feature A1 and original voiceprint feature B, the similarity between voiceprint sub-feature A1 and original voiceprint feature C, the similarity between voiceprint sub-feature B1 and original voiceprint feature A, the similarity between voiceprint sub-feature B1 and original voiceprint feature C, the similarity between voiceprint sub-feature C1 and original voiceprint feature A, and the similarity between voiceprint sub-feature C1 and original voiceprint feature B. Of course, the similarity between multiple voiceprint sub-features and multiple original voiceprint features may also include other contents, which are not limited in the embodiment of the present application. For example, it may also include the similarity between each voiceprint sub-feature and the original voiceprint feature corresponding to the voiceprint sub-feature.

[0068] In some embodiments, the error statistical processing includes at least one of the following: root mean square error (RMSE) statistical processing and mean absolute percentage error (MAPE) statistical processing. Of course, error statistical processing can also be implemented in other ways, which are not limited in the embodiments of the present application. Error statistical processing can be used to quantify the size of the consistency difference between different audio data. In some embodiments, taking the error statistical processing including root mean square error statistical processing and mean absolute percentage error statistical processing as an example, the above step S133 includes: performing root mean square error statistical processing on the similarities between multiple voiceprint sub-features and multiple original voiceprint features to obtain a first deviation value; performing mean absolute percentage error statistical processing on the similarities between multiple voiceprint sub-features and multiple original voiceprint features to obtain a second deviation value; and determining a voiceprint deviation value based on the first deviation value and the second deviation value. The first deviation value and the second deviation value can be averaged or weighted summed to obtain the voiceprint deviation value. By combining multiple error statistical processing, the accuracy of the voiceprint deviation value can be further improved.

[0069] In summary, the voiceprint processing method provided in the embodiment of the present application determines the standard voiceprint features and voiceprint deviation values of the reference speaker based on multiple audio data of the reference speaker. Among them, the standard voiceprint features of the reference speaker include multiple voiceprint features, and the voiceprint deviation value is used to indicate the deviation between the multiple voiceprint features. The standard voiceprint features and voiceprint deviation value can be used for voiceprint matching to achieve voiceprint recognition. For each speaker, the embodiment of the present application obtains multiple voiceprint features and voiceprint deviation values of the speaker based on the multiple audio data of the speaker. On the one hand, it can cover the various states of the speaker as much as possible, avoiding voiceprint recognition errors due to changes in the speaker's state; on the other hand, it can distinguish different speakers more accurately, avoiding voiceprint recognition errors due to the similarity of the speakers' voices. The embodiment of the present application improves the accuracy of voiceprint recognition by obtaining multiple voiceprint features and voiceprint deviation values of the speaker, and helps to achieve more accurate distinction between different speakers.

[0070] See also Figure 2 , Figure 2 This is a flow chart of a voiceprint processing method provided by an embodiment of the present application. Figure 2 As shown, the voiceprint processing method may include the following steps:

[0071] Step S200: performing voiceprint matching on the speaker to be identified and the reference speaker based on the voiceprint features of the speaker to be identified and the standard voiceprint features of the reference speaker.

[0072] The speaker to be identified can be any speaker. In order to confirm the identity of the speaker to be identified, the voiceprint features of the speaker to be identified can be obtained first. In some embodiments, the voiceprint features of the speaker to be identified can be extracted from the audio data of the speaker to be identified, so that the voiceprint features of the speaker to be identified accurately reflect the voice characteristics of the speaker to be identified. For the method of extracting the voiceprint features of the speaker to be identified, please refer to the above-mentioned method of extracting the original voiceprint features and voiceprint sub-features, which will not be described in detail here. For example, the above method also includes: inputting the audio data of the speaker to be identified into the voiceprint recognition model, so that the voiceprint recognition model outputs the voiceprint features of the speaker to be identified. Among them, the voiceprint recognition model can be an artificial intelligence model, for example, it can be a voiceprint recognition model based on deep learning, such as ResNet34, ResNet221, ECAPA, etc. The voiceprint features of the speaker to be identified can be embedded vectors.

[0073] In addition, a standard voiceprint feature of at least one reference speaker may be obtained. For example, a pre-stored standard voiceprint feature of at least one reference speaker may be obtained from a voiceprint database. The standard voiceprint feature of the reference speaker may include multiple voiceprint features, for example, multiple original voiceprint features and one fused voiceprint feature.

[0074] In the embodiments of the present application, voiceprint matching is performed between the speaker to be identified and a reference speaker based on the voiceprint features of the speaker to be identified and the standard voiceprint features of a reference speaker. Because the standard voiceprint features of a reference speaker include multiple voiceprint features, which can cover various states of the reference speaker and facilitate more refined differentiation between different speakers, matching the voiceprints of the speaker to be identified and the reference speaker based on the multiple voiceprint features of the reference speaker can improve the accuracy and precision of voiceprint recognition. For example, similarity calculations and matching score calculations can be performed based on the voiceprint features of the speaker to be identified and the standard voiceprint features of the reference speaker, and voiceprint matching can then be performed based on the calculation results. Similarity calculations include, but are not limited to, cosine similarity and Euclidean distance.

[0075] In some embodiments, the above step S200 may include the following steps S210 to S220:

[0076] Step S210: determining a voiceprint matching score between the speaker to be identified and the reference speaker based on the voiceprint features of the speaker to be identified and the standard voiceprint features of the reference speaker;

[0077] Step S220: performing voiceprint matching on the speaker to be identified and the reference speaker according to the voiceprint matching scores between the speaker to be identified and the reference speaker.

[0078] The voiceprint matching score is used to indicate the degree of voiceprint matching between the speaker to be identified and the reference speaker. Performing voiceprint matching based on the voiceprint matching score can make voiceprint matching more intuitive. The voiceprint matching score between the speaker to be identified and the reference speaker can be determined based on the feature similarity between the voiceprint features of the speaker to be identified and the standard voiceprint features of the reference speaker. In some embodiments, to improve the accuracy of the voiceprint matching score, the above step S210 may include the following steps:

[0079] Step S211: determining the voiceprint similarity between the speaker to be identified and the reference speaker based on the feature similarity between the voiceprint features of the speaker to be identified and the standard voiceprint features of the reference speaker;

[0080] Step S212: determining a voiceprint matching score between the speaker to be identified and the reference speaker based on the voiceprint similarity between the speaker to be identified and the reference speaker.

[0081] Because the standard voiceprint features include multiple voiceprint features, the voiceprint similarity between the speaker to be identified and the reference speaker can be determined based on the feature similarity between the voiceprint features of the speaker to be identified and the multiple voiceprint features of the reference speaker, making the voiceprint similarity more intuitive and accurate. For example, the feature similarities between the voiceprint features of the speaker to be identified and the multiple voiceprint features of the reference speaker can be averaged or weighted summed to determine the voiceprint similarity. In some embodiments, step S211 may include determining the voiceprint similarity between the speaker to be identified and the reference speaker based on a first similarity threshold and the feature similarity between the voiceprint features of the speaker to be identified and the standard voiceprint features of the reference speaker. A voiceprint matching score may be further determined based on the voiceprint similarity. In some embodiments, step S212 may include determining the voiceprint matching score between the speaker to be identified and the reference speaker based on a second similarity threshold and the voiceprint similarity between the speaker to be identified and the reference speaker.

[0082] The first similarity threshold and the second similarity threshold can be preset according to actual needs or determined based on test results. For example, voiceprint recognition can be performed in advance using a verification data set, and the false rejection rate and false acceptance rate of the verification data set can be calculated. When the false rejection rate is too high, the first similarity threshold and the second similarity threshold can be lowered; when the false acceptance rate is too high, the first similarity threshold and the second similarity threshold can be increased. In addition, the first similarity threshold and the second similarity threshold can be the same or different. For example, the first similarity threshold can be 0.20, 0.30, or 0.35, and the second similarity threshold can be 0.58, 0.6, or 0.65, etc. The embodiment of the present application not only considers the best match on a single dimension, but also introduces statistical principles to optimize the decision-making process, which helps to reduce the false acceptance rate while maintaining a good user experience. By adjusting the values of the first similarity threshold and the second similarity threshold, it can adapt to changing needs in various application scenarios.

[0083] Based on the first similarity threshold, the feature similarity between the voiceprint features of the speaker to be identified and the standard voiceprint features of the reference speaker can be screened to eliminate potential outliers or misidentifications, thereby improving the reliability of voiceprint recognition. In some embodiments, the above step S211 may include the following steps S2111 to S2112:

[0084] Step S2111: screening the feature similarities between the voiceprint features of the speaker to be identified and multiple voiceprint features of reference speakers based on a first similarity threshold;

[0085] Step S2112: Determine the voiceprint similarity between the speaker to be identified and the reference speaker based on the feature similarity retained after screening.

[0086] In step S2111, a first similarity threshold can be used to filter out feature similarities that are significantly too high or significantly too low to avoid misidentification. Thus, in some embodiments, step S2111 includes: determining a statistical number of feature similarities with values greater than the first similarity threshold based on the feature similarities between the voiceprint features of the speaker to be identified and multiple voiceprint features of the reference speaker; and filtering the feature similarities between the voiceprint features of the speaker to be identified and multiple voiceprint features of the reference speaker based on the statistical number. The filtering of the feature similarities between the voiceprint features of the speaker to be identified and multiple voiceprint features of the reference speaker based on the statistical number includes: if the statistical number is greater than the first number threshold, eliminating the feature similarity with the smallest value from the feature similarities between the voiceprint features of the speaker to be identified and the multiple voiceprint features of the reference speaker; and if the statistical number is less than the first number threshold, eliminating the feature similarity with the largest value from the feature similarities between the voiceprint features of the speaker to be identified and the multiple voiceprint features of the reference speaker. If the statistical quantity is equal to the first quantity threshold, no feature similarity may be eliminated, the feature similarity with the largest value or the feature similarity with the smallest value may be eliminated, or any one or more feature similarities may be eliminated, which is not limited in this embodiment of the present application.

[0087] The first quantity threshold may be preset or determined based on actual conditions. For example, the first quantity threshold may be half, one-third, or two-thirds of the number of voiceprint features of the reference speaker. For example, if the first quantity threshold is half of the number of voiceprint features of the reference speaker, if more than half of the feature similarities are greater than the first similarity threshold, the feature similarity with the smallest value is eliminated; if more than half of the feature similarities are less than the first similarity threshold, the feature similarity with the largest value is eliminated.

[0088] In step S2112, based on the feature similarities retained after screening, the voiceprint similarity between the speaker to be identified and the reference speaker can be determined. The voiceprint similarity can indicate the overall matching level between the speaker to be identified and the reference speaker. In some embodiments, the above step S2112 includes: averaging the feature similarities retained after screening to obtain the voiceprint similarity between the speaker to be identified and the reference speaker. Of course, the embodiments of the present application do not exclude other ways of determining voiceprint similarity, for example, arbitrarily selecting a feature similarity from the feature similarities retained after screening as the voiceprint similarity; or performing weighted summation on the feature similarities retained after screening to obtain the voiceprint similarity.

[0089] According to the second similarity threshold, the voiceprint matching score between the speaker to be identified and the reference speaker can be determined. In some embodiments, the above step S212 may include the following steps S2121 to S2122:

[0090] Step S2121: If the voiceprint similarity between the speaker to be identified and the reference speaker is greater than a second similarity threshold, a voiceprint matching score between the speaker to be identified and the reference speaker is determined based on the voiceprint deviation value of the reference speaker and the voiceprint similarity between the speaker to be identified and the reference speaker;

[0091] Step S2122: If the voiceprint similarity between the speaker to be identified and the reference speaker is less than the second similarity threshold, the voiceprint similarity between the speaker to be identified and the reference speaker is used as the voiceprint matching score between the speaker to be identified and the reference speaker.

[0092] The voiceprint deviation value of the reference speaker is used to indicate the deviation between the multiple voiceprint features of the reference speaker. If the voiceprint similarity between the speaker to be identified and the reference speaker is greater than the second similarity threshold, it means that the reference speaker is more matched with the speaker to be identified. At this time, the voiceprint matching score between the speaker to be identified and the reference speaker can be determined in combination with the voiceprint deviation value of the reference speaker to increase the importance of the reference speaker. In some embodiments, the above step S2121 may include: summing the product of the voiceprint deviation value of the reference speaker and the voiceprint similarity between the speaker to be identified and the reference speaker, and the voiceprint similarity between the speaker to be identified and the reference speaker to obtain the voiceprint matching score between the speaker to be identified and the reference speaker. That is, the voiceprint matching score between the speaker to be identified and the reference speaker = the voiceprint similarity between the speaker to be identified and the reference speaker × (1 + the voiceprint deviation value of the reference speaker).

[0093] In step S220, voiceprint matching is performed between the speaker to be identified and the reference speakers based on the voiceprint matching scores between the speaker to be identified and the reference speakers. For example, for any reference speaker, the voiceprint matching score between the speaker to be identified and the reference speaker can be compared with a set threshold, and based on the comparison result, whether the voiceprints of the speaker to be identified and the reference speakers match is determined. In some embodiments, there may be multiple reference speakers, and thus voiceprint matching can be performed between the speaker to be identified and the reference speakers based on the voiceprint matching scores between the speaker to be identified and the reference speakers. In some embodiments, step S220 may include selecting the reference speaker with the highest voiceprint matching score among the voiceprint matching scores between the speaker to be identified and the reference speakers as the reference speaker that matches the speaker to be identified. Of course, in embodiments of the present application, the reference speakers with the highest voiceprint matching scores may also be selected as the reference speakers that match the speaker to be identified. A match between the speaker to be identified and the reference speakers means that the speaker to be identified and the reference speakers are the same person, and their identities may be the same.

[0094] In some embodiments, the embodiments of the present application may also set a matching threshold, such as a matching threshold of 0.8, 0.88, or 0.9. If the voiceprint matching score between the speaker to be identified and one or more reference speakers is greater than or equal to the matching threshold, the reference speaker corresponding to the largest voiceprint matching score is used as the reference speaker matching the speaker to be identified; if the voiceprint matching scores between the speaker to be identified and all reference speakers are less than the matching threshold, the speaker to be identified fails to match the multiple reference speakers, and there is no reference speaker among the multiple reference speakers that matches the speaker to be identified.

[0095] To ensure that the standard voiceprint features of the reference speakers in the voiceprint database can adapt to changes in the reference speakers' voices, if the speaker to be identified successfully matches the voiceprints of multiple reference speakers, the standard voiceprint features of the matched reference speakers can be updated. Of course, if the speaker to be identified does not match any of the reference speakers, the standard voiceprint features of the reference speakers will not be updated.

[0096] In some embodiments, the above method further comprises the following steps:

[0097] Step S300: updating the standard voiceprint features of a reference speaker that matches the speaker to be identified based on the voiceprint features of the speaker to be identified.

[0098] If the voiceprint of the speaker to be identified successfully matches that of a reference speaker, the standard voiceprint features of the reference speaker that matches the speaker to be identified can be updated based on the voiceprint features of the speaker to be identified in step S200. The standard voiceprint features include multiple voiceprint features. In this embodiment of the application, the voiceprint features of the speaker to be identified can be added to the standard voiceprint features of the reference speaker, or any one of the standard voiceprint features of the reference speaker can be replaced with the voiceprint features of the speaker to be identified.

[0099] In some embodiments, the above step S300 may include the following steps S310 to S320:

[0100] Step S310: Determine the number of target voiceprint features in the standard voiceprint features of the reference speaker that matches the speaker to be identified; wherein the target voiceprint features refer to the voiceprint features added during the updating process;

[0101] Step S320: updating the standard voiceprint features of the reference speaker that matches the speaker to be identified according to the number of target voiceprint features and the voiceprint features of the speaker to be identified.

[0102] The standard voiceprint features of the reference speaker in the voiceprint database may have been updated multiple times. For example, the reference speaker may have been successfully matched multiple times. In this case, some voiceprint features may have been added to the standard voiceprint features of the reference speaker during the update process. In this embodiment of the application, the voiceprint features added during the update process are referred to as target voiceprint features. For example, the standard voiceprint features of the reference speaker in the voiceprint database can be updated actively and passively. Active update refers to updating the voiceprint features and voiceprint deviation values of the reference speaker when the reference speaker registers his voiceprint, and passive update refers to updating the voiceprint features of the reference speaker when the reference speaker is successfully matched. Among them, the target voiceprint features are the voiceprint features added during the passive update. To ensure the stability and robustness of the voiceprint data, the standard voiceprint features and voiceprint deviation values added during the active update can be fixed, while the target voiceprint features added during the passive update can be deleted or replaced.

[0103] In some embodiments, the above step S320 may include: if the number of target voiceprint features is less than a second quantity threshold, then the voiceprint features of the speaker to be identified are added to the standard voiceprint features of the reference speaker that matches the speaker to be identified; if the number of target voiceprint features is greater than the second quantity threshold, then any one of the target voiceprint features of the reference speaker that matches the speaker to be identified is replaced with the voiceprint features of the speaker to be identified. That is, if the number of target voiceprint features of the reference speaker is less than the second quantity threshold, then the voiceprint features of the speaker to be identified are directly added to the standard voiceprint features of the reference speaker; if the number of target voiceprint features of the reference speaker is greater than or equal to the second quantity threshold, then any one of the target voiceprint features of the reference speaker can be replaced with the voiceprint features of the speaker to be identified. The second quantity threshold can be set based on actual conditions and is not limited in this embodiment of the present application.

[0104] In summary, the voiceprint processing method provided in the embodiments of the present application performs voiceprint matching between a speaker to be identified and a reference speaker based on the voiceprint characteristics of the speaker to be identified and the standard voiceprint characteristics of a reference speaker. During the voiceprint matching process, the voiceprint similarity between the speaker to be identified and the reference speaker is determined using a first similarity threshold and the feature similarity between the voiceprint characteristics of the speaker to be identified and the standard voiceprint characteristics of the reference speaker. The voiceprint matching score between the speaker to be identified and the reference speaker is determined using a second similarity threshold and the voiceprint similarity between the speaker to be identified and the reference speaker. By setting these two similarity thresholds, the accuracy of voiceprint recognition is improved. Among them, by screening the feature similarity between the speaker to be identified and the reference speaker through the first similarity threshold, potential outliers or misidentification situations can be eliminated; by correcting the voiceprint similarity between the speaker to be identified and the reference speaker through the second similarity threshold, the importance of the reference speaker matching the speaker to be identified is increased; by combining the first similarity threshold and the second similarity threshold, the accuracy of voiceprint recognition is improved; by adjusting the values of the first similarity threshold and the second similarity threshold, it can adapt to changes in needs in various application scenarios.

[0105] In addition, the voiceprint processing method provided in the embodiment of the present application updates the standard voiceprint features of the reference speaker that matches the speaker to be identified based on the voiceprint features of the speaker to be identified. This ensures that the standard voiceprint features of the reference speaker adapt to changes in the reference speaker's voice, reduces the problem of increased false recognition rate due to natural aging or other factors, and increases the diversity of voiceprint features, which helps to improve the accuracy of subsequent voiceprint recognition. In addition, the voiceprint processing method provided in the embodiment of the present application can add the voiceprint features of the speaker to be identified to the standard voiceprint features of the reference speaker, or use the voiceprint features of the speaker to be identified to replace one of the standard voiceprint features of the reference speaker. This not only achieves the update of the standard voiceprint features of the reference speaker, but also helps to reduce storage pressure. In particular, for large-scale deployment scenarios, it reasonably balances the relationship between data storage capacity and voiceprint recognition efficiency.

[0106] The following describes the voiceprint processing method provided in the embodiments of the present application with several examples.

[0107] See also Figure 3 , Figure 3 This is a schematic diagram of a voiceprint processing process provided by an embodiment of the present application. Figure 3 As shown, after obtaining the speaker's audio data, front-end processing such as noise reduction and signal processing can be performed; if the speaker is registering for the first time, the voiceprint registration process can be entered. After the registration is completed, the speaker's standard voiceprint features and voiceprint deviation values are saved in the voiceprint database; if the speaker is matched, the voiceprint matching process is entered. If the voiceprint match is successful, the standard voiceprint features of the matching speaker in the voiceprint database are updated. If the voiceprint match fails, the voiceprint features of the speaker are directly discarded.

[0108] See also Figure 4 , Figure 4 This is a schematic diagram of a voiceprint registration process provided by an embodiment of the present application. Figure 4 As shown, the voiceprint registration process may include the following steps:

[0109] Step S401: inputting multiple audio data of the speaker into a voiceprint recognition model, so that the voiceprint recognition model outputs multiple original voiceprint features;

[0110] Step S402: fusing multiple original voiceprint features to obtain a fused voiceprint feature;

[0111] Step S403: extracting multiple audio sub-data from the multiple audio data of the speaker;

[0112] Step S404: inputting the speaker's multiple audio sub-data into the voiceprint recognition model, so that the voiceprint recognition model outputs multiple voiceprint sub-features;

[0113] Step S405: determining similarities between the plurality of original voiceprint features and the plurality of voiceprint sub-features;

[0114] Step S406: performing error statistical processing on the similarities between the multiple original voiceprint features and the multiple voiceprint sub-features to obtain a voiceprint deviation value;

[0115] Step S407: The speaker's multiple original voiceprint features, fused voiceprint features, and voiceprint deviation values are stored in the voiceprint database. The speaker's standard voiceprint features include the speaker's multiple original voiceprint features and fused voiceprint features, i.e., the speaker's standard voiceprint features include multiple voiceprint features.

[0116] For the convenience of description, in the voiceprint matching process, the speaker to be identified is called the speaker to be identified, and the speaker in the voiceprint database is called the reference speaker.

[0117] See also Figure 5 , Figure 5 This is a schematic diagram of a voiceprint matching process provided by an embodiment of the present application. Figure 5 As shown, the voiceprint matching process may include the following steps:

[0118] Step S501: inputting audio data of a speaker to be identified into a voiceprint recognition model, so that the voiceprint recognition model outputs the voiceprint features of the speaker to be identified;

[0119] Step S502: determining feature similarities between the voiceprint features of the speaker to be identified and multiple voiceprint features of multiple reference speakers;

[0120] Step S503: for each reference speaker, determining a statistical number of feature similarities having values greater than a first similarity threshold based on feature similarities between the voiceprint feature of the speaker to be identified and multiple voiceprint features of the reference speaker;

[0121] Step S504: Determine whether the statistical quantity is greater than or equal to a first quantity threshold; if it is greater than or equal to the first quantity threshold, execute the following step S505; otherwise, execute the following step S506;

[0122] Step S505: Eliminate the feature similarity with the smallest value from the feature similarities between the voiceprint feature of the speaker to be identified and multiple voiceprint features of the reference speaker;

[0123] Step S506: Eliminate the feature similarity with the largest value from the feature similarities between the voiceprint feature of the speaker to be identified and multiple voiceprint features of the reference speaker;

[0124] Step S507: determining the voiceprint similarity between the speaker to be identified and the reference speaker based on the feature similarity retained after screening;

[0125] Step S508: for each reference speaker, determine whether the voiceprint similarity between the speaker to be identified and the reference speaker is greater than a second similarity threshold; if it is greater than the second similarity threshold, execute the following step S509; otherwise, execute the following step S510;

[0126] Step S509: summing the product of the reference speaker's voiceprint deviation value and the voiceprint similarity between the speaker to be identified and the reference speaker, and the voiceprint similarity between the speaker to be identified and the reference speaker, to obtain a voiceprint matching score between the speaker to be identified and the reference speaker;

[0127] Step S510: using the voiceprint similarity between the speaker to be identified and the reference speaker as a voiceprint matching score between the speaker to be identified and the reference speaker;

[0128] Step S511: The reference speaker having the largest voiceprint matching score among the voiceprint matching scores between the speaker to be identified and multiple reference speakers is used as the reference speaker that matches the speaker to be identified.

[0129] See also Figure 6 , Figure 6 This is a schematic diagram of the update process of a voiceprint database provided by an embodiment of the present application. Figure 6 As shown, the updating process of the voiceprint database may include the following steps:

[0130] Step S601: when the voiceprint of the speaker to be identified matches the voiceprint of the reference speaker successfully, determining the number of target voiceprint features in the standard voiceprint features of the reference speaker;

[0131] Step S602: Determine whether the number of target voiceprint features is greater than a second number threshold; if so, execute the following step S603; otherwise, execute the following step S604;

[0132] Step S603: replacing any one of the target voiceprint features of the reference speaker with the voiceprint feature of the speaker to be identified;

[0133] Step S604: adding the voiceprint features of the speaker to be identified to the standard voiceprint features of the reference speaker.

[0134] According to a second aspect of the present application, embodiments of the present application further provide a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the above-described voiceprint processing method. This non-transitory computer-readable storage medium has all the beneficial effects of the above-described voiceprint processing method, and this application will not further elaborate on them.

[0135] According to the third aspect of the present application, an embodiment of the present application further provides a computer program product, including a computer program. When the computer program is executed by a processor, it implements the above-mentioned voiceprint processing method and has all the beneficial effects of the above-mentioned voiceprint processing method. This application will not go into details here.

[0136] According to a fourth aspect of the present application, an embodiment of the present application further provides an electronic device comprising: a memory and a processor, wherein the memory stores a computer program; the processor is configured to execute the computer program in the memory to implement the steps of the above-described voiceprint processing method. This electronic device has all the beneficial effects of the above-described voiceprint processing method, and this application will not further elaborate on them.

[0137] The computer-readable storage medium may be, for example, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or any combination thereof, and this application does not specifically limit this. More specific examples of computer-readable storage media may include, but are not limited to, an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.

[0138] In some embodiments of the present application, a computer-readable storage medium may be any tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device.

[0139] The computer-readable storage medium may be included in the electronic device or may exist independently without being incorporated into the electronic device. The computer-readable storage medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device:

[0140] Determining a standard voiceprint feature and a voiceprint deviation value of the reference speaker based on multiple audio data of the reference speaker;

[0141] The standard voiceprint feature includes multiple voiceprint features, and the voiceprint deviation value is used to indicate the deviation between the multiple voiceprint features; the standard voiceprint feature and the voiceprint deviation value are used for voiceprint matching.

[0142] Computer program code for performing the operations of some embodiments of the present application can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer can be connected to the user's computer through any type of network (including a local area network (LAN) or a wide area network (WAN)), or can be connected to an external computer (for example, using an Internet service provider to connect via the Internet).

[0143] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of the systems, methods, and computer program products according to various embodiments of the present application. In this regard, each box in the flowchart or block diagram may represent a module, program segment, or portion of code, which contains one or more executable instructions for implementing a specified logical function.

[0144] It should also be noted that, in some alternative implementations, the functions noted in the block may occur out of the order noted in the figures.

[0145] For example, two blocks shown in succession may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functionality involved. It should also be noted that each block in the block diagrams and / or flow charts, and combinations of blocks in the block diagrams and / or flow charts, may be implemented using a dedicated hardware-based system that performs the specified functions or operations, or may be implemented using a combination of dedicated hardware and computer instructions.

[0146] The units described in some embodiments of the present application may be implemented in software or hardware, and may also be provided in a processor.

[0147] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: Field Programmable Gate Array (FPGA), Application Specific Integrated Circuit (ASIC), Application Specific Standard Parts (ASSP), System on Chip (SOC), Complex Programmable Logic Device (CPLD), and the like.

[0148] According to the fifth aspect of this application, Figure 7 As shown, the embodiment of the present application further provides a vehicle 10, which includes the above-mentioned electronic device. The vehicle has all the beneficial effects of the above-mentioned electronic device, etc., which will not be described in detail in this application.

[0149] The vehicle may be a fuel vehicle, a plug-in hybrid vehicle or a new energy vehicle, etc., and this application does not make any specific restrictions on this.

[0150] In the description of this application, the terms "first" and "second" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more features. In the description of this application, "plurality" means two or more, unless otherwise specifically defined.

[0151] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0152] The embodiments, implementation methods and related technical features of the present application can be combined and replaced with each other without conflict.

[0153] The above are merely preferred embodiments of the present application and do not constitute any form of limitation to the present application. Although the descriptions of each embodiment in the embodiments of the present application have different focuses, for parts that are not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments. However, any simple modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of the present application without departing from the content of the technical solution of the present application are still within the scope of the technical solution of the present application.

Claims

1. A voiceprint processing method, characterized in that: The method comprises: Determining a standard voiceprint feature and a voiceprint deviation value of the reference speaker based on multiple audio data of the reference speaker; The standard voiceprint feature includes multiple voiceprint features, and the voiceprint deviation value is used to indicate the deviation between the multiple voiceprint features; the standard voiceprint feature and the voiceprint deviation value are used for voiceprint matching.

2. The voiceprint processing method according to claim 1, characterized in that: The determining of the standard voiceprint features of the reference speaker includes: Extracting a plurality of original voiceprint features of a reference speaker from a plurality of audio data of the reference speaker; wherein the standard voiceprint features include a plurality of the original voiceprint features.

3. The voiceprint processing method according to claim 2, characterized in that: The method further comprises: The multiple original voiceprint features of the reference speaker are fused to obtain a fused voiceprint feature of the reference speaker; wherein the standard voiceprint feature includes the fused voiceprint feature.

4. The voiceprint processing method according to claim 3, characterized in that: The fusing of the plurality of original voiceprint features of the reference speaker comprises: Performing weighted summation processing on the multiple original voiceprint features of the reference speaker to obtain a fused voiceprint feature of the reference speaker; or, A plurality of the original voiceprint features of the reference speaker are averaged to obtain a fused voiceprint feature of the reference speaker.

5. The voiceprint processing method according to claim 2, characterized in that: Determining the voiceprint deviation value of the reference speaker includes: Extracting a plurality of audio sub-data from a plurality of audio data of a reference speaker; Extracting a plurality of voiceprint sub-features from the plurality of audio sub-data; Determine a voiceprint deviation value of the reference speaker based on the plurality of voiceprint sub-features.

6. The voiceprint processing method according to claim 5, characterized in that: Determining the voiceprint deviation value of the reference speaker includes: A voiceprint deviation value of the reference speaker is determined according to the plurality of voiceprint sub-features and the plurality of original voiceprint features.

7. The voiceprint processing method according to claim 6, characterized in that: Determining the voiceprint deviation value of the reference speaker includes: Error statistical processing is performed on the similarities between the plurality of voiceprint sub-features and the plurality of original voiceprint features to obtain a voiceprint deviation value of the reference speaker.

8. The voiceprint processing method according to claim 7, characterized in that: The error statistical processing includes at least one of the following: root mean square error statistical processing, mean absolute percentage error statistical processing.

9. The voiceprint processing method according to claim 8, characterized in that: The performing error statistical processing on the similarities between the plurality of voiceprint sub-features and the plurality of original voiceprint features includes: performing the root mean square error statistical processing on the similarities between the plurality of the voiceprint sub-features and the plurality of the original voiceprint features to obtain a first deviation value; performing mean absolute percentage error statistical processing on the similarities between the plurality of voiceprint sub-features and the plurality of original voiceprint features to obtain a second deviation value; The voiceprint deviation value is determined according to the first deviation value and the second deviation value.

10. The voiceprint processing method according to claim 1, characterized in that: The method further comprises: According to the voiceprint features of the speaker to be identified and the standard voiceprint features of the reference speaker, voiceprint matching is performed on the speaker to be identified and the reference speaker.

11. The voiceprint processing method according to claim 10, characterized in that: The performing voiceprint matching on the speaker to be identified and the reference speaker includes: determining a voiceprint matching score between the speaker to be identified and the reference speaker based on the voiceprint features of the speaker to be identified and the standard voiceprint features of the reference speaker; Voiceprint matching is performed on the speaker to be identified and the reference speaker according to the voiceprint matching score between the speaker to be identified and the reference speaker.

12. The voiceprint processing method according to claim 11, characterized in that: The determining of the voiceprint matching score between the speaker to be identified and the reference speaker includes: determining the voiceprint similarity between the speaker to be identified and the reference speaker based on the feature similarity between the voiceprint feature of the speaker to be identified and the standard voiceprint feature of the reference speaker; A voiceprint matching score between the speaker to be identified and the reference speaker is determined according to the voiceprint similarity between the speaker to be identified and the reference speaker.

13. The voiceprint processing method according to claim 12, characterized in that: The determining of the voiceprint similarity between the speaker to be identified and the reference speaker includes: The voiceprint similarity between the speaker to be identified and the reference speaker is determined based on a first similarity threshold and feature similarity between the voiceprint feature of the speaker to be identified and the standard voiceprint feature of the reference speaker.

14. The voiceprint processing method according to claim 13, characterized in that: The determining of the voiceprint similarity between the speaker to be identified and the reference speaker includes: screening, according to a first similarity threshold, feature similarities between the voiceprint feature of the speaker to be identified and the plurality of voiceprint features of the reference speaker; The voiceprint similarity between the speaker to be identified and the reference speaker is determined based on the feature similarity retained after screening.

15. The voiceprint processing method according to claim 14, characterized in that: The screening of feature similarities between the voiceprint feature of the speaker to be identified and the plurality of voiceprint features of the reference speaker includes: Determining a statistical number of feature similarities having a value greater than a first similarity threshold based on feature similarities between a voiceprint feature of the speaker to be identified and a plurality of voiceprint features of the reference speaker; The feature similarities between the voiceprint feature of the speaker to be identified and the plurality of voiceprint features of the reference speaker are screened according to the statistical quantity.

16. The voiceprint processing method according to claim 15, characterized in that: The screening of feature similarities between the voiceprint feature of the speaker to be identified and the plurality of voiceprint features of the reference speaker according to the statistical quantity includes: If the statistical quantity is greater than a first quantity threshold, the feature similarity with the smallest value is eliminated from the feature similarities between the voiceprint feature of the speaker to be identified and the multiple voiceprint features of the reference speaker.

17. The voiceprint processing method according to claim 15, characterized in that: The screening of feature similarities between the voiceprint feature of the speaker to be identified and the plurality of voiceprint features of the reference speaker according to the statistical quantity includes: If the statistical quantity is less than a first quantity threshold, the feature similarity with the largest value is eliminated from the feature similarities between the voiceprint feature of the speaker to be identified and the multiple voiceprint features of the reference speaker.

18. The voiceprint processing method according to claim 14, characterized in that: The determining of the voiceprint similarity between the speaker to be identified and the reference speaker includes: The feature similarities retained after screening are averaged to obtain the voiceprint similarity between the speaker to be identified and the reference speaker.

19. The voiceprint processing method according to claim 12, characterized in that: The determining of the voiceprint matching score between the speaker to be identified and the reference speaker includes: A voiceprint matching score between the speaker to be identified and the reference speaker is determined according to a second similarity threshold and the voiceprint similarity between the speaker to be identified and the reference speaker.

20. The voiceprint processing method according to claim 19, characterized in that: The determining of the voiceprint matching score between the speaker to be identified and the reference speaker includes: If the voiceprint similarity between the speaker to be identified and the reference speaker is greater than a second similarity threshold, a voiceprint matching score between the speaker to be identified and the reference speaker is determined based on the voiceprint deviation value of the reference speaker and the voiceprint similarity between the speaker to be identified and the reference speaker.

21. The voiceprint processing method according to claim 20, characterized in that: The determining of the voiceprint matching score between the speaker to be identified and the reference speaker includes: The product of the voiceprint deviation value of the reference speaker and the voiceprint similarity between the speaker to be identified and the reference speaker is summed, and the voiceprint similarity between the speaker to be identified and the reference speaker is summed to obtain a voiceprint matching score between the speaker to be identified and the reference speaker.

22. The voiceprint processing method according to claim 19, characterized in that: The determining of the voiceprint matching score between the speaker to be identified and the reference speaker includes: If the voiceprint similarity between the speaker to be identified and the reference speaker is less than the second similarity threshold, the voiceprint similarity between the speaker to be identified and the reference speaker is used as the voiceprint matching score between the speaker to be identified and the reference speaker.

23. The voiceprint processing method according to claim 11, characterized in that: There are multiple reference speakers; and performing voiceprint matching on the speaker to be identified and the reference speakers includes: The reference speaker corresponding to the largest voiceprint matching score among the voiceprint matching scores between the speaker to be identified and the multiple reference speakers is used as the reference speaker that matches the speaker to be identified.

24. The voiceprint processing method according to claim 10, characterized in that: The method further comprises: Extracting the voiceprint features of the speaker to be identified from the audio data of the speaker to be identified.

25. The voiceprint processing method according to claim 10, characterized in that: The method further comprises: The standard voiceprint feature of the reference speaker that matches the speaker to be identified is updated according to the voiceprint feature of the speaker to be identified.

26. The voiceprint processing method according to claim 25, characterized in that: The updating of the standard voiceprint feature of the reference speaker that matches the speaker to be identified includes: Determining the number of target voiceprint features in the standard voiceprint features of the reference speaker that matches the speaker to be identified; wherein the target voiceprint features refer to the voiceprint features added during the updating process; The standard voiceprint features of the reference speaker that matches the speaker to be identified are updated according to the number of the target voiceprint features and the voiceprint features of the speaker to be identified.

27. The voiceprint processing method according to claim 26, characterized in that: The updating of the standard voiceprint feature of the reference speaker that matches the speaker to be identified includes: If the number of the target voiceprint features is less than a second number threshold, the voiceprint features of the speaker to be identified are added to the standard voiceprint features of the reference speaker that matches the speaker to be identified.

28. The voiceprint processing method according to claim 26, characterized in that: The updating of the standard voiceprint feature of the reference speaker that matches the speaker to be identified includes: If the number of the target voiceprint features is greater than a second number threshold, any one of the target voiceprint features of the reference speaker that matches the speaker to be identified is replaced with the voiceprint feature of the speaker to be identified.

29. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the voiceprint processing method according to any one of claims 1 to 28 is implemented.

30. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the voiceprint processing method according to any one of claims 1 to 28 is implemented.

31. An electronic device, characterized in that: include: a memory having a computer program stored thereon; A processor, configured to execute the computer program in the memory to implement the voiceprint processing method according to any one of claims 1 to 28.

32. A vehicle, characterized in that: Comprising the electronic device as claimed in claim 31.