Speaker confirmation method, device, equipment and storage medium

By segmenting the effective audio of the speaker to be identified and inputting the preset speaker confirmation model, the model is constructed using high-quality sample audio and audio implicit vector information, the problem of low recognition accuracy of speakers under low-quality channels is solved, and a higher recognition accuracy is achieved.

CN119495306BActive Publication Date: 2025-05-30中邮消费金融有限公司
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510073943.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-17
Publication Date
2025-05-30
Estimated Expiration
2045-01-17

AI Technical Summary

Technical Problem

The speaker confirmation model in the prior art is not very accurate when performing speaker recognition in low-quality channels in practical application scenarios.

Method used

By dividing the effective audio of the speaker to be identified into several clause audio and inputting it into the preset speaker confirmation model for processing, the preset model is constructed based on high-quality sample audio and audio implicit vector information, and finally identity identification is based on the speaker average feature vector and standard feature vector.

Benefits of technology

The recognition accuracy of the speaker confirmation model under low quality channels is improved, and the problem of low accuracy in the prior art is solved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119495306B_ABST
    Figure CN119495306B_ABST
Patent Text Reader

Abstract

The present application discloses a speaker verification method, apparatus, device, and storage medium, relating to the technical field of information processing. The method includes: segmenting the valid audio corresponding to the speaker to be recognized into a plurality of clause audios; inputting each clause audio into a preset speaker verification model to obtain a plurality of speaker feature vectors, where the preset speaker verification model is constructed based on high-quality sample audio and audio implicit vector information; determining a speaker average feature vector according to each speaker feature vector; and performing identity recognition on the speaker to be recognized based on the speaker average feature vector and the standard speaker feature vector. By applying the above technical solution, the technical problem that the accuracy of the speaker verification model in the prior art is not high when performing speaker recognition in a low-quality channel in an actual application scenario is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of information processing technologies, and in particular, to a speaker verification method, apparatus, device, and storage medium. Background Art

[0002] In the consumer finance scenario, for reasons of protecting property safety and anti-fraud, it is necessary to verify the identity of users. Among them, one of the identity verification steps is speaker verification to confirm whether the user's voice is the same as the speaker when the user registered through the user's call voice.

[0003] In the existing solutions, a phased speaker verification model can be used to compare the registered audio and the audio to be verified, so as to determine whether the two audio segments are spoken by the same speaker. Specifically, the audio segments to be compared can be input into the feature extraction model to extract audio vectors first, and then a similarity calculation method such as cosine distance is used to calculate the similarity score of a pair of audio feature vectors, and speaker recognition is achieved based on the similarity score. However, the training and test sets used by most models are lossless audio in the original wav format with a sampling rate of 16k. In low-quality channels in actual application scenarios, such as telephone call channels, it is more common to use audio with a lower sampling rate or compressed audio for convenient storage. At this time, the accuracy of model recognition is not high and the recognition effect is not good.

[0004] The above content is only used to assist in understanding the technical solution of the present invention, and does not represent an admission that the above content is prior art. Summary of the Invention

[0005] The main purpose of this application is to provide a speaker verification method, apparatus, device, and storage medium, aiming to solve the technical problem of low accuracy in speaker recognition by the existing speaker verification model in low-quality channels in actual application scenarios.

[0006] To achieve the above purpose, this application proposes a speaker verification method, and the method includes:

[0007] Segment the valid audio corresponding to the speaker to be recognized into several clause audio segments;

[0008] Input each clause audio segment into a preset speaker verification model to obtain several speaker feature vectors, and the preset speaker verification model is constructed based on high-quality sample audio and audio latent vector information;

[0009] Determine the speaker average feature vector according to each speaker feature vector;

[0010] Perform identity recognition on the speaker to be recognized based on the speaker average feature vector and the standard speaker feature vector.

[0011] In one embodiment, before the step of splitting the effective audio corresponding to the speaker to be recognized into a plurality of clause audios, the method further includes:

[0012] Performing audio restoration processing and implicit vector extraction processing on the low-quality audio through a preset audio restoration model to obtain high-quality audio and implicit vector information, where the implicit vector information is the audio information during audio restoration of the low-quality audio;

[0013] Training an initial speaker verification model based on the high-quality audio and the implicit vector information to obtain a preset speaker verification model.

[0014] In one embodiment, an audio feature extraction module, a long and short speech statistical aggregation module, and a feature projection module are provided in the initial speaker verification model; the step of training the initial speaker verification model based on the high-quality audio and the implicit vector information to obtain a preset speaker verification model includes:

[0015] Adding an implicit vector to the high-quality audio by the audio feature extraction module based on the implicit vector information to obtain processed high-quality audio;

[0016] Performing attention transformation on the processed high-quality audio by the audio feature extraction module in the channel dimension and the time series dimension to obtain a plurality of input audio features with different channels;

[0017] Splitting the features of the input audio features by the long and short speech statistical aggregation module to obtain short audio features corresponding to the input audio features;

[0018] Generating an input audio feature matrix corresponding to the input audio features and a short audio feature matrix corresponding to the short audio features;

[0019] Performing feature splicing and fusion processing on the input audio feature matrix and the short audio feature matrix to obtain an aggregated audio feature matrix;

[0020] Projecting the aggregated audio feature matrix into a low-dimensional space by the feature projection module to obtain projected audio features;

[0021] Training the initial speaker verification model based on the projected audio features to obtain a preset speaker verification model.

[0022] In one embodiment, the step of training the initial speaker verification model based on the projected audio features to obtain a preset speaker verification model includes:

[0023] Determine the feature similarity between the projected audio features and the audio features to be compared through the initial speaker verification model;

[0024] Obtain the speaker recognition result according to the feature similarity;

[0025] Determine the false rejection rate and false acceptance rate of the initial speaker verification model based on the speaker recognition result;

[0026] Optimize the initial speaker verification model based on the false rejection rate and the false acceptance rate to obtain a preset speaker verification model.

[0027] In one embodiment, before the step of performing audio restoration processing and implicit vector extraction processing on the low-quality audio through the preset audio restoration model to obtain high-quality audio and implicit vector information, it further includes:

[0028] Collect a number of high-quality sample audios;

[0029] Perform audio compression processing on the high-quality sample audios to obtain low-quality sample audios;

[0030] Train the initial audio restoration model based on the low-quality sample audios to obtain a preset audio restoration model.

[0031] In one embodiment, the initial audio restoration model is provided with: an audio encoder, a diffusion restoration module, and an audio decoder; the step of training the initial audio restoration model based on the low-quality sample audios to obtain a preset audio restoration model includes:

[0032] Input the low-quality sample audios into the initial audio restoration model;

[0033] Perform compression encoding processing on the low-quality sample audios through the audio encoder to obtain low-quality audio implicit vectors;

[0034] Perform audio restoration on the low-quality audio implicit vectors through the diffusion restoration module to obtain high-quality audio implicit vectors;

[0035] Perform decoding on the high-quality audio implicit vectors through the audio decoder to obtain high-quality predicted audios;

[0036] Optimize the initial audio restoration model based on the high-quality sample audios and the high-quality predicted audios to obtain a preset audio restoration model.

[0037] In one embodiment, the step of optimizing the initial audio restoration model based on the high-quality sample audios and the high-quality predicted audios to obtain a preset audio restoration model includes:

[0038] Obtain a first audio spectrum corresponding to the high-quality sample audio and a second audio spectrum corresponding to the high-quality predicted audio;

[0039] Determine the spectral convergence degree between the first audio spectrum and the second audio spectrum;

[0040] Optimize the initial audio restoration model based on the spectral convergence degree to obtain a preset audio restoration model.

[0041] In addition, to achieve the above object, the present application also proposes a speaker verification device, which includes:

[0042] An audio segmentation module, configured to segment the valid audio corresponding to the speaker to be recognized into a plurality of clause audios;

[0043] A feature extraction module, configured to input each clause audio into a preset speaker verification model to obtain a plurality of speaker feature vectors, and the preset speaker verification model is constructed based on high-quality sample audio and audio latent vector information;

[0044] A feature vector determination module, configured to determine a speaker average feature vector according to each speaker feature vector;

[0045] A speaker recognition module, configured to perform identity recognition on the speaker to be recognized based on the speaker average feature vector and the standard speaker feature vector.

[0046] In addition, to achieve the above object, the present application also proposes a speaker verification device, which includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, and the computer program is configured to implement the steps of the speaker verification method as described above.

[0047] In addition, to achieve the above object, the present application also proposes a storage medium, which is a computer-readable storage medium, and a computer program is stored on the storage medium, and when the computer program is executed by a processor, it implements the steps of the speaker verification method as described above.

[0048] The present application provides a speaker verification method. It discloses that the valid audio corresponding to the speaker to be recognized is segmented into several clause audios; each clause audio is input into a preset speaker verification model to obtain several speaker feature vectors, and the preset speaker verification model is constructed based on high-quality sample audio and audio implicit vector information; the speaker average feature vector is determined according to each speaker feature vector; the speaker to be recognized is identified based on the speaker average feature vector and the standard speaker feature vector; compared with the prior art, both the training set and the test set used in the speaker verification model are lossless audio in the original wav format with a sampling rate of 16k, but when recognizing audio with a lower sampling rate or compressed audio in a low-quality channel in the actual application scenario, the accuracy is not high. Since the present invention can pre-construct a preset speaker verification model based on high-quality sample audio and audio implicit vectors, then obtain the speaker feature vectors corresponding to the valid audio of the speaker to be recognized based on the preset speaker verification model, and finally perform speaker recognition based on the corresponding speaker average feature vector and the standard speaker feature vector, thereby solving the technical problem that the accuracy of the speaker verification model in the prior art is not high when performing speaker recognition in a low-quality channel in the actual application scenario. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] The accompanying drawings herein are incorporated into and constitute a part of this specification, showing embodiments consistent with the present application and, together with the specification, are used to explain the principles of the present application.

[0050] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0051] Figure 1 It is a schematic flowchart provided for Embodiment 1 of the speaker verification method of the present application;

[0052] Figure 2 It is a schematic overall flowchart of the speaker verification method of the present application;

[0053] Figure 3 It is a schematic flowchart provided for Embodiment 2 of the speaker verification method of the present application;

[0054] Figure 4 It is a schematic structural diagram of the speaker verification model in the speaker verification method of the present application;

[0055] Figure 5 It is a schematic flowchart provided for Embodiment 3 of the speaker verification method of the present application;

[0056] Figure 6Schematic structural diagram of the audio restoration model in the speaker verification method of the present application;

[0057] Figure 7 Schematic module structure diagram of the speaker verification device according to an embodiment of the present application;

[0058] Figure 8 Schematic device structure diagram of the hardware operating environment involved in the speaker verification method according to an embodiment of the present application.

[0059] The implementation, functional features, and advantages of the present application will be further described with reference to the embodiments and the accompanying drawings. Detailed implementation manners

[0060] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of the present application and are not used to limit the present application.

[0061] To better understand the technical solutions of the present application, the following will be described in detail with reference to the accompanying drawings of the specification and specific implementation manners.

[0062] The main solution of the embodiment of the present application is: dividing the effective audio corresponding to the speaker to be recognized into several clause audios; inputting each clause audio into a preset speaker verification model to obtain several speaker feature vectors, where the preset speaker verification model is constructed based on high-quality sample audio and audio implicit vector information; determining the average speaker feature vector according to each speaker feature vector; and performing identity recognition on the speaker to be recognized based on the average speaker feature vector and the standard speaker feature vector.

[0063] Since most of the existing speaker verification models use the original wav format lossless audio with a sampling rate of 16k for training and testing sets, in low-quality channels in actual application scenarios, such as telephone call channels, in order to facilitate storage, it is more common to use audio with a lower sampling rate or compressed audio. At this time, the accuracy of model recognition is not high.

[0064] The present application provides a solution. A preset speaker verification model can be constructed in advance based on high-quality sample audio and audio implicit vectors, and then the speaker feature vectors corresponding to the effective audio of the speaker to be recognized can be obtained based on the preset speaker verification model. Finally, speaker recognition is performed based on the corresponding average speaker feature vector and the standard speaker feature vector, thereby solving the technical problem that the accuracy of the existing speaker verification model is not high when performing speaker recognition in low-quality channels in actual application scenarios.

[0065] It should be noted that the execution subject of this embodiment can be a computing service device with data processing, network communication, and program running functions, such as a tablet computer, a personal computer, a mobile phone, etc., or an electronic device, a speaker verification device, etc. that can implement the above functions. Hereinafter, taking the speaker verification device as an example (hereinafter referred to as the device), this embodiment and the following embodiments will be described.

[0066] Based on this, an embodiment of the present application provides a speaker verification method. Referring to Figure 1 , Figure 1 is a schematic flowchart of the first embodiment of the speaker verification method of the present application.

[0067] In this embodiment, the speaker verification method includes steps S10 to S40:

[0068] Step S10: Segment the valid audio corresponding to the speaker to be recognized into a plurality of clause audio.

[0069] It should be noted that the above valid audio can be the audio obtained by removing the silent part and non-speech part in the audio of the speaker to be recognized collected by using valid voice detection. In practical applications, this solution can be applied but is not limited to speaker recognition in financial call scenarios.

[0070] It should be understood that the above clause audio can be the audio obtained by segmenting the valid audio. In this embodiment, for the financial call scenario, the device can obtain the vocal tract audio of the speaker to be recognized in real time during the call, and use valid voice detection to remove the silent part and non-speech part in the audio to obtain the valid audio of the speaker to be recognized, and then use valid voice detection to segment it into multiple audio clause samples to obtain the above clause audio.

[0071] Step S20: Input each clause audio into a preset speaker verification model to obtain a plurality of speaker feature vectors, and the preset speaker verification model is constructed based on high-quality sample audio and audio implicit vector information.

[0072] It should be noted that the above high-quality sample audio can be high-sampling-rate and lossless sample audio used for training the speaker verification model; the above audio implicit vector information can be audio information that helps to train the speaker verification model during the process of restoring low-quality audio to high-quality audio.

[0073] In practical applications, before training a preset speaker verification model, audio restoration processing can be performed on low-quality sample audio collected from an actual scenario to restore it to high-quality sample audio, and audio information that helps train the speaker verification model during the audio restoration process of the low-quality sample audio can be collected to obtain audio latent vector information, so as to jointly improve the recognition effect of the speaker verification model on low-quality channel audio through the high-quality sample audio and the audio latent vector information. Among them, the latent vector information is generally one or a group of vectors of the latent audio features obtained through training in the audio restoration model.

[0074] It should be noted that the above-mentioned preset speaker verification model can be constructed and trained using a deep learning-based model. In this embodiment, the device can input the high-quality sample audio and the audio latent vector information into the preset speaker verification model. At this time, the model can respectively extract features that can be used to compare and confirm the speaker for identifying the speaker, that is, the above-mentioned speaker feature vectors.

[0075] Step S30: Determine the speaker average feature vector according to each speaker feature vector.

[0076] It can be understood that the above-mentioned speaker average feature vector is the average vector of the speaker feature vectors. In this embodiment, using the speaker average feature vector to identify the speaker can reduce the differences between the features of different audio samples, thereby improving the accuracy of speaker recognition.

[0077] Step S40: Perform identity recognition on the to-be-identified speaker based on the speaker average feature vector and the standard speaker feature vector.

[0078] It should be noted that the above-mentioned standard speaker feature vector can be the feature vectors corresponding to the comparison audio pre-recorded by all speakers during registration.

[0079] In practical applications, the device can compare the speaker average feature vector of the to-be-identified speaker with the feature vector corresponding to the audio for comparison to determine whether the valid audio input to the preset speaker verification model and the audio to be compared belong to the same speaker, thereby realizing speaker recognition.

[0080] In specific implementation, refer to Figure 2 , Figure 2 is the schematic diagram of the overall flowchart of the speaker verification method of this application. As Figure 2As shown, this solution can first collect high-quality audio (i.e., high-sampling-rate, lossless audio), and generate low-quality audio (i.e., low-sampling-rate, lossy audio) through downsampling and audio compression processing. Then, an audio restoration model can be trained based on the high-quality audio and the low-quality audio. Before constructing the speaker verification model, the low-quality audio can be restored to high-quality audio through the audio restoration model to obtain audio latent vector information, and then the high-quality audio and the audio latent vector information are jointly used as inputs to construct and train the speaker verification model. After training, the above-mentioned preset speaker verification model is obtained. It should be noted that since high-quality audio and the speaker verification model are sensitive to noise (including stationary noise, transient noise, other speech noises, etc.), before inputting the audio into the speaker verification model, noise can be reduced through an adaptive filter or a noise reduction model. Among them, stationary noise can be estimated in advance through the adaptive filter during the period without speech and applied to cancel the noise during the period with speech; transient noise cannot be estimated and canceled through the filter, and usually a noise reduction model is used to estimate and eliminate the transient noise in the audio segment. Thereafter, the input speech of the speaker to be recognized (i.e., the above-mentioned valid audio) can be first segmented into multiple clause audios, and each clause audio is input into the preset speaker verification model. At this time, the preset speaker verification model can output a speaker feature vector. Then, the speaker average feature vector of the speaker feature vector can be calculated. Finally, the speaker average feature vector is compared with the standard speaker feature vector corresponding to the speech to be compared to confirm whether the speaker to be recognized and the speaker to which the speech to be compared belongs are the same person, thereby realizing the recognition of the speaker to be recognized.

[0081] In this embodiment, for the financial phone scenario, first, high-quality speech of multiple speakers can be collected with 44k wav format audio at the sampling rate as the standard. Among them, the number of speakers is more than a thousand, and the target is to have at least 10 minutes of valid speech for each person. The collection channels of the speech can include network video / live data, open-source data, self-recorded data, etc. For the collected high-quality speech audio, it can be compressed with the target of 8k sampling rate and MP3 format to obtain corresponding high- and low-quality audio datasets. Then, the device can use the high-quality audio dataset as the input and the low-quality audio dataset as the target to train an audio restoration model. Since the audio restoration model uses a reversible diffusion process, during the training process, the transformation from high to low quality will be learned, and when applied, the model will be reversed to restore the low-quality audio to high quality. When performing financial phone identity verification, due to phone channel limitations and storage costs, the dialogue audio format in the financial phone scenario is basically in MP3 format with an 8k sampling rate. In this embodiment, the dialogue audio can be classified under each speaker's name with the speaker's ID card / mobile phone number as the main key. In addition, for stereo audio, the channel where the speaker is located can be selected, and the silence and non-speech parts in the audio can be removed using valid speech detection to obtain the valid audio of each speaker. Among them, the target number of data in the financial phone identity verification scenario dataset is more than ten thousand, and the target is to have at least 30 seconds of valid speech for each person. The collection channel can be the actual phone identity verification scenario. After collecting the financial phone identity verification scenario dataset, a speaker verification model can be trained based on this dataset. The training objective is to find a transformation model that maximizes the distance between the clusters of multiple speakers. The evaluation metrics are that the false positive rate and false negative rate of the model are as low as possible. Specifically, this solution can select the valid audio of multiple speakers, cross-construct a test dataset for pairwise comparison between two speakers, and use the model to predict the speaker feature vectors, and select a threshold to measure the evaluation metrics. After training the above audio restoration model and speaker verification model, the two models can be used jointly and applied to speaker verification in the financial phone scenario. Specifically, during the call, the channel audio of the speaker is obtained in real time and segmented into multiple audio clause samples using valid speech detection. Then, the model is applied to obtain multiple feature vectors respectively, and the average vector is calculated for the multiple feature vectors to reduce the differences between the features of different audio samples. Speaker verification is divided into two stages: registration and verification. Among them, in the registration stage, after the customer's first call ends, the above average vector is associated and saved with the customer identity identifier; in the verification stage, during each subsequent call of the customer, the obtained average vector is compared with the registered average vector to determine whether the calling customer is the same person.

[0082] This embodiment provides a speaker verification method, which discloses that the valid audio corresponding to the speaker to be recognized is segmented into several clause audio; each clause audio is input into a preset speaker verification model to obtain several speaker feature vectors, and the preset speaker verification model is constructed based on high-quality sample audio and audio implicit vector information; the speaker average feature vector is determined according to each speaker feature vector; the identity of the speaker to be recognized is identified based on the speaker average feature vector and the standard speaker feature vector; compared with the prior art, both the training set and the test set used in the speaker verification model are lossless audio in the original wav format with a sampling rate of 16k, and when recognizing lower sampling rate audio or compressed audio in a low-quality channel in an actual application scenario, the accuracy is not high. Since this embodiment can pre-construct a preset speaker verification model based on high-quality sample audio and audio implicit vector information, then obtain the speaker feature vectors corresponding to the valid audio of the speaker to be recognized based on the preset speaker verification model, and finally perform speaker recognition based on the corresponding speaker average feature vector and the standard speaker feature vector, thus solving the technical problem of low accuracy when the speaker verification model in the prior art performs speaker recognition in a low-quality channel in an actual application scenario.

[0083] Based on the first embodiment of the present application, in the second embodiment of the present application, the same or similar content as in the above-mentioned first embodiment can be referred to the above introduction and will not be repeated hereinafter. On this basis, please refer to Figure 3 , Figure 3 which is a schematic flowchart provided for the second embodiment of the speaker verification method of the present application.

[0084] In this embodiment, before step S10, the speaker verification method further includes steps S01 to S02:

[0085] Step S01: The low-quality audio is respectively subjected to audio restoration processing and implicit vector extraction processing through a preset audio restoration model to obtain high-quality audio and implicit vector information, and the implicit vector information is the audio information during the audio restoration of the low-quality audio.

[0086] It should be understood that the above-mentioned low-quality audio can be lossy audio with a low sampling rate; correspondingly, the above-mentioned high-quality audio can be high-sampling-rate lossless audio obtained by restoring the low-quality audio. During the process of restoring the low-quality audio to high-quality audio through the preset audio restoration model, the audio restoration model can be spliced to obtain implicit vector information, so as to jointly improve the effect of the speaker verification model on the low-quality channel audio through the high-quality audio and the implicit vector information.

[0087] Step S02: The initial speaker verification model is trained based on the high-quality audio and the implicit vector information to obtain a preset speaker verification model.

[0088] Further, an audio feature extraction module, a long and short speech statistics aggregation module, and a feature projection module are provided in the initial speaker confirmation model; the step S02 includes:

[0089] Step S021: The audio feature extraction module adds an implicit vector to the high-quality audio based on the implicit vector information to obtain the processed high-quality audio.

[0090] It should be noted that the above audio feature extraction module can be used to extract audio features and implicit vector features that can distinguish speakers, so as to be used for subsequent comparison and confirmation of speakers.

[0091] Step S022: The audio feature extraction module performs an attention transformation on the processed high-quality audio in the channel dimension and the time sequence dimension to obtain a plurality of input audio features of different channels.

[0092] In this embodiment, an audio feature extraction module, a long and short speech statistics aggregation module, and a feature projection module are provided in the initial speaker confirmation model. Refer to Figure 4 , Figure 4 is a schematic structural diagram of the speaker confirmation model in the speaker confirmation method of the present application. As Figure 4 shown, the audio feature extraction module may include a convolutional part and a channel attention part. Among them, the convolutional part has 4 layers, and each layer includes 3, 4, 6, and 3 convolutional modules respectively. Standardization and implicit vector addition are performed on the output between layers to obtain the processed high-quality audio; the channel attention part performs an attention transformation on the input high-quality audio in both the channel and time sequence dimensions, captures the correlation of the high-quality audio in different channels and time sequences, and finally obtains a plurality of input audio features of different channels.

[0093] Step S023: The long and short speech statistics aggregation module performs feature segmentation on the input audio features to obtain short audio features corresponding to the input audio features.

[0094] It should be noted that the long and short speech statistics aggregation module can be used to unify audio features. In practical applications, after the audio is processed by the audio feature extraction module, the obtained features still have different time lengths. At this time, they need to be aggregated through aggregation into vectors of a fixed dimension before they can be used to compare and confirm speakers. Therefore, in this embodiment, the long and short speech statistics aggregation module can further process the input audio features. Specifically, the long and short speech statistics aggregation module can perform fixed segmentation on the input long speech, cut it into short speeches, so as to obtain short audio features corresponding to the input audio features.

[0095] Step S024: Generate an input audio feature matrix corresponding to the input audio features and a short audio feature matrix corresponding to the short audio features.

[0096] It should be noted that the above input audio feature matrix can be a feature matrix corresponding to the input audio features; correspondingly, the above short audio feature matrix can be a feature matrix corresponding to the short audio features. In this embodiment, after the long speech is segmented into short speeches by the long and short speech statistical aggregation module, the average values, variances, period estimations of 20, 50, and 100 millisecond windows, and six statistics of the maximum fluctuation amount of multiple different channels of the long speech and short speeches can be calculated, so that each audio feature is converted into a feature matrix with the same length and the number of features equal to 6, and at this time, the above input audio feature matrix and short audio feature matrix are obtained.

[0097] Step S025: Perform feature splicing and fusion processing on the input audio feature matrix and the short audio feature matrix to obtain an aggregated audio feature matrix.

[0098] It should be noted that the above aggregated audio feature matrix can be a feature matrix obtained by splicing and fusing the input audio feature matrix and the short audio feature matrix. In practical applications, although long speech contains more speaker information, it is more easily affected by the content; although the voiceprint features of short speech are more pure, the information is less. Therefore, in order to improve the accuracy of speaker recognition, the matrices obtained from long and short speeches respectively can be spliced and fused to obtain a more comprehensive and stable aggregated audio feature matrix.

[0099] Step S026: Project the aggregated audio feature matrix into a low-dimensional space through the feature projection module to obtain the projected audio features.

[0100] It should be understood that since the aggregated feature matrix usually has a high dimension and is difficult to discriminate, therefore, in the training process, if it is necessary to discriminate a large number of speakers, it is necessary to project the high-dimensional features into a low-dimensional space that is easy to discriminate. Among them, the projection methods usually include full connection, cosine similarity projection, inter-class angle margin projection, etc. As Figure 4 shown, this embodiment can use a projection method of hypersphere projection plus inter-class angle, and its function is to increase the inter-class distance and decrease the intra-class distance during the training process. The projected features are finally adjusted to the required feature vector size using a fully connected layer, thereby reducing the calculation amount during speaker comparison and further improving the speaker recognition efficiency.

[0101] Step S027: Train the initial speaker verification model based on the projected audio features to obtain a preset speaker verification model.

[0102] Further, the step S027 includes: determining a feature similarity between the projected audio feature and the audio feature to be compared through the initial speaker verification model; obtaining a speaker recognition result according to the feature similarity; determining a false rejection rate and a false acceptance rate of the initial speaker verification model based on the speaker recognition result; and optimizing the initial speaker verification model based on the false rejection rate and the false acceptance rate to obtain a preset speaker verification model.

[0103] It can be understood that the above audio feature to be compared can be the feature corresponding to the audio for comparing with the audio of the speaker to be recognized, such as the feature corresponding to the audio provided by the speaker during registration; the above feature similarity can be a value used to represent the similarity degree between the projected audio feature and the audio feature to be compared.

[0104] It should be understood that the device can compare the audio of the speaker to be recognized with the audio to be compared to determine whether these two pieces of audio belong to the same speaker, and then obtain a speaker recognition result. Specifically, if the feature similarity between the projected audio feature and the audio feature to be compared exceeds a certain set threshold (such as 95%), it can be determined that the audio of the speaker to be recognized and the audio to be compared belong to the same person; otherwise, it is determined that the audio of the speaker to be recognized and the audio to be compared do not belong to the same person.

[0105] It should be noted that the above false rejection rate can be the ratio of the speaker verification model misidentifying two voices that should belong to the same speaker as different people (false rejection), where the false rejection rate = the number of test pairs of false rejection / the total number of test pairs * 100%; the above false acceptance rate can be the ratio of the speaker verification model misidentifying two voices that should not belong to the same speaker as the same person (false acceptance), where the false acceptance rate = the number of test pairs of false acceptance / the total number of test pairs × 100%.

[0106] In practical applications, the current false rejection rate and false acceptance rate of the model can be calculated respectively based on the speaker recognition result, and the model effect of the speaker verification model can be evaluated based on the false rejection rate and the false acceptance rate. Generally speaking, the lower the values of these two indicators, the better the model effect. In this embodiment, when the false rejection rate and the false acceptance rate reach the expected values, the training of the initial speaker verification model can be completed to obtain a preset speaker verification model.

[0107] In this embodiment, it is disclosed that the low-quality audio is respectively subjected to audio restoration processing and implicit vector extraction processing through a preset audio restoration model to obtain high-quality audio and implicit vector information. The implicit vector information is the audio information when restoring the low-quality audio. Based on the high-quality audio and the implicit vector information, the initial speaker verification model is trained to obtain a preset speaker verification model. Since this embodiment can jointly train the speaker verification model based on the high-quality audio and the implicit vector information, the recognition effect of the speaker verification model for low-quality channel audio is improved.

[0108] Based on the first embodiment and / or the second embodiment of the present application, in the third embodiment of the present application, the same or similar content as the above embodiments can be referred to the above introduction and will not be repeated hereinafter. On this basis, please refer to Figure 5 , Figure 5 which is a schematic flow chart provided for the third embodiment of the speaker verification method of the present application.

[0109] In this embodiment, before step S10, the speaker verification method further includes steps S11 to S13:

[0110] Step S11: Collect a plurality of high-quality sample audios.

[0111] It should be understood that the above high-quality sample audios can be high-sampling-rate lossless audios used for training the audio restoration model. The high-quality sample audios in this embodiment can use 16k sampling rate audio, or lossless audio with a maximum sampling rate of 44.1k or 48k.

[0112] Step S12: Perform audio compression processing on the high-quality sample audios to obtain low-quality sample audios.

[0113] It can be understood that the above low-quality sample audios can be audios obtained by downsampling and compressing high-quality sample audios. In this embodiment, the above audio compression processing can be a process of obtaining an audio with a lower sampling rate, quality, and size by downsampling the high-quality sample audio and compressing it with an audio compression algorithm. Among them, the audio downsampling method is a method of converting a high-sampling-rate audio (such as 48k, 44k, 16k, etc.) to a low-sampling-rate audio such as 8k audio using a downsampling algorithm. The downsampling algorithm includes linear interpolation, sine interpolation, average method, convolution method, etc.; the audio compression method is a method of performing lossy compression on lossless audio using methods such as mp3 compression and acc compression.

[0114] Step S13: Train the initial audio restoration model based on the low-quality sample audios to obtain a preset audio restoration model.

[0115] Further, the initial audio restoration model is provided with: an audio encoder, a diffusion restoration module, and an audio decoder; the step S13 includes:

[0116] Step S131: Input the low-quality sample audio into the initial audio restoration model.

[0117] It can be understood that the initial audio restoration model in this embodiment can be a model based on deep learning, and its model structure can use convolution and deconvolution, diffusion models, variational autoencoders, etc., and this embodiment does not limit this.

[0118] Step S132: Perform compression encoding processing on the low-quality sample audio through the audio encoder to obtain a low-quality audio latent vector.

[0119] It should be noted that the above low-quality audio latent vector can be a set of latent vectors obtained by performing convolutional compression encoding on the low-quality sample audio. In this embodiment, the audio encoder can be used to obtain a set of latent vectors by performing convolutional compression encoding on the input original low-quality audio waveform, so as to facilitate the processing of each subsequent module.

[0120] In this embodiment, the audio encoder includes 4 bidirectional dilated convolutional layers, and each layer has 64 convolutional kernels. Among them, "bidirectional" means that the convolutional kernels slide on the audio waveform in two temporal directions, namely from the start to the end of the audio and from the end to the start, and its function is to capture the forward and backward information at a certain time point of the audio, that is, to obtain the context dependence; "dilated" means that the size of each convolutional kernel is the same (all 3), but the interval between the sampling points of different convolutional kernels gradually increases from 0 to 63, and its function is to increase the receptive field (range) of the convolution without increasing the convolutional kernel size or resolution, thereby reducing the model calculation amount and improving the audio restoration efficiency.

[0121] Step S133: Perform audio restoration on the low-quality audio latent vector through the diffusion restoration module to obtain a high-quality audio latent vector.

[0122] In this embodiment, the diffusion restoration module can perform convolutional and deconvolutional processing on the low-quality audio latent vector to gradually restore it to the latent vector of high-quality audio (i.e., the above high-quality latent vector).

[0123] In practical applications, refer to Figure 6 , Figure 6 is a schematic structural diagram of the audio restoration model in the speaker verification method of this application. As Figure 6As shown, after inputting the low-quality sample audio into the initial audio restoration model, the audio encoder in the initial audio restoration model can first perform convolutional compression encoding on the low-quality audio waveform to obtain a set of latent vectors z. Then, the diffusion restoration module can gradually restore the latent vector z of the low-quality audio to the latent vector Zt of the high-quality audio through multiple convolutional and transposed convolutional modules. Among them, "diffusion" in the diffusion restoration module means that the process of audio restoration is gradual. Each layer of the module predicts the audio difference, and the difference between the latent vectors of the low-quality and high-quality audios gradually decreases until the latent vector of the high-quality audio is finally obtained. Each convolutional and transposed convolutional module in the diffusion restoration module contains 4 convolutional kernels with sizes of 3, 7, 11, and 17 respectively. During the convolutional process, the convolutional kernels are used from large to small, and during the transposed convolutional process, they are used from small to large to ensure that the sizes of the input and output latent vectors remain unchanged.

[0124] Step S134: Decode the high-quality audio latent vector through the audio decoder to obtain high-quality predicted audio.

[0125] It should be noted that, as Figure 6 shown, the audio decoder can restore the latent vector Zt of the high-quality audio output by the diffusion restoration module into a high-quality audio waveform through the transformers module and finally output the above high-quality predicted audio. In this embodiment, the audio decoder contains 2 attention layers and 2 fully connected layers, whose function is to learn the latent features and restoration parameters at different audio time positions to help better reconstruct the high-quality audio.

[0126] Step S135: Optimize the initial audio restoration model based on the high-quality sample audio and the high-quality predicted audio to obtain a preset audio restoration model.

[0127] Further, the step S135 includes: obtaining a first audio spectrum corresponding to the high-quality sample audio and a second audio spectrum corresponding to the high-quality predicted audio; determining the spectral convergence degree between the first audio spectrum and the second audio spectrum; optimizing the initial audio restoration model based on the spectral convergence degree to obtain a preset audio restoration model.

[0128] It should be noted that the above spectral convergence degree can be a value used to measure the error between the original audio spectrum and the restored audio spectrum (i.e., the error between the above first audio spectrum and the above second audio spectrum).

[0129] In specific implementation, this embodiment can use the spectral convergence degree to evaluate the effect of the audio restoration model. It evaluates the model restoration effect by measuring the error between the original audio spectrum and the restored audio spectrum. The calculation method is as follows:

[0130] First, the restored spectrum It is divided into two parts: the target spectrum and the interference error e;

[0131] Secondly, the following calculations are performed through the original spectrum x and the restored spectrum:

[0132]

[0133] Among them, denotes the transpose of, denotes the square of the L2 norm of x (i.e., the signal energy);

[0134] Then calculate , and calculate the energy corresponding to e and ;

[0135] Finally, calculate the value of SDR: , where the smaller the SDR, the closer the original and restored audio are, and the better the model effect. In this embodiment, when the above spectrum convergence reaches the expected value, the training of the initial audio restoration model can be completed to obtain the preset audio restoration model.

[0136] In this embodiment, a number of high-quality sample audios are collected; the high-quality sample audios are subjected to audio compression processing to obtain low-quality sample audios; the initial audio restoration model is trained based on the low-quality sample audios to obtain the preset audio restoration model, so that subsequently, the audio restoration from the low-quality channel to the high-quality can be realized through the preset audio restoration model, thereby avoiding the problem of poor model recognition effect of the speaker verification model in the low-quality channel.

[0137] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the speaker verification method of this application. Based on this technical concept, more forms of simple transformation are within the protection scope of this application.

[0138] This application also provides a speaker verification device. Please refer to Figure 7 , the speaker verification device includes:

[0139] An audio segmentation module 10 for segmenting the valid audio corresponding to the speaker to be identified into a number of clause audios;

[0140] A feature extraction module 20 for inputting each clause audio into a preset speaker verification model to obtain a number of speaker feature vectors, and the preset speaker verification model is constructed based on high-quality sample audios and audio implicit vector information;

[0141] The feature vector determination module 30 is configured to determine the speaker average feature vector according to the feature vectors of each speaker;

[0142] The speaker recognition module 40 is configured to perform identity recognition on the speaker to be recognized based on the speaker average feature vector and the standard speaker feature vector.

[0143] The speaker verification device provided by the present application adopts the speaker verification method in the above embodiment, and can solve the technical problem that the accuracy is not high when the speaker verification model in the prior art performs speaker recognition on a low-quality channel in an actual application scenario. Compared with the prior art, the beneficial effects of the speaker verification device provided by the present application are the same as those of the speaker verification method provided by the above embodiment, and other technical features in the speaker verification device are the same as those disclosed in the method of the above embodiment, and will not be elaborated herein.

[0144] The present application provides a speaker verification device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the speaker verification method in the first embodiment above.

[0145] Reference is made below Figure 8 , which shows a schematic structural diagram of a speaker verification device suitable for implementing the embodiments of the present application. The speaker verification device in the embodiments of the present application may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), PMPs (Portable Media Players), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 8 The speaker verification device shown is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present application.

[0146] As Figure 8As shown, the speaker verification device may include a processing device 1001 (such as a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM: Read Only Memory) 1002 or a program loaded from a storage device 1003 into a random access memory (RAM: Random Access Memory) 1004. In the RAM 1004, various programs and data required for the operation of the speaker verification device are also stored. The processing device 1001, the ROM 1002, and the RAM 1004 are connected to each other through a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Generally, the following systems may be connected to the I / O interface 1006: an input device 1007 including, for example, a touch screen, a touchpad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD: Liquid Crystal Display), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 may allow the speaker verification device to communicate with other devices wirelessly or wiredly to exchange data. Although the figure shows a speaker verification device having various systems, it should be understood that it is not required to implement or have all the shown systems. More or fewer systems may be alternatively implemented or had.

[0147] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program codes for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network through the communication device, or installed from the storage device 1003, or installed from the ROM 1002. When the computer program is executed by the processing device 1001, the above functions defined in the methods of the embodiments disclosed in the present application are executed.

[0148] The speaker verification device provided by the present application adopts the speaker verification method in the above embodiments and can solve the technical problems of speaker verification. Compared with the prior art, the beneficial effects of the speaker verification device provided by the present application are the same as those of the speaker verification method provided by the above embodiments, and the other technical features in the speaker verification device are the same as those disclosed in the method of the previous embodiment, and will not be elaborated here.

[0149] It should be understood that each part disclosed in this application can be implemented by hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in a suitable manner in any one or more embodiments or examples.

[0150] As described above, the above are only specific embodiments of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in this application, and all should be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.

[0151] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., computer programs) stored thereon, and the computer-readable program instructions are used to execute the speaker recognition method in the above embodiments.

[0152] The computer-readable storage medium provided by this application can be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM) or flash memory, optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In this embodiment, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or combined with an instruction execution system, device, or device. The program code contained on the computer-readable storage medium can be transmitted by any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination of the above.

[0153] The above computer-readable storage medium can be included in the speaker recognition device; it can also exist separately and not be assembled into the speaker recognition device.

[0154] The above computer-readable storage medium carries one or more programs, which, when executed by the speaker verification device, cause the speaker verification device to: segment the valid audio corresponding to the speaker to be recognized into a plurality of clause audios; input each clause audio into a preset speaker verification model to obtain a plurality of speaker feature vectors, where the preset speaker verification model is constructed based on high-quality sample audio and audio implicit vector information; determine a speaker average feature vector according to each speaker feature vector; and perform identity verification on the speaker to be recognized based on the speaker average feature vector and the standard speaker feature vector.

[0155] Computer program code for performing the operations of the present application may be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., by connecting through an Internet service provider using the Internet).

[0156] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.

[0157] The modules involved in the embodiments of the present application can be implemented in software or in hardware. Among them, the name of the module does not constitute a limitation to the unit itself in some cases.

[0158] The readable storage medium provided in the present application is a computer-readable storage medium, and the computer-readable storage medium stores computer-readable program instructions (i.e., computer programs) for executing the above-mentioned speaker recognition method, which can solve the technical problem that the accuracy of the speaker recognition model in the prior art is not high when performing speaker recognition in a low-quality channel in an actual application scenario. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in the present application are the same as those of the speaker recognition method provided in the above embodiments, and will not be elaborated here.

[0159] The above are only some embodiments of the present application, and thus do not limit the patent scope of the present application. Any equivalent structural transformation made by using the content of the specification and drawings of the present application under the technical concept of the present application, or any direct / indirect application in other related technical fields, is included in the patent protection scope of the present application.

Claims

1. A method for speaker confirmation, characterized in that: The method includes: Divide the valid audio corresponding to the speaker to be identified into a number of sentence audios; Inputting each sentence audio into a preset speaker verification model to obtain a plurality of speaker feature vectors, wherein the preset speaker verification model is constructed based on high-quality sample audio and audio latent vector information; Determine a speaker average feature vector based on the feature vectors of each speaker; Performing identity recognition on the speaker to be recognized based on the speaker average feature vector and the standard speaker feature vector; Before the step of dividing the valid audio corresponding to the speaker to be identified into a plurality of sentence audios, the method further includes: Performing audio restoration processing and implicit vector extraction processing on low-quality audio respectively through a preset audio restoration model to obtain high-quality audio and implicit vector information, wherein the implicit vector information is audio information when audio restoration is performed on the low-quality audio; Training an initial speaker confirmation model based on the high-quality audio and the implicit vector information to obtain a preset speaker confirmation model, wherein the initial speaker confirmation model is provided with: an audio feature extraction module, a long and short speech statistics aggregation module, and a feature projection module; The step of training the initial speaker verification model based on the high-quality audio and the implicit vector information to obtain a preset speaker verification model includes: Adding an implicit vector to the high-quality audio based on the implicit vector information by the audio feature extraction module to obtain processed high-quality audio; The audio feature extraction module performs attention transformation on the processed high-quality audio in the channel dimension and the time dimension to obtain input audio features of several different channels; Performing feature segmentation on the input audio feature by the long and short speech statistical aggregation module to obtain short audio features corresponding to the input audio feature; Generate an input audio feature matrix corresponding to the input audio feature and a short audio feature matrix corresponding to the short audio feature; Performing feature splicing and fusion processing on the input audio feature matrix and the short audio feature matrix to obtain an aggregated audio feature matrix; Performing concatenation processing on the aggregated audio feature matrix and the implicit vector information to obtain a target aggregated audio feature matrix; Projecting the aggregated audio feature matrix into a low-dimensional space through the feature projection module to obtain projected audio features; The initial speaker verification model is trained based on the projected audio features to obtain a preset speaker verification model.

2. The method according to claim 1, characterized in that The step of training the initial speaker confirmation model based on the projected audio features to obtain a preset speaker confirmation model includes: Determining the feature similarity between the projected audio feature and the audio feature to be compared by using the initial speaker verification model; Obtaining a speaker recognition result according to the feature similarity; determining a false rejection rate and a false acceptance rate of the initial speaker verification model based on the speaker recognition result; The initial speaker verification model is optimized based on the false rejection rate and the false acceptance rate to obtain a preset speaker verification model.

3. The method according to claim 1, characterized in that Before the step of performing audio restoration processing and latent vector extraction processing on the low-quality audio by using a preset audio restoration model to obtain high-quality audio and latent vector information, the step further includes: Collect several high-quality sample audios; Performing audio compression processing on the high-quality sample audio to obtain low-quality sample audio; An initial audio restoration model is trained based on the low-quality sample audio to obtain a preset audio restoration model.

4. The method according to claim 3, characterized in that The initial audio restoration model is provided with: an audio encoder, a diffusion restoration module and an audio decoder; the step of training the initial audio restoration model based on the low-quality sample audio to obtain a preset audio restoration model includes: Inputting the low-quality sample audio into an initial audio restoration model; Performing compression encoding processing on the low-quality sample audio by the audio encoder to obtain a low-quality audio latent vector; Performing audio restoration on the low-quality audio latent vector by the diffusion restoration module to obtain a high-quality audio latent vector; Decoding the high-quality audio latent vector by the audio decoder to obtain high-quality predicted audio; The initial audio restoration model is optimized based on the high-quality sample audio and the high-quality predicted audio to obtain a preset audio restoration model.

5. The method according to claim 4, characterized in that The step of optimizing the initial audio restoration model based on the high-quality sample audio and the high-quality predicted audio to obtain a preset audio restoration model comprises: Acquire a first audio spectrum corresponding to the high-quality sample audio and a second audio spectrum corresponding to the high-quality predicted audio; determining a degree of spectral convergence between the first audio spectrum and the second audio spectrum; The initial audio restoration model is optimized based on the frequency spectrum convergence to obtain a preset audio restoration model.

6. A speaker verification device, characterized in that: The device comprises: An audio segmentation module, used to segment the valid audio corresponding to the speaker to be identified into a number of sentence audios; A feature extraction module, used to input each sentence audio into a preset speaker verification model to obtain a number of speaker feature vectors, wherein the preset speaker verification model is constructed based on high-quality sample audio and audio latent vector information; A feature vector determination module, used to determine a speaker average feature vector based on the feature vectors of each speaker; A speaker recognition module, used for identifying the speaker to be identified based on the speaker average feature vector and the standard speaker feature vector; The audio segmentation module is further used to perform audio restoration processing and implicit vector extraction processing on low-quality audio through a preset audio restoration model to obtain high-quality audio and implicit vector information, wherein the implicit vector information is audio information when the low-quality audio is restored; based on the high-quality audio and the implicit vector information, an initial speaker confirmation model is trained to obtain a preset speaker confirmation model, wherein the initial speaker confirmation model is provided with: an audio feature extraction module, a long and short speech statistics aggregation module, and a feature projection module; The audio segmentation module is also used to add implicit vectors to the high-quality audio based on the implicit vector information through the audio feature extraction module to obtain processed high-quality audio; perform attention transformation on the processed high-quality audio in the channel dimension and the time dimension through the audio feature extraction module to obtain input audio features of several different channels; perform feature segmentation on the input audio features through the long and short speech statistical aggregation module to obtain short audio features corresponding to the input audio features; generate an input audio feature matrix corresponding to the input audio features and a short audio feature matrix corresponding to the short audio features; perform feature splicing and fusion processing on the input audio feature matrix and the short audio feature matrix to obtain an aggregated audio feature matrix; perform splicing processing on the aggregated audio feature matrix and the implicit vector information to obtain a target aggregated audio feature matrix; project the aggregated audio feature matrix to a low-dimensional space through the feature projection module to obtain projected audio features; train an initial speaker confirmation model based on the projected audio features to obtain a preset speaker confirmation model.

7. A speaker verification device, characterized in that: The device comprises: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the speaker confirmation method according to any one of claims 1 to 5.

8. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the speaker identification method according to any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Voiceprint recognition method and device, electronic equipment and storage medium

    CN114822558A

  • Speaker recognition method and device, equipment and storage medium

    CN115083419A

  • Audio restoration method and device, medium and computing equipment

    CN118942468A