Voiceprint extraction model training method, extraction method and electronic device

By performing covariance matrix dimensionality reduction and feature splicing methods on voiceprint information, the accuracy of voiceprint extraction model is improved, and the problem of low accuracy of voiceprint extraction in the prior art is solved.

CN115394303BActive Publication Date: 2025-05-06VOICEAI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210907483.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-29
Publication Date
2025-05-06
Estimated Expiration
2042-07-29

AI Technical Summary

Technical Problem

In the prior art, when using deep neural networks for voiceprint extraction, the extracted voiceprint information has a low accuracy.

Method used

The sample speech information is feature extracted through the initial model, the covariance matrix, variance and mean are calculated, and the covariance matrix is ​​reduced by a one-dimensional convolution layer to obtain one-dimensional preprocessed data, and then the variance, mean and one-dimensional preprocessed data are spliced ​​to train the voiceprint extraction model.

Benefits of technology

The accuracy of voiceprint extraction of the voiceprint extraction model is improved, and the covariance matrix accurately represents the time and frequency dimension characteristics of speech information, enhancing the representation ability of splicing results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115394303B_ABST
    Figure CN115394303B_ABST
Patent Text Reader

Abstract

The present application discloses a training method for a voiceprint extraction model, a voiceprint extraction method, an electronic device and a storage medium. The training method for the voiceprint extraction model includes: extracting features from multiple audio frames through an initial model to obtain multiple frame features; obtaining a covariance matrix, variance and mean based on multiple frame features; reducing the dimension of the covariance matrix through a one-dimensional convolutional layer in the initial model, and vectorizing the result after dimensionality reduction to obtain one-dimensional preprocessed data; splicing the variance, mean and one-dimensional preprocessed data to obtain a splicing result; training the initial model according to the splicing result to obtain a voiceprint extraction model. In the present application, the covariance matrix accurately represents the feature information of the time dimension and frequency dimension of the sample voice information, so that the splicing result can accurately represent the voiceprint features of the sample voice information, thereby improving the voiceprint extraction accuracy of the voiceprint extraction model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of audio processing technology, and more specifically, to a training method for a voiceprint extraction model, a voiceprint extraction method, an electronic device, and a storage medium. Background Art

[0002] Voiceprint recognition is a technology that uses sound to identify the identity of a voice user, and is one of the important research directions in the field of speech. With the continuous development of computer technology, voiceprint recognition has made great progress in recent years. With its convenient and effective characteristics, it has become an efficient identity recognition method and is widely used in public security, banking, and smart homes.

[0003] At present, the deep neural network can be trained through samples to obtain a voiceprint extraction model, and then the voiceprint extraction model is used to extract the voiceprint of the voice information to be extracted. However, when this method is used to extract the voiceprint of the voice information to be extracted, the accuracy of the extracted voiceprint information is low. Summary of the invention

[0004] In view of this, the embodiments of the present application propose a training method for a voiceprint extraction model, a voiceprint extraction method, an electronic device, and a storage medium.

[0005] In a first aspect, an embodiment of the present application provides a method for training a voiceprint extraction model, the method comprising: performing feature extraction on multiple audio frames included in sample voice information through an initial model to obtain multiple frame features corresponding one-to-one to the multiple audio frames; obtaining a covariance matrix, variance and mean corresponding to the sample voice information based on the multiple frame features; performing dimensionality reduction processing on the covariance matrix through a one-dimensional convolutional layer in the initial model to obtain a result after dimensionality reduction, and performing a vectorization operation on the result after dimensionality reduction to obtain one-dimensional preprocessed data; performing a splicing operation on the variance, the mean and the one-dimensional preprocessed data to obtain a splicing result; training the initial model according to the splicing result to obtain a voiceprint extraction model.

[0006] In a second aspect, an embodiment of the present application provides a voiceprint extraction method, the method comprising: performing feature extraction on a plurality of audio frames to be extracted included in voice information to be extracted through a voiceprint extraction model to obtain a plurality of target frame features corresponding one to one to the plurality of audio frames to be extracted, the voiceprint extraction model being trained by the method described in the first aspect; obtaining a covariance matrix, variance and mean corresponding to the voice information to be extracted according to the plurality of target frame features; performing dimensionality reduction processing on the covariance matrix of the voice information to be extracted through a one-dimensional convolutional layer in the voiceprint extraction model to obtain a result after dimensionality reduction, and performing a vectorization operation on the result after dimensionality reduction to obtain target one-dimensional preprocessed data; performing a splicing operation on the variance of the voice information to be extracted, the mean of the voice information to be extracted and the target one-dimensional preprocessed data to obtain a target splicing result; processing the target splicing result through the voiceprint extraction model to obtain voiceprint information corresponding to the voice information to be extracted.

[0007] In a third aspect, an embodiment of the present application provides a training device for a voiceprint extraction model, the device comprising: a first extraction module, used to perform feature extraction on multiple audio frames included in sample voice information through an initial model, and obtain multiple frame features corresponding to the multiple audio frames one by one; a first acquisition module, used to obtain a covariance matrix, variance and mean corresponding to the sample voice information according to the multiple frame features; a first dimensionality reduction module, used to perform dimensionality reduction processing on the covariance matrix through a one-dimensional convolutional layer in the initial model, obtain a result after dimensionality reduction, and perform vectorization operation on the result after dimensionality reduction to obtain one-dimensional preprocessed data; a first splicing module, used to perform splicing operation on the variance, the mean and the one-dimensional preprocessed data to obtain a splicing result; a model acquisition module, used to train the initial model according to the splicing result to obtain a voiceprint extraction model.

[0008] In a fourth aspect, an embodiment of the present application provides a voiceprint extraction device, the device comprising: a second extraction module, used to perform feature extraction on a plurality of audio frames to be extracted included in the voice information to be extracted through a voiceprint extraction model, and obtain a plurality of target frame features corresponding to the plurality of audio frames to be extracted, wherein the voiceprint extraction model is trained by the method described in the first aspect; a second acquisition module, used to obtain a covariance matrix, variance and mean corresponding to the voice information to be extracted according to the plurality of target frame features; a second dimensionality reduction module, used to perform dimensionality reduction processing on the covariance matrix of the voice information to be extracted through a one-dimensional convolutional layer in the voiceprint extraction model, obtain a result after dimensionality reduction, and perform vectorization operation on the result after dimensionality reduction to obtain target one-dimensional preprocessed data; a second splicing module, used to perform splicing operation on the variance of the voice information to be extracted, the mean of the voice information to be extracted and the target one-dimensional preprocessed data to obtain a target splicing result; a voiceprint acquisition module, used to process the target splicing result through the voiceprint extraction model to obtain voiceprint information corresponding to the voice information to be extracted.

[0009] In a fifth aspect, an embodiment of the present application provides an electronic device, comprising a processor and a memory; one or more programs are stored in the memory and configured to be executed by the processor to implement the above method.

[0010] In a sixth aspect, an embodiment of the present application provides a computer-readable storage medium, in which a program code is stored, wherein the above method is executed when the program code is executed by a processor.

[0011] In a seventh aspect, an embodiment of the present application provides a computer program product or a computer program, the computer program product or the computer program including computer instructions, the computer instructions being stored in a computer-readable storage medium. A processor of an electronic device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the electronic device performs the above method.

[0012] The embodiments of the present application provide a training method for a voiceprint extraction model, a voiceprint extraction method, an electronic device, and a storage medium. The initial model can process sample voice information to obtain a covariance matrix, a variance, and a mean, and splice the one-dimensional preprocessed data, the variance, and the mean corresponding to the covariance matrix. The initial model is then trained using the splicing result to obtain a voiceprint extraction model. The covariance matrix accurately represents the feature information of the time dimension and the frequency dimension of the sample voice information, so that the splicing result can accurately represent the voiceprint features of the sample voice information, thereby improving the voiceprint extraction accuracy of the voiceprint extraction model. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application, and those skilled in the art can also refer to the drawings without creative work.

[0014] Figure 1 A flow chart showing a method for training a voiceprint extraction model according to an embodiment of the present application is shown;

[0015] Figure 2 A schematic diagram of the structure of the initial model in an embodiment of the present application is shown;

[0016] Figure 3 A flowchart of a method for training a voiceprint extraction model proposed in another embodiment of the present application is shown;

[0017] Figure 4 A schematic diagram of the structure of the second-order statistical pooling layer in an embodiment of the present application is shown;

[0018] Figure 5 A flow chart of a voiceprint extraction method proposed in one embodiment of the present application is shown;

[0019] Figure 6 A block diagram of a training device for a voiceprint extraction model proposed in one embodiment of the present application is shown;

[0020] Figure 7 A block diagram of a voiceprint extraction device proposed in one embodiment of the present application is shown;

[0021] Figure 8 A structural block diagram of an electronic device proposed in one embodiment of the present application is shown. DETAILED DESCRIPTION

[0022] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. According to the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0023] In the following description, the terms "first\second" involved are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It can be understood that "first\second" can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.

[0024] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.

[0025] See also Figure 1 , Figure 1 A flow chart of a method for training a voiceprint extraction model according to an embodiment of the present application is shown. The method can be used in an electronic device, and the method includes:

[0026] S110. Perform feature extraction on multiple audio frames included in the sample speech information through an initial model to obtain multiple frame features corresponding to the multiple audio frames.

[0027] The sample voice information may refer to the voice information used to train the initial model, which may be audio in any format, such as MP3, AAC, etc. The sample voice information may be recorded by an electronic device, or obtained by the electronic device from the Internet. The sample voice information may be one or more pieces. When the sample voice information includes multiple pieces, each piece of sample voice information is used as a batch of training samples for training the initial model according to the method of the present application.

[0028] The sample voice information may include multiple audio frames. For example, 1s sample voice information may generally include 50 audio frames or 100 audio frames. The user may set a time shift (a time shift refers to the duration of an audio frame) based on demand, and determine the total number of audio frames included in the sample voice information according to the set time shift. For example, if the time shift is 20ms and the sample voice is 2s, the total number of audio frames obtained is 100.

[0029] The initial model refers to the basic model used to obtain the voiceprint extraction model, such as Figure 2 As shown, the initial model may include an input layer, which may be a general neural network input layer. The input layer extracts features of multiple audio frames to obtain multiple frame features corresponding to the multiple audio frames. Each frame feature may include values ​​of multiple dimensions. Accordingly, the number of dimensions of the input layer is the same as the number of dimensions of the frame features. One dimension of the input layer is used to process one dimension of the frame features. For example, assuming that the input sample speech information has T frames, the obtained frame feature is x i ,i=1,2,…,T, each frame feature includes 20 dimensions, and the input layer also includes 20 dimensions.

[0030] As an implementation method, after obtaining the sample voice information, the sample voice information can be formatted and converted into a standard sample voice in wav format, and the standard sample voice can be Fourier transformed to obtain voice features (the voice features can be data in MSCC format, and the voice features include voice features corresponding to multiple audio frames), and then the voice features are input into the input layer of the initial model to obtain multiple frame features corresponding to multiple audio frames.

[0031] S120. Obtain a covariance matrix, a variance, and a mean corresponding to the sample speech information according to the multiple frame features.

[0032] After obtaining multiple frame features corresponding to the sample speech information, the covariance matrix, variance and mean are calculated according to the multiple frame features.

[0033] As an implementation method, the covariance matrix may be calculated according to multiple frame features through Formula 3; Formula 3 is as follows:

[0034]

[0035]

[0036]

[0037] Among them, C (k,m) is the value of the kth row and mth column in the covariance matrix, T is the total number of features of multiple frames, x ik 、x im are the values ​​corresponding to any two dimensions of the i-th frame feature.

[0038] As an implementation method, the mean value may be calculated according to multiple frame features by using Formula 4, where Formula 4 is:

[0039]

[0040] Among them, m and is the mean, x i is any dimension of the i-th frame feature.

[0041] As an implementation method, the variance may be calculated according to multiple frame features by using Formula 5, where Formula 5 is:

[0042]

[0043] Where v is the variance.

[0044] It can be understood that the mean and variance include the mean and variance of each dimension respectively. For example, if the frame feature includes values ​​of 20 dimensions, the variance includes the variances corresponding to the 20 dimensions respectively, and the mean includes the means corresponding to the 20 dimensions respectively.

[0045] S130, performing dimensionality reduction processing on the covariance matrix through the one-dimensional convolution layer in the initial model to obtain a reduced-dimensional result, and performing vectorization operation on the reduced-dimensional result to obtain one-dimensional preprocessed data.

[0046] S140, performing a splicing operation on the variance, the mean, and the one-dimensional preprocessed data to obtain a splicing result.

[0047] The variance and mean are one-dimensional data, and the covariance matrix cannot be directly concatenated with the variance and mean. Therefore, it is necessary to perform dimensionality reduction on the covariance matrix to obtain one-dimensional preprocessed data, and then the one-dimensional preprocessed data can be concatenated with the variance and mean to obtain the concatenation result.

[0048] As an implementation manner, the variance, the mean, and the one-dimensional preprocessed data may be concatenated according to Formula 6 to obtain a concatenated result. Formula 6 is:

[0049] y=cat{m,v,vec}

[0050] Among them, y is the concatenation result, vec is the one-dimensional preprocessed data, and cat is the concatenation function.

[0051] The initial model may include a one-dimensional convolution layer, which is used to perform dimensionality reduction processing on the covariance matrix to obtain one-dimensional preprocessed data.

[0052] Vectorization operation may refer to converting the result after dimensionality reduction into a one-dimensional vector, that is, the one-dimensional preprocessed data is one-dimensional vector data. Through vectorization operation, the result after dimensionality reduction may be converted into one-dimensional vector data that can effectively describe the delay changes of speech features, thereby improving the accuracy of the one-dimensional preprocessed data.

[0053] The one-dimensional convolution layer may be obtained by configuring the initial one-dimensional convolution layer in the initial model, and the initial one-dimensional convolution layer configuration process may include: obtaining the dimension of the covariance matrix; determining the convolution kernel size, padding size, dilation size and stride size according to the dimension; configuring the initial one-dimensional convolution layer in the initial model according to the convolution kernel size, the padding size, the dilation size and the stride size to obtain the one-dimensional convolution layer.

[0054] Wherein, determining the convolution kernel size, padding size, expansion size and stride size according to the dimension includes: based on the dimension, calculating the convolution kernel size, padding size, expansion size and stride size according to formula 1; the formula 1 is:

[0055]

[0056] Wherein, L is the dimension, padding is the padding size, kernel is the convolution kernel size, stride is the stride size, and dilation is the expansion size.

[0057] like Figure 2 As shown, the initial model may include a second-order statistical pooling layer, which includes the above-mentioned one-dimensional convolutional layer. The second-order statistical pooling layer is also used to vectorize the reduced dimensionality results corresponding to the covariance matrix to obtain one-dimensional preprocessed data, and to splice the variance, mean and one-dimensional preprocessed data after vectorization to obtain a spliced ​​result.

[0058] S150: training the initial model according to the splicing result to obtain a voiceprint extraction model.

[0059] After obtaining the splicing result, the loss value can be obtained according to the splicing result, and the parameters of the initial model can be adjusted according to the loss value until the number of iterations meets the requirements, and the voiceprint extraction model is obtained.

[0060] like Figure 2 As shown, the initial model may further include a segment level layer and an output layer. The segment level layer may include at least one hidden layer for processing the splicing result, and the output layer may be a general neural network output layer for processing the output of the segment level layer to obtain an output result.

[0061] After obtaining the output result, the loss value is determined according to the output result, and then the parameters of the input layer, second-order statistical pooling layer, segment level layer and output layer in the initial model are adjusted according to the determined loss value to obtain the voiceprint extraction model.

[0062] In this embodiment, the sample voice information can be processed through the initial model to obtain the covariance matrix, variance and mean, and the one-dimensional preprocessed data, variance and mean corresponding to the covariance matrix are spliced, and then the initial model is trained through the splicing result to obtain the voiceprint extraction model. The covariance matrix accurately represents the characteristic information of the time dimension and frequency dimension of the sample voice information, so that the splicing result can accurately represent the voiceprint characteristics of the sample voice information, thereby improving the voiceprint extraction accuracy of the voiceprint extraction model.

[0063] At the same time, covariance can take into account both time dimension and frequency dimension information, integrate independent and unrelated dimensions of the second-order statistical pooling layer, and provide new information for segment-level network modeling, thereby improving the voiceprint extraction accuracy of the voiceprint extraction model. Through vectorization operations, the vectorization conversion of the dimensionality reduction results can effectively describe the delay changes of speech features and improve the accuracy of one-dimensional preprocessed data.

[0064] See also Figure 3 , Figure 3 A flowchart of a method for training a voiceprint extraction model according to another embodiment of the present application is shown. The method can be used in an electronic device, and the method includes:

[0065] S210. Perform feature extraction on multiple audio frames included in the sample speech information through an initial model to obtain multiple frame features corresponding to the multiple audio frames.

[0066] S220. Obtain a covariance matrix, a variance, and a mean corresponding to the sample speech information according to the multiple frame features.

[0067] S230, performing dimensionality reduction processing on the covariance matrix through the one-dimensional convolution layer in the initial model to obtain a result after dimensionality reduction, and performing vectorization operation on the result after dimensionality reduction to obtain one-dimensional preprocessed data.

[0068] S240, performing a splicing operation on the variance, the mean, and the one-dimensional preprocessed data to obtain a splicing result.

[0069] The description of S210 - S240 refers to the description of S110 - S140 above, which will not be repeated here.

[0070] S250, training the initial model according to the splicing result; whenever the number of iterations reaches an integer multiple of the first preset number, determining whether the number of iterations reaches a second preset number. If so, executing S260, if not, executing S270.

[0071] S260: Obtain a voiceprint extraction model.

[0072] S270, adjusting the convolution layer parameters of the one-dimensional convolution layer to obtain an adjusted one-dimensional convolution layer; if the convolution layer parameters of the adjusted one-dimensional convolution layer meet preset constraints, obtaining new sample speech information including multiple audio frames.

[0073] After obtaining the splicing result, the splicing result is input into the segment level layer of the initial model to obtain the output result of the segment level layer, and then the output result of the segment level layer is input into the output layer to obtain the output result of the output layer. Then, the loss value is determined according to the output result, and the initial model is trained through the loss value.

[0074] The first preset number of times may be a value set according to the scenario and requirements, such as 4 times. The second preset number of times refers to the maximum number of iterations for training the voiceprint extraction model, such as 1000 times. When the number of iterations reaches the second preset number of times, the training ends and the voiceprint extraction model is obtained. When the number of iterations does not reach the second preset number of times, the training does not end and the iterative training continues.

[0075] Whenever the number of iterations reaches an integer multiple of the first preset number of times and does not reach the second preset number of times, it is determined that the convolution layer parameters of the one-dimensional convolution layer need to be adjusted, and the convolution layer parameters of the one-dimensional convolution layer are adjusted to obtain the adjusted one-dimensional convolution layer, and then the initial model including the adjusted one-dimensional convolution layer is continued to be trained. When the number of iterations does not reach an integer multiple of the first preset number of times and does not reach the second preset number of times, the convolution layer parameters of the one-dimensional convolution layer are not adjusted, and the initial model including the one-dimensional convolution layer is continued to be trained.

[0076] As an implementation manner, the adjusting the convolution layer parameters of the one-dimensional convolution layer to obtain the adjusted one-dimensional convolution layer includes: obtaining a preset floating-point coefficient; based on the preset floating-point coefficient, adjusting the convolution layer parameters of the one-dimensional convolution layer according to Formula 2 to obtain the adjusted one-dimensional convolution layer;

[0077] The formula 2 is:

[0078]

[0079] Among them, M′ is the convolution layer parameter of the adjusted one-dimensional convolution layer, M is the convolution layer parameter of the one-dimensional convolution layer, α is the preset floating point coefficient, and M T is the transposed matrix of M, and I is the identity matrix. The preset floating-point coefficient can be 1.

[0080] As an implementation method, the preset constraint condition includes that the gap between the covariance matrix of the sample speech information and the product result is less than a preset gap, and the product result is the product of the convolution layer parameters of the adjusted one-dimensional convolution layer and the one-dimensional preprocessed data, wherein the preset gap can be a value set based on demand, which is not limited in this application. The mathematical formula of the preset constraint condition is as follows:

[0081] C→Mvec

[0082] When the convolution layer parameters of the adjusted one-dimensional convolution layer meet the preset constraints and the preset constraints converge, then: C = Mvec ≈ MM T, that is, the output one-dimensional preprocessed data vec is approximately equal to the convolution layer parameter M of the one-dimensional convolution layer. The information of the input covariance matrix can be retained to the maximum extent in vectorization, thereby making the accuracy of the one-dimensional preprocessed data higher, and further making the voiceprint extraction model trained according to the one-dimensional preprocessed data more accurate.

[0083] When the convolution layer parameters of the adjusted one-dimensional convolution layer do not meet the preset constraints, the training of the initial model is stopped and a prompt message is output, wherein the prompt message is used to prompt that there is a fault in the training process, so that the user can reconfigure the initial model based on the prompt message.

[0084] The new sample voice information also includes multiple audio frames. The description of the new sample voice information refers to the description of the sample voice information above and will not be repeated. After obtaining the new sample voice information, return to step S210 until the number of iterations reaches the second preset number to obtain a voiceprint extraction model.

[0085] In this embodiment, the structure of the second-order statistical pooling layer in the initial model can refer to Figure 4 .like Figure 4 As shown, in the second-order statistical pooling layer, the covariance matrix, variance and mean of the corresponding sample speech information are obtained according to the multiple frame features output by the input layer, and the covariance matrix is ​​reduced in dimension through the one-dimensional convolution layer in the second-order statistical pooling layer to obtain the reduced-dimensional result, and the reduced-dimensional result is vectorized to obtain one-dimensional preprocessed data, and then the variance, mean and one-dimensional preprocessed data are spliced ​​to obtain the spliced ​​result, and then the spliced ​​result is input into the segment level layer for processing.

[0086] Wherein, whenever the number of iterations reaches an integer multiple of the first preset number of times and the number of iterations does not reach the second preset number of times, the convolution layer parameters of the one-dimensional convolution layer are adjusted by semi-orthogonal constraints to obtain an adjusted one-dimensional convolution layer, and the adjusted one-dimensional convolution layer is applied to the subsequent training process. Wherein, the semi-orthogonal constraint refers to adjusting the convolution layer parameters of the one-dimensional convolution layer according to formula 2, and determining the convolution layer parameters of the adjusted one-dimensional convolution layer according to the preset constraint conditions.

[0087] In this embodiment, by constraining the convolution layer parameters of the one-dimensional convolution layer to be semi-orthogonal, the features after the vectorization operation can retain the information of the covariance matrix to the maximum extent, approach the effect of the eigenvalue features, and play a role in characterizing important components and minor components; at the same time, the second-order statistical pooling layer based on the one-dimensional convolution layer and the semi-orthogonal constraint can significantly improve the recognition accuracy of the voiceprint recognition model.

[0088] See also Figure 5 , Figure 5A flow chart of a voiceprint extraction method proposed in one embodiment of the present application is shown. The method can be used in an electronic device, and the method includes:

[0089] S310, performing feature extraction on a plurality of audio frames to be extracted included in the voice information to be extracted through a voiceprint extraction model to obtain a plurality of target frame features corresponding one-to-one to the plurality of audio frames to be extracted, wherein the voiceprint extraction model is trained by the method described in any of the above embodiments.

[0090] The voice information to be extracted refers to the voice information for voiceprint extraction, which is similar to the description of the sample voice information and will not be repeated here.

[0091] The audio frame in the voice information to be extracted is the audio frame to be extracted. Feature extraction is performed on the audio frame to be extracted, and the obtained frame feature is the target frame feature.

[0092] The voice information to be extracted can be converted into a standard voice to be extracted in wav format, and the standard voice to be extracted can be Fourier transformed to obtain the voice features to be extracted (the voice features to be extracted include the voice features to be extracted corresponding to each of the multiple audio frames to be extracted), and then the voice features to be extracted are input into the input layer of the voiceprint extraction model to obtain multiple target frame features corresponding to the multiple audio frames to be extracted.

[0093] S320. Obtain a covariance matrix, variance, and mean corresponding to the speech information to be extracted according to the multiple target frame features.

[0094] Among them, the process of obtaining the covariance matrix, variance and mean corresponding to the voice information to be extracted according to the target frame features is similar to the process of obtaining the covariance matrix, variance and mean corresponding to the sample voice information according to the frame features above, and will not be repeated here.

[0095] S330, performing dimensionality reduction processing on the covariance matrix of the voice information to be extracted through the one-dimensional convolution layer in the voiceprint extraction model to obtain a reduced-dimensional result, and performing a vectorization operation on the reduced-dimensional result to obtain target one-dimensional preprocessed data.

[0096] S340, performing a splicing operation on the variance of the speech information to be extracted, the mean of the speech information to be extracted, and the target one-dimensional preprocessed data to obtain a target splicing result.

[0097] Among them, the description of S330-S340 refers to the description of S130-S140 above, and will not be repeated here.

[0098] The target one-dimensional preprocessed data refers to the result obtained after dimensionality reduction and vectorization operations are performed on the covariance matrix of the voice information to be extracted through the one-dimensional convolution layer in the voiceprint extraction model, and the target splicing result refers to the result obtained after splicing the variance of the voice information to be extracted, the mean of the voice information to be extracted and the target one-dimensional preprocessed data.

[0099] S350: Process the target concatenation result by using the voiceprint extraction model to obtain voiceprint information corresponding to the voice information to be extracted.

[0100] After obtaining the concatenation result, the concatenation result is processed by the segment level layer of the voiceprint extraction model to obtain the processed result, and then the voiceprint information of the speech information to be extracted is obtained according to the processed result output by the segment level layer.

[0101] In this embodiment, the voiceprint extraction model trained by the method of the above embodiment of the present application has a better extraction effect, and accurate voiceprint information can be extracted through the voiceprint extraction model, thereby improving the accuracy of the voiceprint information.

[0102] See also Figure 6 , Figure 6 A block diagram of a training device for a voiceprint extraction model proposed in an embodiment of the present application is shown. The device 500 includes:

[0103] A first extraction module 510 is used to extract features from a plurality of audio frames included in the sample speech information through an initial model to obtain a plurality of frame features corresponding to the plurality of audio frames one by one;

[0104] A first obtaining module 520, configured to obtain a covariance matrix, a variance and a mean corresponding to the sample speech information according to the plurality of frame features;

[0105] A first dimensionality reduction module 530 is used to perform dimensionality reduction processing on the covariance matrix through the one-dimensional convolution layer in the initial model to obtain a result after dimensionality reduction, and perform vectorization operation on the result after dimensionality reduction to obtain one-dimensional preprocessed data;

[0106] A first splicing module 540 is used to perform a splicing operation on the variance, the mean and the one-dimensional pre-processed data to obtain a splicing result;

[0107] The model acquisition module 550 is used to train the initial model according to the splicing result to obtain a voiceprint extraction model.

[0108] Optionally, the first dimensionality reduction module 530 is further used to perform dimensionality reduction processing on the covariance matrix through the one-dimensional convolution layer to obtain a reduced-dimensional result; and perform vectorization operation on the reduced-dimensional result to obtain one-dimensional preprocessed data.

[0109] Optionally, the device 500 also includes a configuration module for obtaining the dimension of the covariance matrix; determining the convolution kernel size, padding size, dilation size and stride size according to the dimension; configuring the initial one-dimensional convolution layer in the initial model according to the convolution kernel size, the padding size, the dilation size and the stride size to obtain the one-dimensional convolution layer.

[0110] Optionally, the configuration module is further configured to calculate the convolution kernel size, padding size, dilation size, and stride size based on the dimension according to Formula 1;

[0111] The formula 1 is:

[0112]

[0113] Wherein, L is the dimension, padding is the padding size, kernel is the convolution kernel size, stride is the stride size, and dilation is the expansion size.

[0114] Optionally, the model acquisition module 550 is also used to train the initial model according to the splicing result; whenever the number of iterations reaches an integer multiple of a first preset number of times and the number of iterations has not reached a second preset number of times, the convolution layer parameters of the one-dimensional convolution layer are adjusted to obtain an adjusted one-dimensional convolution layer; if the convolution layer parameters of the adjusted one-dimensional convolution layer meet preset constraints, new sample speech information including multiple audio frames is obtained; return to execute the step of extracting features from the multiple audio frames included in the sample speech information through the initial model until the number of iterations reaches the second preset number of times, and the voiceprint extraction model is obtained.

[0115] Optionally, the model acquisition module 550 is further used to obtain a preset floating-point coefficient; based on the preset floating-point coefficient, according to Formula 2, the convolution layer parameters of the one-dimensional convolution layer are adjusted to obtain an adjusted one-dimensional convolution layer;

[0116] The formula 2 is:

[0117]

[0118] Among them, M′ is the convolution layer parameter of the adjusted one-dimensional convolution layer, M is the convolution layer parameter of the one-dimensional convolution layer, α is the preset floating point coefficient, and M T is the transposed matrix of M, and I is the identity matrix.

[0119] Optionally, the device also includes a prompt module, which is used to stop training the initial model and output a prompt message if the convolution layer parameters of the adjusted one-dimensional convolution layer do not meet the preset constraints, and the prompt message is used to indicate that there is a fault in the training process.

[0120] See also Figure 7 , Figure 7 A block diagram of a voiceprint extraction device proposed in an embodiment of the present application is shown. The device 600 includes:

[0121] A second extraction module 610 is used to extract features of a plurality of audio frames to be extracted included in the voice information to be extracted by using a voiceprint extraction model to obtain a plurality of target frame features corresponding to the plurality of audio frames to be extracted, wherein the voiceprint extraction model is trained by the method described in any of the above embodiments;

[0122] A second obtaining module 620 is used to obtain a covariance matrix, a variance and a mean corresponding to the speech information to be extracted according to the multiple target frame features;

[0123] The second dimensionality reduction module 630 is used to perform dimensionality reduction processing on the covariance matrix of the voice information to be extracted through the one-dimensional convolution layer in the voiceprint extraction model to obtain a reduced-dimensional result, and perform a vectorization operation on the reduced-dimensional result to obtain target one-dimensional preprocessed data;

[0124] A second splicing module 640 is used to perform a splicing operation on the variance of the speech information to be extracted, the mean of the speech information to be extracted, and the target one-dimensional pre-processed data to obtain a target splicing result;

[0125] The voiceprint obtaining module 650 is used to process the target splicing result through the voiceprint extraction model to obtain the voiceprint information corresponding to the voice information to be extracted.

[0126] It should be noted that the device embodiments in the present application correspond to the aforementioned method embodiments. The specific principles in the device embodiments can be found in the contents of the aforementioned method embodiments and will not be repeated here.

[0127] Figure 8 The structure block diagram of an electronic device proposed in one embodiment of the present application is shown, and the electronic device is used to execute the training method of the voiceprint extraction model and the voiceprint extraction method according to the embodiment of the present application. Figure 8As shown, the electronic device 1200 includes a central processing unit (CPU) 1201, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 1202 or the program loaded from the storage part 1208 to the random access memory (RAM) 1203, such as executing the method in the above embodiment. In RAM 1203, various programs and data required for system operation are also stored. CPU 1201, ROM 1202 and RAM 1203 are connected to each other through bus 1204. Input / output (I / O) interface 1205 is also connected to bus 1204.

[0128] The following components are connected to the I / O interface 1205: an input section 1206 including a keyboard, a mouse, etc.; an output section 1207 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 1208 including a hard disk, etc.; and a communication section 1209 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication section 1209 performs communication processing via a network such as the Internet. A drive 1210 is also connected to the I / O interface 1205 as needed. A removable medium 1211, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 1210 as needed so that a computer program read therefrom is installed into the storage section 1208 as needed.

[0129] In particular, according to an embodiment of the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through a communication section 1209, and / or installed from a removable medium 1211. When the computer program is executed by a central processing unit (CPU) 1201, various functions defined in the system of the present application are executed.

[0130] It should be noted that the computer-readable medium shown in the embodiment of the present application may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, - but not limited to - an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, a computer-readable storage medium may be any tangible medium containing or storing a program, which may be used by an instruction execution system, device or device or used in combination with it. In the present application, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, wherein a computer-readable program code is carried. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. Computer-readable signal media may also be any computer-readable medium other than computer-readable storage media, which may send, propagate, or transmit programs for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any appropriate medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.

[0131] The flowchart and block diagram in the accompanying drawings illustrate the possible architecture, functions and operations of the system, method and computer program product according to various embodiments of the present application. Wherein, each box in the flowchart or block diagram can represent a module, a program segment, or a part of the code, and the above-mentioned module, program segment, or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, and the combination of boxes in the block diagram or flowchart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0132] The units involved in the embodiments described in this application may be implemented by software or hardware, and the units described may also be set in a processor. The names of these units do not, in some cases, constitute limitations on the units themselves.

[0133] As another aspect, the present application further provides a computer-readable storage medium, which may be included in the electronic device described in the above embodiments; or may exist independently without being assembled into the electronic device. The above computer-readable storage medium carries computer-readable instructions, and when the computer-readable storage instructions are executed by a processor, the method in any of the above embodiments is implemented.

[0134] According to one aspect of the embodiments of the present application, a computer program product or a computer program is provided, the computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of an electronic device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the electronic device executes the method in any of the above embodiments.

[0135] It should be noted that, although several modules or units of the equipment for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present application, the features and functions of two or more modules or units described above can be embodied in one module or unit. On the contrary, the features and functions of one module or unit described above can be further divided into being embodied by multiple modules or units.

[0136] Through the description of the above implementation methods, it is easy for those skilled in the art to understand that the example implementation methods described here can be implemented by software or by combining software with necessary hardware. Therefore, the technical solution according to the implementation methods of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, and includes several instructions to enable a computing device (which can be a personal computer, a server, a touch terminal, or a network device, etc.) to execute the method according to the implementation methods of the present application.

[0137] Those skilled in the art will readily appreciate other embodiments of the present application after considering the specification and practicing the embodiments disclosed herein. The present application is intended to cover any variations, uses or adaptations of the present application, which follow the general principles of the present application and include common knowledge or customary techniques in the art that are not disclosed in the present application. It should be understood that the present application is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from the scope thereof. The scope of the present application is limited only by the appended claims.

[0138] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for training a voiceprint extraction model, characterized in that: The method comprises: Extracting features of multiple audio frames included in the sample speech information through the initial model to obtain multiple frame features corresponding to the multiple audio frames one by one; According to the multiple frame features, a covariance matrix, a variance and a mean corresponding to the sample speech information are obtained; Performing dimensionality reduction processing on the covariance matrix through the one-dimensional convolution layer in the initial model to obtain a reduced-dimensional result, and performing vectorization operation on the reduced-dimensional result to obtain one-dimensional preprocessed data; Performing a splicing operation on the variance, the mean, and the one-dimensional preprocessed data to obtain a splicing result; The initial model is trained according to the splicing result to obtain a voiceprint extraction model.

2. The method according to claim 1, characterized in that Before performing dimensionality reduction processing on the covariance matrix through the one-dimensional convolution layer to obtain a result after dimensionality reduction, the method further includes: Obtaining the dimension of the covariance matrix; Determine the convolution kernel size, padding size, dilation size, and stride size according to the dimension; According to the convolution kernel size, the padding size, the expansion size, and the stride size, the initial one-dimensional convolution layer in the initial model is configured to obtain the one-dimensional convolution layer.

3. The method according to claim 2, characterized in that The determining of the convolution kernel size, the padding size, the dilation size, and the stride size according to the dimension includes: Based on the dimensions, calculate the convolution kernel size, padding size, dilation size, and stride size according to Formula 1; The formula 1 is: Wherein, L is the dimension, padding is the padding size, kernel is the convolution kernel size, stride is the stride size, and dilation is the expansion size.

4. The method according to claim 1, characterized in that The initial model is trained according to the splicing result to obtain a voiceprint extraction model, including: Training the initial model according to the splicing result; Whenever the number of iterations reaches an integral multiple of a first preset number of times and the number of iterations does not reach a second preset number of times, adjusting the convolution layer parameters of the one-dimensional convolution layer to obtain an adjusted one-dimensional convolution layer; If the convolution layer parameters of the adjusted one-dimensional convolution layer meet the preset constraint conditions, obtaining new sample speech information including multiple audio frames; Return to the step of extracting features from multiple audio frames included in the sample voice information using the initial model until the number of iterations reaches a second preset number, thereby obtaining the voiceprint extraction model.

5. The method according to claim 4, characterized in that The step of adjusting the convolution layer parameters of the one-dimensional convolution layer to obtain an adjusted one-dimensional convolution layer includes: Get the preset floating point coefficient; Based on the preset floating-point coefficient, according to Formula 2, the convolution layer parameters of the one-dimensional convolution layer are adjusted to obtain an adjusted one-dimensional convolution layer; The formula 2 is: Among them, M ′ is the convolution layer parameter of the adjusted one-dimensional convolution layer, M is the convolution layer parameter of the one-dimensional convolution layer, α is the preset floating point coefficient, and M T is the transposed matrix of M, and I is the identity matrix.

6. The method according to claim 4, characterized in that The preset constraint condition includes that the gap between the covariance matrix of the sample speech information and the product result is less than the preset gap, and the product result is the product of the convolution layer parameters of the adjusted one-dimensional convolution layer and the one-dimensional preprocessed data.

7. The method according to claim 4, characterized in that Whenever the number of iterations reaches an integer multiple of the first preset number and the number of iterations does not reach the second preset number, the convolution layer parameters of the one-dimensional convolution layer are adjusted to obtain the adjusted one-dimensional convolution layer, the method further includes: If the convolution layer parameters of the adjusted one-dimensional convolution layer do not meet the preset constraint conditions, the training of the initial model is stopped and a prompt message is output, where the prompt message is used to indicate that there is a fault in the training process.

8. A voiceprint extraction method, characterized in that: The method comprises: Performing feature extraction on a plurality of audio frames to be extracted included in the voice information to be extracted by a voiceprint extraction model, obtaining a plurality of target frame features corresponding one-to-one to the plurality of audio frames to be extracted, wherein the voiceprint extraction model is trained by the method according to any one of claims 1 to 7; According to the multiple target frame features, a covariance matrix, a variance and a mean corresponding to the speech information to be extracted are obtained; Performing dimensionality reduction processing on the covariance matrix of the voice information to be extracted through the one-dimensional convolution layer in the voiceprint extraction model to obtain a reduced-dimensional result, and performing a vectorization operation on the reduced-dimensional result to obtain target one-dimensional preprocessed data; Performing a splicing operation on the variance of the speech information to be extracted, the mean of the speech information to be extracted, and the target one-dimensional preprocessed data to obtain a target splicing result; The target splicing result is processed by the voiceprint extraction model to obtain voiceprint information corresponding to the voice information to be extracted.

9. An electronic device, characterized in that: include: one or more processors; Memory; One or more application programs, wherein the one or more application programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs being configured to execute the method according to any one of claims 1-8.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores program codes, which can be called by a processor to execute the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Vocal print feature fusion method and device

    CN104183240A

  • Voiceprint recognition method and device, computer equipment and storage medium

    CN112562691A