Voiceprint information extraction method, device, electronic device and storage medium
Through the voiceprint extraction model, the voiceprint information is processed, the dimensionality reduction and splicing feature information is reduced, and the problem of low accuracy of voiceprint extraction in the existing technology is solved, and a higher accuracy of voiceprint information extraction is achieved.
Patent Information
- Application Number
- CN202210907481.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-29
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2042-07-29
AI Technical Summary
In the prior art, when the voiceprint extracted through deep neural networks, the accuracy of the extracted voiceprint information is low.
The target voice information is processed through the voiceprint extraction model, and the target covariance, target variance and target mean are obtained. The target covariance is reduced through the bilinear parameter layer to obtain the target one-dimensional data. Then the target variance, target mean and target one-dimensional data are spliced to obtain the target splicing results. Finally, the target splicing results are processed through the voiceprint extraction model to obtain the voiceprint information.
By accurately characterizing the characteristic information of the time and frequency dimensions of the target speech information, the accuracy of the extracted voiceprint information is improved.
Smart Images

Figure CN115394302B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of audio processing technology, and more specifically, to a voiceprint information extraction method, device, electronic device and storage medium. Background Art
[0002] Voiceprint recognition is a technology that uses sound to identify the identity of a voice user, and is one of the important research directions in the field of speech. With the continuous development of computer technology, voiceprint recognition has made great progress in recent years. With its convenient and effective characteristics, it has become an efficient identity recognition method and is widely used in public security, banking, and smart homes.
[0003] At present, the deep neural network can be trained through samples to obtain a voiceprint extraction model, and then the voiceprint extraction model is used to extract the voiceprint of the voice information to be extracted. However, when this method is used to extract the voiceprint of the voice information to be extracted, the accuracy of the extracted voiceprint information is low. Summary of the invention
[0004] In view of this, the embodiments of the present application propose a voiceprint information extraction method, device, electronic device and storage medium.
[0005] In a first aspect, an embodiment of the present application provides a method for extracting voiceprint information, the method comprising: processing target voice information through a voiceprint extraction model to obtain a target covariance, target variance and target mean corresponding to the target voice information; performing dimensionality reduction processing on the target covariance through a bilinear parameter layer in the voiceprint extraction model to obtain target one-dimensional data; performing a splicing operation on the target variance, the target mean and the target one-dimensional data to obtain a target splicing result; processing the target splicing result through the voiceprint extraction model to obtain voiceprint information corresponding to the target voice information.
[0006] In a second aspect, an embodiment of the present application provides a voiceprint information extraction device, the device comprising: a speech processing module, used to process the target speech information through a voiceprint extraction model to obtain a target covariance, a target variance and a target mean corresponding to the target speech information; a dimension reduction module, used to perform dimension reduction processing on the target covariance through a bilinear parameter layer in the voiceprint extraction model to obtain target one-dimensional data; a splicing module, used to perform a splicing operation on the target variance, the target mean and the target one-dimensional data to obtain a target splicing result; a voiceprint acquisition module, used to process the target splicing result through the voiceprint extraction model to obtain voiceprint information corresponding to the target speech information.
[0007] In a third aspect, an embodiment of the present application provides an electronic device, including a processor and a memory; one or more programs are stored in the memory and configured to be executed by the processor to implement the above method.
[0008] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, in which a program code is stored, wherein the above method is executed when the program code is executed by a processor.
[0009] In a fifth aspect, an embodiment of the present application provides a computer program product or a computer program, the computer program product or the computer program including computer instructions, the computer instructions being stored in a computer-readable storage medium. A processor of an electronic device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the electronic device performs the above method.
[0010] The embodiments of the present application provide a voiceprint information extraction method, device, electronic device and storage medium. The target voice information is processed by a voiceprint extraction model to obtain a target covariance, a target variance and a target mean, and the target one-dimensional data, the target variance and the target mean corresponding to the target covariance are spliced to obtain a target splicing result, and then the voiceprint information is obtained according to the target splicing result. The target covariance accurately represents the characteristic information of the time dimension and the frequency dimension of the target voice information, so that the target splicing result can accurately represent the voiceprint characteristics of the target voice information, thereby improving the accuracy of the extracted voiceprint information. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.
[0012] Figure 1 A flowchart showing a method for training a voiceprint extraction model in an embodiment of the present application is shown;
[0013] Figure 2 A schematic diagram of the structure of the target model in an embodiment of the present application is shown;
[0014] Figure 3 Shows Figure 1 A flowchart of an implementation of step S140;
[0015] Figure 4 A schematic diagram of the structure of the deep statistical pooling layer in an embodiment of the present application is shown;
[0016] Figure 5 A flow chart of a voiceprint information extraction method proposed in one embodiment of the present application is shown;
[0017] Figure 6 A block diagram of a voiceprint information extraction device proposed in one embodiment of the present application is shown;
[0018] Figure 7 A structural block diagram of an electronic device proposed in one embodiment of the present application is shown. DETAILED DESCRIPTION
[0019] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. According to the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0020] In the following description, the terms "first\second" involved are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It can be understood that "first\second" can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.
[0021] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0022] See also Figure 1 , Figure 1 A flow chart of a method for training a voiceprint extraction model in an embodiment of the present application is shown. The method can be used in an electronic device, and the method includes:
[0023] S110 , processing the sample speech information through the target model to obtain the covariance, variance and mean corresponding to the sample speech information.
[0024] The sample voice information may refer to the voice information used to train the target model, which may be voice in any format, such as MP3, AAC, etc. The sample voice information may be recorded by an electronic device or obtained by the electronic device from the Internet. The sample voice information may be one or more pieces. When the sample voice information includes multiple pieces, each piece of sample voice information is used as a batch of training samples.
[0025] The target model refers to the basic model used to obtain the voiceprint extraction model. The sample voice information is processed by the target model to obtain the covariance, variance and mean of the sample voice information.
[0026] As an implementation mode, the sample speech information includes multiple audio frames, and S110 may include: extracting features from the multiple audio frames through the target model to obtain multiple audio frame features corresponding to the multiple audio frames; and obtaining the covariance, variance and mean corresponding to the sample speech information according to the multiple audio frame features. In this application, covariance refers to a covariance matrix.
[0027] The sample voice information may include multiple audio frames, for example, 1s sample voice information may generally include 30 audio frames or 60 audio frames, etc. The user may set a time shift (a time shift refers to the duration of an audio frame), and the electronic device determines the total number of audio frames included in the sample voice information according to the set time shift, for example, if the time shift is 10ms and the sample voice information is 2s, the total number of audio frames obtained is 200.
[0028] The target model may include an input layer, which may be a neural network input layer. Feature extraction is performed on multiple audio frames through the input layer to obtain multiple audio frame features corresponding to the multiple audio frames. After feature extraction is performed on each audio frame, the obtained feature is used as an audio frame feature, and each audio frame feature may include numerical values of multiple dimensions.
[0029] Correspondingly, the number of dimensions of the input layer is the same as the number of dimensions of the audio frame features, and one dimension of the input layer is used to process one dimension of the audio frame features. For example, the input sample speech information has T frames, and the obtained audio frame feature is x i ,i=1,2,…,T, each audio frame feature includes 20 dimensions, and the input layer also includes 20 dimensions.
[0030] As an implementation method, after obtaining the sample voice information, the sample voice information can be formatted and converted into a standard sample voice in wav format, and the standard sample voice can be Fourier transformed to obtain voice features (the voice features can be data in MSCC format, and the voice features include voice features corresponding to multiple audio frames), and then the voice features are input into the input layer of the target model to obtain multiple audio frame features corresponding to the multiple audio frames.
[0031] After obtaining multiple audio frame features corresponding to the sample speech information, the covariance, variance and mean are calculated according to the multiple audio frame features.
[0032] As an implementation method, the covariance may be calculated according to multiple audio frame features by using Formula 3; Formula 3 is as follows:
[0033]
[0034]
[0035]
[0036] Among them, C (k,m) is the value of the kth row and mth column in the covariance, T is the total number of features of multiple audio frames, x ik 、x im are the values corresponding to any two dimensions of the i-th audio frame feature.
[0037] As an implementation method, the mean value may be calculated according to the features of multiple audio frames by using Formula 4, where Formula 4 is:
[0038]
[0039] Among them, m and is the mean, x i is any dimension of the feature of the i-th audio frame.
[0040] As an implementation method, the variance may be calculated according to multiple audio frame features by using Formula 5, where Formula 5 is:
[0041]
[0042] Where v is the variance.
[0043] It can be understood that the mean and variance include the mean and variance of each dimension respectively. For example, if the audio frame feature includes values of 10 dimensions, the variance includes the variances corresponding to the 10 dimensions respectively, and the mean includes the means corresponding to the 10 dimensions respectively.
[0044] S120, performing dimensionality reduction processing on the covariance through the bilinear parameter layer in the target model to obtain a reduced-dimensional result, and performing a square root regularization operation on the reduced-dimensional result to obtain one-dimensional data.
[0045] Variance and mean are one-dimensional data, while covariance is multi-dimensional data. Therefore, covariance cannot be directly concatenated with variance and mean. It is necessary to perform dimensionality reduction on covariance to obtain one-dimensional data, and then the one-dimensional data can be concatenated with variance and mean to obtain the concatenated result.
[0046] The target model may include a bilinear parameter layer, which is used to reduce the dimension of the covariance to obtain one-dimensional data.
[0047] As an implementation mode, the bilinear parameter layer in the target model includes a parameter matrix; the covariance is subjected to dimensionality reduction processing through the bilinear parameter layer in the target model to obtain a reduced-dimensionality result, including: transforming each column of data in the covariance through the parameter matrix to obtain a corresponding conversion result for each column of data in the covariance; and concatenating the corresponding conversion results for each column of data in the covariance to obtain the reduced-dimensionality result.
[0048] The method of converting each column of data in the covariance by using the parameter matrix to obtain a conversion result corresponding to each column of data in the covariance comprises: converting each column of data in the covariance according to formula 1 by using the parameter matrix to obtain a conversion result corresponding to each column of data in the covariance;
[0049] The formula 1 is:
[0050]
[0051] Among them, w j is the j-th column of the parameter matrix, w j The transposed matrix, C is the covariance, z j is the conversion result corresponding to the j-th column of the covariance.
[0052] Get z j After that, you can stitch all the z j , thus obtaining the result after dimensionality reduction, wherein the splicing process can refer to Formula 6, which is:
[0053]
[0054] Where L is the dimension of the covariance.
[0055] As an implementation method, the square root normalization operation can be performed on the result after dimensionality reduction through Formula 7 to obtain one-dimensional data. Formula 7 is as follows:
[0056]
[0057] Among them, vec is one-dimensional data, and z is the result after dimensionality reduction.
[0058] After performing a square root normalization operation on the reduced-dimensional result, the obtained one-dimensional data is one-dimensional vector data. Through the square root normalization operation, the reduced-dimensional result can be converted into one-dimensional vector data that can effectively describe the delay changes of speech features, so that the accuracy of the one-dimensional data is higher.
[0059] The parameter matrix of the initial state of the bilinear parameter layer of the target model may be determined by using algorithms such as uniform distribution or Gaussian distribution. Then, the initial bilinear parameter layer in the target model is configured by using the parameter matrix of the initial state, and the configured bilinear parameter layer is obtained as the bilinear parameter layer of the target model.
[0060] S130, performing a splicing operation on the variance, the mean, and the one-dimensional data to obtain a splicing result.
[0061] According to Formula 8, the variance, the mean and the one-dimensional data may be concatenated to obtain a concatenation result. Formula 8 is:
[0062] y=cat{m,v,vec}
[0063] Among them, y is the concatenation result, vec is one-dimensional data, and cat is the concatenation function.
[0064] like Figure 2 As shown, the target model may include a deep statistical pooling layer, the deep statistical pooling layer includes a bilinear parameter layer, and the deep statistical pooling layer is also used to perform a square root regularization operation on the dimensionality reduction result corresponding to the covariance to obtain one-dimensional data, and to splice the variance, mean and one-dimensional data after vectorization to obtain a splicing result.
[0065] S140. Training the target model according to the splicing result to obtain a voiceprint extraction model.
[0066] After obtaining the splicing result, the loss value can be obtained according to the splicing result, and the parameters of the target model can be adjusted according to the loss value until the number of iterations meets the requirements to obtain the voiceprint extraction model.
[0067] like Figure 2 As shown, the target model may further include a segment level layer and an output layer. The segment level layer may include at least one hidden layer for processing the splicing result, and the output layer may be a general neural network output layer for processing the output of the segment level layer to obtain an output result.
[0068] After obtaining the output result, the loss value is determined according to the output result, and then the parameters of the input layer, deep statistical pooling layer, segment level layer and output layer in the target model are adjusted according to the determined loss value to obtain the voiceprint extraction model.
[0069] In this embodiment, the sample voice information can be processed by the target model to obtain the covariance, variance and mean, and the one-dimensional data, variance and mean corresponding to the covariance are spliced, and then the target model is trained through the splicing result to obtain the voiceprint extraction model. The covariance accurately represents the characteristic information of the time dimension and frequency dimension of the sample voice information, so that the splicing result can accurately represent the voiceprint characteristics of the sample voice information, thereby improving the voiceprint extraction accuracy of the voiceprint extraction model.
[0070] At the same time, covariance can take into account both time dimension and frequency dimension information, integrate independent and unrelated dimensions of the deep statistical pooling layer, and provide new information for segment-level network modeling, thereby improving the voiceprint extraction accuracy of the voiceprint extraction model. Through the square root regularization operation, the vectorized conversion of the dimensionality reduction result can effectively describe the delay changes of voice features and improve the accuracy of one-dimensional data.
[0071] See also Figure 3 , Figure 3 Shows Figure 1 The method can be used in an electronic device, and the method includes:
[0072] S210, training the target model according to the splicing result; whenever the number of iterations reaches an integer multiple of the first target number, determining whether the number of iterations reaches the second target number. If so, executing S220, if not, executing S230.
[0073] S220: Obtain a voiceprint extraction model.
[0074] S230, adjusting the parameter matrix of the bilinear parameter layer in the target model to obtain an adjusted bilinear parameter layer; if the parameter matrix of the adjusted bilinear parameter layer satisfies a preset constraint condition, obtaining new sample speech information, and returning to execute step S110.
[0075] After obtaining the splicing result, the splicing result is input into the segment level layer of the target model to obtain the output result of the segment level layer, and then the output result of the segment level layer is input into the output layer to obtain the output result of the output layer. Then, the loss value is determined according to the output result, and the target model is trained through the loss value.
[0076] The first target number can be a value set according to the scenario and requirements, such as 3 times. The second target number refers to the maximum number of iterations for training the voiceprint extraction model, such as 10,000 times. When the number of iterations reaches the second target number, the training ends and the voiceprint extraction model is obtained. When the number of iterations does not reach the second target number, the training does not end and iterative training continues.
[0077] Whenever the number of iterations reaches an integer multiple of the first target number and does not reach the second target number, it is determined that the parameter matrix of the bilinear parameter layer needs to be adjusted, and the parameter matrix of the bilinear parameter layer is adjusted to obtain the adjusted bilinear parameter layer, and then the target model including the adjusted bilinear parameter layer is continued to be trained. When the number of iterations does not reach an integer multiple of the first target number and does not reach the second target number, the parameter matrix of the bilinear parameter layer is not adjusted, and the target model including the bilinear parameter layer is continued to be trained.
[0078] As an implementation manner, the adjusting the parameter matrix of the bilinear parameter layer in the target model to obtain the adjusted bilinear parameter layer includes: obtaining a preset floating-point coefficient; based on the preset floating-point coefficient, according to Formula 2, adjusting the parameter matrix of the bilinear parameter layer in the target model to obtain the adjusted bilinear parameter layer;
[0079] The formula 2 is:
[0080]
[0081] Where W′ is the parameter matrix of the adjusted bilinear parameter layer, W is the parameter matrix of the bilinear parameter layer in the target model, α is the preset floating point coefficient, and W T is the transposed matrix of W, and I is the identity matrix. The preset floating-point coefficient can be 1.
[0082] As an implementation mode, the preset constraint condition includes that the gap between the covariance of the sample speech information and the product result is less than a preset gap, and the product result is the product of the parameter matrix of the adjusted bilinear parameter layer and the one-dimensional data, wherein the preset gap can be a value set based on demand, which is not limited in this application. The mathematical formula of the preset constraint condition is as follows:
[0083] C→Wvec
[0084] When the parameter matrix of the adjusted bilinear parameter layer satisfies the preset constraints and the preset constraints converge, then: C = Wvec ≈ WW T , that is, the output one-dimensional data vec is approximately equal to the parameter matrix W of the bilinear parameter layer. The input covariance information can be retained to the maximum extent in the dimensionality reduction operation and the square root normalization operation, so that the accuracy of the one-dimensional data is higher, and then the accuracy of the voiceprint extraction model obtained by training based on the one-dimensional data is higher.
[0085] When the parameter matrix of the adjusted bilinear parameter layer does not meet the preset constraints, the training of the target model is stopped and a prompt message is output to prompt the user that an error has occurred in the training process, so that the user can reconfigure the target model.
[0086] The new sample voice information may also include multiple audio frames. The description of the new sample voice information refers to the description of the sample voice information above and will not be repeated. After obtaining the new sample voice information, return to step S110 until the number of iterations reaches the second target number to obtain a voiceprint extraction model.
[0087] In this embodiment, the structure of the deep statistical pooling layer in the target model can refer to Figure 4 .like Figure 4 As shown, in the deep statistical pooling layer, the covariance, variance and mean of the corresponding sample speech information are obtained according to the multiple audio frame features output by the input layer, and the covariance is reduced in dimension through the bilinear parameter layer in the deep statistical pooling layer to obtain the reduced-dimensional result, and the reduced-dimensional result is square-rooted to obtain one-dimensional data, and then the variance, mean and one-dimensional data are concatenated to obtain the concatenated result, and then the concatenated result is input into the segment-level layer for processing, and finally the output layer processes the result output by the segment-level layer to obtain the output result.
[0088] Wherein, whenever the number of iterations reaches an integer multiple of the first target number and the number of iterations does not reach the second target number, the parameter matrix of the bilinear parameter layer is adjusted by semi-orthogonal constraint to obtain an adjusted bilinear parameter layer, and the adjusted bilinear parameter layer is applied to the subsequent training process. Wherein, the semi-orthogonal constraint refers to adjusting the parameter matrix of the bilinear parameter layer according to Formula 2, and determining the parameter matrix of the adjusted bilinear parameter layer according to the preset constraint condition.
[0089] In this embodiment, the parameter matrix of the bilinear parameter layer is semi-orthogonal by semi-orthogonal constraints, so that the features after the square root normalization operation can retain the covariance information to the maximum extent, approach the effect of the eigenvalue features, and play a role in characterizing important components and minor components; at the same time, the deep statistical pooling layer based on the bilinear parameter layer and the semi-orthogonal constraints can significantly improve the recognition accuracy of the voiceprint recognition model.
[0090] See also Figure 5 , Figure 5 A flow chart of a voiceprint extraction method proposed in one embodiment of the present application is shown. The method can be used in an electronic device, and the method includes:
[0091] S310. Process the target voice information through a voiceprint extraction model to obtain a target covariance, a target variance, and a target mean corresponding to the target voice information.
[0092] The training process of the voiceprint extraction model is as described in the above embodiment and will not be repeated here.
[0093] The target voice information refers to the voice information to be extracted from the voiceprint information, which is similar to the description of the sample voice information and will not be repeated here. The target voice information may include multiple audio frames, which are used as multiple audio frames to be extracted. The voiceprint extraction model can be used to extract features of the multiple audio frames to be extracted to obtain multiple target audio frame features corresponding to the multiple audio frames to be extracted.
[0094] The audio frame in the target speech information is the audio frame to be extracted. Feature extraction is performed on the audio frame to be extracted, and the obtained audio frame feature is the target audio frame feature.
[0095] The target speech information can be converted into a standard speech to be extracted in wav format, and the standard speech to be extracted is subjected to Fourier transform to obtain the speech features to be extracted (the speech features to be extracted include the speech features to be extracted corresponding to each of the multiple audio frames to be extracted), and then the speech features to be extracted are input into the input layer of the voiceprint extraction model to obtain multiple target audio frame features corresponding to the multiple audio frames to be extracted.
[0096] Afterwards, the covariance of the corresponding target speech information can be obtained as the target covariance based on the multiple target audio frame features, the variance of the corresponding target speech information can be obtained as the target variance based on the multiple target audio frame features, and the mean of the corresponding target speech information can be obtained as the target mean based on the multiple target audio frame features.
[0097] Among them, the process of obtaining the target covariance, target variance and target mean based on multiple target audio frame features is similar to the process of obtaining the covariance, variance and mean of the corresponding sample speech information based on multiple audio frame features mentioned above, and will not be repeated here.
[0098] S320, performing dimensionality reduction processing on the target covariance through the bilinear parameter layer in the voiceprint extraction model to obtain target one-dimensional data.
[0099] The bilinear parameter layer target covariance in the voiceprint extraction model can be used to perform dimensionality reduction processing to obtain the reduced-dimensional result, and the reduced-dimensional result can be subjected to a square root regularization operation to obtain the target one-dimensional data.
[0100] The specific process of obtaining the target one-dimensional data according to the target covariance can refer to the process of obtaining the one-dimensional data according to the covariance of the sample speech information mentioned above, which will not be repeated here.
[0101] S330, performing a splicing operation on the target variance, the target mean, and the target one-dimensional data to obtain a target splicing result.
[0102] The description of S330 refers to the description of S130 above and will not be repeated here.
[0103] The target one-dimensional data refers to the result obtained after performing dimensionality reduction processing and square root regularization operation on the target covariance of the target speech information through the bilinear parameter layer in the voiceprint extraction model, and the target splicing result refers to the result obtained after splicing the target variance, the target mean and the target one-dimensional data.
[0104] S340: Process the target concatenation result by using the voiceprint extraction model to obtain voiceprint information corresponding to the target voice information.
[0105] After obtaining the concatenation result, the concatenation result is processed by the segment level layer of the voiceprint extraction model to obtain the processed result, and then the voiceprint information of the target voice information is obtained according to the processed result output by the segment level layer.
[0106] In this embodiment, the target speech information is processed by a voiceprint extraction model to obtain a target covariance, a target variance and a target mean, and the target one-dimensional data, the target variance and the target mean corresponding to the target covariance are spliced to obtain a target splicing result, and then the voiceprint information is obtained based on the target splicing result. The target covariance accurately represents the characteristic information of the time dimension and the frequency dimension of the target speech information, so that the target splicing result can accurately represent the voiceprint characteristics of the target speech information, thereby improving the accuracy of the extracted voiceprint information.
[0107] See also Figure 6 , Figure 6 A block diagram of a voiceprint information extraction device proposed in an embodiment of the present application is shown. The device 500 includes:
[0108] The speech processing module 510 is used to process the target speech information through the voiceprint extraction model to obtain the target covariance, target variance and target mean corresponding to the target speech information;
[0109] A dimension reduction module 520, configured to perform dimension reduction processing on the target covariance through a bilinear parameter layer in the voiceprint extraction model to obtain target one-dimensional data;
[0110] A splicing module 530 is used to perform a splicing operation on the target variance, the target mean and the target one-dimensional data to obtain a target splicing result;
[0111] The voiceprint obtaining module 540 is used to process the target splicing result through the voiceprint extraction model to obtain voiceprint information corresponding to the target voice information.
[0112] Optionally, the device 500 also includes a training module, which is used to process the sample voice information through the target model to obtain the covariance, variance and mean corresponding to the sample voice information; perform dimensionality reduction processing on the covariance through the bilinear parameter layer in the target model to obtain the reduced dimension result, and perform a square root regularization operation on the reduced dimension result to obtain one-dimensional data; perform a splicing operation on the variance, the mean and the one-dimensional data to obtain a splicing result; train the target model according to the splicing result to obtain a voiceprint extraction model.
[0113] Optionally, the sample speech information includes multiple audio frames; the training module is also used to extract features from the multiple audio frames through the target model to obtain multiple audio frame features corresponding one to one to the multiple audio frames; based on the multiple audio frame features, obtain the covariance, variance and mean corresponding to the sample speech information.
[0114] Optionally, the bilinear parameter layer in the target model includes a parameter matrix; the training module is also used to transform each column of data in the covariance through the parameter matrix to obtain a conversion result corresponding to each column of data in the covariance; and perform a splicing operation on the conversion results corresponding to each column of data in the covariance to obtain the result after dimensionality reduction.
[0115] Optionally, the training module is further used to transform each column of data in the covariance according to Formula 1 through the parameter matrix to obtain a transformation result corresponding to each column of data in the covariance;
[0116] The formula 1 is:
[0117]
[0118] Among them, w j is the j-th column of the parameter matrix, w j The transposed matrix, C is the covariance, z j is the conversion result corresponding to the j-th column of the covariance.
[0119] Optionally, the training module is also used to train the target model according to the splicing result; whenever the number of iterations reaches an integer multiple of the first target number and the number of iterations has not reached the second target number, the parameter matrix of the bilinear parameter layer in the target model is adjusted to obtain an adjusted bilinear parameter layer; if the parameter matrix of the adjusted bilinear parameter layer satisfies preset constraints, new sample speech information is obtained; return to execute the step of processing the sample speech information through the target model until the number of iterations reaches the second target number to obtain the voiceprint extraction model.
[0120] Optionally, the training module is further used to obtain preset floating-point coefficients; based on the preset floating-point coefficients, according to Formula 2, the parameter matrix of the bilinear parameter layer in the target model is adjusted to obtain an adjusted bilinear parameter layer;
[0121] The formula 2 is:
[0122]
[0123] Where W′ is the parameter matrix of the adjusted bilinear parameter layer, W is the parameter matrix of the bilinear parameter layer in the target model, α is the preset floating point coefficient, and W T is the transposed matrix of W, and I is the identity matrix.
[0124] It should be noted that the device embodiments in the present application correspond to the aforementioned method embodiments. The specific principles in the device embodiments can be found in the contents of the aforementioned method embodiments and will not be repeated here.
[0125] Figure 7 The structure block diagram of an electronic device proposed in one embodiment of the present application is shown, and the electronic device is used to execute the training method of the voiceprint extraction model and the voiceprint extraction method according to the embodiment of the present application. Figure 7 As shown, the electronic device 1200 includes a central processing unit (CPU) 1201, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 1202 or the program loaded from the storage part 1208 to the random access memory (RAM) 1203, such as executing the method in the above embodiment. In RAM 1203, various programs and data required for system operation are also stored. CPU 1201, ROM 1202 and RAM 1203 are connected to each other through bus 1204. Input / output (I / O) interface 1205 is also connected to bus 1204.
[0126] The following components are connected to the I / O interface 1205: an input section 1206 including a keyboard, a mouse, etc.; an output section 1207 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 1208 including a hard disk, etc.; and a communication section 1209 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication section 1209 performs communication processing via a network such as the Internet. A drive 1210 is also connected to the I / O interface 1205 as needed. A removable medium 1211, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 1210 as needed so that a computer program read therefrom is installed into the storage section 1208 as needed.
[0127] In particular, according to an embodiment of the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through a communication section 1209, and / or installed from a removable medium 1211. When the computer program is executed by a central processing unit (CPU) 1201, various functions defined in the system of the present application are executed.
[0128] It should be noted that the computer-readable medium shown in the embodiment of the present application may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, - but not limited to - an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, a computer-readable storage medium may be any tangible medium containing or storing a program, which may be used by an instruction execution system, device or device or used in combination with it. In the present application, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, wherein a computer-readable program code is carried. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. Computer-readable signal media may also be any computer-readable medium other than computer-readable storage media, which may send, propagate, or transmit programs for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any appropriate medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.
[0129] The flowchart and block diagram in the accompanying drawings illustrate the possible architecture, functions and operations of the system, method and computer program product according to various embodiments of the present application. Wherein, each box in the flowchart or block diagram can represent a module, a program segment, or a part of the code, and the above-mentioned module, program segment, or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, and the combination of boxes in the block diagram or flowchart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0130] The units involved in the embodiments described in this application may be implemented by software or hardware, and the units described may also be set in a processor. The names of these units do not, in some cases, constitute limitations on the units themselves.
[0131] As another aspect, the present application further provides a computer-readable storage medium, which may be included in the electronic device described in the above embodiments; or may exist independently without being assembled into the electronic device. The above computer-readable storage medium carries computer-readable instructions, and when the computer-readable storage instructions are executed by a processor, the method in any of the above embodiments is implemented.
[0132] According to one aspect of the embodiments of the present application, a computer program product or a computer program is provided, the computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of an electronic device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the electronic device executes the method in any of the above embodiments.
[0133] It should be noted that, although several modules or units of the equipment for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present application, the features and functions of two or more modules or units described above can be embodied in one module or unit. On the contrary, the features and functions of one module or unit described above can be further divided into being embodied by multiple modules or units.
[0134] Through the description of the above implementation methods, it is easy for those skilled in the art to understand that the example implementation methods described here can be implemented by software or by combining software with necessary hardware. Therefore, the technical solution according to the implementation methods of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, and includes several instructions to enable a computing device (which can be a personal computer, a server, a touch terminal, or a network device, etc.) to execute the method according to the implementation methods of the present application.
[0135] Those skilled in the art will readily appreciate other embodiments of the present application after considering the specification and practicing the embodiments disclosed herein. The present application is intended to cover any variations, uses or adaptations of the present application, which follow the general principles of the present application and include common knowledge or customary techniques in the art that are not disclosed in the present application. It should be understood that the present application is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from the scope thereof. The scope of the present application is limited only by the appended claims.
[0136] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A method for extracting voiceprint information, characterized in that: The method comprises: Processing the target voice information through the voiceprint extraction model to obtain a target covariance, a target variance, and a target mean corresponding to the target voice information; Performing dimensionality reduction processing on the target covariance through the bilinear parameter layer in the voiceprint extraction model to obtain target one-dimensional data; Performing a splicing operation on the target variance, the target mean, and the target one-dimensional data to obtain a target splicing result; The target splicing result is processed by the voiceprint extraction model to obtain voiceprint information corresponding to the target voice information.
2. The method according to claim 1, characterized in that: The training method of the voiceprint extraction model comprises: Processing the sample speech information through the target model to obtain the covariance, variance and mean corresponding to the sample speech information; Performing dimensionality reduction processing on the covariance through a bilinear parameter layer in the target model to obtain a dimensionality reduced result, and performing a square root regularization operation on the dimensionality reduced result to obtain one-dimensional data; Performing a splicing operation on the variance, the mean, and the one-dimensional data to obtain a splicing result; The target model is trained according to the splicing result to obtain a voiceprint extraction model.
3. The method according to claim 2, characterized in that The sample speech information includes a plurality of audio frames; the sample speech information is processed by the target model to obtain the covariance, variance and mean corresponding to the sample speech information, including: Extracting features from the multiple audio frames using the target model to obtain multiple audio frame features corresponding to the multiple audio frames one by one; According to the multiple audio frame features, the covariance, variance and mean corresponding to the sample speech information are obtained.
4. The method according to claim 2, characterized in that: The bilinear parameter layer in the target model includes a parameter matrix; the dimensionality reduction processing is performed on the covariance through the bilinear parameter layer in the target model to obtain a result after dimensionality reduction, including: By using the parameter matrix, each column of data in the covariance is transformed to obtain a transformation result corresponding to each column of data in the covariance; The conversion results corresponding to each column of data in the covariance are concatenated to obtain the result after dimensionality reduction.
5. The method according to claim 4, characterized in that The method of converting each column of data in the covariance by using the parameter matrix to obtain a conversion result corresponding to each column of data in the covariance includes: By using the parameter matrix, according to formula 1, each column of data in the covariance is transformed to obtain a corresponding transformation result of each column of data in the covariance; The formula 1 is: Among them, w j is the j-th column of the parameter matrix, w j The transposed matrix, C is the covariance, z j is the conversion result corresponding to the j-th column of the covariance.
6. The method according to claim 2, characterized in that The target model is trained according to the splicing result to obtain a voiceprint extraction model, including: Training the target model according to the splicing result; Whenever the number of iterations reaches an integer multiple of the first target number and the number of iterations does not reach the second target number, the parameter matrix of the bilinear parameter layer in the target model is adjusted to obtain an adjusted bilinear parameter layer; the first target number is less than the second target number; the second target number is the maximum number of iterations of the training process of the voiceprint extraction model; If the parameter matrix of the adjusted bilinear parameter layer satisfies a preset constraint condition, new sample speech information is obtained; the preset constraint condition includes that the gap between the covariance of the sample speech information and the product result is less than a preset gap, and the product result is the product of the parameter matrix of the adjusted bilinear parameter layer and the one-dimensional data; Return to the step of processing the sample speech information through the target model until the number of iterations reaches the second target number, and obtain the voiceprint extraction model.
7. The method according to claim 6, characterized in that The step of adjusting the parameter matrix of the bilinear parameter layer in the target model to obtain an adjusted bilinear parameter layer includes: Get the preset floating point coefficient; Based on the preset floating-point coefficients, according to Formula 2, the parameter matrix of the bilinear parameter layer in the target model is adjusted to obtain an adjusted bilinear parameter layer; The formula 2 is: Where W′ is the parameter matrix of the adjusted bilinear parameter layer, W is the parameter matrix of the bilinear parameter layer in the target model, α is the preset floating point coefficient, and W T is the transposed matrix of W, and I is the identity matrix.
8. A voiceprint information extraction device, characterized in that: The device comprises: A speech processing module, used to process the target speech information through a voiceprint extraction model to obtain a target covariance, a target variance and a target mean corresponding to the target speech information; A dimensionality reduction module, used for performing dimensionality reduction processing on the target covariance through a bilinear parameter layer in the voiceprint extraction model to obtain target one-dimensional data; A splicing module, used for performing a splicing operation on the target variance, the target mean and the target one-dimensional data to obtain a target splicing result; The voiceprint acquisition module is used to process the target splicing result through the voiceprint extraction model to obtain voiceprint information corresponding to the target voice information.
9. An electronic device, characterized in that: include: one or more processors; Memory; One or more application programs, wherein the one or more application programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs being configured to execute the method according to any one of claims 1-7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores program codes, which can be called by a processor to execute the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Method and device for data evaluation based on voice
CN107452385A
Voice model training method, speaker identification method, apparatus and device, and medium
CN108777146A