A method, device, and storage medium for voiceprint recognition

By integrating channel compression-excitation and frequency compression-excitation in the voiceprint recognition model, the problem of limited resolution of frequency information in the prior art is solved, and more efficient voiceprint recognition effect is achieved, and excellent performance in cross-language applications.

CN114446310BActive Publication Date: 2025-05-30XIAMEN KUAISHANGTONG TECH CORP LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210079352.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-24
Publication Date
2025-05-30
Estimated Expiration
2042-01-24

AI Technical Summary

Technical Problem

The existing voiceprint recognition system mixes the spatial dimension and channel dimension together in the process of feature abstraction, resulting in limited resolution of frequency information, and poor results when used alone.

Method used

The SEfwSE module that combines the channel compression-excitation SE submodule and the frequency compression-excitation fwSE submodule is adopted to improve the resolution of the channel and frequency dimensions by compressing the voiceprint recognition model of the excitation channel and frequency dimensions.

Benefits of technology

On the premise of increasing the amount of calculation, the effect of voiceprint recognition is improved, especially in cross-language voiceprint recognition application scenarios, and the robustness of the model is improved through attention mechanism and data enhancement.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114446310B_ABST
    Figure CN114446310B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of artificial intelligence technology. Specifically, the present invention discloses a method for voiceprint recognition, which uses a voiceprint recognition model trained through the following steps to perform voiceprint recognition: obtaining a training set, the training set containing multiple audio data; extracting the audio features of the audio data contained in the training set; performing a slicing operation on the audio features to obtain multiple audio slice features of the same length; randomly obtaining a fixed number of audio slice features each time and inputting them into the voiceprint recognition model for training, and iteratively training multiple times to obtain a trained voiceprint recognition model; wherein, the voiceprint recognition model is implemented based on a neural network and includes a processing module that integrates at least two squeeze-and-excitation sub-modules. The method and device for voiceprint recognition provided by the present invention can perform excitation on the channel dimension and frequency dimension of voiceprint recognition, add the excitation results, improve the resolution of both the channel dimension and frequency dimension at the same time, and improve the effect of voiceprint recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of artificial intelligence, in particular to the technical field of voiceprint recognition, and more particularly to a method, device and storage medium for voiceprint recognition. Background Art

[0002] With the development of deep learning, deep neural networks have also been applied to the field of voiceprint recognition. Currently, mainstream voiceprint recognition systems generally consist of two parts, namely the front-end embedding extraction part and the back-end loss function calculation and similarity calculation part. In the training stage, the front-end embedding extraction network is used to extract the embedding and input it into the back-end loss function calculation part, and the network parameters are updated through backpropagation; in the test stage, the back-end loss function calculation part is replaced by the similarity calculation part, and the embedding is extracted through forward propagation to calculate the similarity. The front-end embedding extraction part adopts a neural network structure. In order to speed up the training speed, convolutional operators are currently mainly stacked, and both one-dimensional convolution and two-dimensional convolution have been successfully applied. In order to reduce the training difficulty and deepen the network layer, several convolutional stacking structures are usually changed into residual structures, and the most successful network structure is ResNet.

[0003] Aiming at the problem that almost all network structures mix the spatial dimension and the channel dimension together for feature abstraction, Document 1 (Squeeze-and-Excitation Networks, SE) proposes a method of separating the spatial dimension and the channel dimension, controlling the spatial dimension (compression), and improving the resolution of the channel dimension (excitation). For speech, frequency information is very important. For one-dimensional convolution, the frequency information will be completely compressed, and for two-dimensional convolution, it will be mixed with time information, restricting the resolution of the frequency dimension. Aiming at this problem, Document 2 (The IDLAB VoxCeleb Speaker Recognition Challenge 2021 System Description) proposes a method (fwSE) of compressing channel information and time dimension information to improve the resolution of the frequency dimension. In order to further improve the resolution of the frequency dimension, the author also proposes a learnable frequency position encoding method.

[0004] Although the above methods have improved the voiceprint recognition effect to a certain extent, channel compression-excitation and frequency compression-excitation are used separately. Summary of the Invention

[0005] To overcome the above-mentioned technical problems, the present invention proposes a method for voiceprint recognition, which uses a voiceprint recognition model trained by the following technical solutions to perform voiceprint recognition:

[0006] S1. Obtain a training set, where the training set includes multiple audio data;

[0007] S2. Use the training set to train the voiceprint recognition model. The voiceprint recognition model is implemented based on a neural network and includes a processing module that integrates at least two squeeze-and-excitation sub-modules;

[0008] The step S2 includes:

[0009] S21. Extract the audio features of the audio data included in the training set;

[0010] S22. Perform a slicing operation on the audio features to obtain multiple audio slice features of the same length;

[0011] S23. Randomly obtain a fixed number of the audio slice features each time, input them into the voiceprint recognition model for training, and perform iterative training multiple times to obtain a trained voiceprint recognition model.

[0012] Furthermore, the neural network is a residual network, and the processing module is an SEfwSE module that integrates a channel squeeze-and-excitation SE sub-module and a frequency squeeze-and-excitation fwSE sub-module; the SE sub-module is used to compress the time dimension and frequency dimension of the audio slice features and stimulate the channel dimension of the audio slice features; the fwSE sub-module is used to compress the time dimension and channel dimension of the audio slice features and stimulate the frequency dimension of the audio slice features.

[0013] Furthermore, the compression function Fsq and the excitation function Fex of the channel squeeze-and-excitation SE sub-module are respectively:

[0014]

[0015] F ex (z, W) = σ(g(z, W)) = σ(W 2 δ(W 1 z));

[0016] where x c is the audio slice feature, T is the number of frames of the audio slice feature, F is the dimension of the audio slice feature, i is a positive integer, j is a positive integer, z is a two-dimensional tensor, W represents a linear transformation matrix, δ represents the ReLU activation function, σ represents the sigmoid function, g represents an intermediate function, and W 1 represents the first linear layer that compresses the number of channels to reduce the computational amount, and W2 The second linear layer that restores the number of compressed channels to the size before compression.

[0017] Furthermore, the formulas of the compression function Fsq and the excitation function Fex of the frequency compression-excitation fwSE sub-module are respectively:

[0018]

[0019] F ex (z, W) = σ(g(z, W)) = σ(W 2 δ(W 1 z));

[0020] Where, x F is the audio slice feature, T is the number of frames of the audio slice feature, C is the number of channels of the frequency compression-excitation fwSE sub-module, i is a positive integer, j is a positive integer, z is a two-dimensional tensor, W represents a linear transformation matrix, δ represents the ReLU activation function, σ represents the sigmoid function, g represents an intermediate function, W1 represents the third linear layer that compresses the frequency dimension to reduce the computational amount, and W2 represents the fourth linear layer that restores the compressed frequency dimension to the size before compression.

[0021] Furthermore, the training set also includes speaker information. The number of the audio data included in the same speaker in the training set is not less than 8, and the duration of each piece of the audio data is not less than 2 seconds.

[0022] Furthermore, the audio features are extracted in a frame-by-frame manner. The frame length is 25 milliseconds, the frame shift is 10 milliseconds, and the audio features are 80-dimensional FBank features.

[0023] Furthermore, the slice length of the slicing operation is 200 frames, the length of the overlapping part of the slices is 20 frames, and the fixed quantity is 16.

[0024] Furthermore, the number of layers of the residual network is 34 layers. The convolutional layer of the residual network uses two-dimensional convolution, and the residual network also includes an attention mechanism layer.

[0025] Furthermore, before the step S21, it also includes performing a data augmentation operation on at least part of the audio data in the training set, and adding the audio data after the data augmentation to the training set. The data augmentation operation includes at least one of the following: adding noise or adding reverberation.

[0026] Furthermore, the data augmentation operation is performed in an online manner, and the audio features are extracted in an online manner.

[0027] Further, before the step S23, it further includes performing cepstral mean normalization (CMN) operation on the audio slice features.

[0028] The present invention also provides a voiceprint recognition device, which stores computer instructions; the computer instructions cause the voiceprint recognition device to execute the voiceprint recognition method as described in any one of the above.

[0029] The present invention also provides a computer-readable storage medium, which stores computer instructions, and the computer instructions cause the computer to execute the voiceprint recognition method as described in any one of the above.

[0030] The beneficial effects brought by the technical solution provided by the present invention are:

[0031] The voiceprint recognition method and device according to the embodiments of the present invention can excite the channel dimension and frequency dimension of voiceprint recognition, and add the excitation results. On the premise that the increased computational complexity is negligible, it can improve the resolution of both the channel dimension and frequency dimension, and enhance the voiceprint recognition effect. In a further aspect of the present invention, two-dimensional convolution is used in the voiceprint recognition process, which has a good effect on the application scenario of cross-lingual voiceprint recognition; online feature extraction and data augmentation are adopted. On the one hand, it reduces the demand for the hard disk, and on the other hand, it also performs data augmentation more flexibly; attention mechanism is used for pooling, which not only converts the frame-level features of different lengths of input into fixed-length segment-level embedding features, but also makes the features more robust to noise by using the attention mechanism. Description of the Drawings

[0032] Figure 1 and Figure 2 is a flowchart for training a voiceprint recognition model in the voiceprint recognition method according to the embodiments of the present invention;

[0033] Figure 3 is a schematic diagram of the basic component module of a ResNet34SEfwSE model according to the embodiments of the present invention;

[0034] Figure 4 is a schematic diagram of a SEfwSE module according to the embodiments of the present invention;

[0035] Figure 5 is a schematic diagram of the structure of a voiceprint recognition device according to the embodiments of the present invention. Detailed Embodiments

[0036] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention, and are not used to limit the present invention.

[0037] Example 1:

[0038] As Figure 1 shown in the flowchart of training a voiceprint recognition model in the voiceprint recognition method of the embodiment of the present invention, the specific implementation steps of the method are shown, including:

[0039] S1. Obtain a training set, where the training set includes multiple pieces of audio data;

[0040] S2. Use the training set to train the voiceprint recognition model. The voiceprint recognition model is implemented based on a neural network and includes a processing module that integrates at least two squeeze-and-excitation sub-modules.

[0041] Specifically, as Figure 2 shown, step S2 includes:

[0042] S21. Extract the audio features of the audio data included in the training set;

[0043] S22. Perform a slicing operation on the audio features to obtain multiple audio slice features of the same length;

[0044] S23. Randomly obtain a fixed number of the audio slice features each time, input them into the voiceprint recognition model for training, and perform iterative training multiple times to obtain a trained voiceprint recognition model.

[0045] Specifically, the neural network is a residual network, and the processing module is an SEfwSE module that integrates a channel squeeze-and-excitation SE sub-module and a frequency squeeze-and-excitation fwSE sub-module; the SE sub-module is used to compress the time dimension and frequency dimension of the audio slice features and stimulate the channel dimension of the audio slice features; the fwSE sub-module is used to compress the time dimension and channel dimension of the audio slice features and stimulate the frequency dimension of the audio slice features.

[0046] Specifically, the compression function Fsq and excitation function Fex of the channel squeeze-and-excitation SE sub-module are respectively:

[0047]

[0048] F ex (z, W) = σ(g(z, W)) = σ(W 2 δ(W 1 z));

[0049] where x cFor the audio slice feature, T is the number of frames of the audio slice feature, F is the dimension of the audio slice feature, i is a positive integer, j is a positive integer, z is a two-dimensional tensor, W represents a linear transformation matrix, δ represents the ReLU activation function, σ represents the sigmoid function, g represents an intermediate function, W 1 represents the first linear layer that compresses the number of channels to reduce the computational amount, W 2 represents the second linear layer that restores the compressed number of channels to the size before compression.

[0050] Specifically, the formulas of the compression function Fsq and the excitation function Fex of the frequency compression-excitation fwSE sub-module are respectively:

[0051]

[0052] F ex (z, W) = σ(g(z, W)) = σ(W 2 δ(W 1 z));

[0053] where, x F is the audio slice feature, T is the number of frames of the audio slice feature, C is the number of channels of the frequency compression-excitation fwSE sub-module, i is a positive integer, j is a positive integer, z is a two-dimensional tensor, W represents a linear transformation matrix, δ represents the ReLU activation function, σ represents the sigmoid function, g represents an intermediate function, W1 represents the third linear layer that compresses the frequency dimension to reduce the computational amount, and W2 represents the fourth linear layer that restores the compressed frequency dimension to the size before compression.

[0054] Specifically, the training set also includes speaker information. The number of the audio data contained in the same speaker in the training set is not less than 8, and the duration of each piece of audio data is not less than 2 seconds.

[0055] Specifically, the audio features are extracted in a frame-by-frame manner. The frame length is 25 milliseconds, the frame shift is 10 milliseconds, and the audio features are 80-dimensional FBank features.

[0056] Specifically, the slice length of the slicing operation is 200 frames, the length of the overlapping part of the slices is 20 frames, and the fixed number is 16.

[0057] Specifically, the number of layers of the residual network is 34 layers. The convolutional layer of the residual network uses two-dimensional convolution, and the residual network also includes an attention mechanism layer.

[0058] Specifically, before the step S21, it further includes performing data augmentation operations on at least part of the audio data in the training set, adding the augmented audio data to the training set, and the data augmentation operations include at least one of the following: adding noise or adding reverberation.

[0059] Specifically, the data augmentation operations are performed in an online manner, and the audio features are extracted in an online manner.

[0060] Specifically, before the step S23, it further includes performing cepstral mean normalization (CMN) operations on the audio slice features.

[0061] Embodiment 2:

[0062] The present invention proposes a method for speaker recognition based on channel squeeze-and-excitation and frequency squeeze-and-excitation, and names the model ResNet34SEfwSE model. The ResNet34SEfwSE model includes an SEfwSE module that fuses channel squeeze-and-excitation and frequency squeeze-and-excitation. The method includes the following steps:

[0063] 1) Collect the speech audio of speakers according to the usage scope and scenarios to construct a training set, ensure that the number of speech audio of each speaker is not less than 8, and the duration of a single speech audio is not less than 2 seconds;

[0064] 2) Input the training set into the ResNet34SEfwSE model for training, and save the model files obtained from each epoch of training to obtain a trained ResNet34SEfwSE model;

[0065] Among them, step 2 includes the following steps:

[0066] 21) Perform random data augmentation operations such as adding noise and adding reverberation to the speech audio in an online manner;

[0067] 22) Extract 80-dimensional FBank features of the speech audio in an online manner, with a frame length of 25 ms and a frame shift of 10 ms;

[0068] 23) Perform overlapping slicing on the 80-dimensional FBank features, slice them into chunks with a length of 200 frames, with an overlapping length of 20 frames, and perform cepstral mean normalization (CMN) operations;

[0069] 24) Randomly take a batch of chunks each time and input them into the ResNet34SEfwSE model for training, where the batch size of a single graphics card is set to 16.

[0070] Among them, in step 1, the number of voice audio files of each speaker is not less than 8, and the duration of a single voice audio file is not less than 2 seconds, which is to ensure that each speaker has a long enough audio file, and at least one chunk can be obtained from a single audio file, so as to extract reliable embeddings. In step 21, online data augmentation operations such as randomly adding noise and reverberation to the voice audio files are adopted, which can increase data diversity and improve anti-interference ability. In step 23, the 80-dimensional FBank features are sliced with overlap, cut into chunks with a length of 200 frames, the overlapping part has a length of 20 frames, and the CMN (cepstral mean normalization) operation is performed. On the one hand, it is convenient for batch training, and on the other hand, it increases the training data of each speaker, so that when randomly obtaining data in each batch, as many speakers as possible can be covered as much as possible.

[0071] It should be noted that the number of voice audio files of each speaker not less than 8, the duration of a single voice audio file not less than 2 seconds, the 80-dimensional FBank features, the frame length of 25 ms, the frame shift of 10 ms, the chunk with a length of 200 frames, the length of the overlapping part of 20 frames, and the batch size of a single graphics card set to 16 are only the preferred parameters used in the embodiments of the present invention and do not limit the present invention. In other embodiments, adjustments can be made according to the actual application scenario.

[0072] As Figure 3 The figure shows a schematic diagram of the basic component module of a ResNet34SEfwSE model according to an embodiment of the present invention, showing the basic component module of the ResNet34SEfwSE model. Among them, weight represents a 3x3 convolution, BN represents two-dimensional Batch Normalization, ReLU represents a non-linear activation function, and SEfwSE represents a module that fuses channel squeeze-excitation and frequency squeeze-excitation.

[0073] To make the network structure more suitable for the needs of voiceprint recognition, partial adjustments are made to the original ResNet34 network. First, for the first convolutional layer, the kernel size is changed to 3 and the stride is changed to 1. The max-pooling layer in the second layer is removed. The stride of the first block is changed to 1. The final global average pooling is changed to pooling based on the attention mechanism, and the other parts remain the same. The input format of the ResNet34 network is converted into a four-dimensional tensor. The first dimension is the batch size, the second dimension is the number of channels, the third dimension is the dimension of the features, and the fourth dimension is the number of frames. The two-dimensional convolution of the ResNet34 network is performed on the third and fourth dimensions. Specifically, four blocks are connected after the first convolutional layer. The first block includes 3 basic building blocks with 64 channels. The second block includes 4 basic building blocks with 128 channels. The third block includes 6 basic building blocks with 256 channels. The fourth block includes 3 basic building blocks with 512 channels. After the fourth block, pooling based on the attention mechanism is connected, and finally it is input into the classifier for classification. Since the number of channels does not match between different blocks, a 1x1 convolution is required to adapt the channel dimension between the last basic building block of the previous block and the first basic building block of the next block.

[0074] As Figure 4 shown in the following figure is a schematic diagram of an SEfwSE module according to an embodiment of the present invention, showing the composition structure of the SEfwSE module, which includes a channel squeeze-and-excitation SE sub-module and a frequency squeeze-and-excitation fwSE sub-module.

[0075] The SE sub-module is as shown in the upper half of the appendix Figure 4 First, the time dimension and the frequency dimension are squeezed into 1x1, only the channel dimension is retained (Fsq function), and the channel information is statistically calculated as a vector. In actual use, to reduce the computational complexity, the dimension of the vector is usually set to a fraction of the number of channels, such as 1 / 4 or 1 / 8. After excitation (Fex function), the excited vector is expanded to be the same as the dimension of the input, and then added to the input element by element to obtain the feature map after squeeze-and-excitation.

[0076] The formulas for the squeeze function (Fsq) and the excitation function (Fex) are as follows:

[0077]

[0078] F ex (z, W) = σ(g(z, W)) = σ(W 2 δ(W 1 z));

[0079] where, xc For the audio slice feature, T is the number of frames of the audio slice feature, F is the dimension of the audio slice feature, i is a positive integer, j is a positive integer, z is a two-dimensional tensor, W represents a linear transformation matrix, δ represents the ReLU activation function, σ represents the sigmoid function, g represents an intermediate function, W 1 represents the first linear layer that compresses the number of channels to reduce the computational amount, W 2 represents the second linear layer that restores the compressed number of channels to the size before compression.

[0080] The fwSE sub-module is as shown in the lower part of the appendix Figure 4 First, the position encoding is initialized as a vector with a length equal to the frequency dimension, and this vector is extended to be consistent with the input dimension and then added to the input, and is updated together with the network parameters through backpropagation. Similar to the processing process of the SE sub-module, except that the time dimension and the channel dimension are compressed, and the frequency dimension is excited.

[0081] It can be seen that the dimensions are exactly the same before and after compression-excitation. Therefore, this module can be embedded in any part of the ResNet34 network and is very convenient to use.

[0082] Embodiment 3:

[0083] The present invention also provides a voiceprint recognition device, as Figure 5 shown. The device includes a processor 501, a memory 502, a bus 503, and a computer program stored in the memory 502 and executable on the processor 501. The processor 501 includes one or more processing cores. The memory 502 is connected to the processor 501 through the bus 503. The memory 502 is used to store program instructions. When the processor executes the computer program, the steps in the above method embodiments of the present invention are implemented.

[0084] Further, as an executable solution, the voiceprint recognition device may be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The system / electronic device may include, but is not limited to, a processor and a memory. Those skilled in the art can understand that the above composition structure of the system / electronic device is only an example of the system / electronic device and does not constitute a limitation on the system / electronic device. It may include more or fewer components than the above, or combine some components, or different components. For example, the system / electronic device may further include input / output devices, network access devices, a bus, etc. The embodiments of the present invention do not make limitations in this regard.

[0085] Further, as an executable solution, the so-called processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The processor is the control center of the system / electronic device, and connects various parts of the entire system / electronic device through various interfaces and circuits.

[0086] The memory can be used to store computer programs and / or modules. The processor realizes various functions of the system / electronic device by running or executing the computer programs and / or modules stored in the memory, and by calling the data stored in the memory. The memory mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the mobile phone, etc. In addition, the memory may include high-speed random access memory, and may also include non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices.

[0087] Embodiment 4:

[0088] The present invention also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the method in the above embodiments of the present invention are realized.

[0089] If the modules / units integrated in the system / electronic device are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-described embodiment methods of the present invention, it can also be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-described various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice within the jurisdiction.

[0090] Although the present invention has been specifically shown and described in conjunction with the preferred embodiments, those skilled in the art should understand that various changes can be made to the present invention in terms of form and details without departing from the spirit and scope of the present invention defined by the appended claims. All such changes are within the protection scope of the present invention.

Claims

1. A method for voiceprint recognition, characterized in that, a voiceprint recognition model trained through the following steps is used for voiceprint recognition: S1. Obtain a training set, where the training set includes multiple audio data; S2. Use the training set to train the voiceprint recognition model. The voiceprint recognition model is implemented based on a neural network and includes a processing module that integrates at least two squeeze-and-excitation sub-modules; The step S2 includes: S21. Extract the audio features of the audio data included in the training set; S22. Perform a slicing operation on the audio features to obtain multiple audio slice features of the same length; S23. Randomly obtain a fixed number of the audio slice features each time, input them into the voiceprint recognition model for training, and iterate multiple times to obtain a trained voiceprint recognition model; wherein, the neural network is a residual network, and the processing module is an SEfwSE module that integrates a channel squeeze-and-excitation SE sub-module and a frequency squeeze-and-excitation fwSE sub-module; the SE sub-module is used to squeeze the time dimension and frequency dimension of the audio slice features and excite the channel dimension of the audio slice features; the fwSE sub-module is used to squeeze the time dimension and channel dimension of the audio slice features and excite the frequency dimension of the audio slice features; wherein, the formulas of the compression function Fsq and the excitation function Fex of the channel squeeze-and-excitation SE sub-module are respectively: F ex (z, W) = σ(g(z, W)) = σ(W 2 δ(W 1 z)); where x c is the audio slice feature, T is the number of frames of the audio slice feature, F is the dimension of the audio slice feature, i is a positive integer, j is a positive integer, z is a two-dimensional tensor, W represents a linear transformation matrix, δ represents the ReLU activation function, σ represents the sigmoid function, g represents an intermediate function, W 1 represents the first linear layer that compresses the number of channels to reduce the computational amount, W 2 represents the second linear layer that restores the compressed number of channels to the size before compression; wherein, the formulas of the compression function Fsq and the excitation function Fex of the frequency squeeze-and-excitation fwSE sub-module are respectively: F ex (z, W) = σ(g(z, W)) = σ(W 2 δ(W 1 z)); where x F is the audio slice feature, T is the number of frames of the audio slice feature, C is the number of channels of the frequency compression-excitation fwSE sub-module, i is a positive integer, j is a positive integer, z is a two-dimensional tensor, W represents a linear transformation matrix, δ represents the ReLU activation function, σ represents the sigmoid function, g represents an intermediate function, W1 represents the third linear layer that compresses the frequency dimension to reduce the computational amount, and W2 represents the fourth linear layer that restores the compressed frequency dimension to the size before compression.

2. The method according to claim 1, characterized in that, the training set further includes speaker information. The number of the audio data included for the same speaker in the training set is not less than 8, and the duration of each audio data is not less than 2 seconds.

3. The method according to claim 1, characterized in that, the audio features are extracted in a frame-by-frame manner, the frame length is 25 milliseconds, the frame shift is 10 milliseconds, and the audio features are 80-dimensional FBank features.

4. The method according to claim 1, characterized in that, the slicing length of the slicing operation is 200 frames, the overlapping length of the slices is 20 frames, and the fixed number is 16.

5. The method according to claim 1, characterized in that, the number of layers of the residual network is 34 layers. The convolutional layer of the residual network uses two-dimensional convolution, and the residual network further includes an attention mechanism layer.

6. The method according to claim 1, characterized in that, before the step S21, it further includes performing a data augmentation operation on at least part of the audio data in the training set, adding the augmented audio data to the training set. The data augmentation operation includes at least one of the following: adding noise or adding reverberation.

7. The method according to claim 6, characterized in that, the data augmentation operation is performed in an online manner, and the audio features are extracted in an online manner.

8. The method according to claim 1, characterized in that, before the step S23, it further includes performing a cepstral mean normalization CMN operation on the audio slice features.

9. An apparatus for voiceprint recognition, characterized in that, it includes a memory and a processor, and the memory stores at least one program, and the at least one program is executed by the processor to implement the voiceprint recognition method according to any one of claims 1 to 8.

10. A computer-readable storage medium, characterized in that, the storage medium stores at least one program, and the at least one program is executed by a processor to implement the voiceprint recognition method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Voice spoofing attack detection method based on voice signal spectrum characteristics and deep learning

    CN112201255A

  • Optical remote sensing image ground object classification method based on double-path sparse hierarchical network

    CN112464732A

  • Data identification method and device, electronic equipment and storage medium

    CN113408539A