Voiceprint Feature Extraction Method, Speaker Recognition Method, Model Training Method and Device
Through the voiceprint extraction model with dense connection delay network and multiple two-dimensional convolution processing, the problem of insufficient accuracy of voiceprint feature extraction is solved, and higher accuracy voiceprint feature extraction and speaker recognition are achieved.
Patent Information
- Application Number
- CN202310157038.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-20
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2043-02-20
AI Technical Summary
The accuracy of voiceprint feature extraction in the prior art has insufficient impact on the effectiveness of application scenarios such as security and information security.
The voiceprint extraction model of densely connected delay network is adopted, combining multiple two-dimensional convolution processing and pooling processing of different particle sizes to extract spectrum features, fuse context information of different scales, and improve the accuracy of voiceprint features.
It improves the accuracy of vocalprint feature extraction, reduces the number of parameters and calculations of the model, and at the same time enhances the local translation invariance of the frequency spectrum, and improves the accuracy of speaker recognition.
Smart Images

Figure CN116246636B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical fields of artificial intelligence and audio processing, and particularly to a method for extracting voiceprint features, a speaker recognition method, a model training method, and an apparatus therefor. Background Art
[0002] Voiceprint features refer to the voice features contained in speech that can represent the identity of a speaker. Due to multiple factors such as a person's physiology and personality, each person will have their own different pronunciation characteristics during the speaking process. Therefore, voiceprint feature extraction technology is widely applied in scenarios such as security, information security, and multi-person conversations. Regardless of the scenario used, improving the accuracy of voiceprint feature extraction is crucial. Summary of the Invention
[0003] In view of this, this application provides a method for extracting voiceprint features, a speaker recognition method, a model training method, and an apparatus therefor, so as to improve the accuracy of voiceprint feature extraction.
[0004] This application provides the following solutions:
[0005] In a first aspect, a method for extracting voiceprint features is provided. The method includes:
[0006] Obtain an audio segment containing speech;
[0007] Extract the spectral features of the audio segment;
[0008] Input the spectral features of the audio segment into a voiceprint extraction model, and obtain the voiceprint features output by the voiceprint extraction model; wherein, the voiceprint extraction model includes a first convolutional processing layer and a pooling layer;
[0009] The first convolutional processing layer includes one or more serially connected first convolutional processing modules, and each first convolutional processing module includes a plurality of serially connected basic modules;
[0010] The basic module performs dimensionality reduction processing on the features input to the basic module to obtain first features; performs pooling processing on the first features with at least one granularity, and performs convolutional processing on the result of the pooling processing to obtain second features; and performs convolutional processing on the first features to obtain third features; fuses the second features and the third features to obtain the features output by the basic module;
[0011] The pooling layer performs pooling processing on the features output by the first convolutional processing layer to obtain the voiceprint features.
[0012] According to an implementable manner in an embodiment of this application, the multiple serially connected basic modules are densely connected.
[0013] According to an implementable manner in an embodiment of the present application, the convolution processing includes: performing one-dimensional convolution processing using a time-delay neural network.
[0014] According to an implementable manner in an embodiment of the present application, the pooling processing of the at least one granularity includes at least one of global average pooling, segment average pooling, or max pooling.
[0015] According to an implementable manner in an embodiment of the present application, the voiceprint extraction model further includes: a second convolution processing layer, and the second convolution processing layer includes more than one cascaded second convolution processing modules;
[0016] The second convolution processing layer outputs the obtained features to the first convolution processing layer after performing two-dimensional convolution processing on the spectral features of the audio segment via the more than one cascaded second convolution processing modules.
[0017] In a second aspect, a speaker recognition method is provided, and the method includes:
[0018] Segmenting an audio containing speech into multiple audio segments;
[0019] Using the method described in the first aspect above to extract the voiceprint features of each audio segment for the multiple audio segments respectively;
[0020] Determining the similarity between the voiceprint features of adjacent audio segments;
[0021] Determining whether the corresponding adjacent audio segments belong to the same speaker according to the similarity.
[0022] In a third aspect, a method for training a voiceprint extraction model is provided, and the method includes:
[0023] Obtaining training data including a plurality of training samples, where the training samples include audio segment samples and their corresponding speaker labels;
[0024] Using the training data to train a voiceprint extraction model and a classification model, where after extracting the spectral features of the audio segment samples, inputting them into the voiceprint extraction model, the voiceprint extraction model obtains voiceprint features using the spectral features and outputs them to the classification model, and the classification model performs classification using the voiceprint features to obtain speaker information; the objective of the training includes: minimizing the difference between the speaker information obtained by the classification model and the corresponding speaker labels;
[0025] After the training is completed, obtaining the trained voiceprint extraction model; where the voiceprint extraction model includes a first convolution processing layer and a pooling layer;
[0026] The first convolutional processing layer includes one or more serially connected first convolutional processing modules, and each first convolutional processing module includes a plurality of serially connected basic modules;
[0027] Each basic module performs dimensionality reduction processing on the features input to the basic module to obtain first features; performs pooling processing of at least one granularity on the first features, and performs convolutional processing on the result of the pooling processing to obtain second features; and performs convolutional processing on the first features to obtain third features; and fuses the second features and the third features to obtain the features output by the basic module;
[0028] The pooling layer performs pooling processing on the features output by the first convolutional processing layer to obtain voiceprint features.
[0029] According to one implementable manner in the embodiments of the present application, the pooling processing of at least one granularity includes at least one of global pooling, average pooling, or max pooling.
[0030] According to one implementable manner in the embodiments of the present application, the voiceprint extraction model further includes: a second convolutional processing layer, and the second convolutional processing layer includes one or more serially connected second convolutional processing modules;
[0031] The second convolutional processing layer outputs the obtained features to the first convolutional processing layer after performing two-dimensional convolutional processing on the spectral features of the audio segment through the one or more serially connected second convolutional processing modules.
[0032] In a fourth aspect, a voiceprint feature extraction device is provided, and the device includes:
[0033] An audio acquisition unit configured to acquire an audio segment containing speech;
[0034] A spectral extraction unit configured to extract the spectral features of the audio segment;
[0035] A voiceprint extraction unit configured to input the spectral features of the audio segment into a voiceprint extraction model to obtain the voiceprint features output by the voiceprint extraction model; wherein, the voiceprint extraction model includes a first convolutional processing layer and a pooling layer;
[0036] The first convolutional processing layer includes one or more serially connected first convolutional processing modules, and each first convolutional processing module includes a plurality of serially connected basic modules;
[0037] The basic module performs dimensionality reduction on the features input to the basic module to obtain first features; performs pooling processing on the first features with at least one granularity, and performs convolution processing on the result of the pooling processing to obtain second features; and performs convolution processing on the first features to obtain third features; performs fusion processing on the second features and the third features to obtain the features output by the basic module.
[0038] The pooling layer performs pooling processing on the features output by the first convolution processing layer to obtain the voiceprint features.
[0039] In a fifth aspect, a speaker recognition device is provided. The device includes: an audio segmentation unit, a similarity calculation unit, and a voiceprint feature extraction device as described in the fourth aspect above.
[0040] The audio segmentation unit is configured to segment an audio containing speech into a plurality of audio segments.
[0041] The voiceprint feature extraction device is configured to extract the voiceprint features of each audio segment for the plurality of audio segments respectively.
[0042] The similarity calculation unit is configured to determine the similarity between the voiceprint features of adjacent audio segments, and determine whether the corresponding adjacent audio segments belong to the same speaker according to the similarity.
[0043] In a sixth aspect, a device for training a voiceprint extraction model is provided. The device includes:
[0044] A sample acquisition unit is configured to acquire training data including a plurality of training samples. The training samples include audio segment samples and their corresponding speaker labels.
[0045] A model training unit is configured to train a voiceprint extraction model and a classification model using the training data. After extracting the spectral features of the audio segment samples and inputting them into the voiceprint extraction model, the voiceprint extraction model uses the spectral features to obtain voiceprint features and outputs them to the classification model. The classification model uses the voiceprint features to perform classification to obtain speaker information. The objective of the training includes: minimizing the difference between the speaker information obtained by the classification model and the corresponding speaker labels. After the training is completed, the trained voiceprint extraction model is obtained. Among them, the voiceprint extraction model includes a first convolution processing layer and a pooling layer.
[0046] The first convolution processing layer includes more than one cascaded first convolution processing module, and the first convolution processing module includes a plurality of cascaded basic modules.
[0047] The basic module performs dimensionality reduction on the features input to the basic module to obtain first features; performs pooling processing on the first features at at least one granularity, and performs convolutional processing on the result of the pooling processing to obtain second features; and performs convolutional processing on the first features to obtain third features; fuses the second features and the third features to obtain the features output by the basic module.
[0048] The pooling layer performs pooling processing on the features output by the first convolutional processing layer to obtain voiceprint features.
[0049] According to a seventh aspect, there is provided a computer-readable storage medium having stored thereon a computer program, which when executed by a processor implements the steps of the method described in any one of the first, second, and third aspects above.
[0050] According to an eighth aspect, there is provided an electronic device, including:
[0051] One or more processors; and
[0052] A memory associated with the one or more processors, the memory being used to store program instructions, and the program instructions, when read and executed by the one or more processors, execute the steps of the method described in any one of the first, second, and third aspects above.
[0053] According to the specific embodiments provided by the present application, the present application discloses the following technical effects:
[0054] 1) By adopting pooling processing at at least one granularity in the basic module of the voiceprint extraction model and performing convolutional processing on the result of the pooling processing, the present application fuses context information of different scales to promote the ability of the model to extract more discriminative information, so as to improve the accuracy of the extracted voiceprint features.
[0055] 2) By using a densely connected delay network as the backbone network of the voiceprint extraction model, the present application makes full use of the advantages of dense connection and delay network, so that the voiceprint extraction model has better performance while having fewer parameters and less computational complexity.
[0056] 3) In the present application, at the front end of the voiceprint extraction model, the second convolutional processing layer extracts more detailed features from the spectral features through multiple two-dimensional convolutional processings, thereby increasing the frequency local translation invariance of the model to the input spectrum.
[0057] Of course, it is not necessary for any product implementing the present application to achieve all the above-mentioned advantages simultaneously. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required in the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0059] Figure 1 is the system architecture diagram applicable to the embodiments of the present application;
[0060] Figure 2 is the flowchart of the voiceprint feature extraction method provided by the embodiments of the present application;
[0061] Figure 3 is the structural schematic diagram of the voiceprint extraction model provided by the embodiments of the present application;
[0062] Figure 4 is the structural schematic diagram of the basic module provided by the embodiments of the present application;
[0063] Figure 5 is the flowchart of a method for training a voiceprint extraction model provided by the embodiments of the present application;
[0064] Figure 6 is the schematic diagram in the speaker recognition scenario provided by the embodiments of the present application;
[0065] Figure 7 is the schematic block diagram of the voiceprint feature extraction device provided by the embodiments of the present application;
[0066] Figure 8 is the schematic block diagram of the speaker recognition device provided by the embodiments of the present application;
[0067] Figure 9 is the schematic block diagram of the device for training a voiceprint extraction model provided by the embodiments of the present application;
[0068] Figure 10 is the schematic block diagram of the electronic device provided by the embodiments of the present application. Detailed Embodiments
[0069] The following will clearly and completely describe the technical solutions in the embodiments of the present application in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art belong to the scope of protection of the present application.
[0070] The terms used in the embodiments of the present invention are for the purpose of describing specific embodiments only and are not intended to limit the present invention. The singular forms "a", "said", and "the" used in the embodiments of the present invention and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise.
[0071] It should be understood that the term "and / or" used herein is merely a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " herein generally indicates that the associated objects before and after are in an "or" relationship.
[0072] Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining" or "in response to detecting". Similarly, depending on the context, the phrase "if determined" or "if detected (stated condition or event)" may be interpreted as "when determined" or "in response to determining" or "when detected (stated condition or event)" or "in response to detecting (stated condition or event)".
[0073] For the convenience of understanding the present application, the system architecture applicable to the present application is briefly described first. Figure 1 An exemplary system architecture to which the embodiments of the present application can be applied is shown, as Figure 1 shown in the figure, the system architecture may include: a model training device, a voiceprint feature extraction device, and a speaker recognition device.
[0074] Among them, after obtaining the training data composed of multiple training samples in the offline stage, the model training device can train the voiceprint extraction model by using the method provided in the embodiments of the present application.
[0075] The voiceprint feature extraction device can extract the voiceprint features of the input audio segment online by using the method provided in the embodiments of the present application. Based on the extracted voiceprint features, various types of downstream speaker recognition tasks can be performed. For example, the speaker recognition device uses the extracted voiceprint features to identify the speaker information corresponding to the audio segment. Another example is that the speaker recognition device uses the voiceprint features extracted by the voiceprint feature extraction device for multiple audio segments to identify whether the multiple audio segments correspond to the same speaker.
[0076] The model training device, the voiceprint feature extraction device, and the speaker recognition device can be respectively set as independent servers, or can be set in the same server or server group, or can also be set in an independent or the same cloud server. A cloud server, also known as a cloud computing server or a cloud host, is a host product in the cloud computing service system, which is used to solve the defects of large management difficulty and weak service scalability existing in traditional physical hosts and virtual private servers (VPSs, Virtual Private Servers). The model training device, the voiceprint feature extraction device, and the speaker recognition device can also be set in a computer terminal with strong computing power.
[0077] It should be noted that, in addition to extracting voiceprint features and recognizing speakers online, the above-mentioned voiceprint feature extraction device and speaker recognition device can also extract voiceprint features and recognize speakers in an offline manner.
[0078] It should be understood that Figure 1 the numbers of the voiceprint extraction models, model training devices, voiceprint feature extraction devices, and speaker recognition devices in
[0079] Figure 2 is only illustrative. According to the implementation requirements, there can be any number of voiceprint extraction models, model training devices, voiceprint feature extraction devices, and speaker recognition devices. Figure 1 Figure 2
[0080]
[0081]
[0082]
[0083] Step 202: Obtain an audio segment containing speech.
[0081] Step 204: Extract the spectral features of the audio segment.
[0082] Step 206: Input the spectral features of the audio segment into the voiceprint extraction model to obtain the voiceprint features output by the voiceprint extraction model; wherein, the voiceprint extraction model includes a first convolutional processing layer and a pooling layer. The first convolutional processing layer includes more than one cascaded first convolutional processing module, and the first convolutional processing module includes multiple cascaded basic modules; the basic module performs dimensionality reduction processing on the features input to the basic module to obtain first features; performs pooling processing on the first features with at least one granularity, and performs convolutional processing on the result of the pooling processing to obtain second features; and performs convolutional processing on the first features to obtain third features; fuses the second features and the third features to obtain the features output by the basic module; the pooling layer performs pooling processing on the features output by the first convolutional processing layer to obtain the voiceprint features.
[0083] As can be seen from the above process, in this application, at least one granularity of pooling processing is adopted in the basic module of the voiceprint extraction model, and the result of the pooling processing is subjected to convolution processing, so as to fuse context information of different scales to promote the model's ability to extract more discriminative information, thereby improving the accuracy of the extracted voiceprint features.
[0084] It should be noted that the "first", "second", etc. limitations involved in this disclosure do not have limitations in terms of size, order, quantity, etc., and are only used to distinguish by name. For example, the "first convolution processing layer" and the "second convolution processing layer" are used to distinguish two convolution processing layers by name. For another example, "first feature", "second feature", and "third feature", etc. are used to distinguish three feature representations by name.
[0085] The following will describe each step in the above process in detail. The audio segment obtained in step 202 above is usually an audio segment containing the speech of one speaker, so as to extract the voiceprint features of the speaker from the audio segment.
[0086] In step 204 above, the spectral features extracted from the audio segment can be, for example, FBank (Filter Bank) features or MFCC (Mel-scale Frequency Cepstral Coefficients) features, etc.
[0087] Among them, FBank is obtained by performing Fourier transform on each frame after the audio segment is framed. The MFCC features are obtained by performing DCT (Discrete Cosine Transform) on the basis of FBank. Since these spectral features are relatively common features at present, they will not be elaborated here.
[0088] The following will focus on describing step 206 above, that is, "input the spectral features of the audio segment into the voiceprint extraction model to obtain the voiceprint features output by the voiceprint extraction model" in detail in combination with embodiments.
[0089] The embodiment of this application provides a brand-new structure of the voiceprint extraction model, as Figure 3 shown in, the voiceprint extraction model may include a first convolution processing layer and a pooling layer, and may further include a second convolution processing layer.
[0090] Among them, the second convolution processing layer is located at the front end of the first convolution pooling layer, and may include M serially connected second convolution processing modules, where M is a positive integer. The second convolution processing layer is used to perform two-dimensional convolution processing on the spectral features of the audio segment through M serially connected second convolution processing modules, and output the obtained features to the first convolution processing layer. In order to control the amount of calculation, usually the value of M is set to be small, for example, 10 is taken.
[0091] The second convolutional processing module described above can adopt a two-dimensional convolutional module. The input spectral feature is a two-dimensional spectral feature, for example, represented as [F, T], where F represents the spectral feature and T is the duration of the audio segment. First, a dimension can be added to [F, T] to [1, F, T], and then after passing through the second convolutional processing layer, [C, F, T] is obtained. Among them, C is related to the number of channels of the two-dimensional convolutional processing layer. After [C, F, T] is transformed into [CF, T], it is input into the second convolutional processing layer.
[0092] It should be noted that the second convolutional processing layer in the voiceprint extraction model is not necessary. [F, T] can also be directly input into the second convolutional processing layer. The second convolutional processing layer can extract more detailed features from the spectral features through multiple two-dimensional convolutional processes, increasing the frequency local translation invariance of the model to the input spectrum.
[0093] The first convolutional processing layer can include N serially connected first convolutional processing modules, and each first convolutional processing module includes multiple serially connected basic modules, where N is a positive integer. Figure 3 Taking each first convolutional processing module including 4 basic modules as an example, but the actual value is usually set to be larger, and each first convolutional processing module can include different numbers of basic modules. For example, the three first convolutional processing modules include 12, 24, and 16 basic modules respectively.
[0094] As a more preferred way, the above-mentioned multiple basic modules can adopt dense connections on the basis of being serially connected. The so-called dense connection means that the input of each basic module includes the concatenation of the outputs of all previous basic modules. That is to say, feature reuse is achieved by concatenating the features output by all previous basic modules in the channel, so as to reach a deeper number of layers with fewer parameters and computational costs, and thus have a stronger information extraction ability.
[0095] The structure of the basic module can be as Figure 4 shown. First, the features input to this basic module can be dimension-reduced through a convolutional layer to obtain the first feature. For the basic module connected to the second convolutional processing layer, the input feature is [CF, T]. Assuming that the output of the l-th basic module is represented as [D l , T], its input is represented as [D l-1 , T]. The basic module first reduces the dimension of D l-1 , and T remains unchanged.
[0096] Then, perform pooling processing on the first feature at at least one granularity, and perform convolution processing on the result of the pooling processing to obtain a second feature, and perform convolution processing on the first feature to obtain a third feature. The pooling processing may include at least one of global average pooling, segment average pooling, or max pooling.
[0097] For example Figure 4 As shown in. Three calculations can be performed on the first feature in parallel: global average pooling, segment average pooling, and convolution processing through a CNN. Among them, global average pooling refers to performing average pooling on the first feature globally. Segment average pooling refers to performing average pooling on the first feature separately within sub-windows. After global average pooling and segment average pooling, the pooling results are then passed through a CNN and an activation layer to obtain a second feature. The activation layer can adopt, for example, a sigmoid activation function.
[0098] Perform fusion processing on the second feature and the third feature to obtain the feature output by this basic module. The second feature and the third feature can be fused, for example, by multiplication, and the output is [D l ,T]. After being processed by the second convolutional processing layer, the final output is [D,T].
[0099] As one possible implementation, the CNN in the above basic module can perform one-dimensional convolution processing using a TDNN (Time Delay Neural Network). This method of using a densely connected time delay network as the backbone network of the voiceprint extraction model makes full use of the advantages of dense connection and time delay network, enabling the voiceprint extraction model to have better performance while having fewer parameters and less computational complexity.
[0100] Continue to refer to Figure 3 . The pooling layer can be composed of a pooling module and a linear module, which respectively perform pooling processing and linear processing on the features output by the first convolutional processing layer to obtain voiceprint features, denoted as [D].
[0101] Figure 5 The figure is a flowchart of a method for training a voiceprint extraction model provided by an embodiment of the present application. This method can be executed by Figure 1 the model training device in the system shown in. As shown in Figure 5 , this method may include:
[0102] Step 502: Obtain training data including multiple training samples. The training samples include audio segment samples and their corresponding speaker labels.
[0103] When training the voiceprint extraction model in the embodiments of the present application, multiple audio segments each containing only the voice of one speaker can be obtained as audio segment samples, and the audio segment samples can contain the voices of different speakers. Then, speaker labels are respectively marked for each audio segment, for example, the speaker ID is marked.
[0104] Step 504: Train the voiceprint extraction model and the classification model using the training data. After extracting the spectral features of the audio segment samples, input them into the voiceprint extraction model. The voiceprint extraction model uses the spectral features to obtain voiceprint features and outputs them to the classification model. The classification model uses the voiceprint features for classification to obtain speaker information. The training objective includes minimizing the difference between the speaker information obtained by the classification model and the corresponding speaker label.
[0105] In the embodiments of the present application, when training the voiceprint extraction model, a classification task of a speaker is connected downstream. For the structure and processing of the voiceprint extraction model, reference can be made to the relevant records in the embodiments of the voiceprint extraction method for Figure 3 and Figure 4 which will not be elaborated here.
[0106] After extracting the spectral features of the audio segment samples, input them into the voiceprint extraction model. The voiceprint extraction model outputs the voiceprint features to the classification model, and the classification model classifies according to the voiceprint features, that is, maps the voiceprint features to the space corresponding to the speaker ID to obtain the speaker ID.
[0107] A loss function can be constructed according to the above training objective. In each iteration, the values of the loss function are used to update the model parameters (including the parameters of the voiceprint extraction model and the classification model) by means such as gradient descent until the preset training end condition is met. The training end condition can include, for example, that the value of the loss function is less than or equal to the preset loss function threshold, the number of iterations reaches the preset number threshold, etc. The loss function can be, for example, the cross-entropy loss function, etc.
[0108] Step 506: After the training is over, obtain the trained voiceprint extraction model.
[0109] The voiceprint extraction method provided in the above embodiments of the present application can be applied to a variety of application scenarios:
[0110] For example, in the security application scenario, after extracting the voiceprint features from the audio containing the user's voice, compare them with the voiceprint features of the legitimate user to determine whether it is a legitimate user.
[0111] For another example, in audio or video, speaker recognition is performed to determine the speaker transition points therein, that is, to determine the speaker information in a multi-person conversation audio and the corresponding speakers for each time period. Currently, it is widely used in multi-person meeting scenarios, customer service call scenarios, and sales scenarios. Taking the multi-person meeting scenario as an example, the intelligent speech recognition system can more quickly recognize the text information corresponding to the meeting audio and segment it according to different speakers to determine the position of the speaker transition point, which will greatly improve the work efficiency of users. The following describes this application scenario as an example.
[0112] As Figure 6 shown in, first, the audio containing speech is segmented into multiple audio segments. The audio containing speech can be segmented according to a preset duration so that each audio segment is of the preset duration. The specific duration can be set according to actual needs. For example, the audio can be segmented according to a duration of 1.5 seconds to obtain two or more audio segments.
[0113] Then, in the manner provided in the embodiments of the present application, first, the spectral features of each audio segment are extracted respectively for each audio segment, and then the spectral features are input into the voiceprint extraction model, so as to obtain the voiceprint features of each audio segment extracted by the voiceprint extraction model for multiple audio segments respectively.
[0114] Next, the similarity between the voiceprint features is calculated respectively for adjacent audio segments, and based on the similarity, it is determined whether the adjacent audio segments belong to the same speaker. For example, if the similarity between the voiceprint features of adjacent audio segments is greater than or equal to a preset similarity threshold, it is considered that the adjacent audio segments belong to the same speaker; otherwise, it is considered that the adjacent audio segments do not belong to the same speaker, and the segmentation point between the adjacent audio segments can be used as the speaker transition point for annotating the speaker representation for the audio or speech recognition result.
[0115] The above describes specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require the specific order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0116] According to an embodiment of another aspect, a voiceprint feature extraction device is provided. Figure 7 The schematic block diagram of a voiceprint feature extraction device according to an embodiment is shown. As Figure 7 shown, the device 700 includes: an audio acquisition unit 701, a spectral extraction unit 702, and a voiceprint extraction unit 703. The main functions of each component unit are as follows:
[0117] An audio acquisition unit 701, configured to acquire an audio segment containing speech.
[0118] A spectrum extraction unit 702, configured to extract spectral features of the audio segment.
[0119] A voiceprint extraction unit 703, configured to input the spectral features of the audio segment into a voiceprint extraction model and obtain the voiceprint features output by the voiceprint extraction model.
[0120] Wherein, the structure of the voiceprint extraction model can be as Figure 3 shown, mainly including a first convolutional processing layer and a pooling layer, and may further include a second convolutional processing layer.
[0121] The second convolutional processing layer performs two-dimensional convolutional processing on the spectral features of the audio segment through one or more cascaded second convolutional processing modules, and outputs the obtained features to the first convolutional processing layer.
[0122] The first convolutional processing layer includes one or more cascaded first convolutional processing modules, and each first convolutional processing module includes a plurality of cascaded basic modules.
[0123] The basic module performs dimensionality reduction processing on the features input to the basic module to obtain first features; performs pooling processing of at least one granularity on the first features, and performs convolutional processing on the result of the pooling processing to obtain second features; and performs convolutional processing on the first features to obtain third features; fuses the second features and the third features to obtain the features output by the basic module.
[0124] The pooling layer performs pooling processing on the features output by the first convolutional processing layer to obtain voiceprint features.
[0125] As one possible implementation, the multiple cascaded basic modules are densely connected.
[0126] As one possible implementation, the convolutional processing in the basic module can be one-dimensional convolutional processing using a time-delay neural network.
[0127] As one possible implementation, the above-mentioned pooling processing of at least one granularity includes at least one of global average pooling, segment average pooling, or max pooling.
[0128] According to an embodiment of still another aspect, a speaker recognition device is provided. Figure 8 The schematic block diagram of a speaker recognition device according to an embodiment is shown. As Figure 8 shown, the device 800 includes: an audio segmentation unit 801, a voiceprint extraction device 700, and a similarity calculation unit 802. The main functions of each component unit are as follows:
[0129] The audio segmentation unit 801 is configured to segment the audio containing speech into multiple audio segments.
[0130] The voiceprint feature extraction device 700 is configured to extract the voiceprint features of each audio segment respectively for multiple audio segments.
[0131] For the specific structure and functions of the voiceprint feature extraction device 700, reference can be made to the relevant records in the above embodiments for Figure 7 which will not be elaborated here.
[0132] The similarity calculation unit 802 is configured to determine the similarity between the voiceprint features of adjacent audio segments, and determine whether the corresponding adjacent audio segments belong to the same speaker based on the similarity.
[0133] For example, if the similarity between the voiceprint features of adjacent audio segments is greater than or equal to a preset similarity threshold, it is considered that the adjacent audio segments belong to the same speaker; otherwise, it is considered that the adjacent audio segments do not belong to the same speaker, and the segmentation point between the adjacent audio segments can be used as the speaker conversion point for annotating the speaker representation in the audio or speech recognition result.
[0134] According to an embodiment of another aspect, a device for training a voiceprint extraction model is provided. Figure 9 The schematic block diagram of the device for training a voiceprint extraction model according to an embodiment is shown. As Figure 9 shown, the device 900 includes: a sample acquisition unit 901 and a model training unit 902. The main functions of each component unit are as follows:
[0135] The sample acquisition unit 901 is configured to acquire training data including multiple training samples, and the training samples include audio segment samples and their corresponding speaker labels.
[0136] The model training unit 902 is configured to train a voiceprint extraction model and a classification model using the training data. After extracting the spectral features of the audio segment samples and inputting them into the voiceprint extraction model, the voiceprint extraction model uses the spectral features to obtain voiceprint features and outputs them to the classification model, and the classification model uses the voiceprint features for classification to obtain speaker information; the training objective includes: minimizing the difference between the speaker information obtained by the classification model and the corresponding speaker labels; after the training is completed, the trained voiceprint extraction model is obtained;
[0137] Among them, the structure of the voiceprint extraction model can be as Figure 4 shown, including a first convolutional processing layer and a pooling layer, and may further include a second convolutional processing layer.
[0138] The second convolutional processing layer includes one or more cascaded second convolutional processing modules.
[0139] The second convolutional processing layer outputs the obtained features to the first convolutional processing layer after performing two-dimensional convolutional processing on the spectral features of the audio segment through one or more cascaded second convolutional processing modules.
[0140] The first convolutional processing layer includes one or more cascaded first convolutional processing modules, and each first convolutional processing module includes a plurality of cascaded basic modules.
[0141] The basic module performs dimensionality reduction processing on the features input to the basic module to obtain first features; performs pooling processing with at least one granularity on the first features, and performs convolutional processing on the result of the pooling processing to obtain second features; and performs convolutional processing on the first features to obtain third features; fuses the second features and the third features to obtain the features output by the basic module.
[0142] As one preferred implementation manner, the plurality of cascaded basic modules are densely connected.
[0143] As one implementable manner, the above-mentioned convolutional processing performed by the basic module can be one-dimensional convolutional processing using a time-delay neural network.
[0144] As one implementable manner, the above-mentioned pooling processing with at least one granularity includes at least one of global average pooling, segment average pooling, or max pooling.
[0145] The pooling layer performs pooling processing on the features output by the first convolutional processing layer to obtain voiceprint features.
[0146] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the apparatus embodiments, since they are basically similar to the method embodiments, they are described relatively simply. For the related parts, reference can be made to the partial description of the method embodiments. The apparatus embodiments described above are only illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0147] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data that have been authorized by the user or fully authorized by all parties. Moreover, the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or reject.
[0148] In addition, an embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the steps of the method described in any one of the foregoing method embodiments are implemented.
[0149] And an electronic device, including:
[0150] One or more processors; and
[0151] A memory associated with the one or more processors, where the memory is used to store program instructions. When the program instructions are read and executed by the one or more processors, the steps of the method described in any one of the foregoing method embodiments are executed.
[0152] The present application also provides a computer program product, including a computer program. When the computer program is executed by a processor, the steps of the method described in any one of the foregoing method embodiments are implemented.
[0153] Wherein, Figure 10 An exemplary architecture of the electronic device is shown, which may specifically include a processor 1010, a video display adapter 1011, a disk drive 1012, an input / output interface 1013, a network interface 1014, and a memory 1020. The above-mentioned processor 1010, video display adapter 1011, disk drive 1012, input / output interface 1013, network interface 1014, and the memory 1020 can be communicatively connected through a communication bus 1030.
[0154] Wherein, the processor 1010 can be implemented in ways such as a general-purpose CPU, a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solution provided by the present application.
[0155] The memory 1020 can be implemented in the form of ROM (ReadOnlyMemory), RAM (RandomAccessMemory), static storage devices, dynamic storage devices, etc. The memory 1020 can store the operating system 1021 for controlling the operation of the electronic device 1000, and the basic input / output system (BIOS) 1022 for controlling the low-level operations of the electronic device 1000. Additionally, it can also store a web browser 1023, a data storage management system 1024, and a voiceprint extraction device / speaker recognition device / model training device 1025, etc. The above-mentioned voiceprint extraction device / speaker recognition device / model training device 1025 can be the application program that specifically implements the operations of the foregoing steps in the embodiments of the present application. In summary, when implementing the technical solution provided by the present application through software or firmware, the relevant program codes are stored in the memory 1020 and are called and executed by the processor 1010.
[0156] The input / output interface 1013 is used to connect to the input / output module to achieve information input and output. The input / output module can be configured as a component in the device (not shown in the figure) or externally connected to the device to provide corresponding functions. Among them, the input devices can include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output devices can include a display, a speaker, a vibrator, an indicator light, etc.
[0157] The network interface 1014 is used to connect to a communication module (not shown in the figure) to achieve communication interaction between this device and other devices. Among them, the communication module can achieve communication through wired means (such as USB, network cable, etc.) or through wireless means (such as mobile network, WIFI, Bluetooth, etc.).
[0158] The bus 1030 includes a path for transmitting information between various components of the device (such as the processor 1010, the video display adapter 1011, the disk drive 1012, the input / output interface 1013, the network interface 1014, and the memory 1020).
[0159] It should be noted that although the above device only shows the processor 1010, the video display adapter 1011, the disk drive 1012, the input / output interface 1013, the network interface 1014, the memory 1020, the bus 1030, etc., in the specific implementation process, this device may also include other components necessary for normal operation. In addition, those skilled in the art can understand that the above device may also only include the components necessary to implement the solution of the present application and does not necessarily include all the components shown in the figure.
[0160] From the description of the above embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer program product, which can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., including several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of this application.
[0161] The above provides a detailed introduction to the technical solution provided by this application. Specific examples are used in this article to elaborate on the principle and implementation manner of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application; at the same time, for those of ordinary skill in the art, according to the idea of this application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to this application.
Claims
1. A voiceprint feature extraction method, characterized in that, The method includes: Obtaining an audio segment containing speech; Extracting spectral features of the audio segment; Inputting the spectral features of the audio segment into a voiceprint extraction model to obtain voiceprint features output by the voiceprint extraction model; wherein, the voiceprint extraction model includes a first convolutional processing layer and a pooling layer; The first convolutional processing layer includes one or more serially-connected first convolutional processing modules, and each first convolutional processing module includes a plurality of serially-connected basic modules; The basic module performs dimensionality reduction processing on the features input to the basic module to obtain first features; performs pooling processing on the first features with at least one granularity, and performs convolutional processing on the result of the pooling processing to obtain second features; and performs convolutional processing on the first features to obtain third features; performs fusion processing on the second features and the third features to obtain the features output by the basic module; The pooling layer performs pooling processing on the features output by the first convolutional processing layer to obtain the voiceprint features.
2. The method according to claim 1, wherein There is a dense connection between the plurality of serially-connected basic modules.
3. The method according to claim 1, wherein The convolutional processing includes: performing one-dimensional convolutional processing using a time-delay neural network.
4. The method according to claim 1, wherein The at least one granularity of pooling processing includes at least one of global average pooling, segment average pooling, or max pooling.
5. The method according to claim 1, wherein The voiceprint extraction model further includes: a second convolutional processing layer, and the second convolutional processing layer includes one or more serially-connected second convolutional processing modules; The second convolutional processing layer performs two-dimensional convolutional processing on the spectral features of the audio segment through the one or more serially-connected second convolutional processing modules, and outputs the obtained features to the first convolutional processing layer.
6. A speaker recognition method, characterized in that, The method includes: Segmenting an audio containing speech into a plurality of audio segments; Using the method described in claim 1 to extract the voiceprint features of each audio segment for the plurality of audio segments respectively; Determining the similarity between the voiceprint features of adjacent audio segments; Determining whether the corresponding adjacent audio segments belong to the same speaker according to the similarity.
7. A method for training a voiceprint extraction model, characterized in that The method includes: Obtaining training data containing a plurality of training samples, where the training samples include audio segment samples and their corresponding speaker labels; Using the training data to train a voiceprint extraction model and a classification model, wherein after extracting the spectral features of the audio segment samples and inputting them into the voiceprint extraction model, the voiceprint extraction model uses the spectral features to obtain voiceprint features and outputs them to the classification model, and the classification model uses the voiceprint features to perform classification to obtain speaker information; the objective of the training includes: minimizing the difference between the speaker information obtained by the classification model and the corresponding speaker labels; After the training is completed, obtaining the trained voiceprint extraction model; wherein, the voiceprint extraction model includes a first convolutional processing layer and a pooling layer; The first convolutional processing layer includes one or more serially-connected first convolutional processing modules, and each first convolutional processing module includes a plurality of serially-connected basic modules; The basic module performs dimensionality reduction on the features input to the basic module to obtain first features; performs pooling processing on the first features with at least one granularity, and performs convolutional processing on the result of the pooling processing to obtain second features; and performs convolutional processing on the first features to obtain third features; performs fusion processing on the second features and the third features to obtain the features output by the basic module; The pooling layer performs pooling processing on the features output by the first convolutional processing layer to obtain voiceprint features.
8. The method according to claim 7, wherein The at least one granularity of pooling processing includes at least one of global pooling, average pooling, or max pooling.
9. The method according to claim 7, characterized in that The voiceprint extraction model further includes: a second convolutional processing layer, and the second convolutional processing layer includes more than one second convolutional processing module connected in series; The second convolutional processing layer performs two-dimensional convolutional processing on the spectral features of the audio segment through the more than one second convolutional processing module connected in series, and outputs the obtained features to the first convolutional processing layer.
10. A voiceprint feature extraction device, characterized in that, The device includes: An audio acquisition unit configured to acquire an audio segment containing speech; A spectral extraction unit configured to extract the spectral features of the audio segment; A voiceprint extraction unit configured to input the spectral features of the audio segment into a voiceprint extraction model to obtain the voiceprint features output by the voiceprint extraction model; wherein, the voiceprint extraction model includes a first convolutional processing layer and a pooling layer; The first convolutional processing layer includes more than one first convolutional processing module connected in series, and the first convolutional processing module includes a plurality of basic modules connected in series; The basic module performs dimensionality reduction on the features input to the basic module to obtain first features; performs pooling processing on the first features with at least one granularity, and performs convolutional processing on the result of the pooling processing to obtain second features; and performs convolutional processing on the first features to obtain third features; performs fusion processing on the second features and the third features to obtain the features output by the basic module; The pooling layer performs pooling processing on the features output by the first convolutional processing layer to obtain the voiceprint features.
11. A speaker recognition device, characterized in that, The device includes: an audio segmentation unit, a similarity calculation unit, and a voiceprint feature extraction device as described in claim 10; The audio segmentation unit is configured to segment an audio containing speech into a plurality of audio segments; The voiceprint feature extraction device is configured to extract the voiceprint features of each audio segment respectively for the plurality of audio segments; The similarity calculation unit is configured to determine the similarity between the voiceprint features of adjacent audio segments, and determine whether the corresponding adjacent audio segments belong to the same speaker according to the similarity.
12. An apparatus for training a voiceprint extraction model, characterized in that, The device includes: A sample acquisition unit configured to acquire training data including a plurality of training samples, and the training samples include audio segment samples and their corresponding speaker labels; A model training unit configured to train a voiceprint extraction model and a classification model by using the training data, wherein after extracting the spectral features of the audio segment samples, the spectral features are input into the voiceprint extraction model, the voiceprint extraction model uses the spectral features to obtain voiceprint features and outputs them to the classification model, and the classification model uses the voiceprint features to perform classification to obtain speaker information; the objective of the training includes: minimizing the difference between the speaker information obtained by the classification model and the corresponding speaker label; after the training is completed, obtaining the trained voiceprint extraction model; wherein, the voiceprint extraction model includes a first convolutional processing layer and a pooling layer; The first convolutional processing layer includes one or more cascaded first convolutional processing modules, and the first convolutional processing module includes a plurality of cascaded basic modules; The basic module performs dimensionality reduction processing on the features input to the basic module to obtain first features; performs pooling processing on the first features at at least one granularity, and performs convolutional processing on the result of the pooling processing to obtain second features; and performs convolutional processing on the first features to obtain third features; performs fusion processing on the second features and the third features to obtain the features output by the basic module; The pooling layer performs pooling processing on the features output by the first convolutional processing layer to obtain voiceprint features.
13. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 9.
14. An electronic device, characterized in that, Including: One or more processors; And A memory associated with the one or more processors, the memory is used to store program instructions, and when the program instructions are read and executed by the one or more processors, the steps of the method according to any one of claims 1 to 9 are executed.
Citation Information
Patent Citations
Multi-task speech recognition model training method and multi-task speech recognition method
CN112331187A
Speech recognition method and device, equipment and storage medium
CN114333782A