Speech processing methods and devices

By ensuring consistency between the speaker characteristics of the reference speech and the speech characteristics of the mixed speech during speech processing, and by using an attention matrix for feature fusion, the problems of high computational cost and low extraction efficiency are solved, achieving efficient and accurate extraction of the target speaker's speech.

CN114400007BActive Publication Date: 2025-10-28LENOVO (BEIJING) LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111674684.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-31
Publication Date
2025-10-28
Estimated Expiration
2041-12-31

AI Technical Summary

Technical Problem

Existing technologies involve high computational demands in target speaker speech extraction scenarios, resulting in low extraction efficiency and failing to effectively distinguish between the speech of the target speaker and that of non-target speakers.

Method used

By obtaining the first voiceprint features of the reference speech and processing the consistency of the speech features of the mixed speech to be processed, feature fusion is performed using an attention matrix to avoid feature dimensionality amplification. Furthermore, the attention matrix is ​​used to enhance the speech of the target speaker and suppress the speech of non-target speakers.

Benefits of technology

It reduces computational load, improves the efficiency and accuracy of target speaker speech extraction, reduces computational resource consumption, and enhances the extraction effect of target speaker speech.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114400007B_ABST
    Figure CN114400007B_ABST
Patent Text Reader

Abstract

This application discloses a speech processing method and apparatus. The method involves acquiring reference speech (speech of a target object), processing the reference speech based on the speech features of the mixed speech to be processed, obtaining a first voiceprint feature whose feature dimension is consistent with that of the mixed speech to be processed, and performing feature fusion processing on the speech features of the mixed speech to be processed based on the first voiceprint feature of the reference speech, obtaining a fused feature whose dimension is also consistent with that of the mixed speech to be processed, thus avoiding feature dimension amplification during feature fusion processing. Finally, based on the fused feature whose dimension is consistent with that of the mixed speech to be processed, the speech of the target object in the mixed speech to be processed is extracted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of speech processing technology, and in particular relates to a speech processing method and apparatus. Background Technology

[0002] In the scenario of extracting the speech of the target speaker, the current feature fusion method is relatively simple. Generally, the concat method is used to fuse the speech features of the mixed speech with the voiceprint features of the target speaker. The purpose is to extract the speech of the target speaker from the mixed speech by referring to the voiceprint features of the target speaker.

[0003] However, the aforementioned processing methods increase the computational load for speech extraction from the target speaker, resulting in low efficiency in speech extraction. Summary of the Invention

[0004] Therefore, this application discloses the following technical solution:

[0005] A speech processing method, comprising:

[0006] Obtain reference speech, wherein the reference speech is the speech of the target object;

[0007] Based on the speech features of the mixed speech to be processed, the reference speech is processed to obtain the first voiceprint feature of the reference speech, so that the dimension of the first voiceprint feature of the reference speech is consistent with the dimension of the speech features of the mixed speech to be processed.

[0008] Based on the first voiceprint feature of the reference speech, the speech features of the mixed speech to be processed are subjected to feature fusion processing to obtain fused features. The dimension of the fused features is consistent with the dimension of the speech features of the mixed speech to be processed.

[0009] Based on the fusion features, the speech of the target object in the mixed speech to be processed is extracted.

[0010] Optionally, the step of performing feature fusion processing on the speech features of the mixed speech to be processed based on the first voiceprint features of the reference speech to obtain fused features includes:

[0011] The feature vector of the first voiceprint feature of the reference speech is subjected to matrix transformation processing to obtain an attention matrix, wherein the number of rows and columns of the attention matrix are the dimensions of the first voiceprint feature, respectively.

[0012] The attention matrix is ​​used to perform feature processing on the speech features of the mixed speech to be processed, and the fused features are obtained.

[0013] Optionally, the step of performing matrix transformation processing on the feature vector of the first voiceprint feature of the reference speech to obtain the attention matrix includes:

[0014] The feature vector of the first voiceprint feature is subjected to a first convolution process to obtain a first vector matrix;

[0015] The feature vector of the first voiceprint feature is subjected to a second convolution to obtain a second vector matrix, and the second vector matrix is ​​transposed to obtain a third vector matrix.

[0016] The first vector matrix and the third vector matrix are processed, and the processing results are normalized to obtain the attention matrix;

[0017] Wherein, the number of rows of the first vector matrix and the second vector matrix are respectively the speech feature dimension of the mixed speech to be processed, and the number of columns are 1, or the number of rows of the first vector matrix and the second vector matrix are respectively 1, and the number of columns are respectively the speech feature dimension of the mixed speech to be processed.

[0018] Optionally, the step of using the attention matrix to perform feature processing on the speech features of the mixed speech to be processed to obtain the fused features includes:

[0019] Using each row vector of the attention matrix, the speech features of each frame of speech in at least a portion of the mixed speech to be processed are processed to obtain the fusion features corresponding to each frame of speech in at least a portion of the mixed speech.

[0020] Based on the fusion features corresponding to at least some of the speech frames, the fusion features corresponding to the mixed speech to be processed are obtained.

[0021] Optionally, the step of using the attention matrix to perform feature processing on the speech features of the mixed speech to be processed to obtain the fused features includes:

[0022] Using each row vector of the attention matrix, the speech features of each frame of speech in at least a portion of the mixed speech to be processed are processed to obtain a vector processing result corresponding to the number of rows of the attention matrix for each frame of speech in at least a portion of the speech. Each vector processing result corresponding to each frame of speech is used as the fusion feature of each dimension corresponding to each frame of speech, so that the dimension of the fusion feature corresponding to each frame of speech is consistent with the feature dimension of the mixed speech to be processed.

[0023] Based on the fusion features corresponding to at least some of the speech frames, the fusion features corresponding to the mixed speech to be processed are obtained.

[0024] Optionally, extracting the speech of the target object from the mixed speech to be processed based on the fusion features includes:

[0025] Extract the speech features of the target object from the fused features;

[0026] The speech features of the target object are subjected to speech conversion processing to obtain the speech of the target object.

[0027] Optionally, before performing feature fusion processing on the speech features of the mixed speech to be processed based on the first voiceprint feature of the reference speech, the method further includes:

[0028] Extract the voiceprint features of the mixed speech to be processed to obtain the second voiceprint features of the mixed speech to be processed;

[0029] Determine whether the first voiceprint feature and the second voiceprint feature meet the preset similarity conditions;

[0030] If not, skip the steps of performing feature fusion processing on the speech features of the mixed speech to be processed and extracting the speech of the target object in the mixed speech to be processed;

[0031] If so, the step of performing feature fusion processing on the speech features of the mixed speech to be processed and extracting the speech of the target object in the mixed speech to be processed is triggered.

[0032] Optionally, determining whether the first voiceprint feature and the second voiceprint feature satisfy a preset similarity condition includes:

[0033] Using a preset distance algorithm, the feature distance between the first voiceprint feature and the second voiceprint feature is determined, and it is determined whether the feature distance is within a preset distance range;

[0034] If they are within the preset distance range, then the first voiceprint feature and the second voiceprint feature satisfy the similarity condition;

[0035] If they are not within the preset distance range, then the first voiceprint feature and the second voiceprint feature do not satisfy the similarity condition.

[0036] Optionally, the method further includes:

[0037] Based on the fact that the first voiceprint feature and the second voiceprint feature do not meet the similarity condition, quiet speech that meets the volume condition is output as the target speech extraction result of the mixed speech to be processed.

[0038] A voice processing device, comprising:

[0039] The acquisition module is used to acquire reference speech, wherein the reference speech is the speech of the target object;

[0040] The processing module is used to process the reference speech based on the speech features of the mixed speech to be processed, to obtain the first voiceprint feature of the reference speech, so that the dimension of the first voiceprint feature of the reference speech is consistent with the dimension of the speech features of the mixed speech to be processed.

[0041] The feature fusion module is used to perform feature fusion processing on the speech features of the mixed speech to be processed based on the first voiceprint feature of the reference speech, so as to obtain the fused features after fusion processing. The dimension of the fused features is consistent with the dimension of the speech features of the mixed speech to be processed.

[0042] An extraction module is used to extract the speech of the target object in the mixed speech to be processed based on the fusion features.

[0043] As can be seen from the above scheme, the speech processing method and apparatus disclosed in this application acquires reference speech (speech of the target object), processes the reference speech based on the speech features of the mixed speech to be processed, obtains a first voiceprint feature whose feature dimension is consistent with the speech feature dimension of the mixed speech to be processed, and performs feature fusion processing on the speech features of the mixed speech to be processed based on the first voiceprint feature of the reference speech, to obtain a fused feature whose dimension is also consistent with the speech feature dimension of the mixed speech to be processed, thus avoiding the amplification of feature dimension in the feature fusion processing, and finally extracts the speech of the target object in the mixed speech to be processed based on the fused feature whose dimension is consistent with the speech feature dimension of the mixed speech to be processed. Attached Figure Description

[0044] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0045] Figure 1 This is a schematic diagram of the traditional concat feature fusion process;

[0046] Figure 2 This is a flowchart of a speech processing method provided in this application;

[0047] Figure 3 This is another flowchart of the speech processing method provided in this application;

[0048] Figure 4 This is a schematic diagram of the feature fusion method based on the attention matrix provided in this application;

[0049] Figure 5 This is yet another flowchart of the speech processing method provided in this application;

[0050] Figure 6(a) is a network structure diagram of a traditional speech extraction network;

[0051] Figure 6(b) is a network structure diagram of the speech extraction network implemented based on this application;

[0052] Figure 7 This is another flowchart of the speech processing method provided in this application;

[0053] Figure 8 This is a structural diagram of the speech processing device provided in this application;

[0054] Figure 9 This is a structural diagram of the electronic device provided in this application. Detailed Implementation

[0055] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0056] See Figure 1 This paper presents the processing procedure of the traditional concat feature fusion method. In this method, the target speaker's voiceprint feature vector is first copied along the temporal sequence of the mixed speech to be processed. The number of copies corresponds to the number of speech frames in the mixed speech, ensuring a one-to-one correspondence between each frame of the mixed speech and the target speaker's voiceprint feature vector. Then, frame-by-frame, the speech features of each frame of the mixed speech are concatted with their corresponding voiceprint features to achieve feature fusion. However, this feature fusion method introduces feature dimensionality amplification; for example, for a 100-dimensional feature fusion... The mixed speech features are 256 (i.e., the speech features of the mixed speech to be processed), where 100 is the number of speech frames in the mixed speech and 256 is the feature dimension of each speech frame in the mixed speech. Assuming that the target speaker's voiceprint features with a feature dimension of 256 are concatenated with it, a fusion feature of 100*512 can be obtained. Compared with the mixed speech features of 100*256, the feature dimension increases after fusion. The amplification of the feature dimension during the feature fusion process will introduce a large number of parameters and computational load for the subsequent target speaker speech extraction processing, thus resulting in low efficiency of target speaker speech extraction.

[0057] To at least address the aforementioned problems, embodiments of this application disclose a speech processing method, apparatus, and electronic device for more efficiently extracting target speaker speech from mixed speech to be processed in a target speaker speech extraction scenario. This speech processing method can be applied to electronic devices, which may be, but are not limited to, devices in a variety of general-purpose or dedicated computing environments or configurations, such as personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor devices, etc.

[0058] See Figure 2 The speech processing method disclosed in this application includes at least the following processing steps:

[0059] Step 201: Obtain the reference audio.

[0060] The reference speech is the speech of the target object, which is the target speaker.

[0061] Figure 2 Subsequent processing steps will use the reference speech as a basis to extract the target speech, i.e., the speech of the target speaker, from the mixed speech to be processed. The mixed speech to be processed is the combined speech of multiple speakers, or the combined sound of multiple speakers and their environmental noise.

[0062] Optionally, this step may specifically acquire the target speaker's clean speech, that is, the target speaker's speech after noise has been filtered out using noise reduction technology, as a reference speech.

[0063] Step 202: Based on the speech features of the mixed speech to be processed, process the reference speech to obtain the first voiceprint feature of the reference speech, so that the dimension of the first voiceprint feature of the reference speech is consistent with the dimension of the speech features of the mixed speech to be processed.

[0064] The speech features of the mixed speech to be processed can be, but are not limited to, frequency domain features or frequency-like features of the mixed speech to be processed.

[0065] Taking frequency-domain features as an example, the time-domain features of the mixed speech to be processed, such as amplitude, zero-crossing rate (ZCR), and linear predictive cepstral coefficient (LPCC), can be analyzed. The corresponding time-frequency domain feature transformation methods can be simulated to transform the time-domain features of the mixed speech to be processed into frequency-domain features. For example, the STFT (Short-Time Fourier Transform) can be simulated to process the time-domain features of the mixed speech to be processed, generating corresponding spectral features similar to STFT, which can then be used as frequency-domain features of the mixed speech to be processed.

[0066] The speech features of the mixed speech to be processed can be regarded as a vector, and the length of the vector is the feature dimension of the speech features.

[0067] For the acquired reference speech, this embodiment of the application performs voiceprint feature extraction processing on the reference speech based on the speech features of the mixed speech to be processed, so that the dimension of the first voiceprint feature of the processed reference speech is consistent with the dimension of the speech feature of the mixed speech to be processed.

[0068] The following example provides a specific implementation method:

[0069] The temporal features of the reference speech are analyzed, and the temporal features of the reference speech are transformed into frequency-domain-like features by simulating STFT and other methods. Based on this, the voiceprint features of the reference speech are extracted by referring to the speech feature dimensions of the mixed speech to be processed. The frequency-domain-like features of the reference speech are extracted into voiceprint information vectors with feature vector length equal to the speech feature dimensions of the mixed speech to be processed. Thus, the voiceprint features corresponding to the reference speech with the same feature dimensions as the speech feature dimensions of the mixed speech to be processed are obtained, which is the first voiceprint feature.

[0070] Step 203: Based on the first voiceprint feature of the reference speech, perform feature fusion processing on the speech features of the mixed speech to be processed to obtain the fused features. The dimension of the fused features is consistent with the dimension of the speech features of the mixed speech to be processed.

[0071] Specifically, the first voiceprint feature of the reference speech is used to perform feature fusion processing on the speech features of the mixed speech to be processed. This means fusing the first voiceprint feature of the reference speech with the speech features of the mixed speech to be processed, and based on the consistency of the feature dimensions of the first voiceprint feature and the speech features of the mixed speech to be processed, the dimension of the resulting fused feature is consistent with the dimension of the speech features of the mixed speech to be processed.

[0072] See Figure 3 The flowchart of the speech processing method shown in this embodiment demonstrates that, through steps 2031-2032, the first voiceprint feature of the reference speech and the speech features of the mixed speech to be processed are fused.

[0073] Step 2031: Perform matrix transformation on the feature vector of the first voiceprint feature of the reference speech to obtain the attention matrix. The number of rows and columns of the attention matrix are the dimensions of the first voiceprint feature, respectively.

[0074] The process of performing matrix transformation on the feature vector of the first voiceprint feature to obtain the attention matrix can be further implemented as follows:

[0075] 11) Perform a first convolution on the feature vector of the first voiceprint feature of the reference speech to obtain the first vector matrix.

[0076] Specifically, the first convolution processing of the first voiceprint feature can be achieved by performing a 1D (one-dimensional) convolution with a kernel of 1 on the feature vector of the first voiceprint feature, thereby obtaining the first vector matrix.

[0077] like Figure 4 The provided example performs a 1D convolution with a kernel of 1 on the feature vector of size 1×C corresponding to the first voiceprint feature of the reference object (the feature vector size can also be C×1, which is not limited), to obtain a first vector matrix of size 1×C, where C represents the feature vector length of the first voiceprint feature, that is, the feature dimension of the first voiceprint feature.

[0078] 12) Perform a second convolution on the feature vector of the first voiceprint feature of the reference speech to obtain a second vector matrix, and then transpose the second vector matrix to obtain a third vector matrix.

[0079] Similarly, a second convolution process can be performed on the feature vector of the first voiceprint feature by performing a 1D convolution with a kernel of 1, to obtain a second vector matrix, and then the second vector matrix can be transposed to obtain a third vector matrix.

[0080] The second vector matrix obtained after the second convolution process is different from the first vector matrix obtained after the first convolution process. The number of rows in the first vector matrix and the number of columns in the second vector matrix are respectively the speech feature dimension of the mixed speech to be processed, and are respectively 1; or, the number of rows in the first vector matrix and the number of columns in the second vector matrix are respectively 1, and are respectively the speech feature dimension of the mixed speech to be processed.

[0081] like Figure 4 The provided example performs a 1D convolution with a kernel of 1 on the feature vector of size 1×C corresponding to the first voiceprint feature of the reference object, to obtain a second vector matrix of size 1×C, and then transposes the second vector matrix to obtain a third vector matrix of size C×1.

[0082] 13) Process the first and third vector matrices and normalize the processing results to obtain the attention matrix.

[0083] Next, the first vector matrix and the third vector matrix are multiplied to obtain the multiplied matrix, and the multiplied matrix is ​​normalized (softmax) to obtain the attention matrix.

[0084] For example, targeting Figure 4For example, by performing matrix multiplication of the first vector matrix of 1×C with the third vector matrix of C×1, we obtain the C×C multiplication result matrix, and then perform softmax processing on the multiplication result matrix to obtain the C×C attention matrix.

[0085] Step 2032: Using the attention matrix, perform feature processing on the speech features of the mixed speech to be processed to obtain fused features.

[0086] After obtaining the attention matrix, the speech features of the mixed speech to be processed are processed using the attention matrix to obtain the fused features.

[0087] Specifically, each row vector of the attention matrix can be used to process the speech features of each frame of speech in at least a subset of the mixed speech to be processed, thereby obtaining the fusion features corresponding to each frame of speech in the at least a subset of the speech. Based on the fusion features corresponding to the at least a subset of the speech, the fusion features corresponding to the mixed speech to be processed are obtained. (See also...) Figure 4 The fusion feature corresponding to the mixed speech to be processed is the overall fusion feature composed of the fusion features corresponding to each frame of speech in at least some of the frames of speech. This overall fusion feature is also the "fusion feature after fusion processing" mentioned in step 203.

[0088] It should be noted that processing the speech features of each frame of speech in at least a portion of the mixed speech using each row vector of the attention matrix can mean processing the speech features of all frames of the mixed speech in the mixed speech in the mixed speech in the mixed speech in the mixed speech in the mixed speech in the mixed speech in the mixed speech in the mixed speech in the mixed speech in the mixed speech in the mixed speech in the mixed speech in the mixed speech in the mixed speech in the mixed speech in the mixed speech in the mixed speech, for example, processing every 1 frame, etc., without limitation.

[0089] Specifically, by utilizing each row vector of the attention matrix, the speech features of each frame of speech in at least a portion of the mixed speech to be processed are processed, resulting in vector processing results for each frame of speech with a number equal to the number of rows in the attention matrix. Each vector processing result for each frame of speech serves as the fusion feature for each dimension of that frame of speech; that is, the fusion feature dimension for each frame of speech is equal to the number of rows in the attention matrix. Since the number of rows in the attention matrix (in...) Figure 4 In the example, C) essentially represents the speech feature dimension of the mixed speech to be processed. This ensures that the dimension of the fusion feature corresponding to each frame of speech is consistent with the speech feature dimension of the mixed speech to be processed. Based on this, the resulting fusion feature of the mixed speech to be processed also has the same speech feature dimension as the mixed speech to be processed. (See also...) Figure 4 As shown.

[0090] Furthermore, by utilizing each row vector of the attention matrix to process the speech features of the corresponding frame of speech in the mixed speech to be processed, this can be achieved as follows: For the current row vector of the attention matrix, each vector element of the row vector is matched one-to-one with each vector element of the speech feature vector of the corresponding frame of speech according to the order of the elements. Each pair of corresponding elements is multiplied, resulting in a product of the number of elements in the row vector. These product products are then summed, and the sum is the vector processing result corresponding to the current row vector for that frame of speech. The vector processing results corresponding to each row vector of the attention matrix for that frame of speech constitute the fusion features of that frame of speech in various dimensions.

[0091] Step 204: Based on the fusion features, extract the speech of the target object in the mixed speech to be processed.

[0092] Based on this, and using the fusion features after fusion processing, i.e., the fusion features corresponding to the mixed speech to be processed, the speech of the target object in the mixed speech to be processed is extracted. This processing can be specifically implemented as follows:

[0093] 21) Extract the speech features of the target object from the fused features after fusion processing;

[0094] Based on the matrix structure of the attention matrix proposed in this application, and the feature fusion processing performed on the speech features of the mixed speech to be processed using this attention matrix, it is possible to perform target object (target speaker) speech filtering on each frame of the mixed speech to be processed. Specifically, for frames that do not contain the target object's speech, the attention matrix can suppress it; for frames that contain the target object's speech, the attention matrix can enhance its effect. Specifically, in the obtained fused features, the amplitude (sound intensity) corresponding to the target object's speech is increased, while the amplitude (sound intensity) corresponding to the non-target object's speech is decreased. Furthermore, feature fusion does not change the feature dimension of the mixed speech to be processed, maintaining its unchanged feature dimension, and thus avoiding the introduction of a large number of parameters and computational load due to feature dimension expansion.

[0095] To address the aforementioned characteristics, this application embodiment performs noise reduction and denoising processing on the fused features after fusion processing. This filters out the speech features of non-target objects from the fused features as noise, thereby extracting the speech features of the target object from the fused features. Specifically, but not limited to, performing RNN (Recurrent Neural Network)-based convolution processing on the fused features can eliminate noise and obtain the speech features of the target object from the fused features.

[0096] Taking the speech features of the mixed speech to be processed as frequency domain features as an example, the speech features of the target object extracted from the fused features are essentially frequency domain features enhanced based on the attention matrix.

[0097] 22) Perform speech conversion processing on the speech features of the target object to obtain the speech of the target object.

[0098] Finally, the extracted speech features of the target object are converted into time-domain speech, and the playable speech of the target object can be obtained.

[0099] As can be seen from the above scheme, the method of this embodiment processes the reference speech based on the speech features of the mixed speech to be processed, so that the first voiceprint feature of the processed reference speech has the same dimension (same number of dimensions) as the speech features of the mixed speech to be processed. Based on the first voiceprint feature, the speech features of the mixed speech to be processed are subjected to feature fusion processing to further obtain fused features that also have the same dimension as the speech features of the mixed speech to be processed. Thus, the amplification of feature dimensions brought about by feature fusion processing is avoided.

[0100] Compared to traditional techniques that use the concat method for feature fusion, which leads to an increase in feature dimensionality and consequently, a corresponding increase in computational cost, this application maintains the original feature dimensionality of the mixed speech to be processed after feature fusion. This avoids the introduction of a large number of parameters and computational cost due to feature dimensionality amplification. Compared to the concat method, the fused features obtained by this application have a lower dimensionality and less computational cost, reducing the computational burden on the target speech extraction model (such as an RNN model used for noise reduction) and thus improving the speech extraction efficiency of the target object.

[0101] The mixed speech to be processed generally falls into two categories: those containing the target speech and those not containing the target speech. In the case where the target speech is not contained, traditional methods still follow the processing flow for mixed speech containing the target speech, processing the mixed speech without the target speech through the target speech extraction model (such as an extraction network).

[0102] Because the speech extraction model for the target object is complex and computationally intensive, traditional methods will result in a large amount of irrelevant computation and resource consumption when there is no target object speech in the mixed speech. At the same time, the speaker will forcibly introduce the voiceprint information of the reference speech in some frames that do not contain the target speaker's speech, which is equivalent to introducing voiceprint noise. This increases the difficulty of extracting the target speaker's speech, reduces the separation effect between the target speaker's speech and the non-target speaker's speech, and consequently reduces the extraction accuracy of the target speaker's speech.

[0103] To further address the aforementioned technical issues, see [link to relevant documentation]. Figure 5 The provided flowchart of the speech processing method indicates that, prior to step 203, the speech processing method disclosed in this application may further include the following processing:

[0104] Step 501: Extract the voiceprint features of the mixed speech to be processed to obtain the second voiceprint features of the mixed speech to be processed.

[0105] Specifically, the second voiceprint feature of the mixed speech to be processed can be obtained by extracting voiceprint features from the frequency domain features of the mixed speech to be processed.

[0106] Step 502: Determine whether the first voiceprint feature of the reference speech and the second voiceprint feature of the mixed speech to be processed meet the preset similarity conditions. If not, skip the step of performing feature fusion processing on the speech features of the mixed speech to be processed and extracting the speech of the target object in the mixed speech to be processed. If yes, proceed to step 302 to trigger the processing of performing feature fusion processing on the speech features of the mixed speech to be processed and extracting the speech of the target object in the mixed speech to be processed.

[0107] Next, the first voiceprint feature of the reference speech and the second voiceprint feature of the mixed speech to be processed are similar to determine whether they meet the similarity condition, thereby identifying whether the mixed speech to be processed contains the speech of the target object.

[0108] Specifically, a preset distance algorithm can be used to determine the feature distance between the first voiceprint feature and the second voiceprint feature, and to determine whether the feature distance is within a preset distance range. If so, the first voiceprint feature and the second voiceprint feature are determined to meet the similarity condition; if not, the first voiceprint feature and the second voiceprint feature are determined not to meet the similarity condition.

[0109] The preset distance algorithm can be, but is not limited to, the calculation algorithm for cosine distance or Manhattan distance.

[0110] If the first voiceprint feature and the second voiceprint feature satisfy the similarity condition, it indicates that the mixed speech to be processed contains the speech of the target object. In this case, proceed to step 203 to trigger the subsequent processing of the mixed speech to be processed, so as to extract the speech of the target object in the mixed speech to be processed.

[0111] Conversely, if the first and second voiceprint features do not meet the similarity condition, it indicates that the mixed speech to be processed does not contain the speech of the target object. In this case, the subsequent processing such as feature fusion of the mixed speech to be processed and extraction of the target object speech based on the fusion features is skipped, and the extraction result can be directly output, such as outputting empty information.

[0112] Referring to Figures 6(a) and 6(b), Figure 6(a) shows the structure of a traditional speech extraction network, including a weight-sharing feature transformation module, a speaker extraction module, a feature fusion module, a speech segmentation module, and a speech conversion module, wherein:

[0113] The weight-sharing feature transformation module is responsible for receiving the input mixed speech to be processed and the reference speech respectively, and simulating STFT to transform the time-domain features of the mixed speech to be processed and the reference speech into frequency-domain-like features respectively.

[0114] The mixed speech to be processed represents mixed speech containing the target speaker or mixed speech not containing the target speaker, while the reference speech is the extracted clean speech of the target speaker.

[0115] The voiceprint extraction module is responsible for extracting the frequency-domain features of the reference speech into a voiceprint information vector.

[0116] The feature fusion module is responsible for fusing the speaker information vector with the frequency-domain features of the mixed speech to be processed using the concat method.

[0117] The speech segmentation module is responsible for extracting the speech features of the target speaker based on fused features;

[0118] The speech conversion module is responsible for converting the extracted speech features of the target speaker into time-domain speech, resulting in the target speaker's time-domain speech that can be output and played back.

[0119] Unlike traditional speech extraction networks, as shown in Figure 6(b) of the speech extraction network structure implemented in this application, this application further adds a voiceprint discrimination module between the feature fusion module and the voiceprint extraction module of the speech extraction network. This module is used to identify whether the mixed speech to be processed contains the voice of the target speaker by comparing the voiceprint features of the mixed speech to be processed with the voiceprint features of the reference speech. If the target speaker is not included, the subsequent target speaker extraction process of the mixed speech to be processed is skipped, and quiet speech is directly output as the target speaker extraction result, so as to reduce unnecessary computation and resource consumption.

[0120] In addition, this application also improves the feature fusion module in the speech extraction network by changing the feature fusion module from a concat-based feature fusion method to an attention matrix-based feature fusion method.

[0121] In summary, this embodiment determines whether the mixed speech to be processed contains the speech of the target object by performing a similarity judgment on the first voiceprint feature of the reference speech and the second voiceprint feature of the mixed speech to be processed. If the target object is not included, the subsequent processing of the mixed speech to be processed is skipped, which further reduces unnecessary computation and resource consumption in the target speaker speech extraction scenario. In addition, for the case where the target speaker's speech is not included, the voiceprint information of the reference speech is not forcibly introduced, avoiding the increase in extraction difficulty when the target object's speech is extracted due to the introduction of voiceprint noise. Furthermore, based on the proposed attention matrix, the target object's speech is filtered in each frame of the mixed speech to be processed that contains the target object's speech. The feature information of the target object is enhanced and the feature information of non-target objects is suppressed, which greatly enhances the extraction effect of the target object's speech and improves the extraction accuracy.

[0122] In one embodiment, see Figure 7 The provided flowchart of the speech processing method, and the speech processing method disclosed in this application, may further include the following processing:

[0123] Step 701: Based on the determination result that the first voiceprint feature and the second voiceprint feature do not meet the similarity condition, output the quiet speech that meets the volume condition as the speech extraction result of the target object of the mixed speech to be processed.

[0124] Specifically, if the first voiceprint feature and the second voiceprint feature do not meet the similarity condition, it indicates that the mixed speech to be processed does not contain the speech of the target object. In this case, in this embodiment, while skipping the feature fusion of the mixed speech to be processed and the extraction of the target object speech based on the fused features, quiet speech is output as the target object speech extraction result of the mixed speech to be processed.

[0125] The output quiet speech volume meets the volume condition; more specifically, the quiet speech volume is lower than the set volume threshold to avoid including noise information different from the target speech in the output result when the mixed speech to be processed does not contain the target speech.

[0126] Corresponding to the above method, this application also discloses a voice processing device, the structure of which is as follows: Figure 8 As shown, it specifically includes:

[0127] The acquisition module 801 is used to acquire reference speech, which is the speech of the target object;

[0128] The processing module 802 is used to process the reference speech based on the speech features of the mixed speech to be processed, so as to obtain the first voiceprint feature of the reference speech, so that the dimension of the first voiceprint feature of the reference speech is consistent with the dimension of the speech features of the mixed speech to be processed.

[0129] The feature fusion module 803 is used to perform feature fusion processing on the speech features of the mixed speech to be processed based on the first voiceprint features of the reference speech, so as to obtain the fused features after fusion processing. The dimension of the fused features is consistent with the dimension of the speech features of the mixed speech to be processed.

[0130] Extraction module 804 is used to extract the speech of the target object in the mixed speech to be processed based on the above-mentioned fusion features.

[0131] In one embodiment, the feature fusion module 803 is specifically used for:

[0132] The feature vector of the first voiceprint feature of the reference speech is subjected to matrix transformation to obtain the attention matrix. The number of rows and columns of the attention matrix are the dimensions of the first voiceprint feature, respectively.

[0133] The attention matrix is ​​used to process the speech features of the mixed speech to obtain fused features.

[0134] In one embodiment, the feature fusion module 803, when performing matrix transformation processing on the feature vector of the first voiceprint feature of the reference speech, is specifically used for:

[0135] The feature vector of the first voiceprint feature is subjected to a first convolution process to obtain the first vector matrix;

[0136] The feature vector of the first voiceprint feature is subjected to a second convolution to obtain a second vector matrix, and the second vector matrix is ​​transposed to obtain a third vector matrix.

[0137] The first and third vector matrices are processed, and the processing results are normalized to obtain the attention matrix;

[0138] Wherein, the number of rows of the first vector matrix and the second vector matrix are respectively the speech feature dimension of the mixed speech to be processed, and the number of columns are 1, or the number of rows of the first vector matrix and the second vector matrix are respectively 1, and the number of columns are respectively the speech feature dimension of the mixed speech to be processed.

[0139] In one embodiment, the feature fusion module 803, when using the attention matrix to perform feature processing on the speech features of the mixed speech to be processed to obtain fused features, specifically performs the following:

[0140] Using each row vector of the attention matrix, the speech features of each frame of speech in at least some of the mixed speech to be processed are processed to obtain the fusion features corresponding to each frame of speech in the aforementioned at least some of the mixed speech.

[0141] Based on the fusion features corresponding to at least some of the above-mentioned frames of speech, the fusion features corresponding to the mixed speech to be processed are obtained.

[0142] In one embodiment, the feature fusion module 803, when using the attention matrix to perform feature processing on the speech features of the mixed speech to be processed to obtain fused features, specifically performs the following:

[0143] Using each row vector of the attention matrix, the speech features of each frame of speech in at least some of the mixed speech to be processed are processed to obtain the vector processing results corresponding to the number of rows of the attention matrix for each frame of speech in the at least some of the speech. Each vector processing result corresponding to each frame of speech is used as the fusion feature of each dimension corresponding to each frame of speech, so that the dimension of the fusion feature corresponding to each frame of speech is consistent with the feature dimension of the mixed speech to be processed.

[0144] Based on the fusion features corresponding to at least some of the above-mentioned frames of speech, the fusion features corresponding to the mixed speech to be processed are obtained.

[0145] In one embodiment, the extraction module 804 is specifically used for:

[0146] Extract the speech features of the target object from the above fusion features;

[0147] The speech features of the target object are processed by speech conversion to obtain the speech of the target object.

[0148] In one embodiment, the above-described apparatus further includes:

[0149] The voiceprint discrimination module is used to: extract the voiceprint features of the mixed speech to be processed to obtain the second voiceprint features of the mixed speech to be processed; determine whether the first voiceprint features and the second voiceprint features meet the preset similarity conditions; if not, skip the step of performing feature fusion processing on the voice features of the mixed speech to be processed and extracting the voice of the target object in the mixed speech to be processed; if yes, trigger the step of performing feature fusion processing on the voice features of the mixed speech to be processed and extracting the voice of the target object in the mixed speech to be processed.

[0150] In one embodiment, the voiceprint discrimination module, when determining whether the first voiceprint feature and the second voiceprint feature meet preset similarity conditions, is specifically used for:

[0151] Using a preset distance algorithm, the feature distance between the first voiceprint feature and the second voiceprint feature is determined, and it is determined whether the feature distance is within a preset distance range. If it is within the preset distance range, the first voiceprint feature and the second voiceprint feature satisfy the similarity condition. If they are not within the preset distance range, the first voiceprint feature and the second voiceprint feature do not satisfy the similarity condition.

[0152] In one embodiment, the above-mentioned apparatus further includes:

[0153] The output module is used to output quiet speech that meets the volume condition as the speech extraction result of the target object of the mixed speech to be processed, based on the fact that the first voiceprint feature and the second voiceprint feature do not meet the similarity condition.

[0154] The speech processing apparatus disclosed in this application is described simply because it corresponds to the speech processing methods disclosed in the above method embodiments. For any similarities, please refer to the descriptions of the corresponding method embodiments above, which will not be detailed here.

[0155] This application also discloses an electronic device, which may be, but is not limited to, a device in a variety of general or special computing device environments or configurations, such as: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor devices, etc.

[0156] The electronic device is composed of the following structure: Figure 9 As shown, including:

[0157] Memory 901 is used to store the computer instruction set;

[0158] The computer instruction set in memory 901 can be implemented in the form of a computer program.

[0159] Processor 902 is used to implement the speech processing method disclosed in the above method embodiments by executing a computer instruction set.

[0160] The processor 902 can be a central processing unit (CPU), an application-specific integrated circuit (ASIC), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices.

[0161] In addition to these components, electronic devices may also include communication interfaces, communication buses, and other parts. Memory, processor, and communication interface communicate with each other through the communication bus.

[0162] Communication interfaces are used for communication between electronic devices and other devices. Communication buses can be Peripheral Component Interconnect (PCI) buses or Extended Industry Standard Architecture (EISA) buses, and can be categorized into address buses, data buses, control buses, etc.

[0163] It should be noted that the various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same and similar parts between the various embodiments can be referenced to each other.

[0164] For ease of description, the above systems or devices are described separately as various modules or units based on their functions. Of course, in implementing this application, the functions of each unit can be implemented in one or more software and / or hardware components.

[0165] As can be seen from the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware platforms. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of this application.

[0166] Finally, it should be noted that in this document, relational terms such as first, second, third, and fourth are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0167] The above is only a preferred embodiment of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.

Claims

1. A speech processing method, comprising: Obtain reference speech, wherein the reference speech is the speech of the target object; Based on the speech features of the mixed speech to be processed, the reference speech is processed to obtain the first voiceprint feature of the reference speech, so that the dimension of the first voiceprint feature of the reference speech is consistent with the dimension of the speech features of the mixed speech to be processed. The feature vector of the first voiceprint feature of the reference speech is subjected to matrix transformation processing to obtain an attention matrix, wherein the number of rows and columns of the attention matrix are the dimensions of the first voiceprint feature, respectively. Using the attention matrix, the speech features of the mixed speech to be processed are processed to obtain fused features, wherein the dimension of the fused features is consistent with the dimension of the speech features of the mixed speech to be processed. Based on the fusion features, the speech of the target object in the mixed speech to be processed is extracted.

2. The method according to claim 1, wherein performing matrix transformation processing on the feature vector of the first voiceprint feature of the reference speech to obtain the attention matrix includes: The feature vector of the first voiceprint feature is subjected to a first convolution process to obtain a first vector matrix; The feature vector of the first voiceprint feature is subjected to a second convolution to obtain a second vector matrix, and the second vector matrix is ​​transposed to obtain a third vector matrix. The first vector matrix and the third vector matrix are processed, and the processing results are normalized to obtain the attention matrix; Wherein, the number of rows of the first vector matrix and the second vector matrix are respectively the speech feature dimension of the mixed speech to be processed, and the number of columns are 1, or the number of rows of the first vector matrix and the second vector matrix are respectively 1, and the number of columns are respectively the speech feature dimension of the mixed speech to be processed.

3. The method according to claim 1, wherein the step of using the attention matrix to perform feature processing on the speech features of the mixed speech to be processed to obtain the fused features includes: Using each row vector of the attention matrix, the speech features of each frame of speech in at least a portion of the mixed speech to be processed are processed to obtain the fusion features corresponding to each frame of speech in at least a portion of the mixed speech. Based on the fusion features corresponding to at least some of the speech frames, the fusion features corresponding to the mixed speech to be processed are obtained.

4. The method according to claim 1, wherein the step of using the attention matrix to perform feature processing on the speech features of the mixed speech to be processed to obtain the fused features includes: Using each row vector of the attention matrix, the speech features of each frame of speech in at least a portion of the mixed speech to be processed are processed to obtain a vector processing result corresponding to the number of rows of the attention matrix for each frame of speech in at least a portion of the speech. Each vector processing result corresponding to each frame of speech is used as the fusion feature of each dimension corresponding to each frame of speech, so that the dimension of the fusion feature corresponding to each frame of speech is consistent with the feature dimension of the mixed speech to be processed. Based on the fusion features corresponding to at least some of the speech frames, the fusion features corresponding to the mixed speech to be processed are obtained.

5. The method according to claim 1, wherein extracting the speech of the target object in the mixed speech to be processed based on the fusion features comprises: Extract the speech features of the target object from the fused features; The speech features of the target object are subjected to speech conversion processing to obtain the speech of the target object.

6. The method according to claim 1, further comprising: Extract the voiceprint features of the mixed speech to be processed to obtain the second voiceprint features of the mixed speech to be processed; Determine whether the first voiceprint feature and the second voiceprint feature meet the preset similarity conditions; If not, skip the steps of performing feature fusion processing on the speech features of the mixed speech to be processed and extracting the speech of the target object in the mixed speech to be processed; If so, the step of performing feature fusion processing on the speech features of the mixed speech to be processed and extracting the speech of the target object in the mixed speech to be processed is triggered.

7. The method according to claim 6, wherein determining whether the first voiceprint feature and the second voiceprint feature satisfy a preset similarity condition includes: Using a preset distance algorithm, the feature distance between the first voiceprint feature and the second voiceprint feature is determined, and it is determined whether the feature distance is within a preset distance range; If they are within the preset distance range, then the first voiceprint feature and the second voiceprint feature satisfy the similarity condition; If they are not within the preset distance range, then the first voiceprint feature and the second voiceprint feature do not satisfy the similarity condition.

8. The method according to claim 6, further comprising: Based on the fact that the first voiceprint feature and the second voiceprint feature do not meet the similarity condition, quiet speech that meets the volume condition is output as the target speech extraction result of the mixed speech to be processed.

9. A voice processing device, comprising: The acquisition module is used to acquire reference speech, wherein the reference speech is the speech of the target object; The processing module is used to process the reference speech based on the speech features of the mixed speech to be processed, to obtain the first voiceprint feature of the reference speech, so that the dimension of the first voiceprint feature of the reference speech is consistent with the dimension of the speech features of the mixed speech to be processed. The feature fusion module is used to perform matrix transformation processing on the feature vector of the first voiceprint feature of the reference speech to obtain an attention matrix, and to use the attention matrix to perform feature processing on the speech features of the mixed speech to be processed to obtain fused features. The number of rows and columns of the attention matrix are respectively the dimensions of the first voiceprint feature, and the dimensions of the fused feature are consistent with the dimensions of the speech feature of the mixed speech to be processed. An extraction module is used to extract the speech of the target object in the mixed speech to be processed based on the fusion features.

Citation Information

Patent Citations

  • Target voice extraction method, device and equipment, medium and joint training method

    CN111179911A