Audio Processing Method and Apparatus, Electronic Device, and Readable Storage Medium

By combining forced alignment model, feature generation model and deep neural network for audio separation, the problem of poor vocal separation effect under multi-user voice overlap is solved, and higher accuracy and feature distinction are achieved.

CN114067793BActive Publication Date: 2025-07-04BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111302400.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-04
Publication Date
2025-07-04
Estimated Expiration
2041-11-04

AI Technical Summary

Technical Problem

In the prior art, when multiple users' voices overlap, the vocal separation effect is poor, making it difficult to accurately separate and analyze the voices of different users.

Method used

By acquiring the initial audio data of multiple sound sources, using forced alignment models and feature generation models for content recognition, obtaining content vectors and time information, and combining deep neural networks for audio separation, cutting audio clips to retain complete content information, and using audio separation models and auxiliary separation models for end-to-end audio separation.

Benefits of technology

It improves the accuracy of vocal separation and the discrimination of overall characteristics, and can accurately separate and retain complete content information when multiple user voices overlap, improving the separation effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114067793B_ABST
    Figure CN114067793B_ABST
Patent Text Reader

Abstract

The present disclosure provides an audio processing method and apparatus, an electronic device, and a readable storage medium, which relate to the technical field of speech processing, and particularly to the fields of artificial intelligence, speech technology, and deep learning. The specific implementation solution is as follows: Obtain the audio to be processed, where the audio to be processed includes: initial audio data collected from multiple sound sources, and the multiple sound sources correspond to multiple objects; perform content recognition on the audio to be processed to obtain a content vector and time information corresponding to the content vector; separate the audio to be processed based on the content vector and the time information to obtain a separation result, where the separation result is used to determine, from the initial audio data, the target audio data corresponding to each of the multiple objects respectively. Through the above implementation solution, the present disclosure achieves the effects of improving the accuracy of the separation result and increasing the distinguishability of the overall features, and solves the problem of poor separation effect of the voice separation method provided in the related art.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of speech processing technologies, and particularly to the fields of artificial intelligence, speech technology, and deep learning. The present disclosure provides an audio processing method, an apparatus, an electronic device, and a readable storage medium. Background Art

[0002] In scenarios such as intelligent customer service, conference discussions, and interview conversations, voices of multiple users are often collected on a single channel. Therefore, it is necessary to separate the human voices in the recorded audio and then perform targeted analysis and processing on the voices of different users. Currently, the collected audio can be separated by an offline human voice separation method. First, the audio is cut into equal-length small segments, and then the number of speakers in the given audio or a threshold is used for separation. However, if the voices of multiple users collected overlap, the separation effect is poor. Summary of the Invention

[0003] The present disclosure provides an audio processing method, an apparatus, an electronic device, and a readable storage medium.

[0004] According to a first aspect of the present disclosure, there is provided an audio processing method, including: obtaining an audio to be processed, where the audio to be processed includes: initial audio data collected from multiple sound sources, and the multiple sound sources correspond to multiple objects; performing content recognition on the audio to be processed to obtain a content vector and time information corresponding to the content vector; and separating the audio to be processed based on the content vector and the time information to obtain a separation result, where the separation result is used to determine, from the initial audio data, target audio data corresponding to each of the multiple objects.

[0005] According to a second aspect of the present disclosure, there is provided an audio processing apparatus, including: an obtaining module, configured to obtain an audio to be processed, where the audio to be processed includes: initial audio data collected from multiple sound sources, and the multiple sound sources correspond to multiple objects; a recognition module, configured to perform content recognition on the audio to be processed to obtain a content vector and time information corresponding to the content vector; and a separation module, configured to separate the audio to be processed based on the content vector and the time information to obtain a separation result, where the separation result is used to determine, from the initial audio data, target audio data corresponding to each of the multiple objects.

[0006] According to a third aspect of the present disclosure, there is provided an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; where the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method determined according to the above.

[0007] According to a fourth aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the method determined as described above.

[0008] According to a fifth aspect of the present disclosure, there is provided a computer program product including a computer program, which when executed by a processor implements the method determined as described above.

[0009] Through the above embodiments of the present disclosure, after obtaining the audio to be processed, content recognition can be performed on the audio to be processed to obtain a content vector and time information, and the audio to be processed can be separated by combining the content vector and time information, achieving the purpose of voice separation. It is easy to notice that since the content vector and time information are combined simultaneously during the voice separation process, the complete content information can be retained in the cut audio segment, making the feature vector corresponding to the audio segment more distinguishable, thereby achieving the effect of improving the accuracy of the separation result and increasing the distinguishability of the overall features, and solving the problem of poor separation effect of the voice separation method provided in the related art.

[0010] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:

[0012] Figure 1 is a flowchart of the audio processing method according to the present disclosure;

[0013] Figure 2 is a schematic diagram of the audio separation model and the auxiliary separation model according to the present disclosure;

[0014] Figure 3 is a schematic diagram of the audio processing device according to the present disclosure;

[0015] Figure 4 is a block diagram of an electronic device for implementing the audio processing method of the embodiments of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0016] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to assist understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, the description of well-known functions and structures is omitted below.

[0017] Currently, common algorithms for voice separation may include TDNN (Time Delay Neural Networks) - xvector (used to extract user feature vectors) + AHC (agglomerative hierarchical clustering). However, this solution is relatively cumbersome and not an end-to-end implementation solution. The training and testing processes may not match, and the separation effect is not ideal for the case where multiple user voices overlap.

[0018] According to an embodiment of the present disclosure, the present disclosure provides an audio processing method. As Figure 1 shown, the method may include the following steps:

[0019] Step S102, obtain the audio to be processed, where the audio to be processed includes: initial audio data collected from multiple sound sources, and the multiple sound sources correspond to multiple objects.

[0020] In some embodiments, in scenarios of multi-person conversations, such as conference discussions, interview conversations, variety shows, etc., the voices of multiple different users can be collected through a microphone to obtain initial audio data from multiple different sound sources.

[0021] Step S104, perform content recognition on the audio to be processed to obtain a content vector and time information corresponding to the content vector.

[0022] The time information in the above steps may refer to the timestamp of the content vector, including but not limited to: the start time and the duration of the content vector.

[0023] In the scenario of multi-person conversations, there is a certain correlation in the context of the same user's speech. To improve the voice separation effect, the audio to be processed can be subjected to content recognition by means of machine learning to identify the text information of the audio to be processed and the timestamps of pronunciation units at different granularities, that is, to identify the start time and duration of the pronunciation unit in the audio to be processed. Here, the different granularities may be phonemes, characters, words, etc., but are not limited thereto. In addition, the text information can be cut according to specific granularities according to the separation accuracy requirements to obtain multiple texts, and further feature extraction is performed on each text to obtain the feature vector of each text, that is, the above-mentioned content vector, which is called content embedding.

[0024] Step S106, separate the audio to be processed based on the content vector and time information to obtain a separation result, where the separation result is used to determine the target audio data corresponding to each of the multiple objects from the initial audio data.

[0025] In some embodiments, the audio to be processed can be cut based on time information to obtain multiple audio segments, and the lengths of different audio segments are different, which is different from the existing uniform cutting. Further, the multiple audio segments are identified by combining the content vector through machine learning, the user corresponding to each audio segment is determined, and finally the audio segments of the same user are aggregated to obtain the target audio data of each user.

[0026] For example, in a meeting discussion scenario, the audio data during the entire meeting can be collected by a sound collection device such as a microphone as the audio to be processed. Since there are multiple participants speaking during the whole process and the speaking time of each participant is not fixed, after the audio to be processed is collected, the timestamps of pronunciation units with different granularities can be determined by identifying the content of the audio to be processed, and then the audio to be processed is segmented according to the identified timestamps to obtain audio segments. At this time, the granularity of the audio segments is the same as that of the content vector. By combining the content vector to identify the audio segments, the speaking audio of each participant can be accurately determined, achieving the purpose of voice separation.

[0027] Through the above steps, after the audio to be processed is obtained, the content of the audio to be processed can be recognized to obtain the content vector and time information, and the audio to be processed is separated by combining the content vector and time information, achieving the purpose of voice separation. It is easy to notice that since the content vector and time information are combined simultaneously during the voice separation process, the complete content information can be retained in the cut audio segments, making the feature vector corresponding to the audio segment more distinguishable, thereby achieving the effect of improving the accuracy of the separation result and increasing the distinguishability of the overall features, and solving the problem of poor separation effect of the voice separation method provided in the related art.

[0028] Optionally, performing content recognition on the audio to be processed to obtain the content vector and time information includes: using a forced alignment model to recognize the audio to be processed to obtain text information and time information; using a feature generation model to extract features from the text information to obtain the content vector.

[0029] The above-mentioned forced alignment model can be a model pre-trained on common model frameworks such as GMM (Gaussian Mixture Model)-HMM (Hidden Markov Model), LSTM (Long Short Term Memory)-CTC (Connectionist Temporal Classification), Chain, CNN (convolutional neural networks) RNN (Recurrent Neural Network)-T, etc., using open-source data such as Aishell or LibriSpeech. The input of this model can be the Mel spectrum of the audio, and the output is the probability of each predicted pronunciation unit. The loss of training is CE (Cross Entropy Loss). After multiple rounds of iterative convergence, a model with stable performance is obtained. In this way, when the Mel spectrum is input during application, the corresponding text information and time information, that is, the timestamps of phonemes, characters, and words, can be obtained.

[0030] The above-mentioned feature generation model can be a common feature extraction model, which is not specifically limited in this disclosure. The input is pronunciation units such as phonemes, characters, and words, and the output is the corresponding feature vector.

[0031] In some embodiments, after obtaining the audio to be processed, the Mel spectrum of the audio to be processed can be extracted. For example, the audio to be processed can be processed by a Mel-scale filter bank to transform it into the corresponding Mel spectrum, but it is not limited thereto. The Mel spectrum is input into the forced alignment model to obtain the corresponding text information and time information, such as phonemes and the corresponding time information. Then, the text information is mapped to a feature vector through the feature generation model, that is, the above-mentioned content vector is obtained.

[0032] Through the above steps, by pre-constructing a forced alignment model and a feature generation model to perform content recognition on the audio to be processed, the efficiency and accuracy of content recognition are improved, and thus the accuracy of voice separation is enhanced.

[0033] Optionally, the content vector includes: the feature vectors of multiple texts with a preset granularity, and the time information includes: the timestamps of the multiple texts. The audio to be processed is separated based on the content vector and the corresponding time information of the content vector, and the separation result includes: the audio to be processed is cut based on the timestamp of each text to obtain multiple target audios; the audio separation model is used to separate the multiple target audios based on the feature vectors of the multiple texts to obtain the separation result.

[0034] The above-mentioned preset granularity can be a phoneme, a character, a word, etc., but is not limited thereto. The text with the preset granularity can be a pronunciation unit in the audio to be processed.

[0035] The above-mentioned audio separation model can be a voice separation module of a deep neural network based on PIT (Permutation Invariant Train). As Figure 2 shown, the model can be composed of multiple layers of BLSTM (Binary Long-Short Term Memory), a Linear mapping layer, and a sigmod activation layer. The loss function for training is BCE (Binary Cross Entropy), and the strategy is PIT. If the training audio contains 2 users, namely A and B, the predicted output result is shown in the output layer in the figure. Since the corresponding relationship between the upper and lower results in the output layer and the two users is unknown, both situations can be calculated, and the result with a smaller Loss is selected as the final loss. It should be noted that if the voices of the 2 users overlap in time, the probabilities corresponding to the two results are relatively high.

[0036] In some embodiments, the audio to be processed can be cut with the same granularity according to the time information to obtain multiple audio segments (i.e., the above-mentioned multiple target audios). That is, the granularity of the input of the audio separation model is equivalent to that of the content embedding, and both can be phonemes, or characters or words. The Mel spectrum is extracted from each audio segment, and then together with the feature vectors of each pronunciation element output by the feature generation model, it is input into the audio separation model for separation to obtain the final voice separation result.

[0037] Through the above steps, by cutting the audio to be processed based on the timestamp of each text, it is ensured that the target audio after cutting retains the complete content, and the pre-trained audio separation model is used to perform voice separation on the multiple target audios and the feature vectors of the multiple texts, achieving the effect of improving the voice separation efficiency and accuracy.

[0038] Optionally, the audio separation model at least includes: a first-layer bidirectional long short-term memory model and a second-layer bidirectional long short-term memory model. Using the audio separation model, multiple target audios are separated based on the feature vectors of multiple texts, and the separation result includes: inputting the multiple target audios into the first-layer bidirectional long short-term memory model for processing to obtain a first output vector; concatenating the first output vector and the feature vectors of the multiple texts to obtain a concatenated vector; inputting the concatenated vector into the second-layer bidirectional long short-term memory model for processing to obtain the separation result.

[0039] In some embodiments, since the audio separation model needs to process the feature vectors of each text and multiple audio segments, therefore, the Mel spectrum can be extracted from each audio segment and then input into the Stacked BLSTM to obtain the feature vector H of the corresponding granularity. Then, the output vector H of the first Stacked BLSTM is concatenated with the feature vector C of the corresponding text to obtain the vector M, and the vector is input into the second Stacked BLSTM as the input of the subsequent module of the audio separation model, as Figure 2 shown.

[0040] Through the above steps, by concatenating the first output vector and the feature vectors of multiple texts, the audio separation model can fully consider the content of multiple texts during the process of recognizing the first output vector, achieving the effect of making the concatenated features more discriminative and improving the accuracy of voice separation.

[0041] Optionally, the method further includes: obtaining training samples, where the training samples include: training audios and the corresponding annotation results of the training audios. The training audios include: audio data collected from multiple training sound sources, and the multiple training sound sources correspond to multiple training objects; performing content recognition on the training audios to obtain the training vectors corresponding to the training audios and the time information corresponding to the training vectors; separating the training audios based on the training vectors and the time information corresponding to the training vectors to obtain a first prediction result, where the first prediction result is used to represent the probability of the training object corresponding to the training vector; processing the annotation results and the first prediction result to obtain a first loss function; adjusting the model parameters of the audio separation model based on the first loss function.

[0042] The above training samples can be a large amount of audio data of multi-person conversations received and contain a small proportion of aliasing (10% - 20%). The annotation results can be the training objects corresponding to different audio segments at a specific granularity obtained through manual annotation.

[0043] It should be noted that in order to improve the quality of training samples, the training samples can be preprocessed, including removing noises, including environmental noises, busy tones, ringback tones, etc., but not limited to this, to obtain high-quality audio. Additionally, in order to ensure that the quantity of training samples meets the training requirements, data augmentation can be performed on the high-quality audio, including time-domain warping, frequency-domain masking, etc., but not limited to this.

[0044] In some embodiments, for the training audio, the Mel spectrum can be extracted and input into the forced alignment model to obtain multiple texts and timestamps. Then, the multiple texts are passed through the feature generation model to obtain the feature vector C. Then, the training audio is cut into segments of the same granularity according to the time information. As Figure 2 shown, the Mel spectrum is extracted from each audio segment and then input into the first Stacked BLSTM to obtain the high-level feature vector H of the corresponding granularity. H is concatenated with C to obtain M. M is input into the second Stacked BLSTM to obtain the corresponding first prediction result, that is, the probability of the corresponding training object. The corresponding first loss function is obtained through the PIT strategy, and the model parameters of the audio separation model are updated based on this loss function, so as to train a high-performance audio separation model.

[0045] Through the above steps, the audio separation model is trained with the training samples to ensure that a high-performance audio separation model is trained, thereby achieving the effect of improving the accuracy of vocal separation.

[0046] Optionally, after processing the annotation result and the first prediction result to obtain the first loss function, the method further includes: obtaining a target vector, where the target vector is the vector input to the second layer of bidirectional long short-term memory model of the audio separation model; using the auxiliary separation model to predict the target vector to obtain the second prediction result corresponding to the target vector, where the second prediction result is used to represent the training object corresponding to the target vector; generating a second loss function based on the annotation result and the second prediction result; obtaining the total loss function based on the first loss function and the second loss function; adjusting the model parameters of the audio separation model based on the total loss function.

[0047] The above-mentioned auxiliary separation model can be composed of a Linear mapping layer, a Tanh activation layer, and a Normalize regularization layer. As Figure 2As shown in the right model in the middle. The input is the concatenated vector after concatenating the output vector of the first Stacked BLSTM and the content embedding, and the output is the feature vector of the corresponding training object. The feature vector here represents the pronunciation features of the training object, such as physiological structures like vocal cords, oral cavity size, nasal cavity, throat, etc. Through this feature vector, similarity comparison can be performed on the audio of the training object for the second time. The loss function for assisting in the training of the separation model can adopt the 2-norm loss function to calculate the error between the output result of the assisting separation model and the labeled result. The calculation formula is as follows:

[0048]

[0049] Among them, J DC represents the second loss function; V represents the output result, V = [v1,..., v T T T represents the number of audio segments after cutting the training audio; L' represents the labeled result, and the dimension is T * 2 c , C represents the number of training objects. Each row in L' is in the one-hot (only one 1 and the others are 0) form. For example, assuming C is 2, then the audio segments after cutting correspond to 4 situations: 0: non-speech, 1: speaker 1, 2: speaker 2, 3: overlapping. If the first audio segment is silent, then the first row of L is [1 0 0 0]. If the second audio segment is speaker 1, then the second row is [0 1 0 0]. If the third audio segment is speaker 2, then the third row is [0 0 3 0]. If the fourth audio segment is speaker 1 and speaker 2, then the fourth row is [0 0 0 1]; F represents the type of norm, and F = 2 represents the 2-norm.

[0050] In some embodiments, as Figure 2 shown, for the training audio, after obtaining the concatenated vector M through the aforementioned steps, while inputting the concatenated vector M into the second Stacked BLSTM, the concatenated vector can be input into the assisting separation model to obtain the corresponding second prediction result, that is, the feature vector of the corresponding training object. Through the 2-norm loss, the corresponding second loss function can be obtained. By performing a weighted sum on the two loss functions, the total loss function can be obtained. The calculation formula is as follows:

[0051] J MULTI =(1 - α)J PIT +αJ DC .

[0052] Among them, J PIT represents the first loss function, J​MULTI Let \(L\) denote the total loss function, and \(\alpha\) denote a hyperparameter used to adjust the weights of the two loss functions, with a preferred value of \(0.4\).

[0053] Furthermore, based on the total loss function, existing optimization algorithms (such as stochastic gradient descent algorithm, least squares method, etc.) are used to update the model parameters of the audio separation model, so that the first Stacked BLSTM can learn the knowledge of the auxiliary separation model, that is, the splicing model contains feature vectors of different training objects, thereby training a high-performance audio separation model.

[0054] Through the above steps, during the training process of the audio separation model, by combining the output results of the auxiliary separation model to calculate the total loss function, it is ensured that the trained audio separation model can learn the pronunciation features of different training objects, so that in the process of voice separation, the audio segments belonging to the same user can be accurately determined, thus achieving the effect of improving the accuracy of voice separation.

[0055] Optionally, adjusting the model parameters of the audio separation model based on the total loss function includes: adjusting the model parameters of the audio separation model using the stochastic gradient descent algorithm based on the total loss function.

[0056] In some embodiments, after calculating the total loss function, the stochastic gradient descent algorithm (SGD) can be used to calculate the gradient of the loss function, and then update the model parameters of the audio separation model, and iterate repeatedly for multiple rounds until convergence. The implementation process of the stochastic gradient descent algorithm is the same as the prior art and will not be elaborated here.

[0057] Through the above steps, the model parameters are adjusted by the stochastic gradient descent algorithm, thereby achieving the effect of reducing the learning time and improving the training efficiency of the audio separation model.

[0058] Based on the above analysis, it can be seen that in the present disclosure, the pronunciation units and corresponding time information are obtained through the forced alignment model, and the content information of the audio can be cut by combining the time information. Different from the traditional equal-length cutting, this can retain the complete content information, and the content information is added on the basis of the features of the user to whom the audio segment belongs, making the overall features more distinguishable; in addition, an end-to-end speaker separation system is constructed based on the order-independent criterion, which supports variable numbers of speakers (within the maximum number of speakers supported by the network) during use, has a simple overall structure, and also has a good separation effect for the case of overlapping speech; in addition, an auxiliary separation model of deep clustering is introduced, and through the double loss function, the accuracy of voice separation is further improved.

[0059] It should be noted that the audio to be processed in this embodiment is not the audio output for a specific user and does not reflect the personal information of a specific user. Moreover, the acquisition, storage, and application of the audio data involved in this embodiment all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0060] According to an embodiment of the present disclosure, the present disclosure also provides an audio processing device, which is used to implement the above-mentioned embodiments and preferred implementation manners, and those that have been described will not be repeated. As used hereinafter, the term "module" can be a combination of software and / or hardware that can achieve a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.

[0061] Figure 3 is a schematic diagram of the audio processing device according to the present disclosure, as Figure 3 shown, the device includes: an acquisition module 32, configured to acquire the audio to be processed, where the audio to be processed includes: initial audio data collected from multiple sound sources, and the multiple sound sources correspond to multiple objects; an identification module 34, configured to perform content identification on the audio to be processed to obtain a content vector and time information corresponding to the content vector; a separation module 36, configured to separate the audio to be processed based on the content vector and time information to obtain a separation result, where the separation result is used to determine the target audio data corresponding to each object among the multiple objects from the initial audio data.

[0062] Optionally, the identification module includes: an identification unit, configured to use a forced alignment model to identify the audio to be processed to obtain text information and time information; an extraction unit, configured to use a feature generation model to extract features from the text information to obtain a content vector.

[0063] Optionally, the content vector includes: feature vectors of multiple texts with a preset granularity, and the time information includes: timestamps of the multiple texts. The separation module includes: a cutting unit, configured to cut the audio to be processed based on the timestamp of each text to obtain multiple target audios; a separation unit, configured to use an audio separation model to separate the multiple target audios based on the feature vectors of the multiple texts to obtain a separation result.

[0064] Optionally, the audio separation model at least includes: a first-layer bidirectional long short-term memory model and a second-layer bidirectional long short-term memory model. The separation unit is further configured to: input the multiple target audios into the first-layer bidirectional long short-term memory model for processing to obtain a first output vector; splice the first output vector and the feature vectors of the multiple texts to obtain a spliced vector; input the spliced vector into the second-layer bidirectional long short-term memory model for processing to obtain a separation result.

[0065] Optionally, the apparatus further includes: The acquisition module is further configured to acquire training samples, where the training samples include: training audio and the corresponding annotation results of the training audio, and the training audio includes: audio data collected from multiple training sound sources, and the multiple training sound sources correspond to multiple training objects; the recognition module is further configured to perform content recognition on the training audio to obtain a training vector corresponding to the training audio and time information corresponding to the training vector; the separation module is further configured to separate the training audio based on the training vector and the time information corresponding to the training vector to obtain a first prediction result, where the first prediction result is used to represent the probability of the training object corresponding to the training vector; the processing module is configured to process the annotation result and the first prediction result to obtain a first loss function; the adjustment module is configured to adjust the model parameters of the audio separation model based on the first loss function.

[0066] Optionally, the apparatus further includes: The acquisition module is further configured to acquire a target vector, where the target vector is a vector input to the second-layer bidirectional long short-term memory model of the audio separation model; the prediction module is configured to use the auxiliary separation model to predict the target vector to obtain a second prediction result corresponding to the target vector, where the second prediction result is used to represent the training object corresponding to the target vector; the first generation module is configured to generate a second loss function based on the annotation result and the second prediction result; the second generation module is configured to obtain a total loss function based on the first loss function and the second loss function; the adjustment module is further configured to adjust the model parameters of the audio separation model based on the total loss function.

[0067] Optionally, the adjustment module includes: an adjustment unit configured to adjust the model parameters of the audio separation model based on the total loss function by using the stochastic gradient descent algorithm.

[0068] According to an embodiment of the present disclosure, the present disclosure further provides an electronic device, a readable storage medium, and a computer program product.

[0069] Figure 4 FIG. shows a schematic block diagram of an exemplary electronic device 400 that can be used to implement the embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0070] As Figure 4As shown, device 400 includes a computing unit 401, which can perform various appropriate actions and processes according to computer programs stored in a read-only memory (ROM) 402 or computer programs loaded from a storage unit 408 into a random access memory (RAM) 403. In the RAM 403, various programs and data required for the operation of device 400 can also be stored. The computing unit 401, the ROM 402, and the RAM 403 are connected to each other via a bus 404. An input / output (I / O) interface 405 is also connected to the bus 404.

[0071] Multiple components in device 400 are connected to the I / O interface 405, including: an input unit 406, such as a keyboard, a mouse, etc.; an output unit 407, such as various types of displays, speakers, etc.; a storage unit 408, such as a magnetic disk, an optical disc, etc.; and a communication unit 409, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 409 allows device 400 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0072] The computing unit 401 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 401 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 401 executes the various methods and processes described above, such as an audio processing method. For example, in some embodiments, the audio processing method can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 408. In some embodiments, part or all of the computer program can be loaded and / or installed onto device 400 via the ROM 402 and / or the communication unit 409. When the computer program is loaded into the RAM 403 and executed by the computing unit 401, one or more steps of the audio processing method described above can be executed. Alternatively, in other embodiments, the computing unit 401 can be configured to execute the audio processing method in any other appropriate way (e.g., by means of firmware).

[0073] The various embodiments of the systems and techniques described above in this specification can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems-on-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which may be a special-purpose or general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.

[0074] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus, such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The program code may execute entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine, or entirely on the remote machine or server.

[0075] In the context of the present disclosure, a machine-readable medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0076] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).

[0077] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.

[0078] A computer system can include a client and a server. The client and the server are generally far from each other and usually interact through a communication network. The client-server relationship is generated by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, or a server of a distributed system, or a server combined with a blockchain.

[0079] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitation is imposed herein.

[0080] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.

Claims

1. An audio processing method, comprising: Obtaining audio to be processed, wherein the audio to be processed includes: initial audio data collected from multiple sound sources, and the multiple sound sources correspond to multiple objects; Performing content recognition on the audio to be processed to obtain a content vector and time information corresponding to the content vector; Separating the audio to be processed based on the content vector and the time information to obtain a separation result, wherein the separation result is used to determine, from the initial audio data, target audio data corresponding to each of the multiple objects; Wherein the content vector includes: feature vectors of multiple texts with a preset granularity, and the time information includes: timestamps of the multiple texts, and separating the audio to be processed based on the content vector and the time information to obtain a separation result includes: cutting the audio to be processed based on the timestamp of each text to obtain multiple target audios; separating the multiple target audios based on the feature vectors of the multiple texts to obtain the separation result; Wherein the audio separation model at least includes: a first-layer bidirectional long short-term memory model and a second-layer bidirectional long short-term memory model, and separating the multiple target audios based on the feature vectors of the multiple texts to obtain the separation result includes: inputting the multiple target audios into the first-layer bidirectional long short-term memory model for processing to obtain a first output vector; splicing the first output vector and the feature vectors of the multiple texts to obtain a spliced vector; inputting the spliced vector into the second-layer bidirectional long short-term memory model for processing to obtain the separation result.

2. The method according to claim 1, wherein, The performing content recognition on the audio to be processed to obtain a content vector and the time information includes: Using a forced alignment model to recognize the audio to be processed to obtain text information and the time information; Using a feature generation model to extract features from the text information to obtain the content vector.

3. The method according to claim 1 or 2, further comprising: Obtaining a training sample, wherein the training sample includes training audio and an annotation result corresponding to the training audio, and the training audio includes: audio data collected from multiple training sound sources, and the multiple training sound sources correspond to multiple training objects; Performing content recognition on the training audio to obtain a training vector corresponding to the training audio and time information corresponding to the training vector; Separating the training audio based on the training vector and the time information corresponding to the training vector to obtain a first prediction result, wherein the first prediction result is used to represent the probability of the training object corresponding to the training vector; Processing the annotation result and the first prediction result to obtain a first loss function; Adjusting model parameters of the audio separation model based on the first loss function.

4. The method according to claim 3, after processing the annotation result and the first prediction result to obtain a first loss function, further comprising: Obtaining a target vector, wherein the target vector is a vector input into the second-layer bidirectional long short-term memory model of the audio separation model; Predict the target vector using an auxiliary separation model to obtain a second prediction result corresponding to the target vector, where the second prediction result is used to characterize the training object corresponding to the target vector; Generate a second loss function based on the annotation result and the second prediction result; Obtain a total loss function based on the first loss function and the second loss function; Adjust the model parameters of the audio separation model based on the total loss function.

5. The method according to claim 4, wherein, Adjusting the model parameters of the audio separation model based on the total loss function includes: Adjust the model parameters of the audio separation model using the stochastic gradient descent algorithm based on the total loss function.

6. An audio processing device, comprising: An acquisition module, configured to acquire an audio to be processed, where the audio to be processed includes: initial audio data collected from multiple sound sources, and the multiple sound sources correspond to multiple objects; An identification module, configured to perform content identification on the audio to be processed to obtain a content vector and time information corresponding to the content vector; A separation module, configured to separate the audio to be processed based on the content vector and the time information to obtain a separation result, where the separation result is used to determine target audio data corresponding to each of the multiple objects from the initial audio data; Wherein, the content vector includes: feature vectors of multiple texts with a preset granularity, the time information includes: timestamps of the multiple texts, and the separation module includes: a cutting unit, configured to cut the audio to be processed based on the timestamp of each text to obtain multiple target audios; a separation unit, configured to separate the multiple target audios based on the feature vectors of the multiple texts to obtain the separation result; Wherein, the audio separation model at least includes: a first-layer bidirectional long short-term memory model and a second-layer bidirectional long short-term memory model, and the separation unit is further configured to: input the multiple target audios into the first-layer bidirectional long short-term memory model for processing to obtain a first output vector; splice the first output vector and the feature vectors of the multiple texts to obtain a spliced vector; input the spliced vector into the second-layer bidirectional long short-term memory model for processing to obtain the separation result.

7. The apparatus according to claim 6, wherein, The identification module includes: An identification unit, configured to identify the audio to be processed using a forced alignment model to obtain text information and the time information; An extraction unit, configured to extract features from the text information using a feature generation model to obtain the content vector.

8. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1-5.

9. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1-5.

10. A computer program product comprising a computer program which, when executed by a processor, implements the method according to any one of claims 1-5.

Citation Information

Patent Citations

  • Speech enhancement for target speakers and speed enhancement method

    CN107919133A

  • Object spoken language evaluation method and device, storage medium and electronic device

    CN111986680A