Audio processing method, device, electronic device and storage medium
By extracting the center channel from the audio, determining the probability of speech presence and performing weighted mixing, the problem of unsatisfactory audio voice enhancement effect in the existing technology is solved, and the clarity and intelligibility of the audio content are improved.
Patent Information
- Application Number
- CN202210450469.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-24
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2042-04-24
AI Technical Summary
The existing technology is not ideal for enhancing human voice in audio. The human voice and background sound in the audio are difficult to distinguish, and the clarity and intelligibility are not high.
Extract the original center channel from the audio, determine the probability of speech presence in each segment, perform noise reduction on the center channel, and perform weighted mixing based on the speech presence probability, combining it with other channels to form the mixed audio.
By precisely enhancing the center channel vocals, the clarity and intelligibility of audio content are significantly improved.
Smart Images

Figure CN114944162B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing technology, and in particular to an audio processing method, device, electronic device and storage medium. Background Art
[0002] With the development of technology, people have more and more channels for consuming external information, one of which is watching television programs. However, due to the varying quality of recording technology used during TV production, the human voice in the audio is often mixed with background noise, making it difficult to distinguish between the human voice and the background noise during TV broadcasts, which negatively impacts the viewing experience. Furthermore, when watching TV programs in noisy environments, it is also difficult to hear the human voice in the audio. Therefore, there is a need to enhance the human voice in the audio of TV programs.
[0003] Currently, the main method used to enhance human voices in audio is as follows: Since there are multiple channels in the audio, the center channel is first extracted from the multiple channels. Then, the Voice Activity Detection (VAD) algorithm is used to determine whether there is voice in the center channel. If there is voice, dynamic range control (DRC), filters, etc. are used to reduce noise on the center channel. The noise-reduced center channel is then mixed with the other channels in the multi-channel system, and the mixed audio is output for people to listen to. If there is no voice, the center channel is no longer subjected to noise reduction processing, and the original audio is played directly for people to listen to.
[0004] However, when the above method is used to enhance the human voice in the audio, the enhancement effect of the human voice is not very ideal, and the clarity and intelligibility of the spoken content in the audio are not high. Summary of the Invention
[0005] The purpose of the embodiments of the present application is to provide an audio processing method, device, electronic device and storage medium to further enhance the human voice in the audio and improve the clarity and intelligibility of the spoken content in the audio.
[0006] To solve the above technical problems, the embodiments of the present application provide the following technical solutions:
[0007] In a first aspect, the present application provides an audio processing method, comprising: extracting an original center channel from audio having multiple channels; determining the speech existence probability of each segment in the original center channel; performing noise reduction processing on the original center channel to obtain a noise-reduced center channel; weighting the original center channel and the noise-reduced center channel based on the speech existence probability of each segment in the original center channel, wherein a higher speech existence probability of a corresponding segment results in a greater weight of the corresponding segment in the noise-reduced center channel during mixing; and mixing the weighted center channel with other channels in the audio except the original center channel to obtain mixed audio.
[0008] According to a second aspect of the present application, there is provided an audio processing device, comprising: a signal separation unit for extracting an original center channel from audio having multiple channels; a voice activity detection unit for determining the speech existence probability of each segment in the original center channel; a speech noise reduction unit for performing noise reduction processing on the original center channel to obtain a noise-reduced center channel; a center channel processing unit for weighting the original center channel and the noise-reduced center channel based on the speech existence probability of each segment in the original center channel, wherein the higher the speech existence probability of the corresponding segment, the greater the weight of the corresponding segment in the noise-reduced center channel during mixing; and a mixing unit for mixing the weighted center channel with other channels in the audio except the original center channel to obtain mixed audio.
[0009] The third aspect of the present application provides an electronic device, comprising: a processor, a memory and a bus; wherein the processor and the memory communicate with each other through the bus; the processor is used to call program instructions in the memory to execute the method in the first aspect.
[0010] A fourth aspect of the present application provides a computer-readable storage medium having a computer program stored thereon, wherein when the program is executed, the device where the storage medium is located is controlled to execute the method in the first aspect.
[0011] Compared with the prior art, the audio processing method provided in the first aspect of the present application extracts the original center channel from the audio, then determines the probability of speech existence of each segment in the original center channel, and performs noise reduction processing on the original center channel to obtain the noise-reduced center channel. Then, based on the speech existence probability of each segment in the original center channel, the original center channel and the noise-reduced center channel are weighted, wherein the higher the speech existence probability of the corresponding segment, the greater the weight of the corresponding segment in the noise-reduced center channel during mixing. Finally, the weighted center channel is mixed with other channels in the audio except the original center channel to obtain the mixed audio. In this way, the human voice in the center channel can be accurately enhanced, and the clarity and intelligibility of the content in the audio can be improved to a large extent.
[0012] The audio processing device provided in the second aspect, the electronic device provided in the third aspect, and the computer-readable storage medium provided in the fourth aspect of this application have the same or similar beneficial effects as the audio processing method provided in the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] The above and other objects, features and advantages of the exemplary embodiments of the present application will become readily understood by reading the detailed description below with reference to the accompanying drawings. In the accompanying drawings, several embodiments of the present application are shown in an exemplary and non-limiting manner, and the same or corresponding reference numerals represent the same or corresponding parts, wherein:
[0014] Figure 1 The flow diagram of the audio processing method in the embodiment of the present application is as follows Figure 1 ;
[0015] Figure 2 The flow diagram of the audio processing method in the embodiment of the present application is as follows Figure 2 ;
[0016] Figure 3 The flow diagram of the audio processing method in the embodiment of the present application is as follows Figure 3 ;
[0017] Figure 4 This is a structural diagram of an audio processing device in an embodiment of the present application;
[0018] Figure 5 This is a schematic diagram of the structure of an electronic device in an embodiment of the present application. DETAILED DESCRIPTION
[0019] The following describes exemplary embodiments of the present application in more detail with reference to the accompanying drawings. Although exemplary embodiments of the present application are shown in the accompanying drawings, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments described herein. Instead, these embodiments are provided to enable a more thorough understanding of the present application and to fully convey the scope of the present application to those skilled in the art.
[0020] It should be noted that, unless otherwise specified, the technical or scientific terms used in this application should have the common meanings understood by those skilled in the art to which this application belongs.
[0021] Currently, to enhance human voices in audio, a voice activity detection (VAD) algorithm is used to determine whether a human voice is present in the center channel. If so, the entire center channel is filtered to emphasize the human voice. The filtered center channel is then mixed with the other audio channels to enhance the human voice. However, this method of enhancing human voices in audio remains unsatisfactory. When the enhanced audio is played back, it is still difficult to distinguish the human voice from background noise, making it difficult to hear the human voice clearly. Consequently, the clarity and intelligibility of the audio content remain limited.
[0022] After in-depth research, the inventors found that the main reason why the current method of enhancing the human voice in audio is not ideal is that the signal of the center channel in the audio is not further processed in a specific way, that is, the signal of each section of the center channel is not analyzed separately, and different methods are used for processing according to different analysis results.
[0023] In view of this, the embodiments of the present application provide an audio processing method, device, electronic device and storage medium. For the center channel in the audio, the probability of speech existence in each segment in the center channel is first determined. At the same time, the center channel is also subjected to a complete noise reduction process. When obtaining the final center channel, for segments with a high probability of speech existence, the corresponding segments use the center channel that has been noise-reduced, and for segments with a low probability of speech existence, the corresponding segments use the center channel that has been noise-reduced. Finally, the center channel obtained above is mixed with other channels in the audio. By combining the segments in the center channel processed in different ways based on the probability of speech existence in each segment in the center channel, the human voice in the center channel can be effectively enhanced, thereby improving the clarity and intelligibility of the content in the audio.
[0024] Next, the audio processing method provided in the embodiment of the present application is first described in detail.
[0025] Figure 1 The flow diagram of the audio processing method in the embodiment of the present application is as follows Figure 1 , see Figure 1 As shown, the method may include:
[0026] S101: extracting an original center channel from multi-channel audio.
[0027] An audio clip often contains multiple channels. For example, in an audio clip of a teacher lecturing, one channel contains the teacher's voice, another contains the teacher's writing on the blackboard, another contains the sound of students flipping pages, and so on. In this audio clip, people primarily want to hear the teacher's voice. Therefore, the teacher's voice is the central channel, while the teacher's writing on the blackboard and the student's flipping pages are the other two channels.
[0028] In order to achieve more accurate processing of the human voice in the audio, it is necessary to extract the center channel from the audio of multiple channels, which can also be called the original center channel.
[0029] Generally speaking, a single audio channel has only one center channel. This is the channel closest to the audio collection device. However, if two channels are closest to the audio collection device, a prompt can be generated to prompt the user to select one channel as the center channel. Alternatively, a center channel can be automatically determined based on pre-set rules, such as selecting the loudest channel as the center channel. The specific method for extracting the center channel from audio is not specified here.
[0030] S102: Determine the speech existence probability of each segment in the original center channel.
[0031] After extracting the original center channel from the audio, since it's actually a segment of audio signal, it sometimes contains human voices, meaning someone is speaking, while other times it doesn't. Therefore, to accurately process the center channel, we first need to determine the speech presence probability of each segment in the original center channel—that is, the probability of someone speaking in that segment.
[0032] During the specific implementation process, the voice activity detection VAD algorithm can be used to calculate the probability of speech existence in each segment of the original center channel. For example: in the original center channel, the corresponding total duration is 3 minutes, and each 1 minute duration is a segment. The original center channel can be divided into 3 speech segments. The voice activity detection VAD algorithm can be used to calculate that the probability of speech existence in the 0-1 minute segment is a, the probability of speech existence in the 1-2 minute segment is b, and the probability of speech existence in the 2-3 minute segment is c. Among them, the speech existence probabilities a, b, and c can be the same or different, which needs to be derived based on the actual calculation results. In addition, other algorithms can also be used to calculate the probability of speech existence in each segment of the original center channel. As for the specific algorithm, as long as it can calculate the probability of speech existence, it is not limited here.
[0033] S103: Perform noise reduction processing on the original center channel to obtain a center channel after noise reduction.
[0034] After extracting the original center channel from the audio, although the original center channel is mainly human voice, there will still be some interference sounds. In order to enhance the human voice, these interference sounds need to be removed, that is, the original center channel is subjected to noise reduction processing to obtain the noise-reduced center channel.
[0035] In practical applications, the original center channel may be subjected to noise reduction processing by various methods such as dynamic range control (DRC) or various filters, as long as the noise in the original center channel can be removed. The specific noise reduction method used in step S103 is not limited here.
[0036] It should be noted that the above steps S102 and S103 can be performed simultaneously or non-simultaneously, and the order in which the above steps S102 and S103 are performed is not specifically limited. In the subsequent step S104, the results of steps S102 and S103 are used simultaneously.
[0037] S104: Weighting the original central channel and the noise-reduced central channel based on the speech existence probability of each segment in the original central channel.
[0038] The higher the probability of speech in a corresponding segment, the greater the weight of the corresponding segment in the center channel after noise reduction during mixing. The lower the probability of speech in a corresponding segment, the greater the weight of the corresponding segment in the center channel after noise reduction during mixing.
[0039] For example, assume the original center channel is divided into three segments: segment a1, segment b1, and segment c1. Through step S102, it is determined that the probability of speech presence in segment a1 is 90%, the probability of speech presence in segment b1 is 50%, and the probability of speech presence in segment c1 is 20%. Through step S103, the segments corresponding to the noise-reduced center channel are segment a2, segment b2, and segment c2, respectively. In this step, the original center channel and the noise-reduced center channel are mixed to varying degrees according to the speech presence probability of each segment, i.e., weighted. The weighted center channel can now be: [segment a1 × 90% + segment a2 × (1-90%)] + [segment b1 × 50% + segment b2 × (1-50%)] + [segment c1 × 20% + segment c2 × (1-20%)].
[0040] Of course, if the probability of speech in a certain segment is 100%, it means that there is definitely a human voice in the segment, and the corresponding center channel segment used for this segment is the corresponding segment in the center channel after denoising. If the probability of speech in a certain segment is 0%, it means that there is no human voice in the segment, and the corresponding center channel segment used for this segment is the corresponding segment in the original center channel.
[0041] In this way, the part of the center channel where the human voice exists can be denoised more accurately, and the sound of the part where the human voice does not exist can be retained as much as possible, so as to achieve accurate increase of the human voice and improve the clarity and intelligibility of the content in the audio.
[0042] S105: Mixing the weighted center channel with other channels in the audio except the original center channel to obtain mixed audio.
[0043] After precisely enhancing the vocals in the center channel of the audio, generally speaking, the probability of the presence of vocals in channels other than the center channel is low. Even if vocals exist in other channels, the vocal content in other channels is not as important as that in the center channel. Therefore, it is no longer necessary to precisely enhance the vocals in other channels, and the weighted center channel can be directly mixed with the other channels.
[0044] Of course, the above steps S102-S104 can also be performed again for other channels, and then in this step, the weighted center channel is mixed with the weighted other channels. In this way, the various human voices in the audio can be further enhanced, thereby improving the intelligibility of the entire audio content.
[0045] The specific method used when mixing the weighted center channel with other channels can be various current mixing methods, which are not specifically limited here.
[0046] As can be seen from the above content, the audio processing method provided in the embodiment of the present application extracts the original center channel from the audio, then determines the probability of speech existence of each segment in the original center channel, and performs noise reduction processing on the original center channel to obtain the center channel after noise reduction. Then, the original center channel and the center channel after noise reduction are weighted based on the probability of speech existence of each segment in the original center channel. The higher the probability of speech existence of the corresponding segment, the greater the weight of the corresponding segment in the center channel after noise reduction during mixing. Finally, the weighted center channel is mixed with other channels in the audio except the original center channel to obtain the mixed audio. In this way, the human voice in the center channel can be accurately enhanced, and the clarity and intelligibility of the content in the audio can be improved to a large extent.
[0047] Furthermore, as a Figure 1 As a refinement and extension of the method shown, an embodiment of the present application also provides an audio processing method. Figure 2 The flow diagram of the audio processing method in the embodiment of the present application is as follows Figure 2 , see Figure 2 As shown, the method may include:
[0048] S201: extracting an original center channel from multi-channel audio using an adaptive weighting algorithm.
[0049] When extracting the original center channel from audio, whether it's dual or dual channels, an adaptive panning algorithm can be used. This allows for more accurate center channel extraction, facilitating precise vocal enhancement. Of course, other algorithms can also be used, as long as they can extract the center channel from the audio, and are not specifically limited here.
[0050] S202: Determine the speech existence probability of each segment in the original center channel.
[0051] Here, the specific implementation method of step S202 is the same or similar to that of the aforementioned step S102, so it will not be repeated here.
[0052] S203: Smoothing the speech existence probability of each segment in the original center channel.
[0053] After dividing the original center channel into segments and determining the speech probability corresponding to each segment, if the speech probability of adjacent segments differs significantly, the subsequent weighting of the corresponding segments in the original and noise-reduced center channels may result in speech discontinuities in the weighted center channel, which can degrade the processed audio quality. Therefore, after determining the speech probability of each segment in the original center channel, these speech probabilities need to be smoothed.
[0054] For example, suppose the original center channel is divided into five segments: segment a, segment b, segment c, segment d, and segment e, with corresponding speech probabilities of 50%, 60%, 90%, 20%, and 70%, respectively. If the corresponding segments in the original center channel and the noise-reduced center channel are subsequently weighted according to these probabilities, the resulting audio playback will be uneven, resulting in a choppy sound. Therefore, smoothing the speech probabilities for each segment is necessary. During audio playback, the speech probability of each segment in the original center channel changes from 50% to 60% and then to 90%. The second change in probability is significant, necessitating smoothing. Therefore, the 60% probability is adjusted to 70%. Next, the speech probability changes from 90% to 20% and then to 70%. The second change in probability is also significant, necessitating smoothing. Therefore, the 20% probability is adjusted to 50%. Finally, the probability of speech existence in each segment changes from 50%, 60%, 90%, 20%, and 70% to 50%, 70%, 90%, 50%, and 70% after smoothing.
[0055] Of course, the above specific values are only used as examples for clarity of explanation. As for the specific smoothing method of the speech presence probability of each segment, various smoothing processing algorithms can be used for processing, and the smoothing results obtained will also vary, which is not specifically limited here.
[0056] S204: Using a speech enhancement algorithm based on short-time log spectrum estimation to perform noise reduction processing on the original central channel to obtain a noise-reduced central channel.
[0057] To improve the speed of noise reduction processing on the original center channel and thereby enhance audio processing efficiency, a speech enhancement algorithm based on short-time logarithmic spectrum estimation can be employed to reduce the noise of the original center channel. Compared to other noise reduction algorithms, this algorithm converges more quickly during the noise reduction process. Therefore, it can more quickly reduce the noise of the original center channel and obtain the reduced-noise center channel, further enhancing audio processing efficiency.
[0058] It should be noted that the above steps S202-S203 and this step S204 can be performed simultaneously or at different times. The order in which the above steps S202-S203 and this step S204 are performed is not specifically limited here. Only after both steps S203 and S204 are completed can the following step S205 be performed.
[0059] S205: Weighting the original central channel and the noise-reduced central channel based on the speech existence probability of each segment in the original central channel.
[0060] In the process of weighting the two center channels, the weighting may be performed in the specific manner described in the aforementioned step S104.
[0061] To further speed up the weighting process, if the probability of the first speech corresponding to the first segment in the original center channel is greater than 50% and greater than the probability of the second speech corresponding to the second segment, the weighting process can be simplified. That is, in the final center channel, each segment will use either the original center channel segment or the noise-reduced center channel segment. The specific channel segment to use is determined by the speech probability.
[0062] For example, suppose the probability of the first speech corresponding to the first segment in the original center channel is 80%, and the probability of the second speech corresponding to the second segment is 30%. This indicates that the probability of the human voice in the first segment is higher, while the probability of the human voice in the second segment is lower. In this case, to complete the weighting process as quickly as possible, the weighting calculation process can be simplified and the segment corresponding to the first segment in the denoised center channel can be directly used to splice with the second segment in the original center channel.
[0063] In this way, while ensuring the enhancement of the human voice, the weighted calculation speed can also be accelerated, thereby improving the audio processing efficiency.
[0064] S206: Perform smoothing processing on the weighted center channel.
[0065] The original center channel and the noise-reduced center channel are weighted based on the probability of speech in each segment of the original center channel to produce a weighted center channel. To ensure smoother playback of the final processed audio and avoid interruptions, the weighted center channel can be smoothed.
[0066] In a specific implementation, any one or more current smoothing processing methods can be used to smooth the weighted center channel. The specific smoothing processing method used in this step is not limited here.
[0067] After smoothing the weighted center channel, the processed center channel becomes smoother and has no discontinuities, thereby preventing the center channel from causing discontinuities in the final audio and improving the smoothness of the final output audio.
[0068] S207: Mixing the smoothed center channel with other channels in the audio except the original center channel to obtain mixed audio.
[0069] Here, the specific implementation of step S207 is the same or similar to that of the aforementioned step S105, so it will not be repeated here.
[0070] In actual applications, audio is often output by a multi-channel / stereo system, that is, an audio device, and the multi-channel / stereo system divides the audio into a left channel (L) and a right channel (R) and outputs them separately. Therefore, the audio processing method provided in the embodiment of the present application is to extract the center channel from the left channel and the right channel, and then determine the probability of speech existence and reduce noise for the extracted center channel, and then mix it with the left channel and the right channel respectively, and finally output the processed left channel and right channel respectively, that is, the processed audio.
[0071] Figure 3 The flow diagram of the audio processing method in the embodiment of the present application is as follows Figure 3 , see Figure 3 As shown in the figure, after the multi-channel / stereo system outputs the left channel (L) and the right channel (R), the center channel is first extracted from the left channel (L) and the right channel (R), and the center channel (C) is extracted. Then, the center channel (C) is subjected to speech noise reduction to obtain the noise-reduced center channel (C'), and speech endpoint detection is performed on the center channel (C). Finally, based on the results of speech endpoint detection, the noise-reduced center channel (C') and the left channel (L), as well as the noise-reduced center channel (C') and the right channel (R) are mixed, and finally the processed left channel (L') and right channel (R') are output respectively, thereby enhancing the human voice in the audio and improving the clarity and intelligibility of the audio content.
[0072] Based on the same inventive concept, as an implementation of the above method, an embodiment of the present application further provides an audio processing device. Figure 4 This is a structural diagram of the audio processing device in the embodiment of the present application, see Figure 4 As shown, the device may include:
[0073] The signal separation unit 401 is configured to extract an original center channel from the multi-channel audio;
[0074] a voice activity detection unit 402, configured to determine the probability of speech presence in each segment of the original center channel;
[0075] A speech noise reduction unit 403 is configured to perform noise reduction processing on the original central channel to obtain a noise-reduced central channel;
[0076] a center channel processing unit 404 configured to weight the original center channel and the noise-reduced center channel based on the speech presence probability of each segment in the original center channel, wherein a higher speech presence probability of a corresponding segment results in a greater weight of the corresponding segment in the noise-reduced center channel during mixing;
[0077] The mixing unit 405 is configured to mix the weighted center channel with other channels in the audio except the original center channel to obtain mixed audio.
[0078] In other embodiments of the present application, the speech noise reduction unit 403 is specifically configured to perform noise reduction processing on the original center channel using a speech enhancement algorithm based on short-time log spectrum estimation.
[0079] In other embodiments of the present application, the voice activity detection unit 402 is further configured to smooth the speech presence probabilities of each segment in the original center channel; and the center channel processing unit 404 is specifically configured to weight the original center channel and the denoised center channel based on the speech presence probabilities of each segment in the smoothed original center channel.
[0080] In other embodiments of the present application, the original center channel includes a first segment and a second segment; the probability of existence of the first speech corresponding to the first segment is higher than 50%, and higher than the probability of existence of the second speech corresponding to the second segment; the center channel processing unit 404 is specifically used to splice the first segment and the second segment based on the probability of existence of the first speech in the first segment and the probability of existence of the second speech in the second segment.
[0081] In other embodiments of the present application, the mixing unit 405 is further configured to smooth the weighted center channel; and mix the smoothed center channel with other channels in the audio except the original center channel.
[0082] In other embodiments of the present application, the signal separation unit 401 is specifically configured to extract the original center channel from the multi-channel audio using an adaptive weighting algorithm.
[0083] It should be noted that the description of the above device embodiment is similar to the description of the above method embodiment and has similar beneficial effects as the method embodiment. For technical details not disclosed in the device embodiment of this application, please refer to the description of the method embodiment of this application for understanding.
[0084] Based on the same inventive concept, an embodiment of the present application also provides an electronic device. Figure 5 This is a schematic diagram of the structure of the electronic device in the embodiment of the present application, see Figure 5 As shown, the electronic device may include: a processor 501, a memory 502 and a bus 503; wherein the processor 501 and the memory 502 communicate with each other through the bus 503; the processor 501 is used to call the program instructions in the memory 502 to execute the method in one or more of the above embodiments.
[0085] It should be noted that the description of the electronic device embodiment above is similar to the description of the method embodiment above, and has similar beneficial effects as the method embodiment. For technical details not disclosed in the electronic device embodiment of this application, please refer to the description of the method embodiment of this application for understanding.
[0086] Based on the same inventive concept, an embodiment of the present application further provides a computer-readable storage medium on which a computer program is stored, wherein when the program is running, the device where the storage medium is located is controlled to execute the method in one or more of the above embodiments.
[0087] It should be noted that the description of the above storage medium embodiment is similar to the description of the above method embodiment and has similar beneficial effects as the method embodiment. For technical details not disclosed in the storage medium embodiment of this application, please refer to the description of the method embodiment of this application for understanding.
[0088] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
Claims
1. An audio processing method, characterized in that: The method comprises: Extracting the original center channel from multi-channel audio; Determining the probability of speech existence of each segment in the original center channel; Performing noise reduction processing on the original center channel to obtain a noise-reduced center channel; weighting the original center channel and the noise-reduced center channel based on the speech presence probability of each segment in the original center channel, wherein the higher the speech presence probability of a corresponding segment, the greater the weight of the corresponding segment in the noise-reduced center channel during mixing; Mixing the weighted center channel with other channels in the audio except the original center channel to obtain mixed audio; The weighting of the original central channel and the noise-reduced central channel based on the speech existence probability of each segment in the original central channel includes: Multiply each segment in the original center channel by its speech probability, and then add the multiplication result of the same segment and (1-corresponding speech probability) in the denoised center channel to obtain the weighted center channel.
2. The method according to claim 1, characterized in that The performing noise reduction processing on the original center channel includes: A speech enhancement algorithm based on short-time log spectrum estimation is used to perform noise reduction processing on the original central channel.
3. The method according to claim 1, characterized in that After determining the speech existence probability of each segment in the original center channel, the method further includes: Smoothing the speech existence probability of each segment in the original central channel; The weighting of the original central channel and the noise-reduced central channel based on the speech existence probability of each segment in the original central channel includes: The original central channel and the noise-reduced central channel are weighted based on the speech existence probability of each segment in the smoothed original central channel.
4. The method according to claim 1, wherein The original central channel includes a first segment and a second segment; the probability of existence of a first speech corresponding to the first segment is higher than 50%, and higher than the probability of existence of a second speech corresponding to the second segment; and weighting the original central channel and the noise-reduced central channel based on the speech existence probabilities of the segments in the original central channel includes: The first segment and the second segment are spliced based on the first speech existence probability of the first segment and the second speech existence probability of the second segment.
5. The method according to claim 1, wherein After weighting the original central channel and the noise-reduced central channel based on the speech presence probability of each segment in the original central channel, the method further includes: Smoothing the weighted center channel; Mixing the weighted center channel with other channels in the audio except the original center channel includes: The smoothed center channel is mixed with other channels of the audio except the original center channel.
6. The method according to any one of claims 1 to 5, characterized in that The method of extracting an original center channel from multi-channel audio includes: An adaptive weighting algorithm is used to extract the original center channel from multi-channel audio.
7. An audio processing device, characterized in that: The device comprises: A signal separation unit, configured to extract an original center channel from the multi-channel audio; a voice activity detection unit, configured to determine a probability of speech presence in each segment of the original center channel; a speech noise reduction unit, configured to perform noise reduction processing on the original central channel to obtain a noise-reduced central channel; a center channel processing unit, configured to weight the original center channel and the noise-reduced center channel based on the speech presence probability of each segment in the original center channel, wherein the higher the speech presence probability of a corresponding segment, the greater the weight of the corresponding segment in the noise-reduced center channel during mixing; a mixing unit, configured to mix the weighted center channel with other channels in the audio except the original center channel to obtain mixed audio; The center channel processing unit is specifically configured to multiply each segment in the original center channel by its speech presence probability, and then add the multiplication result of the same segment in the denoised center channel and (1-corresponding speech presence probability) to obtain a weighted center channel.
8. The device according to claim 7, characterized in that The speech noise reduction unit is specifically configured to perform noise reduction processing on the original central channel by using a speech enhancement algorithm based on short-time log spectrum estimation.
9. An electronic device, characterized in that: The electronic device includes: a processor, a memory and a bus; wherein the processor and the memory communicate with each other through the bus; the processor is used to call program instructions in the memory to execute the method according to any one of claims 1 to 6.
10. A computer-readable storage medium, characterized in that A computer program is stored thereon, wherein when the program is run, the device where the storage medium is located is controlled to execute the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Voice recognition method and device, electronic equipment and storage medium
CN111696532A
Apparatus and a method for signal enhancement
US20200286501A1