Voice processing method and device, equipment, medium and program product

By using low-frequency band frequency domain mask vectors to correct high-frequency band frequency domain mask vectors in the audio zoom method, the amplitude spectrum accuracy of the target channel is improved, thus enhancing the audio zoom effect.

CN121600943APending Publication Date: 2026-03-03XIAOMI TECH (WUHAN) CO LTD +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411155986.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-08-21
Publication Date
2026-03-03

AI Technical Summary

Technical Problem

The existing audio zoom methods have room for improvement in speaker separation at a specified angle.

Method used

By acquiring speech signals from different speech acquisition channels, a frequency domain mask vector is obtained. The high-frequency band mask vector is then corrected using a more reliable low-frequency band mask vector. The corrected mask vector is used to mask the initial amplitude spectrum to obtain the amplitude spectrum of the target channel, and finally, the enhanced speech signal is obtained.

Benefits of technology

It improves the reliability of the high-frequency band frequency domain mask vector, enhances the accuracy of the amplitude spectrum of the target channel, and improves the audio zoom effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121600943A_ABST
    Figure CN121600943A_ABST
Patent Text Reader

Abstract

The invention relates to a voice processing method and device, equipment, a medium and a program product, and the method comprises the steps: obtaining a frequency domain mask vector based on the mask extraction processing of to-be-processed voice signals collected by different voice collection channels, and the frequency domain mask vector comprises a first frequency domain mask vector corresponding to a target channel; based on the first frequency domain mask vector of the first frequency band, correcting the first frequency domain mask vector of the second frequency band to obtain a corrected first frequency domain mask vector of the second frequency band, the frequency in the first frequency band being smaller than the frequency in the second frequency band; based on the first frequency domain mask vector of the first frequency band and the corrected first frequency domain mask vector, performing masking processing on an initial amplitude spectrum corresponding to the to-be-processed voice signal to obtain an amplitude spectrum of the target channel; and based on the amplitude spectrum of the target channel, obtaining an enhanced voice signal of the target channel. According to the method, the accuracy of the amplitude spectrum of the determined target channel can be improved, and the accuracy of the enhanced voice signal is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of speech processing technology, and in particular to a speech processing method, apparatus, device, medium, and program product. Background Technology

[0002] Audio zoom technology can capture sound from a target direction while attenuating interfering sound sources from other directions. In related technologies, speaker separation at a specified angle can be achieved by extracting a mask and then applying it to the amplitude spectrum of the input signal.

[0003] However, the audio zoom method in related technologies still needs improvement in its ability to separate speakers at a specified angle. Summary of the Invention

[0004] To overcome the problems existing in related technologies, this disclosure provides a speech processing method, apparatus, device, medium, and program product.

[0005] According to a first aspect of the present disclosure, a speech processing method is provided, the method comprising: Based on the mask extraction processing of the speech signals to be processed acquired from different speech acquisition channels, a frequency domain mask vector is obtained, wherein the frequency domain mask vector includes the first frequency domain mask vector corresponding to the target channel. Based on the first frequency domain mask vector of the first frequency band, the first frequency domain mask vector of the second frequency band is corrected to obtain the corrected first frequency domain mask vector of the second frequency band, wherein the frequency in the first frequency band is less than the frequency in the second frequency band; The initial amplitude spectrum of the speech signal to be processed is masked based on the first frequency domain mask vector of the first frequency band and the modified first frequency domain mask vector to obtain the amplitude spectrum of the target channel. Based on the amplitude spectrum of the target channel, the enhanced speech signal of the target channel is obtained.

[0006] Optionally, the step of modifying the first frequency domain mask vector of the second frequency band based on the first frequency band to obtain the modified first frequency domain mask vector of the second frequency band includes: Obtain the average value of the first frequency domain mask vector for the first frequency band; The average value is determined as the corrected first frequency domain mask vector.

[0007] Optionally, the method further includes: The frequency threshold is determined based on the spacing between different voice acquisition channels and the speed of sound. Based on the frequency threshold, the first frequency band and the second frequency band in the first frequency domain mask vector are determined.

[0008] Optionally, the method further includes: Obtain the first energy corresponding to the target channel and the second energy corresponding to the interference channel; The amplitude spectrum of the target channel is updated based on the first energy and the second energy to obtain the updated amplitude spectrum of the target channel; The process of obtaining the enhanced speech signal of the target channel based on the amplitude spectrum of the target channel includes: Based on the updated amplitude spectrum of the target channel, the enhanced speech signal of the target channel is obtained.

[0009] Optionally, updating the amplitude spectrum of the target channel based on the first energy and the second energy to obtain the updated amplitude spectrum of the target channel includes: If the product of the first energy and the preset coefficient is less than the second energy, the amplitude spectrum of the target channel is suppressed to obtain the updated amplitude spectrum of the target channel, wherein the preset coefficient is greater than 1. If the product of the first energy and the preset coefficient is greater than or equal to the second energy, the amplitude spectrum of the target channel is determined as the updated amplitude spectrum of the target channel.

[0010] Optionally, obtaining the first energy corresponding to the target channel and the second energy corresponding to the interference channel includes: Based on the amplitude spectrum of the candidate channel, the energy corresponding to the candidate channel is obtained; Wherein, when the candidate channel is the target channel, the energy corresponding to the candidate channel is the first energy; when the candidate channel is the interference channel, the energy corresponding to the candidate channel is the second energy.

[0011] Optionally, obtaining the energy corresponding to the candidate channel based on the amplitude spectrum of the candidate channel includes: Based on the amplitude spectrum of the candidate channel in the second audio frame, the initial energy of the candidate channel in the second audio frame is obtained; Based on the initial energy of the candidate channel in the second audio frame and the energy of the candidate channel in the first audio frame, the energy of the candidate channel in the second audio frame is obtained, where the first audio frame is the previous audio frame adjacent to the second audio frame.

[0012] According to a second aspect of the present disclosure, a voice processing apparatus is provided, the apparatus comprising: The extraction module is configured to perform mask extraction processing on the speech signals to be processed acquired from different speech acquisition channels to obtain a frequency domain mask vector, wherein the frequency domain mask vector includes a first frequency domain mask vector corresponding to the target channel. The correction module is configured to correct the first frequency domain mask vector of the second frequency band based on the first frequency domain mask vector of the first frequency band, so as to obtain the corrected first frequency domain mask vector of the second frequency band, wherein the frequency in the first frequency band is less than the frequency in the second frequency band; The masking module is configured to mask the initial amplitude spectrum of the speech signal to be processed based on the first frequency domain mask vector of the first frequency band and the modified first frequency domain mask vector, so as to obtain the amplitude spectrum of the target channel. The generation module is configured to obtain the enhanced speech signal of the target channel based on the amplitude spectrum of the target channel.

[0013] According to a third aspect of the present disclosure, a computer-readable storage medium is provided, having stored thereon computer program instructions that, when executed by a processor, implement the steps of the speech processing method provided in the first aspect of the present disclosure.

[0014] According to a fourth aspect of the present disclosure, an electronic device is provided, comprising: a memory having a computer program stored thereon; and a processor for executing the computer program in the memory to implement the steps of the voice processing method mentioned in the first aspect of the present disclosure.

[0015] According to a fifth aspect of the present disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps of the speech processing method mentioned in the first aspect of the present disclosure.

[0016] The technical solutions provided by the embodiments of this disclosure may include the following beneficial effects: First, a frequency domain mask vector is obtained by mask extraction processing of the speech signals to be processed acquired from different speech acquisition channels. Then, based on the first frequency domain mask vector of the first frequency band, the first frequency domain mask vector of the second frequency band is corrected to obtain the corrected first frequency domain mask vector of the second frequency band. Next, the initial amplitude spectrum corresponding to the speech signal to be processed is masked based on the first frequency domain mask vector of the first frequency band and the corrected first frequency domain mask vector to obtain the amplitude spectrum of the target channel. Finally, the enhanced speech signal of the target channel can be obtained based on the amplitude spectrum of the target channel. By using the more reliable low-frequency band frequency domain mask vector to correct the relatively less reliable high-frequency band frequency domain mask vector, the reliability of the high-frequency band frequency domain mask vector can be improved, thereby improving the accuracy of the amplitude spectrum of the determined target channel, improving the accuracy of the enhanced speech signal, and improving the audio zoom effect.

[0017] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description

[0018] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.

[0019] Figure 1 This is a schematic flowchart illustrating a speech processing method according to an exemplary embodiment of the present disclosure.

[0020] Figure 2 This is a schematic flowchart illustrating a speech processing method according to an exemplary embodiment of the present disclosure.

[0021] Figure 3 This is a structural block diagram of a voice processing apparatus according to an exemplary embodiment of the present disclosure.

[0022] Figure 4 This is a block diagram illustrating an electronic device according to an exemplary embodiment of the present disclosure. Detailed Implementation

[0023] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.

[0024] It should be noted that all actions involving the acquisition of signals, information, or data in this application are carried out in compliance with the relevant data protection laws and policies of the country where the application is located, and with the authorization granted by the owner of the relevant device.

[0025] Figure 1 This is a flowchart illustrating a speech processing method according to an exemplary embodiment. This speech processing method can be used in an electronic device, such as a mobile phone, laptop, or tablet computer. Figure 1 As shown, the speech processing method includes the following steps: S110: Based on the mask extraction processing of the speech signals to be processed acquired from different speech acquisition channels, a frequency domain mask vector is obtained. The frequency domain mask vector includes the first frequency domain mask vector of the corresponding target channel.

[0026] In this embodiment of the disclosure, the electronic device may have at least two voice acquisition devices, each of which may correspond to a voice acquisition channel. Thus, the electronic device can acquire voice signals in the environment through at least two voice acquisition channels, that is, obtain the voice signal to be processed corresponding to each voice acquisition channel.

[0027] Among them, at least two voice acquisition devices can form a voice acquisition device array, and each voice acquisition device in the voice acquisition device array can be called a voice acquisition array element.

[0028] The voice acquisition device is usually a microphone.

[0029] In some implementations, for the speech signal to be processed acquired by each speech acquisition channel, the electronic device can decompose the speech signal in the time and frequency domains using STFT (Short Time Fourier Transform) to obtain the spectrum corresponding to different speech acquisition channels. Then, the spectrum corresponding to different speech acquisition channels is processed by the PDCW (Phase Difference Channel Weighting) algorithm to obtain a continuous mask vector. This continuous mask vector corresponds to multiple different frequency points, where a frequency point refers to a specific frequency value; that is, one frequency point corresponds to one frequency value. Therefore, the obtained continuous mask vector can also be called a frequency domain mask vector. Thus, through the above process, the mask extraction processing of the speech signal to be processed acquired by different speech acquisition channels is realized.

[0030] The core idea of ​​the PDCW algorithm is to calculate the phase difference through the spectrum corresponding to different speech acquisition channels, then obtain the time delay difference based on the phase difference, and then obtain a binary mask vector by whether the time delay difference is within the target threshold. After the binary mask vector is processed by the Gammatone filter, a continuous mask vector is obtained.

[0031] It should be noted that the speech processing method proposed in this embodiment is not only applicable to mask vectors extracted based on the phase difference method, but also applicable to masks obtained by other beamforming methods, such as Minimum Variance Distortionless Response (MVDR) and Generalized Sidelobe Canceller (GSC).

[0032] In some implementations, the frequency domain mask vector may further include a second frequency domain mask vector corresponding to the interference channel.

[0033] The target channel can be understood as the channel corresponding to the audio source within the desired audio acquisition range, while the interference channel can be understood as the channel corresponding to the audio source outside the desired audio acquisition range.

[0034] Here, the audio acquisition range refers to the range in physical space. In this case, for example, the audio source within the desired audio acquisition range is, for example, an audio source within a 60-degree angle range directly in front of the electronic device, and the audio source outside the desired audio acquisition range is, for example, an audio source outside a 60-degree angle range directly in front of the electronic device.

[0035] The mask vector can be viewed as a set of weighting coefficients, which correspond to the spectrum of the speech signal to be processed and are used to adjust different frequency components in the spectrum. When applied to the amplitude spectrum obtained from the spectrum, the mask vector changes the amplitude of each frequency component in the amplitude spectrum, thereby achieving the effect of enhancement or suppression.

[0036] S120, based on the first frequency domain mask vector of the first frequency band, the first frequency domain mask vector of the second frequency band is corrected to obtain the corrected first frequency domain mask vector of the second frequency band, where the frequency in the first frequency band is less than the frequency in the second frequency band.

[0037] In this embodiment, since the frequencies in the first frequency band are lower than those in the second frequency band, the first frequency band can be referred to as the low-frequency band, and the second frequency band as the high-frequency band. Furthermore, the reliability of the first frequency domain mask vector in the first frequency band is greater than the reliability of the first frequency domain mask vector in the second frequency band.

[0038] By using the first frequency domain mask vector of the first frequency band, which has higher reliability, to correct the first frequency domain mask vector of the second frequency band, which has lower reliability, the reliability of the first frequency domain mask vector of the final full frequency band can be enhanced.

[0039] In some implementations, the method of this disclosure may include the steps of determining a first frequency band and a second frequency band: Based on the spacing between different voice acquisition channels and the speed of sound, a frequency threshold is determined; based on the frequency threshold, the first frequency band and the second frequency band in the first frequency domain mask vector are determined.

[0040] For example, suppose the continuous mask vector output by the PDCW is a 256-point frequency domain vector. The value of the k-th frequency point and the m-th channel is represented as mask[k, m], the frequency corresponding to the k-th frequency point is f, the distance between the two voice acquisition channels is d, and the speed of sound is c. According to the spatial sampling theorem, signal aliasing can only be guaranteed when the distance between the voice acquisition channels is no greater than half the wavelength of the signal, i.e., satisfying the condition... .

[0041] Therefore, when the signal frequency f satisfies: f ≤ f max , where f max The signal is reliable only when the frequency threshold is equal to c / 2d. Therefore, in this embodiment of the present disclosure, the frequency threshold, i.e., f, can be determined based on the spacing between different voice acquisition channels and the speed of sound. max Furthermore, frequencies less than or equal to the frequency threshold are determined as frequencies in the first frequency band, and frequencies greater than the frequency threshold are determined as frequencies in the second frequency band.

[0042] In some implementations, modifying the first frequency domain mask vector of the second frequency band based on the first frequency band to obtain the modified first frequency domain mask vector of the second frequency band may include the following steps: Obtain the average value of the first frequency domain mask vector of the first frequency band; determine the average value as the corrected first frequency domain mask vector.

[0043] Based on the foregoing, the first frequency band can include multiple frequency points. Therefore, the average value of the first mask vectors belonging to the multiple frequency points in the first frequency band can be taken, and then the average value can be used to replace the mask vectors of the frequency points included in the second frequency band. It can be understood that when the second frequency band includes multiple frequency points, the mask vectors corresponding to these multiple second frequency points are the same, all being the average value calculated in the previous step.

[0044] In some implementations, the above process can be represented by the following formula:

[0045] Where K1 represents the frequency point in the first frequency band, and K2 represents the frequency point in the second frequency band. This indicates the minimum frequency point in the first frequency band. This indicates the highest frequency point in the first frequency band. This indicates the channel number, used to distinguish different channels, such as the target channel and the interference channel.

[0046] That is, through the above process, the corrected first frequency domain mask vector of the second frequency band and the corrected second frequency domain mask vector of the second frequency band can be obtained according to actual needs.

[0047] S130, based on the first frequency domain mask vector of the first frequency band and the corrected first frequency domain mask vector, the initial amplitude spectrum corresponding to the speech signal to be processed is masked to obtain the amplitude spectrum of the target channel.

[0048] The first frequency domain mask vector of the first frequency band and the corrected first frequency domain mask vector can form the first mask vector of the entire frequency band.

[0049] The amplitude spectrum represents the amplitude distribution of a signal in the frequency domain and can be obtained by calculating the absolute value of the spectrum.

[0050] In some implementations, the electronic device can decompose the speech signals to be processed acquired by different speech acquisition channels in the time domain and frequency domain respectively through STFT (Short Time Fourier Transform) to obtain the spectrum corresponding to different speech acquisition channels. Then, the absolute value of the spectrum corresponding to different speech acquisition channels is taken to obtain the amplitude spectrum corresponding to different speech acquisition channels. Then, the average value of the amplitude spectrum corresponding to different speech acquisition channels is taken to obtain the initial amplitude spectrum corresponding to the speech signal to be processed.

[0051] Furthermore, in this embodiment of the present disclosure, after obtaining the first mask vector of the entire frequency band, the initial amplitude spectrum corresponding to the speech signal to be processed can be masked using the first mask vector of the entire frequency band to obtain the amplitude spectrum of the target channel.

[0052] In some implementations, based on similar operations, after obtaining the second mask vector of the full-band corresponding interference channel, the initial amplitude spectrum corresponding to the speech signal to be processed can be masked using the second mask vector of the full-band, thereby obtaining the amplitude spectrum of the interference channel.

[0053] S140: Based on the amplitude spectrum of the target channel, the enhanced speech signal of the target channel is obtained.

[0054] In some implementations, after obtaining the amplitude spectrum of the target channel, the enhanced speech signal of the target channel can be obtained by passing the amplitude spectrum of the target channel through ISTFT (Inverse Short Time Fourier Transform).

[0055] The above method first involves mask extraction of the speech signals acquired from different speech acquisition channels to obtain frequency domain mask vectors. Then, based on the first frequency domain mask vector of the first frequency band, the first frequency domain mask vector of the second frequency band is corrected to obtain the corrected first frequency domain mask vector of the second frequency band. Next, based on the first frequency domain mask vector of the first frequency band and the corrected first frequency domain mask vector, the initial amplitude spectrum corresponding to the speech signal to be processed is masked to obtain the amplitude spectrum of the target channel. Finally, based on the amplitude spectrum of the target channel, the enhanced speech signal of the target channel can be obtained. By using the more reliable low-frequency band frequency domain mask vector to correct the relatively less reliable high-frequency band frequency domain mask vector, the reliability of the high-frequency band frequency domain mask vector can be improved, thereby improving the accuracy of the determined amplitude spectrum of the target channel, improving the accuracy of the enhanced speech signal, and improving the audio zoom effect.

[0056] Furthermore, although the aforementioned process can improve the accuracy of the amplitude spectrum of the determined target channel to a certain extent, considering the possibility of reverberation interference, which may cause energy from interfering channels to leak into the target channel and affect the accuracy of the amplitude spectrum of the determined target channel, in order to further improve the accuracy of the amplitude spectrum of the determined target channel, in some embodiments, the speech processing method of this disclosure may further include the following steps: Obtain the first energy corresponding to the target channel and the second energy corresponding to the interference channel; update the amplitude spectrum of the target channel based on the first energy and the second energy to obtain the updated amplitude spectrum of the target channel.

[0057] In this case, obtaining the enhanced speech signal of the target channel based on the amplitude spectrum of the target channel may include the following steps: The enhanced speech signal of the target channel is obtained based on the updated amplitude spectrum of the target channel.

[0058] In this embodiment of the disclosure, in order to further improve the accuracy of the amplitude spectrum of the determined target channel, the energy of the target channel (first energy) and the energy of the interference channel (second energy) can be tracked, and the amplitude spectrum of the target channel can be updated based on the first energy and the second energy to obtain the updated amplitude spectrum of the target channel. Subsequently, the enhanced speech signal of the target channel can be obtained based on the updated amplitude spectrum of the target channel.

[0059] In some implementations, updating the amplitude spectrum of the target channel based on the first energy and the second energy to obtain the updated amplitude spectrum of the target channel may include the following steps: If the product of the first energy and the preset coefficient is less than the second energy, the amplitude spectrum of the target channel is suppressed to obtain the updated amplitude spectrum of the target channel, and the preset coefficient is greater than 1; if the product of the first energy and the preset coefficient is greater than or equal to the second energy, the amplitude spectrum of the target channel is determined as the updated amplitude spectrum of the target channel.

[0060] In this embodiment, if the product of the first energy and a preset coefficient greater than 1 is less than the second energy, it indicates that there is no target speaker's voice at the current location, and the energy originates from leakage in the interference channel. Therefore, the amplitude spectrum of the target channel should be suppressed. Otherwise, if the product of the first energy and the preset coefficient is greater than or equal to the second energy, the amplitude spectrum of the target channel can remain unchanged without processing.

[0061] In some implementations, interference can be suppressed by multiplying the amplitude spectrum of the target channel by a small suppression coefficient, such as 0.05.

[0062] In some implementations, the preset coefficient may be, for example, 2.

[0063] In some implementations, the above process can be expressed by the following formula:

[0064] in, This represents the updated amplitude spectrum corresponding to the target channel. This represents the amplitude spectrum before the update corresponding to the target channel. This represents the energy of the target channel. This indicates the energy of the interference channel. Indicates the preset coefficient. This represents the inhibition coefficient.

[0065] In some implementations, obtaining the first energy corresponding to the target channel and the second energy corresponding to the interference channel may include the following steps: Based on the amplitude spectrum of the candidate channel, the energy corresponding to the candidate channel is obtained; In the case where the candidate channel is the target channel, the energy corresponding to the candidate channel is the first energy; in the case where the candidate channel is the interference channel, the energy corresponding to the candidate channel is the second energy.

[0066] In this embodiment of the disclosure, the energy corresponding to the target channel, i.e., the first energy, can be obtained based on the amplitude spectrum of the target channel. Similarly, the energy corresponding to the interference channel, i.e., the second energy, can be obtained based on the amplitude spectrum of the interference channel. The methods for obtaining the first energy corresponding to the target channel and the second energy corresponding to the interference channel are similar.

[0067] In some implementations, the sum of the amplitudes corresponding to each frequency point in the amplitude spectrum of the target channel can be determined as the energy of the target channel. Similarly, the sum of the amplitudes corresponding to each frequency point in the amplitude spectrum of the interference channel can be determined as the energy of the interference channel.

[0068] It should be noted that the method in the above embodiments is applicable to the processing of the speech signal to be processed based on any audio frame, and to obtain the enhanced speech signal of the target channel corresponding to that audio frame.

[0069] In some implementations, to further improve the accuracy of the amplitude spectrum of the target channel corresponding to a certain audio frame, when determining the energy of the target channel or interference channel, in addition to referring to the information of the audio frame, the information of the previous audio frame can also be combined. In this case, obtaining the energy corresponding to the candidate channel based on the amplitude spectrum of the candidate channel may include the following steps: Based on the amplitude spectrum of the candidate channel in the second audio frame, the initial energy of the candidate channel in the second audio frame is obtained. Based on the initial energy of the candidate channel in the second audio frame and the energy of the candidate channel in the first audio frame, the energy of the candidate channel in the second audio frame is obtained. The first audio frame is the previous audio frame adjacent to the second audio frame.

[0070] In this embodiment of the disclosure, when determining the energy of the candidate channel corresponding to the nth audio frame, the calculation can be performed based on the initial energy of the candidate channel corresponding to the nth audio frame and the energy of the candidate channel corresponding to the (n-1)th audio frame.

[0071] In some implementations, the initial energy of the candidate channel corresponding to the audio frame n and the energy of the candidate channel corresponding to the audio frame n-1 can be weighted and summed to obtain the energy of the candidate channel corresponding to the audio frame n.

[0072] In some implementations, the sum of the amplitudes of each frequency point in the amplitude spectrum of the candidate channel corresponding to the nth audio frame can be determined as the initial energy of the candidate channel corresponding to the nth audio frame.

[0073] In some implementations, the process of calculating the energy of the candidate channel corresponding to the nth audio frame can be expressed by the following formula:

[0074] Where n represents the frame index and k represents the k-th frequency point. This is a smoothing factor, which can be a constant such as 0.99. This represents the energy of the candidate channel corresponding to the nth audio frame. This represents the energy of the candidate channel corresponding to the (n-1)th audio frame. This indicates the amplitude corresponding to the k-th frequency point in the amplitude spectrum of the n-th audio frame for the candidate channel, and z represents the number of frequency points.

[0075] The following is combined Figure 2 The flowchart illustrating the speech processing method is provided to illustrate the execution process of the speech processing method according to an embodiment of this disclosure. Figure 2 As shown: Taking the voice acquisition channels as the left and right microphones in the electronic device as an example, the dual-channel voice signals acquired by the left and right microphones are decomposed in the time and frequency domains through STFT (Short Time Fourier Transform). Then, for any audio frame, mask extraction processing is further performed through the PDCW (Phase Difference Channel Weighted) algorithm to obtain the frequency domain mask vector. The frequency domain mask vector includes the first frequency domain mask vector corresponding to the target channel and the second frequency domain mask vector corresponding to the interference channel.

[0076] Based on the distance between the left and right microphones and the speed of sound, a frequency threshold is determined. Frequency points greater than the frequency threshold are assigned to the second frequency band, and frequency points less than or equal to the frequency threshold are assigned to the first frequency band.

[0077] For the first frequency domain mask vector, the average value of the first frequency domain mask vector of the first frequency band is obtained, and the average value is determined as the corrected first frequency domain mask vector. Based on the first frequency domain mask vector of the first frequency band and the corrected first frequency domain mask vector, the initial amplitude spectrum corresponding to the speech signal to be processed is masked to obtain the amplitude spectrum of the target channel.

[0078] Similarly, for the second frequency domain mask vector, the average value of the second frequency domain mask vector of the first frequency band is obtained, and the average value is determined as the corrected second frequency domain mask vector. Based on the second frequency domain mask vector of the first frequency band and the corrected second frequency domain mask vector, the initial amplitude spectrum corresponding to the speech signal to be processed is masked to obtain the amplitude spectrum of the interference channel.

[0079] Through the above process, the amplitude spectrum of the target channel and the amplitude spectrum of the interference channel in any audio frame can be obtained. Subsequently, the amplitude spectrum of the target channel can be further updated based on the first energy and the second energy to obtain the updated amplitude spectrum of the target channel. Based on the updated amplitude spectrum of the target channel, the enhanced speech signal of the target channel in that audio frame can be obtained.

[0080] Taking the enhanced speech signal corresponding to the target channel in the nth audio frame as an example, the initial energy of the target channel in the nth audio frame can be obtained by summing the amplitudes of each frequency point in the amplitude spectrum of the target channel in the nth audio frame. Furthermore, the energy of the target channel in the (n-1)th audio frame can be obtained. Then, the initial energy and the energy of the target channel in the (n-1)th audio frame are weighted and summed to obtain the energy of the target channel in the nth audio frame. Similarly, the energy of the interference channel in the nth audio frame can be obtained.

[0081] Next, the magnitude of the product of the first energy and the preset coefficient is compared with the second energy. If the product of the first energy and the preset coefficient is less than the second energy, the amplitude spectrum of the target channel is suppressed by the suppression coefficient to obtain the updated amplitude spectrum of the target channel in the nth audio frame; if the product of the first energy and the preset coefficient is greater than or equal to the second energy, the amplitude spectrum of the target channel is determined as the updated amplitude spectrum of the target channel in the nth audio frame.

[0082] After obtaining the updated amplitude spectrum of the target channel corresponding to the nth audio frame, an ISTFT (Inverse Short Time Fourier Transform) can be performed on the updated amplitude spectrum of the target channel corresponding to the nth audio frame to obtain the enhanced speech signal of the target channel, that is, to obtain the enhanced speech signal corresponding to the audio source within the desired audio acquisition range.

[0083] Figure 3 This is a structural block diagram of a voice processing device 300 according to an exemplary embodiment, with reference to... Figure 3 The voice processing device 300 includes: The extraction module 310 is configured to perform mask extraction processing on the speech signals to be processed acquired from different speech acquisition channels to obtain a frequency domain mask vector, wherein the frequency domain mask vector includes a first frequency domain mask vector corresponding to the target channel. The correction module 320 is configured to correct the first frequency domain mask vector of the second frequency band based on the first frequency domain mask vector of the first frequency band, so as to obtain the corrected first frequency domain mask vector of the second frequency band, wherein the frequency in the first frequency band is less than the frequency in the second frequency band. The masking module 330 is configured to mask the initial amplitude spectrum of the speech signal to be processed based on the first frequency domain mask vector of the first frequency band and the modified first frequency domain mask vector to obtain the amplitude spectrum of the target channel. The generation module 340 is configured to obtain the enhanced speech signal of the target channel based on the amplitude spectrum of the target channel.

[0084] In some embodiments, the correction module 320 includes: The acquisition submodule is configured to acquire the average value of the first frequency domain mask vector of the first frequency band; The determination submodule is configured to determine the average value as the corrected first frequency domain mask vector.

[0085] In some embodiments, the voice processing device 300 further includes: The first determining module is configured to determine the frequency threshold based on the spacing and sound velocity between different voice acquisition channels; The second determining module is configured to determine the first frequency band and the second frequency band in the first frequency domain mask vector based on the frequency threshold.

[0086] In some embodiments, the voice processing device 300 further includes: The acquisition module is configured to acquire the first energy corresponding to the target channel and the second energy corresponding to the interference channel; The update module is configured to update the amplitude spectrum of the target channel based on the first energy and the second energy to obtain the updated amplitude spectrum of the target channel; Accordingly, the generation module 340 includes: The generation submodule is configured to obtain the enhanced speech signal of the target channel based on the updated amplitude spectrum of the target channel.

[0087] In some embodiments, the update module includes: The first update submodule is configured to suppress the amplitude spectrum of the target channel when the product of the first energy and the preset coefficient is less than the second energy, so as to obtain the updated amplitude spectrum of the target channel, wherein the preset coefficient is greater than 1. The second update submodule is configured to determine the amplitude spectrum of the target channel as the updated amplitude spectrum of the target channel when the product of the first energy and the preset coefficient is greater than or equal to the second energy.

[0088] In some embodiments, the acquisition module includes: An energy determination submodule is configured to obtain the energy corresponding to a candidate channel based on the amplitude spectrum of the candidate channel; wherein, when the candidate channel is the target channel, the energy corresponding to the candidate channel is the first energy, and when the candidate channel is the interference channel, the energy corresponding to the candidate channel is the second energy.

[0089] In some embodiments, the energy determination submodule includes: The first determining unit is configured to obtain the initial energy of the candidate channel in the second audio frame based on the amplitude spectrum of the candidate channel in the second audio frame. The second determining unit is configured to obtain the energy of the candidate channel in the second audio frame based on the initial energy of the candidate channel in the second audio frame and the energy of the candidate channel in the first audio frame, wherein the first audio frame is the previous audio frame adjacent to the second audio frame.

[0090] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.

[0091] This disclosure also provides a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, implement the steps of the speech processing method provided in this disclosure.

[0092] Figure 4 This is a block diagram illustrating an electronic device 400 according to an exemplary embodiment. For example, the electronic device 400 may be a mobile phone, a tablet computer, or a laptop computer.

[0093] Reference Figure 4 The electronic device 400 may include one or more of the following components: processing component 402, memory 404, power supply component 406, multimedia component 408, audio component 410, input / output interface 412, sensor component 414, and communication component 416.

[0094] Processing component 402 typically controls the overall operation of electronic device 400, such as operations associated with display, telephone calls, data communication, camera operation, and recording. Processing component 402 may include one or more processors 420 to execute instructions to complete all or part of the steps of the aforementioned voice processing method. Furthermore, processing component 402 may include one or more modules to facilitate interaction between processing component 402 and other components. For example, processing component 402 may include a multimedia module to facilitate interaction between multimedia component 408 and processing component 402.

[0095] Memory 404 is configured to store various types of data to support the operation of electronic device 400. Examples of such data include instructions for any application or method operating on electronic device 400, contact data, phonebook data, messages, pictures, videos, etc. Memory 404 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0096] Power supply component 406 provides power to various components of electronic device 400. Power supply component 406 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to electronic device 400.

[0097] Multimedia component 408 includes a screen that provides an output interface between the electronic device 400 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touchscreen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors may sense not only the boundaries of the touch or swipe action but also the duration and pressure associated with the touch or swipe operation. In some embodiments, multimedia component 408 includes a front-facing camera and / or a rear-facing camera. When the electronic device 400 is in an operating mode, such as a shooting mode or a video mode, the front-facing camera and / or the rear-facing camera may receive external multimedia data. Each front-facing camera and rear-facing camera may be a fixed optical lens system or have focal length and optical zoom capabilities.

[0098] Audio component 410 is configured to output and / or input audio signals. For example, audio component 410 includes at least two microphones (MICs) configured to receive external audio signals when electronic device 400 is in an operating mode, such as call mode, recording mode, and voice recognition mode. The received audio signals may be further stored in memory 404 or transmitted via communication component 416. In some embodiments, audio component 410 also includes a speaker for outputting audio signals.

[0099] Input / output interface 412 provides an interface between processing component 402 and peripheral interface modules, which may be keyboards, click wheels, buttons, etc. These buttons may include, but are not limited to, home buttons, volume buttons, start buttons, and lock buttons.

[0100] Sensor assembly 414 includes one or more sensors for providing state assessments of various aspects of electronic device 400. For example, sensor assembly 414 may detect the on / off state of electronic device 400, the relative positioning of components such as the display and keypad of electronic device 400, changes in position of electronic device 400 or a component of electronic device 400, the presence or absence of user contact with electronic device 400, orientation or acceleration / deceleration of electronic device 400, and temperature changes of electronic device 400. Sensor assembly 414 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. Sensor assembly 414 may also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, sensor assembly 414 may also include an accelerometer, gyroscope, magnetometer, pressure sensor, or temperature sensor.

[0101] Communication component 416 is configured to facilitate wired or wireless communication between electronic device 400 and other devices. Electronic device 400 can access wireless networks based on communication standards, such as WiFi, 2G, or 3G, or combinations thereof. In one exemplary embodiment, communication component 416 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, communication component 416 also includes a near-field communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on radio frequency identification (RFID) technology, Infrared Data Association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0102] In an exemplary embodiment, the electronic device 400 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the above-described voice processing method.

[0103] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 404 including instructions, which can be executed by a processor 420 of an electronic device 400 to complete the aforementioned voice processing method. For example, the non-transitory computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.

[0104] The aforementioned device can be a standalone electronic device or a part of a standalone electronic device. For example, in one embodiment, the device can be an integrated circuit (IC) or a chip, wherein the integrated circuit can be a single IC or a collection of multiple ICs. The chip can include, but is not limited to, the following types: GPU (Graphics Processing Unit), CPU (Central Processing Unit), FPGA (Field Programmable Gate Array), DSP (Digital Signal Processor), ASIC (Application Specific Integrated Circuit), and SoC (System on Chip). The aforementioned integrated circuit or chip can be used to execute executable instructions (or code) to implement the aforementioned voice processing method. The executable instructions can be stored in the integrated circuit or chip or obtained from other devices or equipment. For example, the integrated circuit or chip includes a processor, memory, and an interface for communicating with other devices. The executable instructions can be stored in the memory, and when the executable instructions are executed by the processor, the above-mentioned voice processing method is implemented; or, the integrated circuit or chip can receive the executable instructions through the interface and transmit them to the processor for execution to implement the above-mentioned voice processing method.

[0105] In another exemplary embodiment, a computer program product is also provided, the computer program product comprising a computer program executable by a programmable device, the computer program having a code portion for performing the above-described speech processing method when executed by the programmable device.

[0106] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of this disclosure. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the following claims.

[0107] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.

Claims

1. A speech processing method, characterized in that, The method includes: Based on the mask extraction processing of the speech signals to be processed acquired from different speech acquisition channels, a frequency domain mask vector is obtained, wherein the frequency domain mask vector includes the first frequency domain mask vector corresponding to the target channel. Based on the first frequency domain mask vector of the first frequency band, the first frequency domain mask vector of the second frequency band is corrected to obtain the corrected first frequency domain mask vector of the second frequency band, wherein the frequency in the first frequency band is less than the frequency in the second frequency band; The initial amplitude spectrum of the speech signal to be processed is masked based on the first frequency domain mask vector of the first frequency band and the modified first frequency domain mask vector to obtain the amplitude spectrum of the target channel. Based on the amplitude spectrum of the target channel, the enhanced speech signal of the target channel is obtained.

2. The method according to claim 1, characterized in that, The process of modifying the first frequency domain mask vector of the second frequency band based on the first frequency band to obtain the modified first frequency domain mask vector of the second frequency band includes: Obtain the average value of the first frequency domain mask vector for the first frequency band; The average value is determined as the corrected first frequency domain mask vector.

3. The method according to claim 1, characterized in that, The method further includes: The frequency threshold is determined based on the spacing between different voice acquisition channels and the speed of sound. Based on the frequency threshold, the first frequency band and the second frequency band in the first frequency domain mask vector are determined.

4. The method according to claim 1, characterized in that, The method further includes: Obtain the first energy corresponding to the target channel and the second energy corresponding to the interference channel; The amplitude spectrum of the target channel is updated based on the first energy and the second energy to obtain the updated amplitude spectrum of the target channel; The process of obtaining the enhanced speech signal of the target channel based on the amplitude spectrum of the target channel includes: Based on the updated amplitude spectrum of the target channel, the enhanced speech signal of the target channel is obtained.

5. The method according to claim 4, characterized in that, The step of updating the amplitude spectrum of the target channel based on the first energy and the second energy to obtain the updated amplitude spectrum of the target channel includes: If the product of the first energy and the preset coefficient is less than the second energy, the amplitude spectrum of the target channel is suppressed to obtain the updated amplitude spectrum of the target channel, wherein the preset coefficient is greater than 1. If the product of the first energy and the preset coefficient is greater than or equal to the second energy, the amplitude spectrum of the target channel is determined as the updated amplitude spectrum of the target channel.

6. The method according to claim 4, characterized in that, The step of obtaining the first energy corresponding to the target channel and the second energy corresponding to the interference channel includes: Based on the amplitude spectrum of the candidate channel, the energy corresponding to the candidate channel is obtained; Wherein, when the candidate channel is the target channel, the energy corresponding to the candidate channel is the first energy; when the candidate channel is the interference channel, the energy corresponding to the candidate channel is the second energy.

7. The method according to claim 6, characterized in that, The process of obtaining the energy corresponding to the candidate channel based on the amplitude spectrum of the candidate channel includes: Based on the amplitude spectrum of the candidate channel in the second audio frame, the initial energy of the candidate channel in the second audio frame is obtained; Based on the initial energy of the candidate channel in the second audio frame and the energy of the candidate channel in the first audio frame, the energy of the candidate channel in the second audio frame is obtained, where the first audio frame is the previous audio frame adjacent to the second audio frame.

8. A voice processing device, characterized in that, The device includes: The extraction module is configured to perform mask extraction processing on the speech signals to be processed acquired from different speech acquisition channels to obtain a frequency domain mask vector, wherein the frequency domain mask vector includes a first frequency domain mask vector corresponding to the target channel. The correction module is configured to correct the first frequency domain mask vector of the second frequency band based on the first frequency domain mask vector of the first frequency band, so as to obtain the corrected first frequency domain mask vector of the second frequency band, wherein the frequency in the first frequency band is less than the frequency in the second frequency band; The masking module is configured to mask the initial amplitude spectrum of the speech signal to be processed based on the first frequency domain mask vector of the first frequency band and the modified first frequency domain mask vector, so as to obtain the amplitude spectrum of the target channel. The generation module is configured to obtain the enhanced speech signal of the target channel based on the amplitude spectrum of the target channel.

9. An electronic device, characterized in that, include: A memory on which computer programs are stored; A processor for executing the computer program in the memory to implement the steps of the method according to any one of claims 1-7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the program implements the steps of the method described in any one of claims 1-7.

11. A computer program product, comprising a computer program, characterized in that, When executed by a processor, the computer program implements the steps of the method described in any one of claims 1-7.