Method and device for enhancing voice signal, processor and computing equipment
By analyzing the statistical results and comparison results of the fundamental frequency signal in the voice signal, filtering and enhancement operations are performed, the problem of identifying and enhancing the speaker's voice signal in a noisy environment is solved, and high-quality voice signal enhancement is achieved.
Patent Information
- Application Number
- CN202510105165.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-06
AI Technical Summary
In noisy environments or multi-person speaking scenarios, how to accurately identify and selectively enhance the speaker's voice signal, eliminate background noise and other speakers' voices.
By analyzing the statistical results of the fundamental frequency signals contained in multiple speech frames in the speech signal, the target fundamental frequency signal is determined, and filtering operations are performed based on the comparison results of the candidate fundamental frequency signal and the target fundamental frequency signal, and finally the unfiltered speech frame and the fundamental frequency signal are enhanced.
It realizes accurate recognition and selective enhancement of the speaker's voice signal, improves the quality and clarity of the voice signal in complex environments, and reduces the consumption of computing and storage resources.
Smart Images

Figure CN119943081A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of audio processing, and more particularly, to a method, apparatus, processor, and computing device for enhancing a speech signal. Background Art
[0002] Voice communication is a natural and efficient way to exchange information. With the development of audio technology, various fields have higher and higher requirements for the quality of voice signals. Among them, speech enhancement technology is an important branch of speech signal processing, which can be applied to noise suppression in noisy environments or scenes where multiple people are talking. When enhancing speech signals, it is necessary to selectively enhance the speech signals. For example, in the scene of extracting the main speaker, it is expected that the voice of the main speaker is extracted and enhanced while the voice of the background person or the environmental noise is filtered out. In this way, how to accurately identify the speech signal that needs to be enhanced and enhance it is the key to improving the quality of the speech signal. Summary of the invention
[0003] The embodiments of this specification provide a method, an apparatus, a processor, and a computing device for enhancing a speech signal.
[0004] In a first aspect of the present specification, a method for enhancing a speech signal is provided. The method comprises determining a target baseband signal based on statistical results corresponding to a plurality of baseband signals contained in a plurality of speech frames in the speech signal. The method further comprises determining a filtering operation corresponding to the target speech frame and the candidate baseband signal contained in the target speech frame based on a comparison result between the candidate baseband signal contained in the target speech frame in the plurality of speech frames and the target baseband signal. The method further comprises enhancing the unfiltered target speech frame and the unfiltered candidate baseband signal in the unfiltered target speech frame based on the filtering operation.
[0005] In a second aspect of the present specification, a device for enhancing a speech signal is provided. The device includes a target baseband signal determination unit, which is configured to determine the target baseband signal based on the statistical results corresponding to multiple baseband signals contained in multiple speech frames in the speech signal. The device also includes a filtering operation determination unit, which is configured to determine the filtering operation corresponding to the target speech frame and the candidate baseband signal contained in the target speech frame based on the comparison result between the candidate baseband signal contained in the target speech frame in the multiple speech frames and the target baseband signal. The device also includes a signal enhancement unit, which is configured to enhance the unfiltered target speech frame and the unfiltered candidate baseband signal in the unfiltered target speech frame based on the filtering operation.
[0006] In a third aspect of the present specification, a processor is provided, which is configured to execute the method according to the first aspect of the present specification.
[0007] In a fourth aspect of the present specification, a computing device is provided. The computing device includes a memory. The computing device also includes a processor according to the third aspect of the present specification.
[0008] It should be understood that the contents described in the Summary of the Invention are not intended to limit the key or important features of the embodiments of this specification, nor are they intended to limit the scope of this specification. Other features of this specification will become easy to understand through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] The above and other features, advantages and aspects of the embodiments of the present specification will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. In the accompanying drawings, the same or similar reference numerals represent the same or similar elements, wherein:
[0010] Figure 1 A schematic diagram illustrating an example environment in which various embodiments of the present specification may be implemented;
[0011] Figure 2 A flowchart of a method for enhancing a speech signal according to some embodiments of the specification is shown;
[0012] Figure 3 A schematic diagram showing a method of determining a spectrogram corresponding to a speech frame according to some embodiments of the present invention is shown;
[0013] Figure 4 A schematic diagram showing enhancement of a speech signal according to some embodiments of the present specification is shown;
[0014] Figure 5 A schematic diagram showing exemplary application scenarios of some embodiments of the present specification; and
[0015] Figure 6 A block diagram of a device that can implement various embodiments of the present specification is shown. DETAILED DESCRIPTION
[0016] The embodiments of the present specification will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present specification are shown in the accompanying drawings, it should be understood that the present specification can be implemented in various forms and should not be construed as being limited to the embodiments described herein, which are instead provided for a more thorough and complete understanding of the present specification. It should be understood that the drawings and embodiments of the present specification are only for exemplary purposes and are not intended to limit the scope of protection of the present specification.
[0017] In the description of the embodiments of this specification, the term "including" and similar terms should be understood as open inclusion, that is, "including but not limited to". The term "based on" should be understood as "based at least in part on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The terms "first", "second", etc. may refer to different or the same objects. Other explicit and implicit definitions may also be included below.
[0018] As mentioned above, in a multi-person speaking scenario or in a noisy environment, the speech signal may contain multiple signals such as the voice of user A, the voice of user B, and background noise, etc. Therefore, when enhancing the speech signal, it is crucial to selectively enhance the multiple signals contained in the speech signal.
[0019] In the related art, a pre-trained speech separation model can be used to separate speech signals, thereby determining the speech signals that are ultimately used for enhancement. For example, the speech separation model can process the input speech data to obtain separated speech signals. However, the performance of the speech separation model is poor in more complex scenarios. Moreover, training the speech separation model and storing the speech classification model require certain processing resources and storage resources.
[0020] To this end, a method for enhancing a speech signal is provided in an embodiment of the present specification, and the fundamental frequency signal corresponding to the speaker can be determined based on the spectrum diagram of the speech signal. On this basis, according to whether a speech frame contains fundamental frequency signals of multiple speakers and whether it contains the fundamental frequency signal corresponding to the speaker, it can be determined which speech frames to filter and which other signals contained in which speech frames to filter, thereby enabling effective selection and selective enhancement of speech signals. In addition, by identifying and enhancing the speech signal corresponding to the speaker, the quality and clarity of the speech signal in a complex environment can be improved. In addition, through such a method, different filtering operations can be performed on different speech frames without relying on a complex speech separation model, thereby reducing the consumption of computing and storage resources.
[0021] Reference below Figures 1 to 6 It should be understood that these exemplary embodiments are provided only to enable those skilled in the art to better understand and implement the embodiments of this specification, and are not intended to limit the scope of this specification in any way.
[0022] Figure 1 1 is a schematic diagram of an example environment 100 in which various embodiments of the present specification may be implemented. Figure 1As shown, in the environment 100, a microphone 102, an audio signal 104, a computing device 106, and a speaker 108 are included. The microphone 102 can be a single microphone or a microphone array. The microphone 102 can be used to collect the audio signal 104, for example, to collect the voice signal of the speaker. It can be understood that the microphone array can sample and process the spatial characteristics of the sound field, so as to determine the angle and distance of the speaker using the audio signal received by the microphone array, so as to achieve tracking of the speaker and directional pickup of the voice signal.
[0023] In some embodiments of this specification, the computing device 106 can be a terminal device or a server. Among them, the server can be an independent server, or a server cluster or distributed system composed of multiple servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, and cloud computing. The terminal device can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, and a smart watch, etc., but is not limited to this. The microphone can be deployed in the computing device 106 such as a terminal device, or it can be directly or indirectly connected to the computing device through a wired or wireless communication method, which is not limited in this specification.
[0024] In reality, when the speaker is speaking to the microphone 102, there may be background sounds of the environment and the voices of other people. For example, in a conference scenario, when the speaker is sharing a work case, there may be conversations among the participants. In order to improve the quality of the speaker's voice signal so that the audience can clearly hear the information conveyed by the speaker, it is necessary to filter out the voices of other people and background noise.
[0025] like Figure 1 As shown, after the microphone 102 acquires the audio signal 104, the audio signal 104 can be sent to the computing device 106. After receiving the audio signal 104, the computing device 106 can analyze the audio signal 104, determine the voice signal contained in the audio signal 104, and determine which voice frames in the voice signal are filtered or which signals in the voice frame are filtered. Since the baseband signals of different users are different, the baseband signal can uniquely characterize a user. Therefore, the voice frame or the signal contained in the voice frame can be filtered based on the baseband signal. It can be understood that in conference scenes, lecture scenes or other scenes, the speaker's voice accounts for a high proportion, so the frequency of the baseband signal corresponding to the speaker in the voice signal is also relatively high. On this basis, the target audio signal corresponding to the speaker can be determined according to the statistical results of the baseband signal.
[0026] Then, when the target baseband signal appears in the speech frame, the speech frame can be retained; and when the speech frame does not contain the target baseband signal, the speech frame can be filtered. If the speech frame contains not only the target baseband signal but also other baseband signals, it is also necessary to filter the other baseband signals contained in the speech frame when retaining the speech frame. In this way, effective speech enhancement can be performed on such a complex audio signal.
[0027] like Figure 1 As shown, after the computing device 106 enhances the voice signal, the enhanced voice signal 112 can be sent to the speaker 108. The speaker 108 is also called a "speaker" and is used to convert the audio electrical signal into a sound signal. The audience 110, such as a participant or a student, can listen to the enhanced voice signal 112 through the speaker 108.
[0028] In this way, the speech signal corresponding to the speaker can be determined based on the statistical results of the baseband signal, and on this basis, the speech signal corresponding to the speaker can be enhanced, thereby improving the clarity and intelligibility of the speaker's voice, allowing the audience to receive and understand the speaker's information more smoothly, and improving the quality and clarity of the speech signal in a complex environment. In addition, by filtering out interfering audio and interfering signals, the workload of subsequent audio processing can be saved, thereby improving the efficiency of audio processing.
[0029] Combined with the above Figure 1 A schematic diagram of an example environment 100 in which embodiments of the present specification can be implemented is described below. Figure 2 200 is a flowchart of a method 200 for enhancement according to some embodiments of the present specification. The method 200 may be Figure 1 Executed by the computing device 106 shown in FIG. Figure 2 As shown, in box 202, method 200 may include determining a target fundamental frequency signal based on statistical results corresponding to multiple fundamental frequency signals contained in multiple speech frames in a speech signal. The speech signal refers to a signal containing human voice, for example, a signal containing the speech sounds of "Zhang San" and "Li Si". The speech signal is composed of multiple speech frames, and a speech signal can be formed by combining multiple speech frames continuously or discontinuously in time. On this basis, a speech frame can be a short time segment of a speech signal, for example, a time segment lasting tens of milliseconds to hundreds of milliseconds.
[0030] In actual applications, the voice signal may contain the voices of multiple users or environmental noise. For example, when the speaker is inputting voice through the microphone, other members may be speaking. Due to the mixing of the voices of other members, the voice signal may be audio data of multiple voices. On this basis, the spectrogram obtained by Fourier transforming the mixed audio data is a mixed spectrum. For example, the time domain waveform signal is converted into a frequency domain signal by short-time Fourier transform (STFT) or fast Fourier transform (FFT) to obtain a spectrogram. The spectrogram can be regarded as a two-dimensional image, the horizontal axis of the spectrogram can represent the frequency, and the vertical axis can represent the amplitude value. It should be noted that the spectrogram can be in units of one speech frame, and one speech frame can correspond to one spectrogram.
[0031] The baseband signal in the spectrum diagram can be the lowest frequency of the periodic signal. The baseband signal is equivalent to the excitation signal emitted by the speaker's vocal cords. The vocal tract modulates the excitation signal to obtain the speaker's voice signal. The baseband signal of the voice is related to the frequency of the vocal cord vibration, and each person's physiological characteristics (such as the length and thickness of the vocal cords) and vocal behavior (such as pronunciation habits) may affect the frequency of the vocal cord vibration. In other words, different users correspond to different baseband signals. In this way, the baseband signal can reflect the user's voice characteristics to a certain extent.
[0032] Since the spectrogram may be a mixed spectrogram, a spectrogram may contain multiple baseband signals. For example, speech frame 1 contains baseband signal m and baseband signal n, and speech frame 2 contains baseband signal a and baseband signal b. This means that if the spectrogram of a speech frame contains multiple baseband signals, it means that there are multiple speakers in the speech frame (for example, speaker a and speaker b).
[0033] On this basis, in order to determine the fundamental frequency signal corresponding to the speaker and provide a reference for the subsequent enhancement of the voice signal, the target fundamental frequency signal can be determined according to the statistical results of the fundamental frequency signal. The target fundamental frequency signal is the fundamental frequency signal corresponding to the speaker. Among them, the statistical results can be the number of times each fundamental frequency signal appears and in which voice frame it appears. In other embodiments, the statistical results can also be the intensity of each fundamental frequency signal (such as the amplitude value), which can be used to characterize the vibration energy size in the voice frame.
[0034] In box 204, method 200 may include determining the target speech frame and the filtering operation corresponding to the candidate baseband signal contained in the target speech frame based on the comparison result of the candidate baseband signal contained in the target speech frame in the multiple speech frames and the target baseband signal. Since the purpose of speech signal enhancement is to enhance the voice of the speaker on the one hand and filter out the speech sounds of other users and various noises on the other hand. On this basis, it can be determined whether to filter out the target speech frame or filter out some baseband signals in the target speech frame (i.e., filter out the voices of other speakers in the target speech frame) based on the comparison result or matching result of the candidate baseband signal contained in the target speech frame and the target baseband signal. In other words, the filtering operation can be used to indicate whether to filter out a certain speech frame in the speech signal or whether to filter out some baseband signals in a certain speech frame.
[0035] In box 206, method 200 may include enhancing the unfiltered target speech frame and the unfiltered candidate baseband signal in the unfiltered target speech frame based on the filtering operation. After determining the filtering operation corresponding to each speech frame in the speech signal in the above manner, the retained speech frame and the retained baseband signal can be determined. On this basis, the retained speech frame can be enhanced and the retained baseband signal can be enhanced. Among them, enhancement refers to the operation used to improve the quality of the speech frame, and the operation of speech enhancement includes but is not limited to noise suppression, echo cancellation, volume adjustment, etc. These operations can be selected and combined according to specific application requirements. In an example, it can include using a filter to enhance the baseband signal of the speech frame and suppress interfering speech, etc.
[0036] In this way, the speech signal corresponding to the speaker can be determined based on the statistical results of the baseband signal, and on this basis, the speech signal corresponding to the speaker can be enhanced, thereby improving the clarity and intelligibility of the speaker's voice, allowing the audience to receive and understand the speaker's information more smoothly, and improving the quality and clarity of the speech signal in a complex environment. In addition, by filtering out interfering audio and interfering signals, the workload of subsequent audio processing can be saved, thereby improving the efficiency of audio processing.
[0037] In some embodiments, an audio signal may be acquired by a microphone or a microphone array. The original audio signal may contain speech segments and non-speech segments. On this basis, speech separation technology may be used to accurately locate the starting point and technical point of the speech signal from the audio signal, thereby determining the speech signal in the audio signal. For example, voice activity detection technology (VAD) may be used to identify speech segments and non-speech segments from the audio signal. The non-speech segment may be a silent segment or a noise segment.
[0038] Generally, speech signals are non-stationary signals, and the vocal cords of users vibrate regularly when speaking, so the speech signals are relatively fixed in a short time range. That is, the speech signal has short-term stability. On this basis, a segment of speech signal can be divided into several frames for short-term analysis. It is generally believed that speech signal segments within the range of tens of milliseconds, such as 10ms-50ms, are quasi-steady-state processes, so the frame length of each frame can be tens of milliseconds, and there may be partial overlap between consecutive frames. In some embodiments, in order to better reduce spectrum leakage and meet the periodicity requirements of Fourier transform, the speech signal can be windowed. For example, the speech signal can be windowed using a window function. Among them, the window function can include but is not limited to a Hanning window, a Hamming window, and a Blackman window.
[0039] After performing frame division and windowing operations on the speech signal, multiple speech frames can be obtained. By performing Fourier transform on multiple speech frames respectively, the spectrogram corresponding to each speech frame can be determined. In this way, the spectrogram corresponding to the speech signal can be obtained by splicing the spectrograms of each frame together along the time axis. For example, the spectrogram corresponding to the speech frame can be determined by performing short-time Fourier transform on the speech frame.
[0040] Figure 3 A schematic diagram of determining a spectrogram corresponding to a speech frame according to some embodiments of the present invention is shown. The speech signal collected by the microphone may be a time domain signal, and the abscissa of the time domain waveform corresponding to the speech signal may be time, and the ordinate may be amplitude. The time domain waveform may be used to reflect the process of speech signal changing over time and the bullying of speech energy. By performing Fourier transform on the speech signal, the time domain signal may be converted into a frequency domain signal to determine the corresponding spectrogram.
[0041] For example, the corresponding spectrogram can be determined by a short-time Fourier transform algorithm. Among them, the short-time Fourier transform can be based on the idea of time-frequency localization, and a window function (the function slides on the time axis) is used to process the signal in segments. In each time window, assuming that the signal is stable (or pseudo-stable), the signal in the time window is Fourier transformed to obtain the spectrum information in the time window. By continuously moving the window function, the spectrum information of the signal in different time periods can be obtained. It can be understood that the horizontal axis of the spectrogram can represent the frequency and the vertical axis can represent the amplitude value. The spectrogram can be used to represent the changes in the amplitude of each frequency component of the speech signal over time. It can be understood that the baseband signal contained in the speech frame can be determined based on the spectrogram of a speech frame. Among them, the baseband signal can be the lowest frequency of the periodic signal.
[0042] In practical applications, the process of collecting audio signals may be affected by background noise. Based on this, the audio signal can be denoised before being processed. For example, echo cancellation can be used to remove echoes or reflected sounds in the audio signal. In other embodiments, a noise gating algorithm can also be used to reduce the impact of noise on the audio signal. Of course, pre-trained deep neural networks or recurrent neural networks can also be used to identify and process noise in audio signals, thereby improving the purity of the audio signal and improving the sound quality of the audio signal, providing reliable data support for subsequent audio signal processing.
[0043] In some embodiments, the number of baseband signals contained in each speech frame is different, and the amplitude values corresponding to the baseband signals are also different. On this basis, there may be multiple baseband signals, so it is necessary to determine the target baseband signal according to the statistical results of the baseband signals in the multiple baseband signals. For example, the baseband signal with the highest frequency of occurrence or the most frequent occurrence can be used as the target baseband signal. Among them, the highest frequency of occurrence can refer to the largest number of frames containing the baseband signal. For example, the speech signal includes 100 speech frames, of which 80 speech frames contain baseband signal a, and 30 speech frames contain baseband signal b. In this case, baseband signal a can be used as the target baseband signal, that is, baseband signal a can be the baseband signal of the main speaker or the main speaker.
[0044] In this way, in a noisy environment with multiple people speaking, by determining the target fundamental frequency signal, the system can more accurately identify the speech content of the main speaker. In addition, the target fundamental frequency signal can also be used to optimize the quality of the speech signal to make it clearer and more understandable.
[0045] It is understandable that in order to determine a more accurate and more representative target baseband signal, the target baseband signal can also be determined by combining information such as the spectrum characteristics or energy distribution of the speech signal. Among them, the spectrum characteristics can reflect the energy distribution of the speech signal at different frequencies to a certain extent. By analyzing the spectrum characteristics, key information such as the harmonic structure and resonance peak of the speech signal can be determined, which helps to more accurately identify the baseband signal. In other embodiments, a more accurate baseband signal can also be determined by combining features such as the duration of the speech signal and the sound intensity of the speech signal.
[0046] Since there are a large number of baseband signals, in order to determine the target baseband signal more clearly and directly, the representative baseband signal in each speech frame can be determined in frame units. For example, when a speech frame contains multiple baseband signals, the baseband signal with an amplitude value greater than a preset amplitude value threshold among the multiple baseband signals can be determined according to the spectrum diagram as the representative baseband signal in the speech frame. According to the statistical results of the representative baseband signals of all speech frames in the speech signal, the target baseband signal can be determined. Table 1 below shows the statistical results of representative baseband signals. As shown in Table 1, the speech signal includes 16 speech frames, the number of frames in which baseband signal a appears is 12, the number of frames in which baseband signal b appears is 2, and the number of frames in which baseband signal c appears is 3. In this way, baseband signal a is the global dominant sound, such as the voice of the speaker.
[0047] Table 1 Statistical results of representative baseband signals
[0048]
[0049]
[0050] Figure 4 FIG. 2 shows a schematic diagram of enhancing a speech signal in some embodiments of the present specification. Figure 4 As shown, the audio signal can be acquired by real-time acquisition or by reading from an audio library. The audio signal may include a speech signal and a non-speech signal. On this basis, the audio signal can be processed by speech separation technology such as VAD technology to determine the speech signal contained in the audio signal and the multiple continuous speech frames contained in the speech signal. The endpoints in VAD refer to the critical points of change between silence and effective speech signals. Using VAD technology, the starting point and the end point corresponding to the speech segment can be found, the speech period and the non-speech period can be distinguished, and silence, noise, etc. can be removed. After determining the speech signal, the spectrum corresponding to the speech frame can be determined by performing Fourier transform on each speech frame of the speech signal. The spectrum can contain one or more baseband signals.
[0051] like Figure 4As shown, the global dominant baseband signal (i.e., the target baseband signal) can be determined in the above manner, for example, the global dominant baseband signal can be determined according to the statistical results of the baseband signal. The global dominant baseband signal can be the baseband signal corresponding to the speaker. For example, in a conference scenario, the speaker can be a leader who presides over the meeting or an employee who is speaking. The existence of a baseband signal can represent a user who is speaking, and the presence of multiple baseband signals in a voice frame can represent that at the moment corresponding to the voice frame, multiple users are speaking. Based on this, after determining the global dominant baseband signal, it can be determined whether to filter the frame, that is, whether to filter out the voice frame, or whether to filter some signals in the voice frame according to the number of baseband signals contained in the voice frame and the amplitude value of the baseband signal. It can be understood that a baseband signal can represent the voice of a user, and the number of baseband signals contained in the voice frame can be used to determine whether the voice frame is in a scene where multiple people are speaking or multiple people are speaking.
[0052] like Figure 4 As shown, when it is determined that there are multiple people speaking in a scene, that is, a voice frame contains multiple baseband signals, it can be determined whether to filter the voice frame according to whether the voice frame contains a global dominant baseband signal. If the voice frame contains a global dominant baseband signal, the voice frame is retained and other baseband signals contained in the voice frame are filtered out. In other words, the voices of users other than the speaker in the voice frame can be filtered out, and only the speaker's voice is retained. And when the voice frame does not contain a global dominant baseband signal, the voice frame can be filtered out. In this way, the voice frame that does not contain the speaker's voice can be filtered out, thereby reducing the interference of background noise, the voices of other participants or other non-target sounds. On the other hand, by filtering out unnecessary voice frames, the amount of data for subsequent audio processing tasks can be reduced.
[0053] like Figure 4 As shown, when it is determined that there is no multi-person speaking scene, that is, when a speech frame contains only one baseband signal, the filtering operation corresponding to the speech frame can be determined according to the amplitude value of the baseband signal. For example, it can be determined whether to filter the speech frame according to the comparison result between the amplitude value of the baseband signal and the amplitude value of the target baseband signal. In one example, if the amplitude value of the baseband signal is the same as the amplitude value of the target baseband signal or the difference between the amplitude value of the baseband signal and the amplitude value of the target baseband signal is less than a preset difference threshold, the speech frame can be retained.
[0054] In practice, there may be a situation where the speaker is replaced within a period of time (for example, in a meeting, the host needs to temporarily add some information). In this case, it is also necessary to retain and enhance the voice of the replaced speaker, that is, the supplementary speech signal of the non-speaker. In this scenario, the user can set a preset baseband signal threshold according to the scenario and the frequency of the human voice, for example, it can be set to 200Hz, 250Hz, etc. If it is greater than the preset baseband signal threshold, it indicates that the user's voice corresponding to the baseband signal is the required voice in this scenario.
[0055] After determining the retained speech frame and the baseband signal retained in the speech frame in the above manner, the enhanced speech signal can be determined by performing an inverse Fourier transform on the speech frame. The type of the inverse Fourier transform can be determined according to the type of the Fourier transform, that is, the inverse Fourier transform corresponds to the Fourier transform.
[0056] In this way, the speech signal corresponding to the speaker can be determined based on the statistical results of the baseband signal, and on this basis, the speech signal corresponding to the speaker can be enhanced, thereby improving the clarity and intelligibility of the speaker's voice, allowing the audience to receive and understand the speaker's information more smoothly, and improving the quality and clarity of the speech signal in a complex environment. In addition, by filtering out interfering audio and interfering signals, the workload of subsequent audio processing can be saved, thereby improving the efficiency of audio processing.
[0057] Generally, different users have different voiceprint features. Voiceprint features can characterize the features of the sound emitted by an object. Voiceprints are not only specific, but also relatively stable. On this basis, the voiceprint features corresponding to the voice signal can also be combined to assist the filtering operation of the voice signal. Among them, the voiceprint features can include physical properties such as sound intensity, sound quality, and pitch, and can also include personal characteristics such as voice speed, speech habits (such as nasal tone, sentence punctuation, speech rhythm), and spoken language. In some embodiments, the voiceprint features corresponding to the voice signal can be extracted using a pre-trained neural network model. Among them, the neural network model can be a network model based on Mel cepstral coefficients or a Gaussian mixture-universal background model. After determining the voiceprint features corresponding to the voice signal, the filtering operations corresponding to different voice frames can be determined in combination with the baseband signal. For example, if the voiceprint features corresponding to a certain voice frame have a high similarity with the voiceprint features corresponding to the speaker (for example, higher than a preset similarity threshold), and the voice frame also contains the target baseband signal corresponding to the speaker, then the voice frame can be retained. In this way, the filtering operation of the speech frame can be determined more comprehensively by combining multiple factors, thereby ensuring the accuracy of the filtering operation.
[0058] Figure 5Schematic diagram showing exemplary application scenarios of some embodiments of this specification. Figure 5 , the speaker 504, such as a teacher, stands on the podium and speaks to the microphone 502. The microphone 502 can collect the corresponding audio signal. Among them, the position of the microphone can be located in the upper right corner or the upper left corner of the podium, and this specification does not limit the position of the microphone. In the teaching scenario, the audio signal not only contains the teacher's voice but also may include background noises such as the class bell, the sound of moving tables and chairs, and may also include the voices of other students. On this basis, the audio signal can be processed by an electronic device connected to the microphone 502 network, for example, the voice signal corresponding to the teacher in the audio signal is enhanced in the above manner, thereby improving the clarity and intelligibility of the teacher's voice, so that students can receive and understand the teacher's speech content more smoothly. It can be understood that the voice signal enhancement method provided in this specification can also be applied to other scenarios, such as business cooperation negotiations, multi-person meetings, shopping mall broadcasts and other scenarios, and this specification does not make specific restrictions here.
[0059] Figure 6 A schematic block diagram of an example device 600 that can be used to implement an embodiment of the present specification is shown. As shown, the device 600 includes a computing unit 601, which can perform various appropriate actions and processes according to computer program instructions stored in a read-only memory (ROM) 602 or loaded from a storage unit 608 into a random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the device 600 can also be stored. The computing unit 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0060] A number of components in the device 600 are connected to the I / O interface 605, including: an input unit 606, such as a keyboard, a mouse, etc.; an output unit 607, such as various types of displays, speakers, etc.; a storage unit 608, such as a disk, an optical disk, etc.; and a communication unit 609, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 606 allows the device 600 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0061] The computing unit 601 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 601 performs the various methods and processes described above, such as method 200. For example, in some embodiments, the method 200 may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 608. In some embodiments, part or all of the computer program may be loaded and / or installed on the device 600 via the ROM 602 and / or the communication unit 609. When the computer program is loaded into the RAM 603 and executed by the computing unit 601, one or more steps of the method 200 described above may be performed. Alternatively, in other embodiments, the computing unit 601 may be configured to perform the method 200 in any other appropriate manner (e.g., by means of firmware).
[0062] The functions described above herein may be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip systems (SOCs), load programmable logic devices (CPLDs), and the like.
[0063] The program code for implementing the method of this specification can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that the program code, when executed by the processor or controller, causes the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code can be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.
[0064] In the context of this specification, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In addition, although each operation is depicted in a specific order, this should be understood as requiring such operations to be performed in the specific order shown or in a sequential order, or requiring that all illustrated operations should be performed to obtain the desired result. In certain environments, multitasking and parallel processing may be advantageous. Likewise, although several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this specification. Certain features described in the context of separate embodiments may also be implemented in a single implementation in combination. Conversely, various features described in the context of a single implementation may also be implemented in multiple implementations individually or in any suitable sub-combination.
[0065] Although the subject matter has been described in language specific to structural features and / or methodological logical actions, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. On the contrary, the specific features and actions described above are merely example forms of implementing the claims.
Claims
1. A method for enhancing a speech signal, comprising: Determine a target fundamental frequency signal based on statistical results corresponding to a plurality of fundamental frequency signals included in a plurality of speech frames in the speech signal; Determine, based on a comparison result between a candidate baseband signal included in a target speech frame among the multiple speech frames and the target baseband signal, a filtering operation corresponding to the target speech frame and the candidate baseband signal included in the target speech frame; as well as Based on the filtering operation, the unfiltered target speech frame and the unfiltered candidate baseband signal in the unfiltered target speech frame are enhanced.
2. The method according to claim 1, wherein based on the comparison result between the candidate baseband signal contained in the target speech frame among the multiple speech frames and the target baseband signal, determining the filtering operation corresponding to the target speech frame and the candidate baseband signal contained in the target speech frame comprises: Determining the number of candidate baseband signals included in the target speech frame; as well as Based on the number and the comparison result between the candidate baseband signal and the target baseband signal, the target speech frame and the filtering operation corresponding to the candidate baseband signal included in the target speech frame are determined.
3. The method according to claim 1, wherein determining the target baseband signal based on statistical results corresponding to multiple baseband signals included in multiple speech frames in the speech signal comprises: Determine a plurality of baseband signals included in the plurality of speech frames based on a plurality of spectrograms corresponding to the plurality of speech frames; as well as Based on the statistical results of the multiple baseband signals, a target baseband signal is determined.
4. The method according to claim 3, wherein determining the target baseband signal based on the statistical results of the plurality of baseband signals comprises: Determining a plurality of baseband signals included in each of the plurality of speech frames; Determine a baseband signal whose amplitude value is greater than a first preset amplitude threshold value among the multiple baseband signals included in each speech frame; as well as The target baseband signal is determined based on the number of frames in which the baseband signal having an amplitude value greater than a first preset amplitude threshold appears.
5. The method according to claim 1, wherein based on the comparison result between the candidate baseband signal contained in the target speech frame among the multiple speech frames and the target baseband signal, determining the filtering operation corresponding to the target speech frame and the candidate baseband signal contained in the target speech frame comprises: Determining whether the target speech frame includes a plurality of candidate baseband signals; In response to the presence of a candidate baseband signal in the target speech frame, determining a comparison result between the candidate baseband signal contained in the target speech frame and a second preset amplitude threshold; as well as In response to the candidate fundamental frequency signal included in the target speech frame being greater than the second preset amplitude threshold, the target speech frame is not filtered.
6. The method according to claim 5, further comprising: Acquire the human voice features corresponding to the target speech frame; as well as Based on the human voice feature, the second preset amplitude threshold is determined.
7. The method according to claim 5, wherein based on the comparison result between the candidate baseband signal contained in the target speech frame among the multiple speech frames and the target baseband signal, determining the target speech frame and the filtering operation corresponding to the candidate baseband signal contained in the target speech frame comprises: In response to the target speech frame including a plurality of candidate baseband signals, determining whether the target baseband signal exists in the plurality of candidate baseband signals; as well as In response to the target baseband signal existing in the plurality of candidate baseband signals, the target speech frame is not filtered and candidate baseband signals other than the target baseband signal in the target speech frame are filtered.
8. The method according to claim 7, wherein based on the comparison result between the candidate baseband signal contained in the target speech frame among the multiple speech frames and the target baseband signal, determining the target speech frame and the filtering operation corresponding to the candidate baseband signal contained in the target speech frame comprises: In response to the target baseband signal not existing in the plurality of candidate baseband signals, the target speech frame is filtered.
9. The method according to claim 1, further comprising: Determining the voice signal contained in the audio signal by performing voice detection on the audio signal; as well as A plurality of speech frames of the speech signal are determined by performing frame processing on the speech signal.
10. The method according to claim 3, further comprising: A plurality of spectrograms corresponding to the plurality of speech frames are determined by performing Fourier transform on the plurality of speech frames.
11. A device for enhancing a speech signal, comprising: A target baseband signal determination unit is configured to determine a target baseband signal based on statistical results corresponding to a plurality of baseband signals included in a plurality of speech frames in a speech signal; a filtering operation determining unit, configured to determine a filtering operation corresponding to the target speech frame and the candidate baseband signal included in the target speech frame based on a comparison result between the candidate baseband signal included in the target speech frame among the multiple speech frames and the target baseband signal; as well as The signal enhancement unit is configured to enhance the unfiltered target speech frame and the unfiltered candidate baseband signal in the unfiltered target speech frame based on the filtering operation.
12. A processor, configured to execute the method according to any one of claims 1 to 10.
13. A computing device comprising: Memory; as well as A processor according to claim 12.