Speech noise reduction method, apparatus, and related apparatus
Patent Information
- Application Number
- PCT/CN2025/072469
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-08
- Filing Date
- 2025-01-15
- Publication Date
- 2025-10-02
AI Technical Summary
In the existing technology, after a voice call is made with a different speaker, the noise reduction performance of the target speaker extraction (TSE) algorithm is severely degraded, resulting in a serious degradation of the call quality and even the inability to continue the call.
By using sensors in the terminal device to detect changes in the distance between the speaker and the microphone, it is determined whether a speaker change has occurred. When a speaker change is detected, the TSE algorithm is stopped or adjusted, the voice features of the current speaker are re-extracted, and the TSE algorithm continues to run to ensure noise reduction performance.
After the voice call is changed, the call quality can be effectively maintained, avoiding the decline in call quality caused by the reduced noise reduction performance of the TSE algorithm, and ensuring the continuity and reliability of personalized noise reduction.
Smart Images

Figure CN2025072469_02102025_PF_FP_ABST
Abstract
Description
Speech noise reduction method, device and related devices
[0001] This application claims priority to Chinese patent application No. 202410266350.6, filed on March 8, 2024, entitled “Speech Noise Reduction Method, Device and Related Devices,” the entire contents of which are incorporated herein by reference. Technical Field
[0002] The present application relates to the field of audio processing technology, and in particular to a speech noise reduction method, device and related devices. Background Art
[0003] Nowadays, mobile phones, computers, and other devices are ubiquitous in our daily lives. Voice calls are a key function of these devices. They can occur in various environments, including homes, offices, train stations, and subways. However, the quality of voice calls is easily affected by ambient noise and human voices, especially in noisy environments. To improve voice call quality, voice noise reduction is needed to reduce or eliminate ambient noise and human voice interference. Summary of the Invention
[0004] This application provides a method, device, and related devices for reducing voice noise, which can ensure the performance of voice noise reduction and thus ensure the quality of voice calls. The technical solution is as follows:
[0005] In a first aspect, a method for reducing speech noise is provided, which is applied to a terminal device, the terminal device including a microphone, and the method comprising:
[0006] A first microphone signal is collected through a microphone, the first microphone signal including a speech signal of a first speaker; a target speaker extraction (TSE) algorithm is executed based on the speech characteristics of the first speaker to reduce noise on the first microphone signal; and the TSE algorithm is stopped when a stop condition for the TSE algorithm is detected. The stop condition for the TSE algorithm includes at least one of the speaker moving away from and then approaching the microphone again, and the speaker changing.
[0007] In this way, if the speaker moves away and then reapproaches the microphone, it's likely that the caller has changed. Stopping the TSE algorithm can avoid the problem of severely corrupting the speaker's voice and negatively impacting call quality caused by continuing the TSE algorithm despite the speaker change. Alternatively, stopping the TSE algorithm after confirming the speaker has changed can still ensure call quality.
[0008] In one possible implementation, after stopping the TSE algorithm, the method further includes: collecting a second microphone signal via the microphone; determining speech characteristics of a second speaker based on the second microphone signal, wherein the second microphone signal includes the speech signal of the second speaker; and continuing to run the TSE algorithm based on the speech characteristics of the second speaker to reduce noise on the signal collected by the microphone. In other words, the terminal device adaptively extracts the speech characteristics of the current speaker and continues to run the TSE algorithm to ensure call quality.
[0009] In one possible implementation, before determining the voice features of the second speaker based on the second microphone signal, the method further includes: determining, based on the second microphone signal, that the second microphone signal contains the voice signal of the speaker. This can save computing resources of the terminal device.
[0010] In one possible implementation, the microphone includes a first microphone and a second microphone, and the first microphone and the second microphone are located at different positions of the terminal device; based on the second microphone signal, determining the presence of a speaker's voice signal in the second microphone signal includes: determining a signal-to-noise ratio of the second microphone signal based on signals collected by the first microphone and / or the second microphone in the second microphone signal; determining an inter-channel level difference (ILD) of the second microphone signal based on signals collected by the first microphone and the second microphone in the second microphone signal; and determining the presence of a speaker's voice signal in the second microphone signal based on the signal-to-noise ratio and the ILD of the second microphone signal.
[0011] In one possible implementation, when the TSE algorithm is stopped, the method further includes: running a non-TSE noise reduction algorithm to reduce noise on the second microphone signal collected by the microphone. In this way, the call quality of the terminal device can continue to be guaranteed by using other noise reduction algorithms.
[0012] Among them, the conditions for stopping the operation of the TSE algorithm include the speaker moving away from and approaching the microphone again and the speaker changing; before stopping the operation of the TSE algorithm, the method also includes: detecting whether the speaker moves away from and approaches the microphone again; when it is detected that the speaker moves away from and approaches the microphone again, collecting a third microphone signal through the microphone; based on the third microphone signal, determining whether the speaker has changed; if the speaker has changed, executing the step of stopping the operation of the TSE algorithm.
[0013] The method further includes: if the speaker has not changed, continuing to run the TSE algorithm, thereby continuously utilizing the TSE algorithm for personalized noise reduction, thereby maintaining noise reduction performance to the greatest extent possible.
[0014] In one possible implementation, the microphone in the terminal device includes a first microphone and a second microphone, and the first microphone and the second microphone are located at different positions of the terminal device; based on a third microphone signal, determining whether the speaker has changed includes: determining the signal-to-noise ratio of the third microphone signal based on the signal collected by the first microphone and / or the second microphone in the third microphone signal; determining the ILD of the third microphone signal based on the signal collected by the first microphone and the second microphone in the third microphone signal; using a TSE algorithm to reduce noise on the third microphone signal to obtain a noise-reduced signal; determining a performance indicator of the TSE algorithm based on the third microphone signal and the noise-reduced signal; and determining whether the speaker has changed based on the signal-to-noise ratio and ILD of the third microphone signal and the performance indicator of the TSE algorithm.
[0015] Determining whether the speaker has changed based on the signal-to-noise ratio and ILD of the third microphone signal and the performance indicator of the TSE algorithm includes: determining whether the third microphone signal contains the speaker's voice signal based on the signal-to-noise ratio and ILD of the third microphone signal; and if the third microphone signal contains the speaker's voice signal, determining whether the speaker has changed based on the performance indicator of the TSE algorithm. In other words, the performance indicator of the TSE algorithm is more meaningful when it is determined that someone is currently speaking, while it may be meaningless when no one is currently speaking. This ensures the reliability of this solution.
[0016] The terminal device further includes a sensor, and the method further includes: using the sensor to detect whether the speaker moves away from and then approaches the microphone again. That is, this solution can be implemented based on hardware.
[0017] In a possible implementation, the sensor in the terminal device includes one or more of a proximity light sensor, an ultrasonic sensor, a visual sensor, and an acoustic sensor.
[0018] In a second aspect, a speech noise reduction device is provided, wherein the speech noise reduction device has the function of implementing the speech noise reduction method of the first aspect. The speech noise reduction device includes one or more modules, wherein the one or more modules are used to implement the speech noise reduction method of the first aspect.
[0019] In a third aspect, a terminal device is provided, comprising a microphone, a processor, and a memory. The microphone is configured to collect a microphone signal. The memory is configured to store a program for executing the speech noise reduction method provided in the first aspect, as well as data used to implement the speech noise reduction method provided in the first aspect. The processor is configured to execute the program stored in the memory.
[0020] In a possible implementation, the terminal device may further include a communication bus, which is used to establish a connection between the processor and the memory.
[0021] In a fourth aspect, a computer-readable storage medium is provided, wherein instructions are stored in the computer-readable storage medium. When the computer-readable storage medium is run on a computer, the computer executes the speech noise reduction method described in the first aspect.
[0022] In a fifth aspect, a computer program product comprising instructions is provided, which, when executed on a computer, enables the computer to execute the speech noise reduction method described in the first aspect.
[0023] The technical effects obtained in the above-mentioned second, third, fourth and fifth aspects are similar to the technical effects obtained by the corresponding technical means in the first aspect, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] FIG1 is a schematic structural diagram of a terminal device provided in an embodiment of the present application;
[0025] FIG2 is a schematic diagram of the structure of another terminal device provided in an embodiment of the present application;
[0026] FIG3 is a schematic structural diagram of another terminal device provided in an embodiment of the present application;
[0027] FIG4 is a schematic structural diagram of another terminal device provided in an embodiment of the present application;
[0028] FIG5 is a schematic structural diagram of another terminal device provided in an embodiment of the present application;
[0029] FIG6 is a schematic structural diagram of another terminal device provided in an embodiment of the present application;
[0030] FIG7 is a schematic structural diagram of another terminal device provided in an embodiment of the present application;
[0031] FIG8 is a schematic structural diagram of another terminal device provided in an embodiment of the present application;
[0032] FIG9 is a flow chart of a method for reducing speech noise provided by an embodiment of the present application;
[0033] FIG10 is a flowchart of another method for speech noise reduction provided in an embodiment of the present application;
[0034] FIG11 is a flowchart of another method for reducing speech noise provided in an embodiment of the present application;
[0035] FIG12 is a flowchart of another method for reducing speech noise provided in an embodiment of the present application;
[0036] FIG13 is a flowchart of another method for reducing speech noise provided in an embodiment of the present application;
[0037] FIG14 is a flowchart of another method for reducing speech noise provided in an embodiment of the present application;
[0038] FIG15 is a flowchart of another method for reducing speech noise provided in an embodiment of the present application;
[0039] FIG16 is a flowchart of another method for reducing speech noise provided in an embodiment of the present application;
[0040] FIG17 is a schematic structural diagram of a speech noise reduction device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0041] In order to make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the implementation methods of the present application will be further described in detail below with reference to the accompanying drawings.
[0042] First, a brief introduction to the speech noise reduction method in related technologies is given.
[0043] With the development of speech processing technology, speech noise reduction technology has gradually evolved from the original traditional noise reduction method based on signal processing to the noise reduction method based on deep neural network (DNN). Among the traditional noise reduction methods based on signal processing (hereinafter referred to as traditional noise reduction methods), the more typical methods include estimating the power spectrum of background noise based on speech features (such as non-stationary features, cepstral features, harmonic features, etc.), and then obtaining the noise-reduced signal based on methods such as spectral subtraction or Wiener filtering. However, traditional noise reduction methods have two major problems: (1) In scenarios with low signal-to-noise ratio, noise is easily over-estimated, resulting in speech damage; (2) many non-stationary noises (such as noises such as closing doors and knocking on tables) are generally unable to be eliminated by traditional noise reduction methods.
[0044] In recent years, DNN-based noise reduction methods have been increasingly used in voice calls. For example, deep neural networks based on the fusion of recurrent neural networks (RNNs) and convolutional neural networks (CNNs), such as convolutional recurrent neural networks (CRNs) and deep complex convolutional recurrent neural networks (DCCRNs), are widely used in end-to-end products such as mobile phones and computers because they require less memory and computing power. Compared with traditional algorithms, DNN-based noise reduction methods can largely solve the two main problems of these traditional noise reduction methods. However, this type of general DNN noise reduction method can usually only eliminate ambient noise and cannot eliminate human voice interference in the environment. In some noisy environments with human voices, call quality may still be low.
[0045] To this end, a noise reduction technology based on the target speaker extraction (TSE) algorithm is proposed. By pre-collecting a segment of the target person's speech signal, the features of the speech signal are extracted as the speech features of the target person. In the speech noise reduction process, based on the speech features of the target person, the noise and human voice interference other than the target person's speech in the noisy speech are suppressed, thereby highlighting the target person's speech. TSE noise reduction technology is also commonly referred to as personalized speech enhancement (PSE) technology or personalized noise reduction (PNR) technology. TSE noise reduction technology can be implemented based on a deep neural network or based on a traditional way of processing signals, which is not limited in the embodiments of the present application. The noise reduction performance of the TSE noise reduction method is usually higher than that of other noise reduction algorithms.
[0046] In related technologies, TSE noise reduction technology is applied to mobile phone calls, pre-extracting the owner's voice. During the call, the human voice and noise in the surrounding environment will be suppressed and basically will not be transmitted to the other end of the call, effectively improving the call quality.
[0047] However, the relevant technology is only suitable for single-person calls. If the owner needs to hand the phone to a family member or other person to continue the call, since the TSE algorithm is always running based on the owner's voice characteristics during the call, the noise reduction performance of the TSE algorithm will be severely reduced after the call is changed, resulting in a serious reduction in call quality and even the inability to continue the call.
[0048] Based on this, the embodiment of the present application provides a method for reducing voice noise, which can ensure the noise reduction performance when the caller changes, thereby ensuring that the call quality is not seriously degraded. The embodiment of the present application will be described in detail below.
[0049] Figure 1 is a schematic diagram of the structure of a terminal device provided in an embodiment of the present application. Referring to Figure 1 , the terminal device 100 includes at least a microphone 101 and a processor 102. Microphone 101 and processor 102 are capable of communicating. Microphone 101 is used to collect voice signals during a call and transmit the collected signals to processor 102. Processor 102 is used to perform noise reduction on the received voice signals according to the voice noise reduction method provided in an embodiment of the present application.
[0050] In order to ensure call quality even after a call is made with a different speaker, the terminal device 100 needs to detect whether the speaker has changed during the voice call. When the speaker changes, the speaker will move away from and approach the microphone again. Therefore, the terminal device 100 needs to detect whether the speaker moves away from and approaches the microphone 101 again. This requires the terminal device 100 to detect the distance between the speaker and the microphone 101. In an embodiment of the present application, referring to FIG2 , the terminal device 100 further includes a sensor 103, and the terminal device 100 uses the sensor 103 for distance measurement, that is, uses the sensor 103 to detect whether the speaker moves away from and approaches the microphone 101 again. Among them, the sensor 103 includes one or more of a proximity light sensor, an ultrasonic sensor, a visual sensor, and an acoustic sensor. In addition, the sensor 103 may also include other sensors that can be used for distance measurement in the embodiment of the present application.
[0051] FIG3 is a schematic diagram of the structure of another terminal device provided in an embodiment of the present application. Referring to FIG3 , the terminal device 100 includes at least a microphone 101, a processor 102, and a proximity light sensor. The proximity light sensor can also communicate with the processor 102. The terminal device 100 uses the proximity light sensor to detect whether a speaker moves away from and then reapproaches the microphone 101.
[0052] In some embodiments, the proximity light sensor is used to determine the distance between the speaker and the microphone 101, and sends the distance to the processor 102. The processor 102 is used to determine whether the speaker has moved away from and then reapproached the microphone 101 based on the received distance. In other embodiments, the proximity light sensor 103 is used to determine the distance between the speaker and the microphone 101, determine a proximity flag based on the distance, and send the proximity flag to the processor 102. The processor 102 determines whether the speaker has moved away from and then reapproached the microphone 101 based on the received proximity flag. The proximity flag indicates whether the speaker has moved away from or approached the microphone 101. In addition to these two embodiments, the proximity light sensor and the processor 102 can also determine whether the speaker has moved away from and then reapproached the microphone 101 through other information exchange methods.
[0053] The proximity light sensor is used to emit a light pulse and determine the time it takes for the light pulse to be reflected by an object, and based on this time, the distance between the object and the proximity light sensor is determined. In the embodiment of the present application, this distance can be used as the distance between the speaker and the microphone 101.
[0054] In the embodiments of the application, the direction of light pulse emission from the proximity light sensor matches the direction of sound reception from microphone 101. For example, if terminal device 100 is a mobile phone, the direction of light pulse emission from the proximity light sensor is toward the direction of the ear when a person answers a call, and the direction of sound reception from microphone 101 is also toward the direction of the ear when a person answers a call. In some embodiments, the proximity light sensor and / or microphone 101 are located at the top or bottom of terminal device 100.
[0055] Taking mobile phones as an example, many phones currently use proximity sensors to detect whether the phone is close to a person's face. Based on this, when the face approaches the phone screen during a call, the screen will turn off, and when the face leaves the screen, the screen will light up.
[0056] FIG4 is a schematic diagram of the structure of another terminal device provided in an embodiment of the present application. Referring to FIG4 , the terminal device 100 includes at least a microphone 101, a processor 102, and an ultrasonic sensor. The ultrasonic sensor can also communicate with the processor 102. The terminal device 100 uses the ultrasonic sensor to detect whether the speaker moves away from and then reapproaches the microphone 101. The principle of ultrasonic sensor ranging is similar to that of proximity light sensor ranging, except that ultrasonic sensors emit ultrasonic waves, while proximity light sensors emit light pulses.
[0057] The use of ultrasonic sensors for distance measurement simulates the bat emitting ultrasonic waves for positioning. In the embodiment of the present application, the ultrasonic sensor is mainly used to detect whether a human face is close to the screen of the terminal device 100.
[0058] In some embodiments, an ultrasonic sensor is configured to alternately transmit and receive ultrasonic waves. For example, after transmitting an ultrasonic wave, the ultrasonic sensor switches from transmit mode to receive mode, thereby receiving the ultrasonic wave reflected from an object. The total time between receiving and transmitting the ultrasonic wave is inversely proportional to the object distance. The object distance here refers to the distance between the object reflecting the ultrasonic wave and the ultrasonic sensor.
[0059] Figure 5 is a schematic diagram of the structure of another terminal device provided in an embodiment of the present application. Referring to Figure 5 , the terminal device 100 includes at least a microphone 101, a processor 102, and a visual sensor. The visual sensor can also communicate with the processor 102. The terminal device 100 uses the visual sensor to detect whether a speaker moves away from and then approaches the microphone 101 again. The principle of visual sensor ranging is to determine the distance between the speaker and the microphone 101 by capturing images. In the embodiment of the present application, ranging based on a visual sensor can be implemented based on the principle of binocular ranging or based on the principle of monocular ranging. When using binocular ranging, the terminal device 100 includes at least two visual sensors. For example, the terminal device 100 includes at least two cameras, each including a visual sensor. The principle of binocular ranging is to achieve ranging by utilizing the parallax between the two visual sensors. Parallax refers to the difference in position of an object in the two images captured by two cameras simultaneously. The specific implementation methods of binocular ranging and monocular ranging can be referred to in related art and will not be elaborated in detail in the present embodiment of the present application.
[0060] The position of the visual sensor in this article is similar to the position of the proximity sensor in the proximity sensor embodiment. Taking a mobile phone as an example, the front camera on the mobile phone can be used to implement the distance measurement in this solution.
[0061] FIG6 is a schematic diagram of the structure of another terminal device provided in an embodiment of the present application. Referring to FIG6 , the terminal device 100 includes at least a microphone 101, a processor 102, and a speaker 104. The speaker 104 is also capable of communicating with the processor 102. The microphone 101 and speaker 104 function as acoustic sensors to detect whether a speaker moves away from and then reapproaches the microphone 101.
[0062] The principle of distance measurement using the speaker 104 and the microphone 101 is as follows: an audio signal is played through the speaker 104, and the audio signal reflected by an object (such as the speaker's head) is collected through the microphone 101; based on the energy of the played audio signal and the energy of the reflected audio signal, the distance between the object and the microphone 101 is determined. For example, when a person's face is close to the mobile phone, the energy of the reflected audio signal is larger, and when the person's face is away from the mobile phone, the energy of the reflected audio signal is smaller. Among them, the audio playback direction of the speaker 104 is consistent with the direction in which the microphone 101 picks up the audio. What is meant here is that the audio played by the speaker 104 should be able to be picked up by the microphone 101 after being reflected by an object, and it does not limit the speaker 104 and the microphone 101 to be directional.
[0063] The audio signal played by the speaker 104 includes the audio signal received from the other end during the call, and in one possible implementation, also includes the ultrasonic signal pre-stored in the terminal device 100. The speaker 104 can mix the pre-stored ultrasonic signal with the audio signal from the other end and play it.
[0064] In a possible implementation, the terminal device 100 further includes a power amplifier (PA), which is configured to amplify the power of the audio signal to be played and output the amplified audio signal to the speaker 104 for playback.
[0065] In some embodiments, the speaker 104 is also used to play signals after noise reduction by the processor 102, such as signals after noise reduction by the processor of the other end of the call. For another example, if the terminal device 100 has a recording function, the processor 102 is also used to reduce noise on the audio signal recorded by the terminal device 100, and the speaker 104 is also used to play the noise-reduced recorded signal. In other embodiments, the speaker 104 is also used to play other audio signals, such as audio signals within a music application.
[0066] The terminal device 100 in the embodiments of FIG. 1 to FIG. 5 may also include a speaker 104 , which is not limited in the embodiments of the present application.
[0067] The speaker 104 is located at any position of the terminal device 100. Taking the terminal device 100 as a mobile phone as an example, the speaker 104 is located at the top or bottom of the mobile phone. In some embodiments, the speaker 104 of the terminal device 100 includes a first speaker and a second speaker, and the first speaker and the second speaker are located at different positions of the terminal device 100. For example, the first speaker is located at the top of the terminal device 100 (e.g., near the first microphone hereinafter), and the second speaker is located at the bottom of the terminal device 100 (e.g., near the second microphone hereinafter).
[0068] FIG7 is a schematic diagram of the structure of another terminal device provided in an embodiment of the present application. Referring to FIG7 , the terminal device 100 includes a first microphone 1011, a second microphone 1012, a speaker 104, and a processor 102. The first microphone 1011 and the second microphone 1012 are both capable of communicating with the processor 102. In some embodiments, the speaker 104 is also capable of communicating with the processor 102. The first microphone 1011 and the second microphone 1012 are both used to collect voice signals during a call and transmit the collected signals to the processor 102. The processor 102 is used to reduce the noise of the received voice signals according to the voice noise reduction method provided in an embodiment of the present application.
[0069] The first microphone 1011 and the second microphone 1012 are located at different positions of the terminal device 100 . For example, the first microphone 1011 is located at the top of the terminal device 100 , and the second microphone 1012 is located at the bottom of the terminal device 100 .
[0070] FIG8 is a schematic diagram of the structure of another terminal device provided in an embodiment of the present application. In one possible implementation, the terminal device is any terminal device 100 in the embodiments of FIG1 to FIG7 . The terminal device includes one or more processors 801, a communication bus 802, a memory 803, and one or more communication interfaces 804.
[0071] The processor 801 is a general-purpose central processing unit (CPU), a digital signal processor (DSP), a network processor (NP), a microprocessor, or one or more integrated circuits for implementing the solution of the present application, such as an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. In one possible implementation, the PLD is a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof. In an embodiment of the present application, the processor is used to execute some or all of the steps in the speech noise reduction method provided in the embodiment of the present application. For example, it is used to execute one or more of steps 901, 902, and 903 in Figure 9 below.
[0072] Communication bus 802 is used to transmit information between the aforementioned components. In one possible implementation, communication bus 802 is divided into an address bus, a data bus, a control bus, etc. For ease of illustration, FIG8 shows only one thick line, but this does not mean that there is only one bus or only one type of bus.
[0073] In one possible implementation, the memory 803 is a read-only memory (ROM), a random access memory (RAM), an electrically erasable programmable read-only memory (EEPROM), an optical disc (including a compact disc read-only memory (CD-ROM), a compact disc, a laser disc, a digital versatile disc, a Blu-ray disc, etc.), a magnetic disk storage medium, or other magnetic storage device, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 803 exists independently and is connected to the processor 801 via the communication bus 802, or the memory 803 is integrated with the processor 801.
[0074] The communication interface 804 uses any transceiver-like device for communicating with other devices or communication networks. The communication interface 804 includes a wired communication interface and, in one possible implementation, also includes a wireless communication interface. Examples of wired communication interfaces include Ethernet interfaces. In one possible implementation, the Ethernet interface is an optical interface, an electrical interface, or a combination thereof. The wireless communication interface is a wireless local area network (WLAN) interface, a cellular network communication interface, or a combination thereof.
[0075] In some embodiments, the terminal device includes multiple processors, such as processor 801 and processor 805 shown in Figure 8. Each of these processors is a single-core processor or a multi-core processor. In one possible implementation, the processor here refers to one or more devices, circuits, and / or processing cores for processing data (such as computer program instructions).
[0076] In a specific implementation, as an embodiment, the terminal device also includes an output device 806 and an input device 807. The output device 806 communicates with the processor 801 and can display information and / or play audio in a variety of ways. For example, the output device 806 includes a speaker, a liquid crystal display (LCD), a light emitting diode (LED) display device, a cathode ray tube (CRT) display device or a projector, etc. The input device 807 communicates with the processor 801 and can receive user input in a variety of ways. For example, the input device 807 includes a microphone, a mouse, a keyboard, a touch screen device or a sensing device, etc. In an embodiment of the present application, the microphone is used to collect a microphone signal to obtain a voice signal to be noise reduced.
[0077] In some embodiments, the memory 803 is used to store program code 810 for executing the solution of the present application, and the processor 801 can execute the program code 810 stored in the memory 803. The program code 810 includes one or more software modules, and the terminal device can implement the speech noise reduction method provided in the embodiment of Figure 9 below through the processor 801 and the program code 810 in the memory 803.
[0078] The embodiments of Figures 1 to 8 above can be implemented individually or in combination. For example, combining the embodiments of Figures 3 and 7 yields a terminal device including a processor, a speaker, a proximity light sensor, and two microphones. Another example is combining the embodiments of Figures 3, 5, and 7 to yield a terminal device including a processor, a speaker, a camera, and two microphones. The various combinations are flexible and not listed here.
[0079] The device or equipment structure and business scenarios described in the embodiments of the present application are intended to more clearly illustrate the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Ordinary technicians in this field can know that with the evolution of the device or equipment structure and the emergence of new business scenarios, the technical solutions provided by the embodiments of the present application are also applicable to similar technical problems.
[0080] FIG9 is a flowchart of a method for reducing speech noise provided in an embodiment of the present application. The method is applied to a terminal device including a microphone. In one possible implementation, the terminal device is any of the terminal devices described in the embodiments of FIG1 through FIG8 above. Referring to FIG9 , the method includes the following steps.
[0081] Step 901: Collect a first microphone signal through a microphone, where the first microphone signal includes a voice signal of a first speaker.
[0082] The first microphone signal is the signal collected by the microphone after the voice call begins but before the speaker changes. Generally speaking, the first speaker after the voice call begins is the owner of the terminal device, and the owner is considered the first speaker.
[0083] Alternatively, the first microphone signal is the signal collected by the microphone after a speaker changes during a voice call. Here, "speaker change" means that the speaker moves away from and then approaches the microphone again. In this case, the first speaker could be anyone, such as the phone owner, or someone else.
[0084] It should be understood that after the voice call starts, the terminal device will execute steps 901 to 903. Each time a person is changed, the terminal device may execute steps 901 to 903 again.
[0085] In some embodiments, the terminal device includes a microphone, and the first microphone signal is a signal collected by the microphone. In other embodiments, the terminal device includes two microphones, which are respectively referred to as a first microphone and a second microphone. The first microphone signal includes a signal collected by the first microphone and a signal collected by the second microphone, and the signals collected by the first microphone and the second microphone both include a voice signal of the first speaker. In still other embodiments, the terminal device includes more than two microphones, and the first microphone signal includes signals collected by these microphones, and these signals both include a voice signal of the first speaker.
[0086] Step 902: Based on the speech characteristics of the first speaker, run a TSE algorithm to reduce noise on the first microphone signal.
[0087] Since the first microphone signal includes the voice signal of the first speaker, the terminal device runs the TSE algorithm based on the voice features of the first speaker to perform noise reduction on the first microphone signal.
[0088] The TSE algorithm refers to an algorithm for suppressing noise and vocal interference other than the target person's voice in noisy speech. The embodiments of this application do not limit the specific implementation of the TSE algorithm. In some embodiments, the TSE algorithm is implemented based on a deep neural network. The deep neural network can be a CNN, CRN, or other network. The embodiments of this application do not limit the architecture, number of network layers, training method, etc. of the deep neural network. In other embodiments, the TSE algorithm is implemented based on signal processing.
[0089] Taking the TSE algorithm implemented based on a deep neural network as an example, the terminal device inputs the speech features of a first speaker and a first microphone signal into the deep neural network to perform noise reduction on the first microphone signal. Alternatively, the terminal device performs a time-frequency transform on the first microphone signal to obtain a frequency domain representation of the first microphone signal, and then inputs the speech features of the first speaker and the frequency domain representation of the first microphone signal into the deep neural network to perform noise reduction on the first microphone signal. The time-frequency transform may include a Fourier transform and / or other time-frequency transforms, which are not limited in this embodiment of the present application.
[0090] After the terminal device runs the TSE algorithm, the algorithm directly outputs the signal after the noise of the first microphone signal, or outputs the frequency domain representation of the signal after the noise of the first microphone signal. The terminal device then performs an inverse time-frequency transform on the frequency domain representation of the signal to obtain the noise-reduced signal.
[0091] As can be seen from the above, the first speaker may be the owner of the device or someone else. In the case where the first speaker is the owner of the device, the voice features of the first speaker may be voice features pre-stored in the terminal device, or voice features extracted in real time by the terminal device. In the case where the first speaker is not the owner of the device, the voice features of the first speaker are voice features extracted in real time by the terminal device. Of course, in some embodiments, in the case where the first speaker is not the owner of the device, the terminal device may also pre-store the voice features of the first speaker. For example, if the first speaker is a family member of the owner of the device, the terminal device may pre-store the voice features of one or more of the owner's family members.
[0092] There are many ways to implement pre-storing a person's voice features in a terminal device, which is not limited in this embodiment of the present application. For example, the owner can input a command word of the intelligent assistant into the terminal device, and the terminal device extracts the voice features of the command word to obtain the owner's voice features.
[0093] If the terminal device has pre-stored voice features of multiple people, the terminal device can determine the person corresponding to the first microphone signal from among the multiple people through voice matching, and identify the determined person as the first speaker. For example, the terminal device extracts the voice features of the first microphone signal, determines the similarity between the extracted voice features and the pre-stored voice features of each person, and determines the first speaker based on the similarity. For example, the terminal device determines the voice features with the highest similarity as the voice features of the first speaker. In addition, the terminal device can also perform voice matching through other methods, which are not limited in this embodiment of the present application.
[0094] In some embodiments, the terminal device automatically starts running the TSE algorithm after a voice call begins.
[0095] In other embodiments, after a voice call starts, the terminal device starts running the TSE algorithm when it detects that there is a speaker in the signal collected by the microphone (that is, the signal collected by the microphone includes the speaker's voice signal) and extracts the voice features of the speaker. The "signal collected by the microphone" mentioned here refers to the signal collected after the voice call starts and before the TSE algorithm is run. An implementation method for the terminal device to detect whether there is a speaker in the signal collected by the microphone includes: the terminal device includes a first microphone and a second microphone, and the first microphone and the second microphone are located at different positions of the terminal device; the terminal device determines the signal-to-noise ratio (SNR) of the signal collected by the first microphone and / or the second microphone, and determines the inter-channel sound pressure difference (ILD) between the signal collected by the first microphone and the signal collected by the second microphone; based on the determined SNR and ILD, it is determined whether there is a speaker in the signal collected by the microphone.
[0096] In one possible implementation, if the SNR exceeds a signal-to-noise ratio threshold and the ILD exceeds a sound pressure difference threshold, the terminal device determines that a speaker is present in the signal collected by the microphone. If the SNR does not exceed the signal-to-noise ratio threshold and / or the ILD does not exceed the sound pressure difference threshold, the terminal device determines that a speaker is not present in the signal collected by the microphone.
[0097] If the terminal device determines two SNRs (i.e., the signal-to-noise ratios of the signals collected by the first microphone and the second microphone), then if both SNRs exceed the signal-to-noise ratio threshold and the ILD exceeds the sound pressure difference threshold, the terminal device determines that a speaker is present in the signals collected by the microphones. Alternatively, if one of the two SNRs exceeds the signal-to-noise ratio threshold and the ILD exceeds the sound pressure difference threshold, the terminal device determines that a speaker is present in the signals collected by the microphones.
[0098] The two SNRs correspond to one signal-to-noise ratio threshold, or the two SNRs correspond to two different signal-to-noise ratio thresholds. For example, the first microphone corresponds to a first signal-to-noise ratio threshold, and the second microphone corresponds to a second signal-to-noise ratio threshold. Then, the SNR of the signal collected by the first microphone is compared with the first signal-to-noise ratio threshold, and the SNR of the signal collected by the second microphone is compared with the second signal-to-noise ratio threshold.
[0099] For example, if the terminal device is a mobile phone, with the first microphone located at the bottom and the second microphone located at the top, the first microphone is closer to the human mouth than the second microphone. Therefore, the energy of the speaker's voice signal collected by the first microphone should be greater, while the energy of the speaker's voice signal collected by the second microphone should be relatively smaller. Therefore, the first signal-to-noise ratio threshold can be greater than the second signal-to-noise ratio threshold. This improves the accuracy of the terminal device's SNR-based detection of the presence of a speaker.
[0100] Alternatively, the terminal device determines the average of the two SNRs to obtain a first SNR average. If the first SNR average exceeds the signal-to-noise ratio threshold and the ILD exceeds the sound pressure difference threshold, the terminal device determines that a speaker is present in the signal collected by the microphone. If the first SNR average does not exceed the signal-to-noise ratio threshold and / or the ILD does not exceed the sound pressure difference threshold, the terminal device determines that no speaker is present in the signal collected by the microphone.
[0101] Here, an implementation method for a terminal device to determine SNR and ILD is introduced.
[0102] For example, consider a terminal device (such as a mobile phone) with the first microphone located at the bottom and the second microphone located at the top. These microphones are referred to as the bottom microphone and the top microphone, respectively. After the terminal device picks up the signal captured by the bottom microphone, it performs a Fourier transform on the signal to obtain its frequency domain representation. Assume that the frequency domain representation of the signal at the kth frequency point in the mth frame is Y(k,m). Y(k,m) serves as an input to any noise reduction algorithm in the terminal device, and the algorithm outputs the frequency domain representation of that frequency point as S(k,m). Then, the SNR(k,m) at that frequency point = S(k,m) / (Y(k,m)–S(k,m)). The terminal device then averages the SNR(k,m) for all frequency points in the mth frame to obtain the SNR for the mth frame. Alternatively, the terminal device can determine the SNR of the signal captured by the top microphone using the same method.
[0103] It should be understood that in this implementation, using any noise reduction algorithm in the terminal device to determine the SNR does not mean that the terminal device is currently running this noise reduction algorithm to perform noise reduction. For example, the terminal device may not be running any noise reduction algorithm to perform noise reduction. Of course, the terminal device may also be running a non-TSE noise reduction algorithm to perform noise reduction.
[0104] In other embodiments, the terminal device may not use a noise reduction algorithm to determine the SNR, but may use other methods to determine the SNR, for example, using any noise estimation method to determine the SNR. This solution does not limit the method for determining the SNR.
[0105] ILD can be calculated using the signals collected by the terminal device's bottom and top microphones. After the terminal device picks up the signals from the bottom and top microphones, it performs Fourier transforms to obtain the frequency domain representations of these two signals. Assuming that the mth frame and kth frequency point of the bottom and top microphones are B(k,m) and T(k,m), respectively, the ILD(k,m) at the kth frequency point is (B(k,m)–T(k,m)) / (B(k,m)+T(k,m)). The terminal device then averages the ILD(k,m) values across all frequency points to obtain the ILD for the mth frame.
[0106] The methods for calculating SNR and ILD introduced in the above two paragraphs are not intended to limit the embodiments of the present application. In some other embodiments, the terminal device may also determine SNR and ILD in other ways, and the specific methods may also refer to related technologies.
[0107] The terminal device determines that a speaker is present in the signal collected by the microphone when it detects that the SNR of one or more consecutive frames in the signal collected by the microphone exceeds the signal-to-noise ratio threshold and the ILD exceeds the sound pressure difference threshold. If the number of consecutive frames is multiple, the total number of frames in the consecutive frames is a preset value, such as 10 or 20 frames, which is not limited in this embodiment of the application.
[0108] Step 903: When a stop condition of the TSE algorithm is detected, stop the TSE algorithm. The stop condition of the TSE algorithm includes at least one of the following: the speaker moves away from and approaches the microphone again, and the speaker changes.
[0109] If the TSE algorithm's termination conditions include the speaker moving away from and then reapproaching the microphone, but not a speaker change, the terminal device detects whether the speaker moves away from and then reapproaches the microphone during a call. If this is detected, it indicates that the speaker has likely changed, and the terminal device stops running the TSE algorithm to prevent continued execution of the algorithm from significantly degrading call quality.
[0110] In some embodiments, the terminal device periodically detects whether the speaker moves away from and then approaches the microphone again, so as to detect a "speaker change" in a timely manner. The detection period can be a default duration configured by the terminal device, such as 1 second or 2 seconds, or other shorter durations.
[0111] In other embodiments, when the terminal device detects a change from someone speaking to no one speaking based on a signal collected by the microphone, or detects a decrease in the speaker's volume, it determines whether the speaker moves away from and approaches the microphone again.
[0112] The terminal device further includes a sensor, and the terminal device uses the sensor to detect whether the speaker moves away from and approaches the microphone again.
[0113] In some embodiments, a sensor detects the distance between a speaker and a microphone and determines whether the speaker has moved away from and then reapproached the microphone based on the detected distance. The sensor periodically detects the distance between the speaker and the microphone and, each time the distance is determined, compares the currently determined distance with a distance threshold to determine whether the speaker is currently approaching or moving away from the microphone. If the currently determined distance does not exceed the distance threshold, the sensor determines that the speaker is currently approaching the microphone. If the currently determined distance exceeds the distance threshold, the sensor determines that the speaker is currently moving away from the microphone. If the speaker was previously away from the microphone and is currently approaching the microphone, the sensor determines that the speaker has moved away from and then reapproached the microphone.
[0114] In one possible implementation, a first flag and a second flag represent the speaker's last and current states, respectively. The value of the first flag is either the first value or the second value. The value of the second flag is also either the first value or the second value. The first value indicates that the speaker is moving away from the microphone, while the second value indicates that the speaker is approaching the microphone. Therefore, if the value of the first flag is the first value and the value of the second flag is the second value, the sensor determines that the speaker has moved away from and then approached the microphone again.
[0115] The first value and the second value are 0 and 1, respectively. Alternatively, the first value and the second value are other different values. The first flag and the second flag can be represented by 'lastAppFlag' and 'curAppFlag', respectively, or by other symbols or character strings.
[0116] In other embodiments, the terminal device further includes a processor, and the sensor detects the distance between the speaker and the microphone, transmits the distance to the processor, and the processor determines whether the speaker has moved away from and then reapproached the microphone based on the received distance. The specific implementation is similar to the process in which the sensor determines whether the speaker has moved away from and then reapproached the microphone based on the distance, and is not repeated here.
[0117] The above-mentioned sensor includes one or more of a proximity light sensor, an ultrasonic sensor, a visual sensor and an acoustic sensor.
[0118] The specific implementation of using a single sensor among these sensors to detect whether the speaker moves away from and approaches the microphone again is described in the embodiments of FIG. 3 to FIG. 6 above, and will not be repeated here.
[0119] The specific implementation method of using two or more of these sensors to detect whether the speaker has moved away from and approached the microphone again includes: using these sensors to respectively determine the distance between the speaker and the microphone to obtain two or more distances; the processor determines the average of these distances as the target speaker distance; if the target speaker distance does not exceed the distance threshold, then the processor determines that the speaker is currently approaching the microphone; if the target speaker distance exceeds the distance threshold, then the processor determines that the speaker is currently away from the microphone; when the speaker was away from the microphone last time and is currently approaching the microphone, the processor determines that the speaker has moved away from and approached the microphone again. That is, the terminal device fuses the distance measurement results of multiple sensors to determine whether the speaker has moved away from and approached the microphone again. The terminal device can also fuse the distance measurement results of multiple sensors in other ways to determine that the speaker has moved away from and approached the microphone again, and the embodiments of the present application are not limited to this.
[0120] From the above, it can be seen that if the stopping condition of the TSE algorithm includes the speaker moving away from and approaching the microphone again, but does not include the speaker changing, then the TSE algorithm is stopped when it is detected that the speaker moves away from and approaches the microphone again.
[0121] Referring to Figure 10, Figure 10 is a flowchart of another voice noise reduction method provided by an embodiment of the present application. After stopping the TSE algorithm, the terminal device does not run any other noise reduction algorithm. It is worth noting that even if no other noise reduction algorithm is run, the call quality is higher than that of continuing to run the TSE algorithm after a real person is replaced. The sensor signal in Figure 10 includes the distance information measured by the sensor or the proximity mark obtained by the sensor based on the distance.
[0122] FIG11 is a flowchart of another method for reducing speech noise provided by an embodiment of the present application. Referring to FIG11 , after stopping the TSE algorithm, the terminal device collects a second microphone signal through the microphone and determines the speech characteristics of the second speaker based on the second microphone signal. The second microphone signal includes the speech signal of the second speaker. Based on the speech characteristics of the second speaker, the terminal device continues to run the TSE algorithm to reduce noise on the signal collected by the microphone. In other words, the terminal device adaptively extracts the speech characteristics of the current speaker, thereby continuing to run the TSE algorithm to ensure call quality.
[0123] In some embodiments, the terminal device includes a microphone, and the second microphone signal is a signal collected by the microphone. In other embodiments, the terminal device includes a first microphone and a second microphone. The second microphone signal includes a signal collected by the first microphone and a signal collected by the second microphone, and the signals collected by the first microphone and the second microphone both include a voice signal of the second speaker. In still other embodiments, the terminal device includes two or more microphones, and the second microphone signal includes signals collected by these microphones, and these signals both include a voice signal of the second speaker.
[0124] To conserve computing resources, before determining the second speaker's voice characteristics based on the second microphone signal, the terminal device determines the presence of the speaker's voice signal in the second microphone signal based on the second microphone signal. If the presence of the speaker's voice signal is determined in the second microphone signal, the terminal device then determines the second speaker's voice characteristics based on the second microphone signal.
[0125] The specific implementation method for the terminal device to determine whether the second microphone signal contains a speaker's voice signal is the same as the specific implementation method for determining whether the signal collected by the microphones contains a speaker, as described above. That is, the terminal device determines the SNR of the second microphone signal based on the signals collected by the first microphone and / or the second microphone in the second microphone signal. The terminal device determines the ILD of the second microphone signal based on the signals collected by the first microphone and the second microphone in the second microphone signal. The terminal device determines the presence of a speaker's voice signal in the second microphone signal based on the SNR and ILD of the second microphone signal.
[0126] In one possible implementation, the terminal device determines that a speaker's voice signal is present in the second microphone signal if the SNR of the second microphone signal exceeds a signal-to-noise ratio threshold and the ILD of the second microphone signal exceeds a sound pressure difference threshold. If the SNR of the second microphone signal does not exceed the signal-to-noise ratio threshold and / or the ILD of the second microphone signal does not exceed the sound pressure difference threshold, the terminal device determines that a speaker's voice signal is not present in the second microphone signal.
[0127] If the terminal device determines two SNRs (i.e., the signal-to-noise ratios of the signals collected by the first microphone and the second microphone in the second microphone signal), then if both SNRs exceed the signal-to-noise ratio threshold and the ILD exceeds the sound pressure difference threshold, the terminal device determines that the speaker's voice signal is not present in the second microphone signal. Alternatively, if one of the two SNRs exceeds the signal-to-noise ratio threshold and the ILD exceeds the sound pressure difference threshold, the terminal device determines that the speaker's voice signal is not present in the second microphone signal.
[0128] The two SNRs correspond to one signal-to-noise ratio threshold, or the two SNRs correspond to two different signal-to-noise ratio thresholds. For example, the first microphone corresponds to a first signal-to-noise ratio threshold, and the second microphone corresponds to a second signal-to-noise ratio threshold. Then, the SNR of the signal collected by the first microphone is compared with the first signal-to-noise ratio threshold, and the SNR of the signal collected by the second microphone is compared with the second signal-to-noise ratio threshold.
[0129] Alternatively, the terminal device determines the average of the two SNRs to obtain a second SNR average. If the second SNR average exceeds the signal-to-noise ratio threshold and the ILD exceeds the sound pressure difference threshold, the terminal device determines that the speaker's voice signal is present in the second microphone signal. If the second SNR average does not exceed the signal-to-noise ratio threshold and / or the ILD does not exceed the sound pressure difference threshold, the terminal device determines that the speaker's voice signal is not present in the second microphone signal.
[0130] Figure 12 is a flowchart of another voice noise reduction method provided by an embodiment of the present application. Referring to Figure 12 , when the TSE algorithm is stopped, the terminal device runs a non-TSE noise reduction algorithm (i.e., a general noise reduction algorithm) to reduce the noise of the second microphone signal collected by the microphone. Since the speaker moves away from and then approaches the microphone again, there is a high probability that the speaker has changed. Therefore, the terminal device can maintain call quality by running a non-TSE noise reduction algorithm.
[0131] Figure 13 is a flowchart of another speech noise reduction method provided by an embodiment of the present application. Referring to Figure 13 , while the TSE algorithm is stopped and the general noise reduction algorithm is executed to reduce noise on the second microphone signal collected by the microphone, the current speaker's speech features are extracted, i.e., the second speaker's speech features are extracted from the second microphone signal. After extracting the current speaker's speech features, the terminal device stops executing the general noise reduction algorithm and resumes executing the TSE algorithm.
[0132] A non-TSE noise reduction algorithm includes any noise reduction algorithm other than the aforementioned TSE algorithm. A non-TSE noise reduction algorithm is a noise reduction algorithm that is not targeted at the target person. Therefore, a non-TSE noise reduction algorithm can also be referred to as a general noise reduction algorithm. In the embodiments of the present application, the general noise reduction algorithm includes a traditional noise reduction algorithm or a general noise reduction algorithm based on a deep neural network.
[0133] As can be seen from the above, after the speaker moves away from and approaches the microphone again, the speaker may have changed, but may not have changed. Based on this, in some embodiments, the running stop condition of the TSE algorithm includes the speaker changing. Since in the process of speaker changing, the speaker usually also moves away from and approaches the microphone again, then the terminal device may also first detect whether the speaker moves away from and approaches the microphone again, and then detect whether the speaker has changed if it is detected that the speaker moves away from and approaches the microphone again. If it is determined that the speaker has changed, the terminal device stops running the TSE algorithm. Then, the running stop condition of the TSE algorithm may also include the speaker moving away from and approaching the microphone again.
[0134] Referring to Figures 14 to 16 , after detecting that the speaker has moved away from and then reapproached the microphone, and before terminating the TSE algorithm, the terminal device collects a third microphone signal through the microphone and, based on this third microphone signal, determines whether the speaker has changed. If the speaker has changed, the terminal device executes the step of terminating the TSE algorithm. If the speaker has not changed, the terminal device continues to execute the TSE algorithm. This ensures that the TSE algorithm is only terminated when a speaker has actually changed, thereby ensuring the reliability of this solution. If the speaker has not changed, the terminal device continues to use the TSE algorithm to ensure call quality.
[0135] The terminal device includes a first microphone and a second microphone, the first microphone and the second microphone being located at different locations on the terminal device. One implementation of the terminal device determining whether the speaker has changed based on the third microphone signal includes: determining the SNR of the third microphone signal based on signals collected by the first microphone and / or the second microphone in the third microphone signal; determining the ILD of the third microphone signal based on signals collected by the first microphone and the second microphone in the third microphone signal; performing noise reduction on the third microphone signal using a TSE algorithm to obtain a noise-reduced signal; determining a performance indicator of the TSE algorithm based on the third microphone signal and the noise-reduced signal; and determining whether the speaker has changed based on the SNR and ILD of the third microphone signal and the performance indicator of the TSE algorithm.
[0136] The specific implementation manner in which the terminal device determines the SNR and ILD of the third microphone signal is similar to the specific implementation manner for determining the SNR and ILD of the signal described above, and will not be repeated here.
[0137] The performance metrics of the TSE algorithm include the output-to-input energy ratio. One implementation method for a terminal device to determine the output-to-input energy ratio of the TSE algorithm is as follows: assuming that the frequency domain representation of the mth frame and the kth frequency point of the microphone signal is Y(k,m), and Y(k,m) serves as an input to the TSE algorithm. The frequency domain representation of the TSE algorithm output at this frequency point is S(k,m). Then, the output-to-input energy ratio feature of this frequency point, R_en(k,m), = S(k,m) / Y(k,m). The output-to-input energy ratio feature of the mth frame is then averaged across all frequency points.
[0138] The performance indicators of the TSE algorithm may also include other indicators in addition to the output-to-input energy ratio feature, which is not limited in this embodiment of the application. The specific implementation method of the terminal device determining other indicators can refer to the relevant technology, and this embodiment of the application will not be described in detail.
[0139] After determining the signal-to-noise ratio (SNR) and ILD of the third microphone signal, as well as the TSE algorithm's performance indicators, the terminal device then determines whether the third microphone signal contains a speaker's voice signal based on the SNR and ILD. If the third microphone signal contains a speaker's voice signal, the terminal device determines whether the speaker has changed based on the TSE algorithm's performance indicators. In other words, the TSE algorithm's performance indicators are more meaningful when someone is currently speaking, while they may be meaningless when no one is currently speaking. This ensures the reliability of this solution.
[0140] In some embodiments, the terminal device does not initially determine the performance indicator of the TSE algorithm. Instead, if the third microphone signal is determined to contain a speaker's voice signal based on the signal-to-noise ratio and ILD of the third microphone signal, the terminal device then determines the performance indicator of the TSE algorithm based on the third microphone signal and the noise-reduced signal. If the third microphone signal is determined to contain no speaker's voice signal, the terminal device does not need to determine the performance indicator of the TSE algorithm. This can conserve computing resources on the terminal device.
[0141] Another implementation method for the terminal device to determine whether the speaker has changed based on the third microphone signal is: based on the third microphone signal, determine the voice characteristics of the third speaker, the third microphone signal includes the voice characteristics of the third speaker; based on the voice characteristics of the first speaker and the voice characteristics of the third speaker, determine whether the speaker has changed.
[0142] One implementation method for the terminal device to determine whether the speaker has changed based on the voice features of the first speaker and the voice features of the third speaker is to determine the similarity between the voice features of the first speaker and the voice features of the third speaker, and determine whether the speaker has changed based on the similarity. For example, if the similarity is greater than a similarity threshold, the terminal device determines that the speaker has changed. If the similarity is not greater than the similarity threshold, the terminal device determines that the speaker has changed.
[0143] In some embodiments, the terminal device first determines whether the third microphone signal contains a speaker's voice signal based on the signal-to-noise ratio and ILD of the third microphone signal. If the third microphone signal contains a speaker's voice signal, the terminal device then determines the voice characteristics of the third speaker based on the third microphone signal.
[0144] The two aforementioned methods for determining whether the speaker has changed based on the third microphone signal can be used independently or in combination. When used in combination, for example, the terminal device uses the two methods to determine whether the speaker has changed, obtaining two results. If at least one of the two results indicates a speaker change, the terminal device determines that the speaker has changed. For example, if both methods determine that the speaker has changed, the terminal device determines that the speaker has changed. If both results indicate that the speaker has not changed, the terminal device determines that the speaker has not changed.
[0145] In some embodiments, the terminal device first uses the first implementation method to determine whether the speaker has changed. If the result obtained using the first implementation method indicates that the speaker has not changed, the second implementation method is used to determine whether the speaker has changed. If the result obtained using the first implementation method indicates that the speaker has changed, there is no need to use the second implementation method. In other embodiments, the terminal device first uses the second implementation method to determine whether the speaker has changed. If the result obtained using the second implementation method indicates that the speaker has not changed, the first implementation method is used to determine whether the speaker has changed. If the result obtained using the second implementation method indicates that the speaker has changed, there is no need to use the first implementation method. In some other embodiments, the terminal device may also use these two implementation methods at the same time to obtain two results.
[0146] In some embodiments, when the stopping condition of the TSE algorithm includes a change of the speaker, the terminal device may also detect in real time whether the speaker has changed, instead of detecting whether the speaker has moved away from and approached the microphone again. For example, the terminal device extracts the voice features of the current speaker in real time, and determines whether the speaker has changed based on the similarity between the reference voice features and the voice features of the current speaker. The reference voice features may include the voice features currently input to the TSE algorithm or the voice features extracted at the previous moment before the current moment. If the similarity between the reference voice features and the voice features of the current speaker does not exceed the similarity threshold, the terminal device determines that the speaker has changed. If the similarity between the reference voice features and the voice features of the current speaker exceeds the similarity threshold, the terminal device determines that the speaker has not changed. For another example, the terminal device determines the SNR and ILD of the signal collected by the microphone in real time and the performance indicators of the TSE algorithm running in real time, and determines whether the speaker has changed based on these data. For specific implementation methods, please refer to the relevant introduction above.
[0147] To sum up, in the embodiment of the present application, during the process of running the TSE algorithm for noise reduction, if it is detected that the speaker moves away from and approaches the microphone again, since there is a high probability that the caller has changed, then the terminal device stops running the TSE algorithm to avoid a serious decline in call quality in time.
[0148] In some embodiments, after stopping the TSE algorithm, the terminal device runs a non-TSE noise reduction algorithm to continue meeting the noise reduction requirements. In other embodiments, while running the non-TSE noise reduction algorithm, the terminal device extracts the voice features of the current speaker and, after extraction, continues to run the TSE algorithm for noise reduction. In still other embodiments, after detecting that the speaker has moved away from and then reapproached the microphone, before stopping the TSE algorithm, the terminal device determines whether the speaker has changed. If so, the TSE algorithm is stopped. If not, the TSE algorithm is continued.
[0149] FIG17 is a schematic diagram of the structure of a speech noise reduction device provided in an embodiment of the present application. The speech noise reduction device can be implemented by software, hardware, or a combination of both to form part or all of an electronic device, and the electronic device can be any terminal device shown in the embodiments of FIG1 to FIG8 . For example, the speech noise reduction device can form part or all of a processor in a terminal device. In an embodiment of the present application, the noise reduction device is included in a terminal device, which also includes a microphone. Referring to FIG17 , the speech noise reduction device includes: an acquisition module 1701 and a processing module 1702.
[0150] An acquisition module 1701 is configured to acquire a first microphone signal collected by a microphone, where the first microphone signal includes a voice signal of a first speaker;
[0151] The processing module 1702 is configured to run a TSE algorithm to reduce noise on a first microphone signal based on speech characteristics of a first speaker, and to stop running the TSE algorithm when a stop condition for the TSE algorithm is detected. The stop condition for the TSE algorithm includes at least one of the speaker moving away from and approaching the microphone again and the speaker changing.
[0152] In a possible implementation, the acquisition module 1701 is further configured to acquire a second microphone signal collected by the microphone after stopping the TSE algorithm;
[0153] The processing module 1702 is further configured to determine a speech feature of a second speaker based on a second microphone signal, where the second microphone signal includes a speech signal of the second speaker;
[0154] The processing module 1702 is further configured to continue running the TSE algorithm based on the speech characteristics of the second speaker to reduce noise on the signal collected by the microphone.
[0155] In a possible implementation, the processing module 1702 is further configured to:
[0156] Before determining the voice feature of the second speaker based on the second microphone signal, it is determined based on the second microphone signal that the voice signal of the speaker exists in the second microphone signal.
[0157] In a possible implementation, the microphone includes a first microphone and a second microphone, and the first microphone and the second microphone are located at different positions of the terminal device;
[0158] The processing module 1702 is specifically configured to:
[0159] determining a signal-to-noise ratio of the second microphone signal based on signals collected by the first microphone and / or the second microphone in the second microphone signal;
[0160] determining an inter-channel sound pressure difference ILD of the second microphone signal based on signals collected by the first microphone and the second microphone in the second microphone signal;
[0161] Based on the signal-to-noise ratio and the ILD of the second microphone signal, it is determined that a speech signal of a speaker exists in the second microphone signal.
[0162] In a possible implementation, the processing module 1702 is further configured to:
[0163] When the TSE algorithm is stopped, a non-TSE noise reduction algorithm is run to reduce noise on the second microphone signal collected by the microphone.
[0164] In a possible implementation, the TSE algorithm may be stopped when the speaker moves away from and then approaches the microphone again, or when the speaker changes. The processing module 1702 is further configured to:
[0165] Detecting if the speaker moves away from and then approaches the microphone again;
[0166] When it is detected that the speaker moves away from the microphone and approaches the microphone again, a third microphone signal collected by the microphone is obtained;
[0167] determining whether the speaker has changed based on the third microphone signal;
[0168] If the speaker changes, the step of stopping the TSE algorithm is performed.
[0169] In a possible implementation, the processing module 1702 is further configured to:
[0170] If the speaker has not changed, the TSE algorithm continues to run.
[0171] In a possible implementation, the microphone in the terminal device includes a first microphone and a second microphone, and the first microphone and the second microphone are located at different positions of the terminal device;
[0172] The processing module 1702 is specifically configured to:
[0173] determining a signal-to-noise ratio of the third microphone signal based on signals collected by the first microphone and / or the second microphone in the third microphone signal;
[0174] determining an ILD of the third microphone signal based on signals collected by the first microphone and the second microphone in the third microphone signal;
[0175] Using the TSE algorithm to reduce noise on the third microphone signal to obtain a noise-reduced signal;
[0176] determining a performance indicator of the TSE algorithm based on the third microphone signal and the noise-reduced signal;
[0177] Based on the signal-to-noise ratio and ILD of the third microphone signal and the performance index of the TSE algorithm, it is determined whether the speaker has changed.
[0178] In a possible implementation, the processing module 1702 is specifically configured to:
[0179] determining whether a speaker's speech signal exists in the third microphone signal based on a signal-to-noise ratio and an ILD of the third microphone signal;
[0180] If the third microphone signal contains a speech signal of the speaker, it is determined whether the speaker has been changed based on the performance indicator of the TSE algorithm.
[0181] In a possible implementation, the terminal device further includes a sensor, and the processing module 1702 is specifically configured to:
[0182] The sensor detects whether the speaker moves away from and then approaches the microphone again.
[0183] In a possible implementation, the sensor includes one or more of a proximity light sensor, an ultrasonic sensor, a visual sensor, and an acoustic sensor.
[0184] In an embodiment of the present application, during the process of running the TSE algorithm for noise reduction, if it is detected that the speaker moves away from and approaches the microphone again, since there is a high probability that the caller has changed, the terminal device stops running the TSE algorithm to avoid a serious decline in call quality in time.
[0185] In some embodiments, after stopping the TSE algorithm, the terminal device runs a non-TSE noise reduction algorithm to continue meeting the noise reduction requirements. In other embodiments, while running the non-TSE noise reduction algorithm, the terminal device extracts the voice features of the current speaker and, after extraction, continues to run the TSE algorithm for noise reduction. In still other embodiments, after detecting that the speaker has moved away from and then reapproached the microphone, before stopping the TSE algorithm, the terminal device determines whether the speaker has changed. If so, the TSE algorithm is stopped. If not, the TSE algorithm is continued.
[0186] It should be noted that the speech noise reduction device provided in the above embodiment only uses the division of the above functional modules as an example to illustrate when reducing the noise of the speech signal. In actual application, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the speech noise reduction device provided in the above embodiment and the speech noise reduction method embodiment are based on the same concept. The specific implementation process is detailed in the method embodiment and will not be repeated here.
[0187] An embodiment of the present application further provides a computer-readable storage medium, which stores instructions. When the computer-readable storage medium is executed on a computer, the computer executes the steps of the speech noise reduction method shown in the above method embodiment.
[0188] The present application also provides a computer program product comprising instructions that, when executed on a computer, causes the computer to perform the steps of the speech noise reduction method described in the above method embodiment. The present application also provides a computer program that, when executed on a computer, causes the computer to perform the steps of the speech noise reduction method described in the above method embodiment.
[0189] In the above embodiments, all or part of the embodiments can be implemented by software, hardware, firmware, or any combination thereof. When implemented using software, all or part of the embodiments can be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that can be accessed by a computer, or a data storage device such as a server or data center that includes one or more available media integrated therein. The available medium may be a magnetic medium (e.g., a floppy disk, a hard disk, or a magnetic tape), an optical medium (e.g., a digital versatile disc (DVD)), or a semiconductor medium (e.g., a solid state disk (SSD)). It is worth noting that the computer-readable storage medium mentioned in the embodiments of the present application may be a non-volatile storage medium, in other words, a non-transient storage medium.
[0190] It should be understood that the "at least one" mentioned herein refers to one or more, and "a plurality of" refers to two or more. In the description of the embodiments of the present application, unless otherwise specified, " / " means or, for example, A / B can mean A or B; "and / or" in this article is merely a description of the association relationship of associated objects, indicating that there can be three relationships, for example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. In addition, in order to facilitate a clear description of the technical solutions of the embodiments of the present application, in the embodiments of the present application, words such as "first" and "second" are used to distinguish between identical or similar items with substantially the same functions and effects. Those skilled in the art will understand that words such as "first" and "second" do not limit the quantity and execution order, and words such as "first" and "second" do not necessarily limit them to be different.
[0191] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.), and signals involved in the embodiments of this application are all authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, the voice signals involved in the embodiments of this application are all obtained with full authorization.
[0192] The above description is an embodiment provided for this application and is not intended to limit this application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of this application should be included in the scope of protection of this application.
Claims
1. A speech noise reduction method, characterized in that: Applied to a terminal device, the terminal device includes a microphone, and the method includes: collecting a first microphone signal through the microphone, wherein the first microphone signal includes a voice signal of a first speaker; running a target speaker extraction (TSE) algorithm based on the speech characteristics of the first speaker to reduce noise on the first microphone signal; When an operation stop condition of the TSE algorithm is detected, the operation of the TSE algorithm is stopped. The operation stop condition of the TSE algorithm includes at least one of a speaker moving away from and approaching the microphone again and a speaker changing.
2. The method according to claim 1, wherein After stopping the TSE algorithm, the method further includes: collecting a second microphone signal through the microphone; determining a speech feature of a second speaker based on the second microphone signal, wherein the second microphone signal includes a speech signal of the second speaker; Based on the speech characteristics of the second speaker, the TSE algorithm is continued to be run to reduce noise on the signal collected by the microphone.
3. The method according to claim 2, wherein Before determining the voice feature of the second speaker based on the second microphone signal, the method further includes: Based on the second microphone signal, it is determined that a speech signal of a speaker exists in the second microphone signal.
4. The method according to claim 3, wherein The microphone includes a first microphone and a second microphone, and the first microphone and the second microphone are located at different positions of the terminal device; The determining, based on the second microphone signal, that a speaker's voice signal exists in the second microphone signal includes: determining a signal-to-noise ratio of the second microphone signal based on signals collected by the first microphone and / or the second microphone in the second microphone signal; determining an inter-channel sound pressure difference ILD of the second microphone signal based on signals collected by the first microphone and the second microphone in the second microphone signal; Based on the signal-to-noise ratio and the ILD of the second microphone signal, it is determined that a speech signal of a speaker exists in the second microphone signal.
5. The method according to any one of claims 1 to 4, characterized in that In the case of stopping the TSE algorithm, the method further includes: A non-TSE noise reduction algorithm is run to reduce noise on the second microphone signal collected by the microphone.
6. The method according to any one of claims 1 to 5, wherein: The stopping conditions of the TSE algorithm include the speaker moving away from and approaching the microphone again and the speaker changing; Before stopping the TSE algorithm, the method further includes: detecting whether the speaker moves away from and approaches the microphone again; When it is detected that the speaker moves away from and approaches the microphone again, collecting a third microphone signal through the microphone; determining whether the speaker has changed based on the third microphone signal; If the speaker changes, the step of stopping the TSE algorithm is performed.
7. The method according to claim 6, wherein The method further comprises: If the speaker has not changed, the TSE algorithm continues to run.
8. The method according to claim 6 or 7, wherein: The microphone includes a first microphone and a second microphone, and the first microphone and the second microphone are located at different positions of the terminal device; The determining whether the speaker has changed based on the third microphone signal includes: determining a signal-to-noise ratio of the third microphone signal based on signals collected by the first microphone and / or the second microphone in the third microphone signal; determining an ILD of the third microphone signal based on signals collected by the first microphone and the second microphone in the third microphone signal; Using the TSE algorithm to reduce noise on the third microphone signal to obtain a noise-reduced signal; determining a performance indicator of the TSE algorithm based on the third microphone signal and the noise-reduced signal; Based on the signal-to-noise ratio and ILD of the third microphone signal and the performance index of the TSE algorithm, it is determined whether the speaker has changed.
9. The method according to claim 8, wherein The determining whether the speaker has changed based on the signal-to-noise ratio and the ILD of the third microphone signal and the performance indicator of the TSE algorithm includes: determining whether a speaker's voice signal exists in the third microphone signal based on a signal-to-noise ratio and an ILD of the third microphone signal; If the third microphone signal contains a speech signal of a speaker, it is determined whether the speaker has changed based on the performance indicator of the TSE algorithm.
10. The method according to any one of claims 1 to 9, wherein The terminal device further includes a sensor, and the method further includes: The sensor is used to detect whether the speaker moves away from and approaches the microphone again.
11. The method according to claim 10, wherein The sensor includes one or more of a proximity light sensor, an ultrasonic sensor, a visual sensor, and an acoustic sensor.
12. A speech noise reduction device, characterized in that: Included in a terminal device, the terminal device includes a microphone, and the device includes: an acquisition module, configured to acquire a first microphone signal collected by the microphone, wherein the first microphone signal includes a voice signal of a first speaker; a processing module for running a target speaker extraction (TSE) algorithm to perform noise reduction on the first microphone signal based on the speech characteristics of the first speaker, and stopping the TSE algorithm when a stop condition for the TSE algorithm is detected, wherein the stop condition for the TSE algorithm includes at least one of the speaker moving away from and approaching the microphone again and the speaker changing.
13. The device according to claim 12, wherein The acquisition module is further configured to acquire a second microphone signal collected by the microphone after stopping the TSE algorithm; The processing module is further configured to determine a speech feature of a second speaker based on the second microphone signal, wherein the second microphone signal includes a speech signal of the second speaker; The processing module is further configured to continue running the TSE algorithm to perform noise reduction on the signal collected by the microphone based on the speech characteristics of the second speaker.
14. The device according to claim 13, wherein The processing module is further configured to: Before determining the speech feature of the second speaker based on the second microphone signal, it is determined based on the second microphone signal that a speech signal of the speaker exists in the second microphone signal.
15. The device according to claim 14, wherein The microphone includes a first microphone and a second microphone, and the first microphone and the second microphone are located at different positions of the terminal device; The processing module is specifically used for: determining a signal-to-noise ratio of the second microphone signal based on signals collected by the first microphone and / or the second microphone in the second microphone signal; determining an inter-channel sound pressure difference ILD of the second microphone signal based on signals collected by the first microphone and the second microphone in the second microphone signal; Based on the signal-to-noise ratio and the ILD of the second microphone signal, it is determined that a speech signal of a speaker exists in the second microphone signal.
16. The device according to any one of claims 12 to 15, characterized in that The processing module is further configured to: When the TSE algorithm is stopped, a non-TSE noise reduction algorithm is run to reduce noise on the second microphone signal collected by the microphone.
17. The device according to any one of claims 12 to 16, characterized in that The TSE algorithm stop conditions include the speaker moving away from and approaching the microphone again and the speaker changing; the processing module is further configured to: detecting whether the speaker moves away from and approaches the microphone again; When it is detected that the speaker moves away from and approaches the microphone again, obtaining a third microphone signal collected by the microphone; determining whether the speaker has changed based on the third microphone signal; If the speaker changes, the step of stopping the TSE algorithm is performed.
18. The device according to claim 17, wherein The processing module is further configured to: If the speaker has not changed, the TSE algorithm continues to run.
19. The device according to claim 17 or 18, characterized in that The microphone includes a first microphone and a second microphone, and the first microphone and the second microphone are located at different positions of the terminal device; The processing module is specifically used for: determining a signal-to-noise ratio of the third microphone signal based on signals collected by the first microphone and / or the second microphone in the third microphone signal; determining an ILD of the third microphone signal based on signals collected by the first microphone and the second microphone in the third microphone signal; Using the TSE algorithm to reduce noise on the third microphone signal to obtain a noise-reduced signal; determining a performance indicator of the TSE algorithm based on the third microphone signal and the noise-reduced signal; Based on the signal-to-noise ratio and ILD of the third microphone signal and the performance index of the TSE algorithm, it is determined whether the speaker has changed.
20. The device according to claim 19, wherein The processing module is specifically used for: determining whether a speaker's voice signal exists in the third microphone signal based on a signal-to-noise ratio and an ILD of the third microphone signal; If the third microphone signal contains a speech signal of a speaker, it is determined whether the speaker has changed based on the performance indicator of the TSE algorithm.
21. The device according to any one of claims 12 to 20, characterized in that The terminal device further includes a sensor, and the processing module is specifically configured to: The sensor is used to detect whether the speaker moves away from and approaches the microphone again.
22. The device according to claim 21, wherein The sensor includes one or more of a proximity light sensor, an ultrasonic sensor, a visual sensor, and an acoustic sensor.
23. A terminal device, characterized in that: The terminal device includes a microphone, a processor and a memory; The microphone is used to collect microphone signals; The memory is used to store computer programs; The processor is configured to execute the computer program to implement the steps of the method according to any one of claims 1 to 11.
24. A computer-readable storage medium, characterized in that The storage medium stores a computer program, which implements the method according to any one of claims 1 to 11 when executed by a processor.
25. A computer program product, characterized in that The computer program product stores computer instructions, and when the computer instructions are executed by a processor, the method according to any one of claims 1 to 11 is implemented.