Audio signal processing method and system for suppressing echo
The audio signal processing method addresses echo suppression challenges by dynamically switching between processing modes based on speaker and noise conditions, enhancing sound quality by effectively using bone and air conduction microphones.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- SHENZHEN SHOKZ CO LTD
- Filing Date
- 2020-12-28
- Publication Date
- 2026-04-17
AI Technical Summary
Conventional echo cancellation algorithms struggle to effectively suppress echoes in audio signals collected by bone conduction microphones when speaker vibrations are strong, leading to poor voice quality due to the dominance of speaker echo signals over human speech signals.
An audio signal processing method that switches between different processing modes based on speaker signal intensity and environmental noise conditions, using bone conduction and air conduction microphones to select and process audio signals for optimal echo suppression and sound quality.
Improves echo cancellation effectiveness and enhances sound quality by dynamically selecting and processing audio signals from bone conduction and air conduction microphones, ensuring high-quality voice communication in varying environments.
Smart Images

Figure 0007847586000004 
Figure 0007847586000005 
Figure 0007847586000006
Abstract
Description
Technical Field
[0001] This specification relates to the field of audio signal processing, and particularly to an audio signal processing method and system for echo suppression.
Background Art
[0002] Currently, vibration sensors are used in electronic products such as earphones, and the application of receiving voice signals as bone conduction microphones is increasing. When a person is speaking, vibrations of the skeleton and skin are simultaneously caused, and these vibrations are bone conduction voice signals, which can be picked up by a bone conduction microphone to generate signals. The system converts the vibration signals collected by the bone conduction microphone into electrical signals or other types of signals and transmits them to an electronic device to implement the pickup function. Currently, more and more electronic devices combine air conduction microphones and bone conduction microphones with different characteristics, use the air conduction microphone to pick up external audio signals, use the bone conduction microphone to pick up vibration signals at the voice generation site, and perform voice enhancement processing and fusion on the picked-up signals. When the bone conduction microphone is arranged in an earphone or other electronic device having a speaker, the bone conduction microphone can receive not only the vibration signals during a person's speech but also the vibration signals generated when the speaker of the earphone or other electronic device is playing sound, thereby generating an echo signal. In this case, echo cancellation algorithm processing is required. Different echo signals of the speaker also affect the voice quality of the microphone. For example, when the input signal of the speaker is strong, the speaker vibration signal received by the bone conduction microphone is relatively large, much larger than the vibration signal generated during a person's speech received by the bone conduction microphone. At this time, it is difficult to cancel the echo in the bone conduction microphone with the conventional echo cancellation algorithm. At this time, the voice quality obtained by using the microphone signals output from the air conduction microphone and the bone conduction microphone as the sound source signal is relatively poor. Therefore, it is not reasonable not to consider the speaker's echo signal when selecting the sound source signal of the microphone.
[0003] Therefore, it is necessary to provide a novel audio signal processing method and system for suppressing echo, which switches the input sound source signal based on different speaker input signals, improves the echo cancellation effect, and enhances sound quality. [Overview of the project]
[0004] This specification provides novel audio signal processing methods and systems for suppressing echoes, which improve the effectiveness of echo cancellation and enhance sound quality. [Means for solving the problem]
[0005] In a first embodiment, the Specified provides an audio signal processing method for echo suppression, comprising the steps of: selecting a target audio processing mode for an electronic device from a plurality of audio processing modes based on at least a speaker signal, wherein the speaker signal is an audio signal transmitted to the electronic device by a control device; reducing echoes in at least the target audio by processing a microphone signal by the target audio processing mode to generate target audio, wherein the microphone signal is an output signal from a microphone module acquired by the electronic device, and the microphone module includes at least one first type microphone and at least one second type microphone; and outputting the target audio signal.
[0006] In some embodiments, the at least one first type microphone outputs a first audio signal, and the at least one second type microphone outputs a second audio signal, wherein the microphone signal includes the first audio signal and the second audio signal.
[0007] In some embodiments, the at least one first type microphone is used to collect human body vibration signals, and the at least one second type microphone is used to collect air vibration signals.
[0008] In some embodiments, the plurality of audio processing modes include at least a first mode for processing the first audio signal and the second audio signal, and a second mode for processing the second audio signal.
[0009] In some embodiments, the step of selecting a target audio processing mode for an electronic device from a plurality of audio processing modes based on at least a speaker signal includes the steps of generating a control signal corresponding to the speaker signal based on at least the speaker signal intensity, wherein the control signal includes a first control signal or a second control signal; and selecting a target audio processing mode corresponding to the control signal based on the control signal, wherein the first mode corresponds to the first control signal and the second mode corresponds to the second control signal.
[0010] In some embodiments, the step of generating a control signal corresponding to the speaker signal based on at least the speaker signal intensity includes the step of determining that the speaker signal intensity is lower than a preset speaker threshold and generating the first control signal, or determining that the speaker signal intensity is higher than the speaker threshold and generating the second control signal.
[0011] In some embodiments, the step of generating a control signal corresponding to the speaker signal based on at least the speaker signal intensity includes the step of generating a corresponding control signal based on the speaker signal intensity and the microphone signal.
[0012] In some embodiments, the step of generating a corresponding control signal based on the speaker signal intensity and the microphone signal includes the step of obtaining evaluation parameters for the microphone signal, wherein the evaluation parameters include environmental noise evaluation parameters, the environmental noise evaluation parameters include at least one of environmental noise level and signal-to-noise ratio, and the step of generating the control signal based on the speaker signal intensity and the evaluation parameters.
[0013] In some embodiments, the step of generating the control signal based on the speaker signal strength and the evaluation parameter includes the steps of: determining that the speaker signal strength is higher than a preset speaker threshold and generating the second control signal; determining that the speaker signal strength is lower than the speaker threshold and the environmental noise evaluation parameter is outside a preset noise evaluation range and generating the first control signal; and determining that the speaker signal strength is lower than the speaker threshold and the environmental noise evaluation parameter is within the noise evaluation range and generating the first or second control signal.
[0014] In some embodiments, the environmental noise evaluation parameter being within the noise evaluation range includes at least one of the following: the environmental noise level being lower than a preset environmental noise threshold, and the signal-to-noise ratio being higher than a preset signal-to-noise ratio threshold.
[0015] In some embodiments, the evaluation parameter further includes human voice signal intensity, and the step of generating the control signal based on the speaker signal intensity and the evaluation parameter determines that the speaker signal intensity is higher than a preset speaker threshold, the human voice signal intensity exceeds a preset human voice threshold, and the environmental noise evaluation parameter is outside a preset noise evaluation range, and generates the first control signal, and the step of determining that the speaker signal intensity is higher than the speaker threshold, the human voice signal intensity exceeds the human voice threshold, and the environmental noise evaluation parameter is within the noise evaluation range The method includes the steps of: determining that the speaker signal strength is higher than the speaker threshold and the human voice signal strength is lower than the human voice threshold, and generating the second control signal; determining that the speaker signal strength is lower than the speaker threshold and the environmental noise evaluation parameter is outside the noise evaluation range, and generating the first control signal; and determining that the speaker signal strength is lower than the speaker threshold and the environmental noise evaluation parameter is within the noise evaluation range, and generating either the first or the second control signal.
[0016] In some embodiments, the environmental noise evaluation parameter being within the noise evaluation range includes at least one of the following: the environmental noise level being lower than a preset environmental noise threshold, and the signal-to-noise ratio being higher than a preset signal-to-noise ratio threshold.
[0017] In some embodiments, the step of generating a target audio includes the step of signal processing the first audio signal and the second audio signal using a first algorithm in the first mode to generate a first target audio, or the step of signal processing the second audio signal using a second algorithm in the second mode to generate a second target audio, wherein the target audio includes one of the first target audio and the second target audio.
[0018] In some embodiments, the step of outputting the target audio includes the steps of performing a smoothing process on the target audio, and when the target audio is switched between the first target audio and the second target audio, performing the smoothing process on the connection point between the first target audio and the second target audio, and outputting the smoothed target audio.
[0019] In some embodiments, the method further includes the step of controlling the intensity of the speaker input signal of the speaker based on the control signal.
[0020] In some embodiments, the step of controlling the intensity of the speaker input signal of the speaker based on the control signal includes determining the control signal as the first control signal and reducing the intensity of the sound output by the speaker by reducing the intensity of the speaker input signal input to the speaker.
[0021] In a second embodiment, the Specified Specified provides an audio signal processing system for echo suppression, comprising: at least one storage medium storing at least one set of instructions for echo suppression audio signal processing; and at least one processor communicatively connected to the at least one storage medium, wherein the at least one processor, when the system is operating, reads the at least one set of instructions and performs the audio signal processing method for echo suppression described in the first embodiment of the Specified Specified based on the instructions of the at least one set of instructions.
[0022] As can be seen from the above technical proposals, the audio signal processing method and system for echo suppression provided herein generates a control signal corresponding to the speaker signal based on the speaker signal strength, and controls or switches the audio processing mode based on the control signal, thereby processing the sound source signal corresponding to the audio processing mode and obtaining better sound quality. The system generates a first control signal, selects a first mode, uses the first audio signal and the second audio signal as the first sound source signal, and processes the first sound source signal to obtain the first target audio. If the speaker signal exceeds the threshold, the speaker echo in the first audio signal is large. At this time, the system generates a second control signal, selects a second mode, uses the second audio signal as the second sound source signal, and processes the second sound source signal to obtain the second target audio. The method and system can switch the sound source signal of the microphone signal by switching different audio processing modes based on the speaker signal, thereby improving sound quality and ensuring that better sound quality can be obtained in different scenes.
[0023] Some of the other features of the audio signal processing methods and systems for echo suppression provided herein are described below. According to this description, the figures and examples presented below will be obvious to those skilled in the art. Creative embodiments of the audio signal processing methods and systems for echo suppression provided herein can be fully illustrated by practicing or using the methods, apparatus, and combinations described in the detailed examples below. [Brief explanation of the drawing]
[0024] To more clearly explain the technical solutions in the embodiments of this specification, the drawings that need to be used in the description of the embodiments are briefly introduced below. However, the drawings in the following description are only some embodiments of this specification, and it is obvious that those skilled in the art can obtain other drawings based on these drawings without creative efforts. [Figure 1] FIG. shows schematic diagrams of some application scenarios of an audio signal processing system for echo suppression provided according to an embodiment of this specification. [Figure 2] FIG. shows schematic diagrams of some electronic devices provided according to an embodiment of this specification. [Figure 3] FIG. shows schematic diagrams of the operation of some first modes provided according to an embodiment of this specification. [Figure 4] FIG. shows schematic diagrams of the operation of some second modes provided according to an embodiment of this specification. [Figure 5] FIG. shows some flowcharts of an audio signal processing method for echo suppression provided according to an embodiment of this specification. [Figure 6] FIG. shows some flowcharts of an audio signal processing method for echo suppression provided according to an embodiment of this specification. [Figure 7] FIG. shows some flowcharts of an audio signal processing method for echo suppression provided according to an embodiment of this specification.
Embodiments for Carrying out the Invention
[0025] The following description provides specific application scenarios and requirements of this specification so that those skilled in the art can manufacture and use the content of this specification. For those skilled in the art, various local corrections to the disclosed embodiments are obvious, and the general principles defined in this specification can be applied to other embodiments and applications without departing from the spirit and scope of this specification. Therefore, this specification is not limited to the disclosed embodiments, but is the broadest scope consistent with the claims.
[0026] The terms used herein are for illustrative purposes only and are not limiting. For example, unless otherwise specified in the context, the singular “one,” “one,” and “the said” as used herein may also include the plural. Where used herein, the terms “include,” “contain,” and / or “contain” mean that there are integers, steps, actions, elements, and / or components related thereto, but not that the existence of one or more other features, integers, steps, actions, elements, components, and / or groups is excluded, or that other features, integers, steps, actions, elements, components, and / or groups may be added to the said system / method.
[0027] Considering the following descriptions, these and other features of this specification, the operation and function of the related structural elements, and the economics of assembly and manufacture of the parts can be significantly improved. Referring to the drawings, they all form part of this specification. However, it should be clearly understood that the drawings are for illustrative and descriptive purposes only and are not intended to limit the scope of this specification. It should also be understood that the drawings are not drawn proportionally.
[0028] The flowcharts used herein illustrate the operations achieved by the systems of some embodiments herein. It should be clearly understood that the operations in the flowchart do not have to be performed in order. Conversely, the operations can be performed in reverse order or simultaneously. One or more other operations can be added to the flowchart. One or more operations can be removed from the flowchart.
[0029] Figure 1 shows a schematic diagram of an application scene of an audio signal processing system 100 for echo suppression (hereinafter simply referred to as System 100) provided according to embodiments of this specification. System 100 may include electronic equipment 200 and control equipment 400.
[0030] The electronic device 200 can store data or commands for performing the audio signal processing method for echo suppression described herein, and can perform the data and / or commands. In some embodiments, the electronic device 200 may be wireless earphones, wired earphones, smart wearable devices, such as smart glasses, smart helmets, smart watches, etc., which have voice collection and voice playback functions. The electronic device 200 may be a mobile device, tablet, laptop computer, in-vehicle device, or similar content, or any combination thereof. In some embodiments, the mobile device may include a smart home device, a smart mobile device, or similar devices, or any combination thereof. For example, the smart mobile device may include a mobile phone, a personal digital assistant, a game device, a navigation device, an ultra-mobile personal computer (UMPC), etc., or any combination thereof. In some embodiments, the smart home device may include a smart TV, a desktop computer, etc., or any combination thereof. In some embodiments, the in-vehicle device may include an in-vehicle computer, an in-vehicle television, etc.
[0031] The control device 400 may be a remote device that communicates with the electronic device 200 via wired and / or wireless audio signal communication. The control device 400 may be a device that is communicatively connected to the electronic device 200. The electronic device 200 can collect local audio signals and output them to the control device 400. The electronic device 200 can also receive and output far-end audio signals transmitted by the control device 400. The far-end audio signals are also called speaker signals. The control device 400 may be a device having voice acquisition and voice playback functions, such as a mobile phone, tablet, laptop computer, earphones, smart wearable device, in-vehicle device or similar content, or any combination thereof. For example, if the electronic device 200 is earphones, the control device 400 may be a terminal device that is communicatively connected to the earphones, such as a mobile phone or personal computer.
[0032] As shown in Figure 1, the electronic device 200 may include a microphone module 240 and a speaker 280. The microphone module 240 may be configured to acquire a local audio signal and output a microphone signal, which is an electronic signal containing audio information. The microphone module 240 may be an extra-ear microphone module or an intra-ear microphone module. For example, the microphone module 240 may be a microphone located outside the ear canal or a microphone located inside the ear canal. The microphone module 240 may include at least one first type microphone 242 and at least one second type microphone 244. The first type microphone 242 is different from the second type microphone 244. The first type microphone 242 may be a microphone that directly collects human body vibration signals, for example, a bone conduction microphone. The second type microphone 244 may be a microphone that directly collects air vibration signals, for example, an air conduction microphone. Of course, the first type of microphone 242 and the second type of microphone 244 may be other types of microphones. For example, the first type of microphone 242 may be an optical microphone, and the second type of microphone 244 may be a microphone that receives electromyographic signals, etc. Because the first type of microphone 242 is different from the second type of microphone 244, the representation of the perceived audio signal will be different, and the noise and echo components in the corresponding audio signal will be different. For convenience of explanation, in the following description, this disclosure will use a bone conduction microphone as an example of the first type of microphone 242 and an air conduction microphone as an example of the second type of microphone 244.
[0033] A bone conduction microphone may include vibration sensors such as optical vibration sensors and acceleration sensors. The vibration sensor can collect mechanical vibration signals (for example, signals from vibrations generated in the skin or skeleton when user 002 is speaking) and convert the mechanical vibration signals into electrical signals. Mechanical vibration signals, as used here, refer mainly to vibrations that propagate through solids. The bone conduction microphone collects vibration signals generated in the skeleton or skin when user 002 makes sound by contacting the user's skin or skeleton with the vibration sensor or a vibration component connected to the vibration sensor, and converts the vibration signals into electrical signals. In some embodiments, the vibration sensor may be sensitive to mechanical vibrations but not to air vibrations (i.e., the vibration sensor's ability to respond to mechanical vibrations exceeds its ability to respond to air vibrations). Because the bone conduction microphone can directly pick up vibrations from the vocal site, it can reduce the influence of ambient noise.
[0034] The air conduction microphone collects air vibration signals generated when user 002 emits sound and converts these air vibration signals into electrical signals. The air conduction microphone may be a single air conduction microphone or a microphone array consisting of two or more air conduction microphones. The microphone array may be a microphone array consisting of beams or other similar microphone arrays. The microphone array can collect sound from different directions or different locations in space.
[0035] A first type microphone 242 can output a first audio signal 243. A second type microphone 244 can output a second audio signal 245. The microphone signal includes the first audio signal 243 and the second audio signal 245. In low-noise scenes, the second audio signal 245 has better sound quality than the first audio signal 243. In scenes with relatively high ambient noise, the first audio signal 243 has higher sound quality in the low-frequency range, and the second audio signal 245 has higher sound quality in the high-frequency range. Therefore, in scenes with relatively high ambient noise, the audio signal characterized by the fusion of the first audio signal 243 and the second audio signal 245 has better sound quality. During actual use, ambient noise is constantly changing and may repeatedly switch between the low-noise and high-noise scenes.
[0036] The speaker 280 can convert electrical signals into audio signals. The speaker 280 may be configured to receive and output the speaker signal from the control device 400. For convenience of explanation, the audio signal input to the speaker 280 is defined as the speaker input signal. In some embodiments, the speaker input signal may be the speaker signal. In some embodiments, the electronic device 200 can process the speaker signal and transmit the processed audio signal to the speaker 280 for output. In this case, the speaker input signal may be the audio signal obtained by the electronic device 200 processing the speaker signal.
[0037] The sound output from the speaker 280 in response to the speaker input signal may be transmitted to the user 002 by air conduction or bone conduction. Speaker 280 may be a speaker that transmits sound by transmitting vibration signals to the human body, for example, a bone conduction speaker, or a speaker that transmits vibration signals through the air, for example, an air conduction speaker. A bone conduction speaker generates mechanical vibrations by a vibration module and conducts the mechanical vibrations into the ear through the skeleton. For example, speaker 280 can contact the user 002's head directly or through a specific medium (e.g., one or more panels) and transmit the audio signal to the user's auditory nerve by vibrating the skull. An air conduction speaker generates vibrations in the air by a vibration module and conducts the air vibrations into the ear through the air. Speaker 280 may be a combination of a bone conduction speaker and an air conduction speaker. Speaker 280 may be other types of speakers. The sound output from the speaker 280 when the speaker input signal is received may be picked up by the microphone module 240 and become an echo. The greater the intensity of the speaker input signal, the greater the sound intensity output by the speaker 280, and the stronger the echo signal becomes.
[0038] The microphone module 240 and speaker 280 may be integrated into the electronic device 200, or they may be external devices to the electronic device 200.
[0039] When the first type microphone 242 and the second type microphone 244 are operating, not only sounds emitted by user 002 but also ambient noise and sounds emitted by speaker 280 can be collected. The electronic device 200 can collect audio signals by microphone module 240 and generate the microphone signal. The microphone signal may include a first audio signal 243 and a second audio signal 245. In different scenes, the audio quality of the first audio signal 243 and the second audio signal 245 will differ. To ensure the quality of voice communication, the electronic device 200 may select a target audio processing mode from a plurality of audio processing modes depending on the different application scene, select an audio signal with better audio quality from the microphone signal as the sound source signal, and process the sound source signal according to the target audio processing mode and output it to control device 400. The sound source signal may also be the input signal of the target audio processing mode. In some embodiments, the signal processing may include noise suppression to reduce noise signals. In some embodiments, the signal processing may include echo suppression to reduce echo signals. In some embodiments, the signal processing may include noise suppression and echo suppression. In some embodiments, the signal processing may also output the sound source signal directly. For convenience of explanation, the following description will assume that the signal processing includes echo suppression. Those skilled in the art will understand that any other signal processing method falls within the scope of this specification.
[0040] The selection of the target audio processing mode by the electronic device 200 is related to the speaker signal as well as the ambient noise. In some scenarios, for example, when the speaker signal is relatively low and the sound output by the speaker 280 is also low, the audio quality of the fused audio signal, which features the first audio signal 243 output by the first type microphone 242 and the second audio signal 245 output by the second type microphone 244, is superior to the audio quality of the second audio signal 245 output by the second type microphone 244.
[0041] In some specific scenarios, such as when the speaker signal is relatively loud and the sound output by speaker 280 is also relatively loud, the influence of the first type microphone 242 on the first audio signal 243 is relatively large, resulting in a relatively large echo in the first audio signal 243. In some embodiments, the echo signal in the first audio signal 243 exceeds the voice signal of user 002. In particular, when speaker 280 is a bone conduction speaker, the echo signal in the first audio signal 243 is more pronounced. Conventional echo cancellation algorithms have difficulty canceling the echo signal in the first audio signal 243 and cannot ensure the effectiveness of echo cancellation. At this time, the audio quality of the second audio signal 245 output by the second type microphone 244 is superior to the audio quality of the audio signal characterized by the fusion of the first audio signal 243 output by the first type microphone 242 and the second audio signal 245 output by the second type microphone 244.
[0042] Therefore, the electronic device 200 can select the target audio processing mode from the plurality of audio processing modes based on the speaker signal and perform the signal processing on the microphone signal. The plurality of audio processing modes may include at least a first mode 1 and a second mode 2.
[0043] The first mode 1 can process the first audio signal 243 and the second audio signal 245. As previously stated, in some embodiments, the signal processing may include noise suppression to reduce noise signals. In some embodiments, the signal processing may include echo suppression to reduce echo signals. In some embodiments, the signal processing may include both noise suppression and echo suppression. For convenience of explanation, the following description will assume that the signal processing includes echo suppression. Those skilled in the art will understand that any other signal processing methods fall within the scope of this specification.
[0044] A second mode 2 can process the second audio signal 245. In some embodiments, the signal processing may include noise suppression to reduce noise signals. In some embodiments, the signal processing may include echo suppression to reduce echo signals. In some embodiments, the signal processing may include both noise suppression and echo suppression. For convenience of explanation, the following description will assume that the signal processing includes echo suppression. Those skilled in the art will understand that any other signal processing methods fall within the scope of this specification.
[0045] The target audio processing mode is one of the first mode 1 and the second mode 2. The plurality of audio processing modes may include other modes, such as a processing mode for signal processing the first audio signal 243.
[0046] Therefore, when the speaker signal is relatively small, in order to ensure that the audio applied to voice communication has relatively high quality, the electronic device 200 selects a first mode 1, uses a first audio signal 243 and a second audio signal 245 as sound source signals, processes the sound source signals to generate and output a first target audio 291, and applies it to voice communication. When the speaker signal is relatively large, in order to ensure that the audio applied to voice communication has relatively high quality, the electronic device 200 selects a second mode 2, uses a second audio signal 245 as a sound source signal, processes the sound source signals to generate and output a second target audio 292, and applies it to voice communication.
[0047] The electronic device 200 can execute data or commands for the audio signal processing method for echo suppression described herein, acquire the microphone signal and the speaker signal, and process the microphone signal by selecting a corresponding target audio processing mode based on the signal intensity of the speaker signal. Specifically, the electronic device 200 selects a target audio processing mode corresponding to the speaker signal intensity from a plurality of audio processing modes based on the speaker signal intensity, selects an audio signal or a combination thereof with better sound quality from the first audio signal 243 and the second audio signal 245 as a sound source signal, processes the sound source signal (e.g., echo cancellation and noise reduction processing) using a corresponding signal processing algorithm, generates and outputs a target audio, and reduces echo in the target audio. The target audio may include one of a first target audio 291 and a second target audio 292. The electronic device 200 can output the target audio to the control device 400.
[0048] As described above, in order to ensure the audio quality of the communication, the electronic device 200 controls and selects a target audio processing mode based on the speaker signal strength, thereby selecting an audio signal with better audio quality as the sound source signal of the electronic device 200, processing the sound source signal, obtaining different target audio for different usage scenarios, and ensuring that the audio quality of the target audio is optimal in each usage scenario.
[0049] Figure 2 shows a schematic diagram of the electronic device 200. The electronic device 200 can perform the audio signal processing method for echo suppression described herein. The audio signal processing method for echo suppression is described in other parts of this specification. For example, the audio signal processing method for echo suppression is described in the description of Figures 5 to 7.
[0050] As shown in Figure 2, the electronic device 200 may include a microphone module 240 and a speaker 280. In some embodiments, the electronic device 200 may also include at least one storage medium 230 and at least one processor 220.
[0051] The storage medium 230 may include a data storage device. The data storage device may be a non-temporary storage medium or a temporary storage medium. For example, the data storage device may include one or more of a magnetic disk, a read-only storage medium (ROM), or a random access storage medium (RAM). The storage medium 230 may also include at least one set of instructions stored in the data storage device for audio signal processing for echo suppression. The instructions are computer program code, and the computer program code may include programs, routines, objects, components, data structures, processes, modules, etc., that perform the audio signal processing method for echo suppression provided herein.
[0052] As shown in Figure 2, the at least one command set may include a control command issued by a control module 231 configured to generate a control signal corresponding to the speaker signal based on the speaker signal or the speaker signal and the microphone signal. The control signal may include a first control signal or a second control signal, where the first control signal corresponds to a first mode 1, and the second control signal corresponds to a second mode 2. The control signal may be any signal; for example, the first control signal may be signal 1, and the second control signal may be signal 2. The control command issued by the control module 231 may generate a corresponding control signal based on the signal strength of the speaker signal or the signal strength of the speaker signal and the evaluation parameters of the microphone. The correspondence between the control signal and the speaker signal or the speaker signal and the microphone signal will be described later. The control module 231 may also select a target audio signal processing mode corresponding to the control signal based on the control signal. If the control signal is the first control signal, the control module 231 selects the first mode 1, and if the control signal is the second control signal, the control module 231 selects the second mode 2.
[0053] In some embodiments, the at least one set of commands may also include echo processing commands issued by an echo processing module 233 configured to process the microphone signal (e.g., echo suppression, noise reduction processing) of the electronic device 200 based on the control signal. The echo processing module 233 processes the microphone signal by employing a first mode 1 if the control signal is the first control signal. The echo processing module 233 processes the microphone signal by employing a second mode 2 if the control signal is the second control signal.
[0054] The echo processing module 233 may include a first algorithm 233-1 and a second algorithm 233-8. The first algorithm 233-1 corresponds to the first control signal and the first mode 1. The second algorithm 233-8 corresponds to the second control signal and the second mode 2.
[0055] In the first mode 1, the electronic device 200 employs the first algorithm 233-1 to process the first audio signal 243 and the second audio signal 245 respectively, features the processed first audio signal 243 and the second audio signal 245, and outputs the first target audio 291.
[0056] Figure 3 shows a schematic diagram of the operation of the first mode 1 provided according to embodiments of this specification. As shown in Figure 3, in the first mode 1, the first algorithm 233-1 can receive the first audio signal 243, the second audio signal 245, and the speaker input signal. The first algorithm 233-1 can echo-cancele the first audio signal 243 based on the speaker input signal using the first echo-canceling module 233-2. The speaker input signal may be a noise-reduced audio signal. The first echo-canceling module 233-2 receives the first audio signal 243 and the speaker input signal and outputs the echo-canceled first audio signal 243. The first echo-canceling module 233-2 may be a single-microphone echo-canceling algorithm.
[0057] In some embodiments, the first algorithm 233-1 can echo-cancele the second audio signal 245 based on the speaker input signal using the second echo-canceling module 233-2. The second echo-canceling module 233-3 receives the second audio signal 245 and the speaker input signal and outputs the echo-canceled second audio signal 245. The second echo-canceling module 233-3 may be a single-microphone echo-canceling algorithm or a multi-microphone echo-canceling algorithm. The first echo-canceling module 233-2 and the second echo-canceling module 233-3 may be the same or different.
[0058] In some embodiments, the first algorithm 233-1 can use a first noise suppression module 233-4 to suppress noise in the first audio signal 243 and the second audio signal 245 after echo cancellation. The first noise suppression module 233-4 is used to suppress noise signals in the first audio signal 243 and the second audio signal 245. The first noise suppression module 233-4 receives the first audio signal 243 and the second audio signal 245 after echo cancellation and outputs the first audio signal 243 and the second audio signal 245 after noise suppression. The first noise suppression module 233-4 may reduce noise in the first audio signal 243 and the second audio signal 245 individually, or it may reduce noise in the first audio signal 243 and the second audio signal 245 simultaneously.
[0059] In some embodiments, the first algorithm 233-1 can perform feature fusion processing on the noise-reduced first audio signal 243 and the second audio signal 245 using a feature fusion module 233-5. The feature fusion module 233-5 receives the noise-reduced first audio signal 243 and the second audio signal 245. The feature fusion module 233-5 can analyze the audio quality of the first audio signal 243 and the second audio signal 245. For example, the feature fusion module 233-5 can analyze the effective audio signal intensity, noise signal intensity, echo signal intensity, and signal-to-noise ratio of the first audio signal 243 and the second audio signal 245 to determine the audio quality of the first audio signal 243 and the second audio signal 245, and then fuse the first audio signal 243 and the second audio signal 245 into the first target audio 291 and output it.
[0060] In some embodiments, the first algorithm 233-1 may also use a second noise suppression module 233-6 to suppress noise in the speaker signal. The second noise suppression module 233-6 is used to suppress noise signals in the speaker signal. The second noise suppression module 233-6 receives the speaker signal transmitted by the control device 400, removes noise signals such as far-end noise, channel noise, and electronic noise in the electronic device 200 from the speaker signal, and outputs a noise-reduced speaker signal.
[0061] Note that Figure 3 is for illustrative purposes only. Those skilled in the art should understand that in some embodiments the first algorithm 233-1 may include the feature fusion module 233-5. In another embodiment, the first algorithm 233-1 may also include any one or any combination thereof of the first echo cancellation module 233-2, the second echo cancellation module 233-3, the first noise suppression module 233-4, and the second noise suppression module 233-6.
[0062] In the second mode 2, the electronic device 200 employs a second algorithm 233-8 to process the second audio signal 245 and outputs the second target audio 292.
[0063] Figure 4 shows a schematic diagram of the operation of a second mode 2 provided according to embodiments of this specification. As shown in Figure 4, in the second mode 2, the second algorithm 233-8 can receive the second audio signal 245 and the speaker input signal. The second algorithm 233-8 can echo-erase the second audio signal 245 based on the speaker input signal using a third echo-erase module 233-9. The third echo-erase module 233-9 receives the second audio signal 245 and the speaker input signal and outputs the echo-erased second audio signal 245. The third echo-erase module 233-9 may be the same as or different from the second echo-erase module 233-3.
[0064] In some embodiments, the second algorithm 233-8 may use a third noise suppression module 233-10 to suppress noise in the second audio signal 245 after echo cancellation. The third noise suppression module 233-10 is used to suppress noise signals in the second audio signal 245. The third noise suppression module 233-10 receives the second audio signal 245 after echo cancellation and outputs the noise-suppressed second audio signal 245 as the second target audio 292. The third noise suppression module 233-10 may be the same as or different from the first noise suppression module 233-4.
[0065] In some embodiments, the second algorithm 233-8 may also use a fourth noise suppression module 233-11 to suppress noise in the speaker signal. The fourth noise suppression module 233-11 is used to suppress noise signals in the speaker signal. The fourth noise suppression module 233-11 receives the speaker signal transmitted by the control device 400, removes noise signals such as far-end noise, channel noise, and electronic noise in the electronic device 200, and outputs a noise-reduced speaker signal. The fourth noise suppression module 233-11 may be the same as or different from the second noise suppression module 233-6.
[0066] Note that Figure 4 is for illustrative purposes only. Those skilled in the art will understand that in some embodiments the second algorithm 233-8 may include any or a combination thereof of the third echo cancellation module 233-9, the third noise suppression module 233-10, and the fourth noise suppression module 233-11. In another embodiment, the second algorithm 233-8 may also directly output the second audio signal 245 without including any of the above signal processing modules.
[0067] The first mode 1 and the second mode 2 can be run at any one time to conserve computational resources. The second mode 2 may be turned off while the first mode 1 is running. The first mode 1 may be turned off while the second mode 2 is running. The first mode 1 and the second mode 2 may be run simultaneously, and when one mode is run, the other mode may update its algorithm parameters. When the electronic device 200 switches between the first mode 1 and the second mode 2, computational resources can be conserved and the calculation results can be made more accurate by commonizing some of the parameters in the first mode 1 and the second mode 2 (e.g., noise parameters obtained by the noise estimation algorithm, human speech parameters obtained by the human speech estimation algorithm, signal-to-noise ratio parameters obtained by the signal-to-noise ratio algorithm, etc.). The first algorithm 233-1 and the second algorithm 233-8 in the first mode 1 and second mode 2 can be made more accurate by sharing some of the parameters in the control commands issued by the control module 231, such as noise parameters obtained by the noise estimation algorithm, human speech parameters obtained by the human speech estimation algorithm, and signal-to-noise ratio parameters obtained by the signal-to-noise ratio algorithm.
[0068] In some embodiments, the at least one set of commands may also include microphone control commands executed by a microphone control module 235, which is configured to perform a smoothing process on the target audio and output the smoothed target audio to a control device 400. The microphone control module 235 receives a control signal and the target audio generated by the control module 231 and can perform the smoothing process on the target audio based on the control signal. If the control signal is the first control signal, a first mode 1 is executed, using the first target audio 291 output by the first algorithm 233-1 as the input signal; if the control signal is the second control signal, a second mode 2 is executed, using the second target audio 292 output by the second algorithm 233-8 as the input signal. When the microphone control module 235 switches the target audio processing mode between the first mode 1 and the second mode 2 by switching the control signal between the first control signal and the second control signal, it can perform smoothing processing on the target audio to avoid signal discontinuities caused by switching between the first target audio 291 and the second target audio 292. Specifically, the microphone control module 235 can adjust the first target audio 291 and the parameters of the first target audio 291 so that the target audio is continuous. The parameters may be stored in advance in the at least one storage medium 230. The parameters may be amplitude, phase, frequency response, etc. The adjustments may include adjustments to the volume of the target audio, adjustments to the EQ balance, adjustments to residual noise, etc. When the microphone control module 235 switches the target audio processing mode between the first mode 1 and the second mode 2, it can make the target audio a continuous signal, making it difficult for the user 002 to perceive the switch between the two.
[0069] In some embodiments, the at least one set of commands may also include speaker control commands executed by a speaker control module 237 configured to adjust the speaker processing signal to obtain the speaker input signal, and output the speaker input signal to speaker 280 to produce sound. The speaker control module 237 can receive the speaker processing signal and the control signals output by the first algorithm 233-1 and the second algorithm 233-8. If the control signal is the first control signal, the speaker control module 237 can reduce the sound output by speaker 280, reduce echoes, and improve the echo cancellation effect of the first algorithm 233-1 by controlling the speaker processing signal output by the first algorithm 233-1 to be reduced or turned off before being output to speaker 280. If the control signal is the second control signal, the speaker control module 237 does not need to adjust the speaker processing signal output by the second algorithm 233-8. When the control signal is switched between the first control signal and the second control signal, the speaker control module 237 can perform smoothing processing on the speaker processing signals output by the first algorithm 233-1 and the second algorithm 233-8 so that the sound output by the speaker 28 does not become discontinuous. When switching between the first control signal and the second control signal, the speaker control module 237 ensures as much continuity of the switching as possible so that the user 002 does not perceive the switching between the two.
[0070] In the first mode 1, the first algorithm 233-1 is biased towards the voice quality of user 002 picked up by the near-end microphone module 240. If the speaker processing signal is too loud, the speaker control module 237 processes the speaker processing signal to reduce the speaker input signal, thereby reducing the sound output by speaker 280, reducing echo, and ensuring near-end voice quality. The second algorithm 233-8 is biased towards the speaker input signal of speaker 280. The first audio signal 243 output by the first type microphone 242 is not used to ensure the voice quality and intelligibility of the speaker input signal of speaker 280.
[0071] At least one processor 220 may be communicatively connected to at least one storage medium 230, a microphone module 240, and a speaker 280. The communicative connection is any form of connection that can receive information directly or indirectly. At least one processor 220 is used to execute the at least one set of commands described above. When the system 100 is running, at least one processor 220 reads the at least one set of commands, retrieves data from the microphone module 240 and speaker 280 based on the instructions of the at least one set of commands, and executes the audio signal processing method for echo suppression provided herein. The processor 220 can execute all the steps included in the audio signal processing method for echo suppression. The processor 220 may take the form of one or more processors, and in some embodiments, the processor 220 may include one or more hardware processors, such as a microcontroller, microprocessor, reduced command set computer (RISC), application-specific integrated circuit (ASIC), application-specific command set processor (ASIP), central processing unit (CPU), graphics processing unit (GPU), physical processing unit (PPU), microcontroller unit, digital signal processor (DSP), field-programmable gate array (FPGA), advanced RISC machine (ARM), programmable logic device (PLD), any circuit or processor capable of performing one or more functions, or any combination thereof. For the purpose of illustrating the issue, only one processor 220 in the electronic device 200 is described herein. However, the electronic device 200 in this specification may also include multiple processors, and therefore the operation and / or method steps disclosed herein may be performed by one processor or jointly by multiple processors as described herein.For example, if the processor 220 of the electronic device 200 performs steps A and B in this specification, it should be understood that steps A and B may be performed jointly by two different processors 220 or separately (for example, the first processor performs step A and the second processor performs step B, or the first and second processors perform steps A and B jointly).
[0072] In some embodiments, the system 100 can select the target audio processing mode of the electronic device 200 based on the signal strength of the speaker signal. In some embodiments, the system 100 can select the target audio processing mode of the electronic device 200 based on the signal strength of the speaker signal and the microphone signal.
[0073] Figure 5 shows a flowchart of an audio signal processing method P100 for echo suppression provided according to embodiments of this specification. The method P100 is a flowchart of how a system 100 selects the target audio processing mode of an electronic device 200 based on the signal intensity of the speaker signal. As shown in Figure 5, the method P100 may include the following steps performed by at least one processor 220.
[0074] S120: Select a target audio processing mode for the electronic device 200 from the first mode 1 and the second mode 2 based on at least the speaker signal. As described above, the target audio processing mode may include one of the first mode 1 and the second mode 2. Specifically, step S120 may include the following steps:
[0075] S121: Acquire the speaker signal.
[0076] S122: Generate a control signal corresponding to the speaker signal based on at least the speaker signal strength. The control signal includes a first control signal or a second control signal. Specifically, the electronic device 200 can receive the speaker signal transmitted by the control device 400, compare the speaker signal strength with a preset speaker threshold, and generate the control signal based on the comparison result. Step S122 may include any of the following cases:
[0077] S122-2: The speaker signal intensity is determined to be lower than the speaker threshold, and the first control signal is generated, or,
[0078] S122-4: It is determined that the speaker signal strength is higher than a preset speaker threshold, and the second control signal is generated.
[0079] Step S120 may also include the following steps:
[0080] S124: Based on the control signal, the target audio processing mode corresponding to the control signal is selected. Here, the first control signal corresponds to the first mode 1. The second control signal corresponds to the second mode 2. If the control signal is the first control signal, the first mode 1 is selected, and if the control signal is the second control signal, the second mode 2 is selected.
[0081] When the speaker signal strength is higher than the speaker threshold, the first audio signal 243 and the second audio signal 245 are processed using the first algorithm 233-1 in the first mode 1. However, it is not possible to remove the echo signal from the signal while retaining a relatively good human voice signal. As a result, the audio quality of the resulting first target audio 291 is relatively poor, while the quality of the second target audio 292 obtained by processing the second audio signal 245 using the second algorithm 233-8 in the second mode 2 is relatively good. Therefore, when the speaker signal strength is higher than the speaker threshold, the electronic device 200 generates the second control signal corresponding to the second mode.
[0082] When the speaker signal strength is lower than the speaker threshold, the first audio signal 243 and the second audio signal 245 are processed using the first algorithm 233-1 in the first mode 1. This allows for the removal of echo signals from the signals while retaining relatively good human voice signals. As a result, the resulting first target audio 291 has relatively good audio quality, and the second target audio 292 obtained by processing the second audio signal 245 using the second algorithm 233-8 in the second mode 2 also has relatively good quality. Therefore, the electronic device 200 can also generate the first control signal corresponding to the first mode 1 and the second control signal corresponding to the second mode 2 when the speaker signal strength is lower than the speaker threshold.
[0083] The control signal is generated by the control module 231. Specifically, the electronic device 200 can monitor the speaker signal strength in real time and compare it with the speaker threshold. The electronic device 200 can also detect the speaker signal strength at regular intervals and compare it with the speaker threshold. If the electronic device 200 monitors that the speaker signal strength has changed significantly and the change exceeds a preset range, it can also compare the speaker signal with the speaker threshold.
[0084] The electronic device 200 generates the second control signal if the speaker signal strength is higher than the speaker threshold. The electronic device generates the first control signal if the speaker signal changes and the speaker signal strength is lower than the speaker threshold. The electronic device 200 generates the first control signal if the speaker signal strength is lower than the speaker threshold, and generates the second control signal if the speaker signal changes and the speaker signal strength is higher than the speaker threshold.
[0085] To ensure that the switching of the control signal is not perceived by user 002, the speaker threshold may be a range. The speaker threshold may be within the range in which a first speaker critical value and a second speaker critical value exist. The first speaker critical value is smaller than the second speaker critical value. The speaker signal strength being higher than the speaker threshold may include the speaker signal strength being higher than the second speaker critical value. The speaker signal strength being lower than the speaker threshold may include the speaker signal strength being lower than the first speaker critical value.
[0086] The electronic device 200 can generate the first control signal or the second control signal when the speaker signal strength is equal to the speaker threshold. The electronic device 200 can generate the second control signal when the speaker signal strength is higher than the second speaker critical value, and the electronic device 200 can generate the second control signal when the speaker signal strength falls between the first speaker critical value and the second speaker critical value. The electronic device 200 can generate the first control signal when the speaker signal strength is lower than the first speaker critical value, and the electronic device 200 can generate the first control signal when the speaker signal strength increases between the first speaker critical value and the second speaker critical value.
[0087] The electronic device 200 can obtain a control model through machine learning and input the speaker signal to the control model that outputs the control signal.
[0088] The method P100 may include the following steps, which are performed by at least one processor 220.
[0089] S140: Process the microphone signal using the target audio processing mode to generate the target audio and reduce echoes in the microphone signal. Specifically, step S140 may include any of the following cases:
[0090] S142: The control signal is determined as the first control signal, and based on the speaker input signal, the first audio signal 243 and the second audio signal 245 are subjected to signal processing and feature fusion using the first algorithm 233-1 in the first mode 1 corresponding to the first control signal to generate the first target audio 291. The specific process is described above and will not be explained in detail here.
[0091] S144: The control signal is determined as the second control signal, and the second audio signal 245 is processed based on the speaker input signal by the second algorithm 233-8 in the second mode 2 corresponding to the second control signal. The specific process is described above and will not be explained in detail here.
[0092] S160: Output the target audio. The electronic device 200 can output the target audio directly. The electronic device 200 can also perform a smoothing process on the target audio so that it is not perceived by the user 002 when the target audio is switched between the first target audio 291 and the second target audio 292. Specifically, step S160 may include the step of performing a smoothing process on the target audio and outputting the smoothed target audio.
[0093] Specifically, the electronic device 200 can perform smoothing processing on the target audio using the microphone control module 235. When the target audio is switched between the first target audio 291 and the second target audio 292, the microphone control module 235 performs the smoothing processing on the connection point between the first target audio 291 and the second target audio 292 so that the connection point transitions smoothly, that is, it can adjust the signals of the first target audio 291 and the second target audio 292.
[0094] The above method P100 may also include the following steps:
[0095] S180: Based on the control signal, the intensity of the speaker input signal of the speaker 280 is controlled. Specifically, step S180 can be performed by the speaker control module 237. Step S180 can be performed by the speaker control module 237 determining the control signal as the first control signal, and the speaker control module 237 can process the speaker processing signal and reduce the intensity of the speaker input signal input to the speaker 280, thereby reducing the intensity of the sound output by the speaker 280, reducing the echo signal in the microphone input signal, and improving the sound quality of the first target audio.
[0096] Table 1 shows the result diagram of the target audio processing mode corresponding to Figure 5. As shown in Table 1, to facilitate matching, the scene is divided into four scenes, each being: Type 1: The near-end audio signal is less than the threshold (e.g., User 002 does not make a sound) and the speaker signal does not exceed the speaker threshold; Type 2: The near-end audio signal is greater than the threshold (e.g., User 002 makes a sound) and the speaker signal does not exceed the speaker threshold; Type 3: The near-end audio signal is less than the threshold (e.g., User 002 does not make a sound) and the speaker signal exceeds the speaker threshold; and Type 4: The near-end audio signal is greater than the threshold (e.g., User 002 makes a sound) and the speaker signal exceeds the speaker threshold. Here, whether or not the near-end audio signal is greater than the threshold can be determined by the control module 231 based on the microphone signal. The fact that the near-end audio signal is greater than the threshold may mean that the audio signal strength emitted by user 002 exceeds a preset threshold. The target audio processing modes corresponding to the four scenes are such that the first and second types correspond to the first mode 1, and the third and fourth types correspond to the second mode 2.
[0097] [Table 1]
[0098] In the method P100 described above, the electronic device 200 can select a target audio processing mode based on the speaker signal in order to ensure that the audio quality processed by the target audio processing mode selected by the electronic device 200 is optimal in any scene and to ensure call quality.
[0099] In some embodiments, the selection of the target audio processing mode is related not only to the echo of the speaker signal but also to ambient noise. The ambient noise can be evaluated by at least one of the ambient noise level and the signal-to-noise ratio in the microphone signal.
[0100] Figure 6 shows a flowchart of an audio signal processing method P200 for echo suppression provided according to embodiments of this specification. The method P200 is a flowchart of how a system 100 selects the target audio processing mode of an electronic device 200 based on the signal intensity of the speaker signal and the microphone signal. Specifically, the method P200 is a flowchart of how a system 100 selects the target audio processing mode based on at least one of the ambient noise level and the signal-to-noise ratio in the speaker signal and the microphone signal. The method P200 may include the following steps performed by at least one processor 220.
[0101] S220: Select a target audio processing mode for the electronic device 200 from a first mode 1 and a second mode 2 based on at least the speaker signal. Specifically, step S220 may include the following steps:
[0102] S222: Generate a control signal corresponding to the speaker signal based on at least the speaker signal strength. The control signal includes a first control signal or a second control signal. Specifically, step S222 may also involve the electronic device 200 generating a corresponding control signal based on the speaker signal strength and noise in the microphone signal. Step S222 may include the following steps:
[0103] S222-2: Evaluation parameters for the speaker signal and the microphone signal are obtained. Here, the evaluation parameters may be environmental noise evaluation parameters for the microphone signal. The environmental noise evaluation parameters may include one of the environmental noise level and the signal-to-noise ratio. The electronic device 200 can obtain the environmental noise evaluation parameters for the microphone signal by the control module 231. Specifically, the electronic device 200 can obtain the environmental noise evaluation parameters based on at least one of the first audio signal 243 and the second audio signal 245. The electronic device 200 can obtain the environmental noise level or the signal-to-noise ratio by a noise estimation algorithm, which is not described in detail here.
[0104] S222-4: Generate the control signal based on the speaker signal strength and the environmental noise evaluation parameter. Specifically, the electronic device 200 can compare the speaker signal strength with a preset speaker threshold and the environmental noise evaluation parameter with a preset noise evaluation range, and generate the control signal based on the comparison result. Step S222-4 may include any of the following cases.
[0105] S222-5: The speaker signal strength is determined to be higher than a preset speaker threshold, and the second control signal is generated.
[0106] S222-6: The speaker signal intensity is determined to be lower than the speaker threshold, and the environmental noise evaluation parameter is outside the preset noise evaluation range, and the first control signal is generated.
[0107] S222-7: It is determined that the speaker signal intensity is lower than the speaker threshold and the environmental noise evaluation parameter is within the noise evaluation range, and the first control signal or the second control signal is generated.
[0108] Here, the state that the environmental noise evaluation parameter is within the noise evaluation range includes at least one of the following: the environmental noise level is lower than a preset environmental noise threshold, and the signal-to-noise ratio is higher than a preset signal-to-noise ratio threshold. In this case, the environmental noise is relatively small. The state that the environmental noise evaluation parameter is outside the noise evaluation range includes at least one of the following: the environmental noise level is higher than a preset environmental noise threshold, and the signal-to-noise ratio is lower than a preset signal-to-noise ratio threshold. In this case, the environmental noise is relatively large. Here, when the environmental noise evaluation parameter is outside the noise evaluation range, i.e., in a noisy environment, the voice quality of the first target audio 291 is better than that of the second target audio 292. When the environmental noise evaluation parameter is within the noise evaluation range, the voice quality of the first target audio 291 is not much different from that of the second target audio 292.
[0109] Step S220 may also include the following steps:
[0110] S224: Based on the control signal, the target audio processing mode corresponding to the control signal is selected. Here, the first control signal corresponds to the first mode 1. The second control signal corresponds to the second mode 2. If the control signal is the first control signal, the first mode 1 is selected, and if the control signal is the second control signal, the second mode 2 is selected.
[0111] When the speaker signal intensity is higher than the speaker threshold, the first algorithm 233-1 in the first mode 1 processes the first audio signal 243 and the second audio signal 245, but it is unable to remove the echo signal from the signal while leaving a relatively good human voice signal. As a result, the voice quality of the obtained first target audio 291 is relatively poor, while the quality of the second target audio 292 obtained by the second algorithm 233-8 in the second mode 2 processing the second audio signal 245 is relatively good. Therefore, when the speaker signal intensity is higher than the speaker threshold, the electronic device 200 generates the second control signal corresponding to the second mode 2, regardless of the range of the ambient noise.
[0112] When the speaker signal intensity is lower than the speaker threshold, the first algorithm 233-1 in the first mode 1 processes the first audio signal 243 and the second audio signal 245, removing the echo signal from the signal while retaining a relatively good human voice signal. As a result, the voice quality of the obtained first target audio 291 is relatively good, and the quality of the second target audio 292 obtained by the second algorithm 233-8 in the second mode 2 processing the second audio signal 245 is also relatively good. Therefore, when the speaker signal intensity is lower than the speaker threshold, the control signal generated by the electronic device 200 is related to environmental noise.
[0113] If the ambient noise level is higher than the ambient noise threshold, or if the signal-to-noise ratio is lower than the signal-to-noise ratio threshold, it indicates that the ambient noise in the microphone signal is relatively large. When the first algorithm 233-1 in the first mode 1 processes the first audio signal 243 and the second audio signal 245, it can reduce noise in the signal while leaving a relatively good human voice signal, so the resulting first target audio 291 has relatively good sound quality, while the second target audio 292 obtained by the second algorithm 233-8 in the second mode 2 processing the second audio signal 245 has worse sound quality than the first target audio 291. Therefore, the electronic device 200 generates the first control signal corresponding to the first mode 1 when the speaker signal strength is lower than the speaker threshold and the ambient noise level is higher than the ambient noise threshold or the signal-to-noise ratio is lower than the signal-to-noise ratio threshold.
[0114] Furthermore, when ambient noise is relatively low, that is, when the ambient noise evaluation parameter is within the noise evaluation range, the voice quality of the first target audio 291 is not significantly different from that of the second target audio 292. In this case, the electronic device 200 always generates the second control signal, selects the second algorithm 233-8 in the second mode 2 to process the second audio signal 245, and can reduce computational complexity and save resources while ensuring the target audio voice quality.
[0115] If the ambient noise level is lower than the ambient noise threshold or the signal-to-noise ratio is higher than the signal-to-noise ratio threshold, it indicates that the ambient noise in the microphone signal is relatively small. The audio quality of both the first target audio 291 obtained when the first algorithm 233-1 in the first mode 1 processes the first audio signal 243 and the second audio signal 245, and the second target audio 292 obtained when the second algorithm 233-8 in the second mode 2 processes the second audio signal 245, is good. Therefore, the electronic device 200 generates the first control signal or the second control signal when the speaker signal strength is lower than the speaker threshold and the ambient noise level is lower than the ambient noise threshold or the signal-to-noise ratio is higher than the signal-to-noise ratio threshold. Specifically, the electronic device 200 can determine the control signal in the current scene based on the control signal in the previous scene. In other words, if an electronic device generates a first control signal in the previous scene, the electronic device also generates a first control signal in the current scene to ensure signal continuity. The reverse is also true.
[0116] The control signal is generated by the control module 231. Specifically, the electronic device 200 can monitor the speaker signal strength and the environmental noise evaluation parameter in real time and compare them with the speaker threshold and the environmental evaluation range. The electronic device 200 can also detect the speaker signal strength and the environmental noise evaluation parameter at regular intervals and compare them with the speaker threshold and the noise evaluation range. If the electronic device 200 monitors that the speaker signal strength or the environmental noise evaluation parameter has changed significantly and the change value exceeds a preset range, it can also compare the speaker signal and the environmental noise evaluation parameter with the speaker threshold and the noise evaluation range.
[0117] To ensure that the user 002 does not perceive the switching of the control signal, the speaker threshold, the ambient noise threshold, and the preset signal-to-noise ratio may be within a range. The speaker threshold, as mentioned above, will not be described in detail here. The ambient noise threshold may be within a range in which a first noise critical value and a second noise critical value exist. The first noise critical value is smaller than the second noise critical value. The ambient noise level being higher than the ambient noise threshold includes the ambient noise level being higher than the second noise critical value. The ambient noise level being lower than the ambient noise threshold includes the ambient noise level being lower than the first noise critical value. The signal-to-noise ratio threshold may be within a range in which a first signal-to-noise ratio critical value and a second signal-to-noise ratio critical value exist. The first signal-to-noise ratio critical value is smaller than the second signal-to-noise ratio critical value. The statement that the signal-to-noise ratio is higher than the signal-to-noise ratio threshold includes the statement that the signal-to-noise ratio is higher than the second critical value of the signal-to-noise ratio. The statement that the signal-to-noise ratio is lower than the signal-to-noise ratio threshold includes the statement that the signal-to-noise ratio is lower than the first critical value of the signal-to-noise ratio.
[0118] The method P200 may include the following steps, which are performed by at least one processor 220.
[0119] S240: Reduce echoes in the microphone signal by processing the microphone signal using the target audio processing mode to generate target audio. Specifically, step S240 may include any of the following cases:
[0120] S242: The control signal is determined as the first control signal, the first mode 1 is selected, the first audio signal 243 and the second audio signal 245 are signal-processed, and the first target audio 291 is generated. Specifically, step S242 may be the same as step S142, but this will not be explained in detail here.
[0121] S244: The control signal is determined to be the second control signal, the second mode 2 is selected, echo suppression is performed on the second audio signal 245, and the second target audio 292 is generated. Specifically, step S244 may be the same as step S144, but this will not be explained in detail here.
[0122] S260: Outputs the target audio. Specifically, step S260 may coincide with step S160, but this will not be explained in detail here.
[0123] The above method P200 may also include the following steps:
[0124] S280: Based on the control signal, the intensity of the speaker input signal of the speaker 280 is controlled. Specifically, step S280 may be the same as step S180, but this will not be explained in detail here.
[0125] Table 2 shows the target audio processing mode result diagram corresponding to Figure 6. As shown in Table 2, to facilitate matching, the scene was divided into eight scenes, each being: Type 1: Near-end audio signal is less than the threshold (e.g., User 002 does not make a sound), the speaker signal does not exceed the speaker threshold, and ambient noise is relatively low; Type 2: Near-end audio signal is greater than the threshold (e.g., User 002 makes a sound), the speaker signal does not exceed the speaker threshold, and ambient noise is relatively low; Type 3: Near-end audio signal is less than the threshold (e.g., User 002 does not make a sound), the speaker signal exceeds the speaker threshold, and ambient noise is relatively low; and Type 4: Near-end audio signal is greater than the threshold (e.g., User 002 makes a sound), the speaker signal exceeds the speaker threshold... Type 5: The speaker signal exceeds the speaker threshold, ambient noise is relatively low, Type 6: The speaker signal exceeds the speaker threshold (for example, user 002 does not make a sound), ambient noise is relatively high, Type 7: The speaker signal exceeds the speaker threshold (for example, user 002 makes a sound), ambient noise is relatively high, Type 8: The speaker signal exceeds the speaker threshold (for example, user 002 makes a sound), ambient noise is relatively high. Here, whether the speaker signal is greater than the threshold can be determined by the control module 231 based on the microphone signal. The near-end audio signal being greater than the threshold may mean that the audio signal intensity emitted by user 002 exceeds a preset threshold. The target audio processing modes corresponding to the eight scenes are such that the 5th and 6th scenes correspond to the first mode 1, the 3rd, 4th, 7th, and 8th scenes correspond to the second mode 2, and the other scenes correspond to either the first mode 1 or the second mode 2.
[0126] [Table 2]
[0127] The method P200 can not only control the target audio processing mode of the electronic device 200 based on the speaker signal, but also control the target audio processing mode based on the nearby ambient noise signal, thereby ensuring optimal audio quality of the audio signal output by the electronic device 200 and maintaining call quality in different scenarios.
[0128] In some embodiments, the selection of the target audio processing mode relates not only to the echo and ambient noise of the speaker signal but also to the speech signal of user 002. The ambient noise signal can be evaluated by at least one of the ambient noise level and the signal-to-noise ratio in the microphone signal. The speech signal of user 002 can be evaluated by the human speech signal intensity in the microphone signal. The human speech signal intensity may be the human speech signal intensity obtained by a noise estimation algorithm, or it may be the intensity of the audio signal obtained after noise reduction.
[0129] Figure 7 shows a flowchart of an audio signal processing method P300 for echo suppression provided according to embodiments of this specification. The method P300 is a flowchart of how a system 100 selects a target audio processing mode for an electronic device 200 based on the signal intensity of the speaker signal and the microphone signal. Specifically, the method P300 is a flowchart of how a system selects the target audio processing mode based on at least one of the following: the human voice signal intensity in the speaker signal and the microphone signal, and the ambient noise level and signal-to-noise ratio. The method P300 may include the following steps performed by at least one processor 220.
[0130] S320: Select a target audio processing mode for the electronic device 2 from the first mode 1 and the second mode 2 based on at least the speaker signal. Specifically, step S320 may include the following steps:
[0131] S322: Generate a control signal corresponding to the speaker signal based on at least the speaker signal strength. The control signal includes a first control signal or a second control signal. Step S320 allows the electronic device 200 to generate a corresponding control signal based on the speaker signal strength, noise in the speaker signal, and human voice signal strength in the microphone signal. Specifically, step S322 may include the following steps:
[0132] S322-2: Evaluation parameters for the speaker signal and the microphone signal are obtained. Here, the evaluation parameters may include environmental noise evaluation parameters in the microphone signal and may also include human voice signal intensity in the microphone signal. The environmental noise evaluation parameters may include at least one of the environmental noise level and the signal-to-noise ratio. The electronic device 200 can obtain the environmental noise evaluation parameters and human voice signal intensity in the microphone signal by the control module 231. Specifically, the electronic device 200 can obtain the evaluation parameters based on at least one of the first audio signal 243 and the second audio signal 245. The electronic device 200 can obtain the human voice signal, the environmental noise level and the signal-to-noise ratio by a noise estimation algorithm, which is not described in detail here.
[0133] S322-4: The control signal is generated based on the speaker signal strength and the evaluation parameters. Specifically, the electronic device 200 can compare the speaker signal strength with a preset speaker threshold, compare the environmental noise evaluation parameters with a preset noise evaluation range, and compare the human voice signal strength with a preset human voice threshold, and generate the control signal based on the comparison results. Step S322-4 may include any of the following cases.
[0134] S322-5: The speaker signal strength is determined to be higher than a preset speaker threshold, the human voice signal strength exceeds the human voice threshold, and the environmental noise evaluation parameter is outside a preset noise evaluation range, and the first control signal is generated.
[0135] S322-6: The speaker signal strength is determined to be higher than the speaker threshold, the human voice signal strength exceeds the human voice threshold, and the environmental noise evaluation parameter is within the noise evaluation range, and the second control signal is generated.
[0136] S322-7: It is determined that the speaker signal strength is higher than the speaker threshold and the human voice signal strength is lower than the human voice threshold, and the second control signal is generated.
[0137] S322-8: It is determined that the speaker signal intensity is lower than the speaker threshold and the environmental noise evaluation parameter is outside the noise evaluation range, and the first control signal is generated.
[0138] S322-9: It is determined that the speaker signal intensity is lower than the speaker threshold and the environmental noise evaluation parameter is within the noise evaluation range, and the first control signal or the second control signal is generated.
[0139] The state that the environmental noise evaluation parameter is within the noise evaluation range includes at least one of the following: the environmental noise level is lower than a preset environmental noise threshold, and the signal-to-noise ratio is higher than a preset signal-to-noise ratio threshold. In this case, the environmental noise is relatively small. The state that the environmental noise evaluation parameter is outside the noise evaluation range includes at least one of the following: the environmental noise level is higher than a preset environmental noise threshold, and the signal-to-noise ratio is lower than a preset signal-to-noise ratio threshold. In this case, the environmental noise is relatively large. Here, when the environmental noise evaluation parameter is outside the noise evaluation range, i.e., in a noisy environment, the voice quality of the first target audio 291 is better than that of the second target audio 292. When the environmental noise evaluation parameter is within the noise evaluation range, the voice quality of the first target audio 291 is not much different from that of the second target audio 292. The speaker threshold, the environmental noise threshold, and the signal-to-noise ratio threshold are described above and will not be explained in detail here.
[0140] Here, the fact that the human voice signal intensity exceeds the human voice threshold means that user 002 is speaking. At this time, in order to ensure the voice quality of user 002, the electronic device 200 can generate the first control signal and reduce the speaker signal to ensure the voice quality of the first target audio 292.
[0141] The speaker signal, the environmental noise threshold, the signal-to-noise ratio, and the human voice threshold may be pre-stored in the electronic device 200.
[0142] Step S320 may also include the following steps:
[0143] S324: Based on the control signal, the target audio processing mode corresponding to the control signal is selected. Here, the first control signal corresponds to the first mode 1. The second control signal corresponds to the second mode 2. If the control signal is the first control signal, the first mode 1 is selected, and if the control signal is the second control signal, the second mode 2 is selected.
[0144] The fact that the speaker signal intensity is higher than the speaker threshold, the human voice signal intensity exceeds the human voice threshold, and the environmental noise evaluation parameter is outside the preset noise evaluation range proves that user 002 is speaking at this time, with significant echo and relatively high noise. To ensure the voice quality and intelligibility of user 002, the electronic device 200 can reduce or turn off the speaker input signal input to speaker 280, thereby reducing the echo in the microphone signal and ensuring the voice quality of the target audio. At this time, the voice quality of the first target audio 291 obtained by the first algorithm 233-1 in first mode 1 processing the first audio signal 243 and the second audio signal 245 is better than the second target audio 292 obtained by the second algorithm 233-8 in second mode 2 processing the second audio signal 245. Therefore, the electronic device 200 generates the first control signal corresponding to the first mode 1 when the speaker signal strength is higher than the speaker threshold, the human voice signal strength exceeds the human voice threshold, and the environmental noise evaluation parameter is outside the preset noise evaluation range. In such a case, the electronic device 200 can ensure the near-end user 002's ability to understand the voice quality. Although a portion of the speaker input signal is missing, the electronic device 200 can improve both voice communication quality by maintaining the voice quality and understanding of most of the speaker input signal.
[0145] When the speaker signal strength is higher than the speaker threshold, and the human voice signal strength is lower than the human voice threshold, or when the human voice signal strength exceeds the human voice threshold, and the environmental noise evaluation parameter is within a preset noise evaluation range, it is proven that user 002 is not speaking at this time, or user 002 is speaking but the noise is relatively low. At this time, the voice quality of the first target audio 291 obtained by the first algorithm 233-1 in the first mode 1 processing the first audio signal 243 and the second audio signal 245 is worse than the second target audio 292 obtained by the second algorithm 233-8 in the second mode 2 processing the second audio signal 245. Therefore, the electronic device 200 generates the second control signal corresponding to the second mode 2 when the speaker signal strength is higher than the speaker threshold, the human voice signal strength is lower than the human voice threshold, or the human voice signal strength exceeds the human voice threshold, and the environmental noise evaluation parameter is within a preset noise evaluation range. Other circumstances in step S322-4 are largely the same as those in step S222-4, so they will not be described in detail here.
[0146] The control signal is generated by the control module 231. Specifically, the electronic device 200 can monitor the speaker signal strength and the evaluation parameter in real time and compare them with the speaker threshold, the noise evaluation range, and the human voice threshold. The electronic device 200 can also detect the speaker signal strength and the evaluation parameter at regular intervals and compare them with the speaker threshold, the noise evaluation range, and the human voice threshold. If the electronic device 200 monitors that the speaker signal strength or the evaluation parameter has changed significantly and the change value exceeds a preset range, it can also compare the speaker signal and the evaluation parameter with the speaker threshold, the evaluation range, and the human voice threshold.
[0147] The method P300 described above may include the following steps performed by at least one processor 220:
[0148] S340: Reduce echoes in the microphone signal by processing the microphone signal with the target audio processing mode to generate the target audio. Specifically, step S340 may include any of the following cases:
[0149] S342: The control signal is determined as the first control signal, the first mode 1 is selected, the first audio signal 243 and the second audio signal 245 are signal-processed, and the first target audio 291 is generated. Specifically, step S342 may be the same as step S142, but this will not be explained in detail here.
[0150] S344: The control signal is determined to be the second control signal, the second mode 2 is selected, the second audio signal 245 is signal-processed, and the second target audio 292 is generated. Specifically, step S344 may be the same as step S144, but this will not be explained in detail here.
[0151] The method P300 described above may include the following steps performed by at least one processor 220:
[0152] S360: Output the target audio. Specifically, step S360 may coincide with step S160, but this will not be explained in detail here.
[0153] The method P300 described above may also include the following steps:
[0154] S380: Based on the control signal, the intensity of the speaker input signal of the speaker 280 is controlled. Specifically, step S380 may be the same as step S180, but this will not be explained in detail here.
[0155] Table 3 shows the target audio processing mode result diagram corresponding to Figure 7. As shown in Table 3, to facilitate matching, the scene was divided into eight scenes, each being: Type 1: Near-end audio signal is less than the threshold (e.g., user 002 does not make a sound), the speaker signal does not exceed the speaker threshold, and ambient noise is relatively low; Type 2: Near-end audio signal is greater than the threshold (e.g., user 002 makes a sound), the speaker signal does not exceed the speaker threshold, and ambient noise is relatively low; Type 3: Near-end audio signal is less than the threshold (e.g., user 002 does not make a sound), the speaker signal exceeds the speaker threshold, and ambient noise is relatively low; and Type 4: Near-end audio signal is greater than the threshold (e.g., user 002 makes a sound), the speaker signal exceeds the speaker threshold... Type 5: The speaker signal exceeds the speaker threshold, ambient noise is relatively low, Type 6: The speaker signal exceeds the speaker threshold (for example, user 002 does not make a sound), ambient noise is relatively high, Type 7: The speaker signal exceeds the speaker threshold (for example, user 002 makes a sound), ambient noise is relatively high, Type 8: The speaker signal exceeds the speaker threshold (for example, user 002 makes a sound), ambient noise is relatively high. Here, whether the speaker signal is greater than the threshold can be determined by the control module 231 based on the microphone signal. The near-end audio signal being greater than the threshold may mean that the audio signal intensity emitted by user 002 exceeds a preset threshold. The target audio processing modes corresponding to the eight scenes are such that the 5th, 6th, and 8th scenes correspond to the first mode 1, the 3rd, 4th, and 7th scenes correspond to the second mode 2, and the other scenes correspond to either the first mode 1 or the second mode 2.
[0156] [Table 3]
[0157] Methods P200 and P300 are applied to different application scenarios. In scenarios where the speaker signal is more important than near-end audio quality, method P200 can be selected to ensure the quality and intelligibility of the speaker signal. In scenarios where near-end audio quality is more important than the speaker signal, method P300 can be selected to ensure the audio quality and intelligibility of the near-end audio.
[0158] As described above, System 100, Method P100, Method P200, and Method P300 can control the sound source signal of the electronic device 200 so that the target audio quality is optimized in any scene and the quality of voice communication is improved, by controlling the target audio processing mode of the electronic device 200 based on the speaker signal for different scenes.
[0159] The signal intensity of the ambient noise differs with frequency. At different frequencies, the audio quality of the first target audio 291 and the second target audio 292 also differs. For example, at the first frequency, the audio quality of the first target audio 291, in which the first audio signal 243 and the second audio signal 245 are processed by the first algorithm 233-1, is superior to the audio quality of the second target audio 292, in which the second audio signal 245 is processed by the second algorithm 233-8. At frequencies other than the first frequency, the audio quality of the first target audio 291, in which the first audio signal 243 and the second audio signal 245 are processed by the first algorithm 233-1, is close to the audio quality of the second target audio 292, in which the second audio signal 245 is processed by the second algorithm 233-8. In this case, the electronic device 200 can also generate the control signal based on the frequency of the ambient noise. A first control signal is generated at the first frequency, and a second control signal is generated at a frequency other than the first frequency.
[0160] If the ambient noise is low-frequency noise (for example, in the case of a subway or bus), the quality of the audio signal at low frequencies of the first target audio 291 obtained by the first audio signal 243 and the second audio signal 245 in the signal processing of the first algorithm 233-1 may be poor, meaning that the understanding of the audio at low frequencies of the first target audio 291 is relatively poor, while the understanding of the audio at high frequencies is relatively high. In this case, the electronic device 200 can control the selection of the target audio processing mode based on the frequency of the ambient noise. For example, in the low-frequency range, the electronic device 200 can select method P300 and control the target audio processing mode to ensure that the voice of the near-end user 002 is picked up and that near-end audio quality is ensured, and in the high-frequency range, the electronic device 200 can select method P200 and control the target audio processing mode to ensure that the near-end user 002 can hear the speaker signal.
[0161] Other aspects of this specification provide a non-temporary storage medium storing at least one group of executable commands based on sound source signal control, wherein, when the executable commands are executed by a processor, the executable commands instruct the processor to carry out steps of the audio signal processing method for echo suppression described herein. In some possible embodiments, each aspect of this specification can also be implemented in the form of a program product including program code. When the program product is executed on an electronic device 200, the program code is used to cause the electronic device 200 to carry out steps based on sound source signal control described herein. A program product for implementing the above method may employ a portable compact disc read-only memory (CD-ROM) containing the program code and be executed on the electronic device 200. However, the program products of this specification are not limited thereto, and the readable storage medium may be any tangible medium that includes or stores a program, the program being used by or in combination with a command execution system (e.g., a processor 220). The program product may employ any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. A readable storage medium may be, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of more than these. More specific examples of readable storage media include electrical connections having one or more wires, portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. The computer-readable storage medium may include data signals propagating in the baseband as part of a carrier wave, and may include readable program code.The data signals propagated in this manner may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The readable storage medium may also be any other readable medium that can transmit, propagate, or transmit programs for use by or in combination with command execution systems, apparatus, or devices. The program code contained in the readable storage medium may be transmitted by any suitable medium, including but not limited to wireless, wired, optical cable, RF, or any suitable combination thereof. The program code for performing the operations described herein may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java and C++, and general procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the electronic device 200, partially on the electronic device 200, as a standalone software package, partially on the electronic device 200, partially on a remote computing device, or fully on a remote computing device.
[0162] The above describes specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the operations or steps described in the claims may be performed in a different order than in the embodiments, and the desired results can still be achieved. Furthermore, the processes shown in the drawings do not necessarily require a specific or sequential order to achieve the desired results. In some embodiments, multitasking and parallel processing may also be possible or advantageous.
[0163] Having read this detailed disclosure, those skilled in the art will understand that the above detailed disclosure may be presented merely as examples and may not be limiting. Although not explicitly stated herein, those skilled in the art will understand that the requirements of this specification encompass a variety of reasonable changes, improvements, and modifications of the embodiments. These changes, improvements, and modifications are intended to be submitted herein and are within the spirit and scope of the exemplary embodiments herein.
[0164] Some terms used herein are used to describe embodiments. For example, “one embodiment,” “embodiment,” and / or “several embodiments” mean that certain features, structures, or characteristics described in conjunction with such embodiments may be included in at least one embodiment of this specification. Therefore, it should be emphasized and understood that two or more references to “embodiment,” “one embodiment,” or “alternative embodiment” in each part of this specification do not necessarily refer to the same embodiment. Certain features, structures, or characteristics may be appropriately combined in one or more embodiments of this specification.
[0165] In the foregoing description of the embodiments herein, for the sake of understanding the features and for the sake of simplification, various features are combined in a single embodiment, drawing, or description thereof. However, these combinations of features are not essential, and those skilled in the art can easily extract some of the features and understand them as individual embodiments when reading this specification. In other words, the embodiments in this specification can be understood as a combination of multiple secondary embodiments. The content of each secondary embodiment may be less than all the features of one disclosed embodiment described above.
[0166] Each patent, patent application, publication of a patent application, and other material referenced herein, such as texts, books, specifications, publications, documents, articles, etc., may be combined herein by reference. All content used for all purposes is excluded from any history of any indictment document relating thereto, any identical document that may not be consistent with or may contradict this document, or any identical indictment document that may have a limited effect on the broadest scope of the claims, whether current or future, relating to this document. For example, terms in this document are used in the event of any inconsistency or conflict between terms relating to any material contained herein, their definitions, and / or terms relating to this document.
[0167] Finally, it should be understood that the embodiments of the application disclosed herein are intended to illustrate the principles of the embodiments herein. Other modified embodiments are also within the scope of this specification. Therefore, the embodiments disclosed herein are merely examples and not limiting. Those skilled in the art can implement the applications herein by taking alternative configurations based on the embodiments herein. Therefore, the embodiments herein are not limited to those strictly described in the application.
Claims
1. An audio signal processing method for suppressing echo in electronic equipment, The steps include acquiring the speaker signal transmitted from the control device, A step of generating a corresponding control signal based on the intensity of the speaker signal and the microphone signal, The control signal includes a first control signal or a second control signal. The electronic device includes a microphone module that outputs the microphone signal, The microphone module includes at least one first type of microphone that outputs a first audio signal, and at least one second type of microphone that outputs a second audio signal and is different from the at least one first type of microphone, A step of selecting a target audio processing mode corresponding to the control signal from a plurality of audio processing modes based on the control signal, The plurality of audio processing modes include a first mode for processing the first audio signal and the second audio signal, and a second mode for processing the second audio signal without processing the first audio signal. A step in which the first mode corresponds to the first control signal and the second mode corresponds to the second control signal, The steps include reducing echoes in the target audio by processing the microphone signal using the target audio processing mode to generate target audio, The step of outputting the target audio is included, An audio signal processing method for suppressing echo, characterized by the following:
2. The microphone signal includes the first audio signal and the second audio signal. The audio signal processing method according to feature 1.
3. The aforementioned at least one first type of microphone is used for collecting human body vibration signals, and, The aforementioned at least one second type of microphone is used for collecting air vibration signals. The audio signal processing method according to feature 2.
4. The step of generating a control signal corresponding to the speaker signal based at least on the intensity of the speaker signal is: The steps include determining that the intensity of the speaker signal is lower than a preset speaker threshold and generating the first control signal, or The process includes the step of determining that the intensity of the speaker signal is higher than the speaker threshold and generating the second control signal, The audio signal processing method according to feature 1.
5. The step of generating a corresponding control signal based on the intensity of the speaker signal and the microphone signal is as follows: A step of obtaining evaluation parameters for the microphone signal, wherein the evaluation parameters include environmental noise evaluation parameters, and the environmental noise evaluation parameters include at least one of environmental noise level and signal-to-noise ratio. The step of generating the control signal based on the intensity of the speaker signal and the evaluation parameter includes: The audio signal processing method according to feature 1.
6. The step of generating the control signal based on the intensity of the speaker signal and the evaluation parameter is: The steps include determining that the intensity of the speaker signal is higher than a preset speaker threshold and generating the second control signal, The steps include determining that the intensity of the speaker signal is lower than the speaker threshold and that the environmental noise evaluation parameter is outside a preset noise evaluation range, and generating the first control signal, The process includes determining that the intensity of the speaker signal is lower than the speaker threshold and that the environmental noise evaluation parameter is within the noise evaluation range, and generating the first control signal or the second control signal. The audio signal processing method according to feature 5.
7. The environmental noise evaluation parameter being within the noise evaluation range means that The aforementioned environmental noise level is lower than a preset environmental noise threshold, The signal-to-noise ratio is higher than a preset signal-to-noise ratio threshold, and includes at least one of the above. The audio signal processing method according to feature 6.
8. The evaluation parameter further includes the human voice signal intensity, and the step of generating the control signal based on the speaker signal intensity and the evaluation parameter is: The steps include determining that the speaker signal intensity is higher than a preset speaker threshold, the human voice signal intensity exceeds a preset human voice threshold, and the environmental noise evaluation parameter is outside a preset noise evaluation range, and generating the first control signal, The steps include determining that the intensity of the speaker signal is higher than the speaker threshold, the intensity of the human voice signal exceeds the human voice threshold, and the environmental noise evaluation parameter is within the noise evaluation range, and generating the second control signal, The steps include determining that the intensity of the speaker signal is higher than the speaker threshold and the intensity of the human voice signal is lower than the human voice threshold, and generating the second control signal, The steps include determining that the intensity of the speaker signal is lower than the speaker threshold and that the environmental noise evaluation parameter is outside the noise evaluation range, and generating the first control signal, The process includes determining that the intensity of the speaker signal is lower than the speaker threshold and that the environmental noise evaluation parameter is within the noise evaluation range, and generating the first control signal or the second control signal. The audio signal processing method according to feature 5.
9. The environmental noise evaluation parameter being within the noise evaluation range means that The aforementioned environmental noise level is lower than a preset environmental noise threshold, The signal-to-noise ratio is higher than a preset signal-to-noise ratio threshold, and includes at least one of the above. The audio signal processing method according to feature 8.
10. The step of generating the target audio is: A step of processing the first audio signal and the second audio signal using a first algorithm in the first mode to generate a first target audio, or, The second mode includes the step of signal processing the second audio signal using a second algorithm to generate a second target audio, Here, the target audio includes one of the first target audio and the second target audio, The step of outputting the target audio is: The steps include: performing a smoothing process on the target audio, and when the target audio is switched between the first target audio and the second target audio, performing the smoothing process on the connection point between the first target audio and the second target audio; The step of outputting the smoothed target audio is included. The audio signal processing method according to feature 1.
11. The aforementioned electronic device further includes a speaker, The aforementioned method, The process further includes controlling the intensity of the speaker input signal of the speaker based on the control signal, specifically, The process includes the steps of determining the aforementioned control signal as the first control signal and reducing the intensity of the speaker input signal input to the speaker, thereby reducing the intensity of the sound output by the speaker. The audio signal processing method according to feature 1.
12. An audio signal processing system for suppressing echoes, A storage medium storing at least one set of commands for audio signal processing to suppress echoes, The system includes at least one processor that is communicatively connected to the at least one storage medium, Here, the at least one processor, when the system is operating, reads the at least one set of instructions and performs the audio signal processing method for echo suppression according to any one of claims 1 to 11 based on the instructions of the at least one set of instructions. An audio signal processing system for suppressing echoes, characterized by the following:
Citation Information
Patent Citations
Echo processor
JP2003101445A
Sound Interference Cancellation
JP2020536273A
Multi-sensor signal optimization for speech communication
US20130070935A1