Audio signal processing method and system for echo suppression
By switching the audio processing mode according to the speaker signal strength, the traditional echo cancellation algorithm solves the difficulty of echo cancellation when the speaker input signal is strong, and high-quality voice communication in different scenarios is achieved.
Patent Information
- Application Number
- CN202080104434.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-12-28
- Publication Date
- 2025-08-15
- Estimated Expiration
- 2040-12-28
AI Technical Summary
Traditional echo cancellation algorithms are difficult to effectively eliminate echoes in the bone-conducting microphone when the speaker input signal is strong, resulting in a decline in the voice quality of the microphone signal.
Switch the audio processing mode according to the speaker signal strength, select a suitable audio processing mode by generating a control signal, and process the first and second types of microphone signals respectively to reduce echoes and improve voice quality.
By switching the audio processing mode in different scenarios, the echo can be effectively reduced, the voice quality of the microphone signal can be improved, and the voice communication quality in different scenarios can be ensured.
Smart Images

Figure CN116158090B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of audio signal processing, and in particular to an audio signal processing method and system for suppressing echo. Background Art
[0002] Vibration sensors are currently being used in electronic products such as headphones, and their use as bone conduction microphones for receiving voice signals is increasing. When a person speaks, their bones and skin vibrate simultaneously. These vibrations are the bone-conducted voice signals, which are picked up by the bone conduction microphone, generating a signal. The system converts the vibration signals collected by the bone conduction microphone into electrical or other types of signals, which are then transmitted to the electronic device to achieve sound pickup. Currently, an increasing number of electronic devices are combining air conduction microphones with bone conduction microphones, each with different characteristics. The air conduction microphone picks up external audio signals, while the bone conduction microphone picks up vibration signals from the vocalization site. These signals are then processed for voice enhancement and fusion. When placed in headphones or other electronic devices with speakers, the bone conduction microphone not only picks up the vibration signals of speech but also the vibration signals generated by the speakers of the headphones or other electronic devices when they play sound, generating an echo signal. This requires echo cancellation algorithms to process these signals. The varying echo signals from the speakers can also affect the microphone's voice quality. For example, when the speaker input signal is strong, the vibration signal received by the bone conduction microphone is larger, much larger than the vibration signal generated by human speech. In this case, traditional echo cancellation algorithms have difficulty eliminating the echo from the bone conduction microphone. In this case, the microphone signals output by the air conduction microphone and the bone conduction microphone as the sound source signal produce poor voice quality. Therefore, it is unreasonable to ignore the speaker echo signal when selecting the microphone's sound source signal.
[0003] Therefore, it is necessary to provide a new audio signal processing method and system for suppressing echo, so as to switch the input sound source signal according to different speaker input signals, improve the effect of echo cancellation, and enhance the voice quality. Summary of the Invention
[0004] This specification provides a new audio signal processing method and system for suppressing echo, so as to improve the effect of echo cancellation and enhance voice quality.
[0005] In a first aspect, the present specification provides an audio signal processing method for suppressing echo, comprising: selecting a target audio processing mode of an electronic device from a plurality of audio processing modes based at least on a speaker signal, wherein the speaker signal is an audio signal sent by a control device to the electronic device; processing a microphone signal through the target audio processing mode to generate target audio, thereby at least reducing the echo in the target audio, wherein the microphone signal is an output signal of a microphone module obtained by the electronic device, the microphone module including at least one first-class microphone and at least one second-class microphone; and outputting the target audio signal.
[0006] In some embodiments, the at least one first type microphone outputs a first audio signal; and the at least one second type microphone outputs a second audio signal, wherein the microphone signals include the first audio signal and the second audio signal.
[0007] In some embodiments, the at least one first type microphone is used to collect human body vibration signals; and the at least one second type microphone is used to collect air vibration signals.
[0008] In some embodiments, the plurality of audio processing modes include at least: a first mode for performing signal processing on the first audio signal and the second audio signal; and a second mode for performing signal processing on the second audio signal.
[0009] In some embodiments, the selecting of a target audio processing mode of an electronic device from a plurality of audio processing modes based at least on a speaker signal includes: generating a control signal corresponding to the speaker signal based at least on the intensity of the speaker signal, the control signal including a first control signal or a second control signal; and selecting a target audio processing mode corresponding to the control signal based on the control signal, wherein the first mode corresponds to the first control signal and the second mode corresponds to the second control signal.
[0010] In some embodiments, generating a control signal corresponding to the speaker signal based at least on the strength of the speaker signal includes: determining that the strength of the speaker signal is lower than a preset speaker threshold, and generating the first control signal; or determining that the strength of the speaker signal is higher than the speaker threshold, and generating the second control signal.
[0011] In some embodiments, generating a control signal corresponding to the speaker signal based at least on the strength of the speaker signal includes generating a corresponding control signal based on the strength of the speaker signal and the microphone signal.
[0012] In some embodiments, generating a corresponding control signal based on the strength of the speaker signal and the microphone signal includes: obtaining evaluation parameters of the microphone signal, the evaluation parameters including environmental noise evaluation parameters, the environmental noise evaluation parameters including at least one of environmental noise level and signal-to-noise ratio; and generating the control signal based on the strength of the speaker signal and the evaluation parameters.
[0013] In some embodiments, the control signal is generated based on the strength of the speaker signal and the evaluation parameter, including one of the following situations: determining that the strength of the speaker signal is higher than a preset speaker threshold, generating the second control signal; determining that the strength of the speaker signal is lower than the speaker threshold, and the environmental noise evaluation parameter is outside the preset noise evaluation range, generating the first control signal; and determining that the strength of the speaker signal is lower than the speaker threshold, and the environmental noise evaluation parameter is within the noise evaluation range, generating the first control signal or the second control signal.
[0014] In some embodiments, the environmental noise evaluation parameter is within the noise evaluation range, including at least one of the following situations: the environmental noise level is lower than a preset environmental noise threshold; and the signal-to-noise ratio is higher than a preset signal-to-noise ratio threshold.
[0015] In some embodiments, the evaluation parameter also includes the strength of the human voice signal, and the generating of the control signal based on the strength of the speaker signal and the evaluation parameter includes one of the following situations: determining that the strength of the speaker signal is higher than a preset speaker threshold, and the strength of the human voice signal exceeds the preset human voice threshold, and the environmental noise evaluation parameter is outside the preset noise evaluation range, thereby generating the first control signal; determining that the strength of the speaker signal is higher than the speaker threshold, and the strength of the human voice signal exceeds the human voice threshold, and the environmental noise evaluation parameter is within the noise evaluation range, thereby generating the second control signal; determining that the strength of the speaker signal is higher than the speaker threshold, and the strength of the human voice signal is lower than the human voice threshold, thereby generating the second control signal; determining that the strength of the speaker signal is lower than the speaker threshold, and the environmental noise evaluation parameter is outside the noise evaluation range, thereby generating the first control signal; and determining that the strength of the speaker signal is lower than the speaker threshold, and the environmental noise evaluation parameter is within the noise evaluation range, thereby generating the first control signal or the second control signal.
[0016] In some embodiments, the environmental noise evaluation parameter is within the noise evaluation range, including at least one of the following situations: the environmental noise level is lower than a preset environmental noise threshold; and the signal-to-noise ratio is higher than a preset signal-to-noise ratio threshold.
[0017] In some embodiments, generating the target audio includes: performing signal processing on the first audio signal and the second audio signal through a first algorithm in the first mode to generate a first target audio; or performing signal processing on the second audio signal through a second algorithm in the second mode to generate a second target audio, wherein the target audio includes one of the first target audio and the second target audio.
[0018] In some embodiments, outputting the target audio includes: smoothing the target audio, and when the target audio switches between the first target audio and the second target audio, performing the smoothing process on the connection between the first target audio and the second target audio; and outputting the target audio after the smoothing process.
[0019] In some embodiments, the method further comprises: controlling the strength of a speaker input signal of the speaker based on the control signal.
[0020] In some embodiments, controlling the intensity of the speaker input signal of the speaker based on the control signal includes: determining that the control signal is the first control signal, reducing the intensity of the speaker input signal input to the speaker, thereby reducing the intensity of the sound output by the speaker.
[0021] In a second aspect, this specification also provides a system for audio signal processing for echo suppression, comprising: at least one storage medium and at least one processor, the at least one storage medium storing at least one instruction set for audio signal processing for echo suppression; the at least one processor being communicatively connected to the at least one storage medium, wherein, when the system is running, the at least one processor reads the at least one instruction set and, according to the instructions of the at least one instruction set, executes the method for audio signal processing for echo suppression described in the first aspect of this specification.
[0022] As can be seen from the above technical solutions, the audio signal processing method and system for echo suppression provided in this specification can generate a control signal corresponding to the speaker signal based on the strength of the speaker signal, and control or switch the audio processing mode based on the control signal, thereby performing signal processing on the sound source signal corresponding to the audio processing mode to obtain better voice quality. When the speaker signal does not exceed the threshold, the system generates a first control signal, selects the first mode, and uses the first audio signal and the second audio signal as the first sound source signal to perform signal processing on the first sound source signal to obtain the first target audio. When the speaker signal exceeds the threshold, the speaker echo in the first audio signal is large. At this time, the system generates a second control signal, selects the second mode, and uses the second audio signal as the second sound source signal to perform signal processing on the second sound source signal to obtain the second target audio. The method and system can switch different audio processing modes based on the speaker signal, thereby switching the sound source signal of the microphone signal to improve voice quality and ensure better voice quality in different scenarios.
[0023] Other features of the audio signal processing method and system for echo suppression provided in this specification are partially outlined in the following description. The following figures and examples will be readily apparent to those skilled in the art based on the description. The inventive aspects of the audio signal processing method and system for echo suppression provided in this specification can be fully explained through practice or use of the methods, devices, and combinations described in the following detailed examples. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] In order to more clearly illustrate the technical solutions in the embodiments of this specification, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of this specification. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0025] Figure 1 Schematic diagrams showing some application scenarios of audio signal processing systems for echo suppression provided according to embodiments of this specification;
[0026] Figure 2 Schematic diagrams of some electronic devices provided according to embodiments of this specification are shown;
[0027] Figure 3 shows some first mode operation schematic diagrams provided according to the embodiments of this specification;
[0028] Figure 4 shows some schematic diagrams of the second mode of operation provided by the embodiments of this specification;
[0029] Figure 5 1 shows a flow chart of some audio signal processing methods for echo suppression provided according to embodiments of this specification;
[0030] Figure 6 shows a flow chart of some audio signal processing methods for suppressing echo provided according to embodiments of this specification; and
[0031] Figure 7 The flowchart of some audio signal processing methods for suppressing echo provided according to the embodiments of this specification is shown. DETAILED DESCRIPTION
[0032] The following description provides specific application scenarios and requirements for this specification, with the goal of enabling those skilled in the art to make and use the contents of this specification. Various modifications to the disclosed embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other embodiments and applications without departing from the spirit and scope of this specification. Therefore, this specification is not limited to the embodiments shown, but is intended to be accorded the broadest scope consistent with the claims.
[0033] The terms used herein are for the purpose of describing specific example embodiments only and are not intended to be limiting. For example, as used herein, the singular forms "a," "an," and "the" may also include the plural forms unless the context clearly indicates otherwise. When used in this specification, the terms "comprise," "include," and / or "contain" are intended to refer to the presence of the associated integers, steps, operations, elements, and / or components, but do not preclude the presence of one or more other features, integers, steps, operations, elements, components, and / or groups or the addition of other features, integers, steps, operations, elements, components, and / or groups in the system / method.
[0034] These and other features of this specification, as well as the operation and function of the associated elements of the structure, and the economical assembly and manufacture of the components, can be significantly improved with consideration of the following description. Reference is made to the accompanying drawings, all of which form a part of this specification. However, it should be expressly understood that the drawings are for illustration and description purposes only and are not intended to limit the scope of this specification. It should also be understood that the drawings are not drawn to scale.
[0035] The flowcharts used in this specification illustrate operations implemented by systems according to some embodiments of the present specification. It should be clearly understood that the operations of the flowcharts may not be implemented in sequence. Rather, the operations may be implemented in reverse order or simultaneously. Furthermore, one or more additional operations may be added to the flowcharts. One or more operations may be removed from the flowcharts.
[0036] Figure 1The following is a schematic diagram showing some application scenarios of an audio signal processing system 100 (hereinafter referred to as system 100 ) for echo suppression provided in accordance with an embodiment of this specification. The system 100 may include an electronic device 200 and a control device 400 .
[0037] The electronic device 200 may store data or instructions for executing the audio signal processing method for echo suppression described in this specification, and may execute the data and / or instructions. In some embodiments, the electronic device 200 may be a wireless headset, a wired headset, a smart wearable device, such as smart glasses, a smart helmet, or a smart watch, etc., which has voice collection and voice playback functions. The electronic device 200 may also be a mobile device, a tablet computer, a laptop computer, a built-in device in a motor vehicle, or the like, or any combination thereof. In some embodiments, the mobile device may include a smart home device, a smart mobile device, or the like, or any combination thereof. For example, the smart mobile device may include a mobile phone, a personal digital assistant, a gaming device, a navigation device, an ultra-mobile personal computer (UMPC), etc., or any combination thereof. In some embodiments, the smart home device may include a smart TV, a desktop computer, etc., or any combination thereof. In some embodiments, the built-in device in a motor vehicle may include an onboard computer, an onboard TV, etc.
[0038] The control device 400 can be a remote device that communicates wired and / or wireless audio signals with the electronic device 200. The control device 400 can also be a local device that is communicatively connected to the electronic device 200. The electronic device 200 can collect local audio signals and output them to the control device 400. The electronic device 200 can also receive and output remote audio signals sent by the control device 400. The remote audio signal can also be called a speaker signal. The control device 400 can also be a device with voice collection and voice playback functions. For example, a mobile phone, a tablet computer, a laptop computer, a headset, a smart wearable device, a built-in device in a motor vehicle or the like, or any combination thereof. For example, when the electronic device 200 is a headset, the control device 400 can be a terminal device that is communicatively connected to the headset, such as a mobile phone, a computer, and the like.
[0039] like Figure 1As shown, electronic device 200 may include a microphone module 240 and a speaker 280. Microphone module 240 may be configured to acquire local audio signals and output microphone signals, that is, electronic signals carrying audio information. Microphone module 240 may be an external or internal microphone module. For example, microphone module 240 may be a microphone located outside or inside the ear canal. Microphone module 240 may include at least one first-type microphone 242 and at least one second-type microphone 244. First-type microphone 242 is different from second-type microphone 244. First-type microphone 242 may be a microphone that directly collects human body vibration signals, such as a bone conduction microphone. Second-type microphone 244 may be a microphone that directly collects air vibration signals, such as an air conduction microphone. Of course, first-type microphone 242 and second-type microphone 244 may also be other types of microphones. For example, first-type microphone 242 may be an optical microphone, while second-type microphone 244 may be a microphone that receives electromyographic signals, and so on. Because the first type of microphone 242 differs from the second type of microphone 244, it perceives audio signals differently, resulting in different noise and echo components in the corresponding audio signals. For ease of illustration, this disclosure uses a bone conduction microphone as an example of the first type of microphone 242 and an air conduction microphone as an example of the second type of microphone 244.
[0040] The bone conduction microphone may include a vibration sensor, such as an optical vibration sensor, an acceleration sensor, etc. The vibration sensor may collect mechanical vibration signals (for example, signals generated by the vibration of the skin or bones when the user 002 speaks) and convert the mechanical vibration signals into electrical signals. The mechanical vibration signals mentioned here mainly refer to vibrations transmitted through solids. The bone conduction microphone contacts the skin or bones of the user 002 through the vibration sensor or a vibration component connected to the vibration sensor, thereby collecting the vibration signals generated by the bones or skin of the user 002 when the user 002 makes a sound, and converting the vibration signals into electrical signals. In some embodiments, the vibration sensor may be a device that is sensitive to mechanical vibrations but insensitive to air vibrations (that is, the response capability of the vibration sensor to mechanical vibrations exceeds the response capability of the vibration sensor to air vibrations). Since the bone conduction microphone can directly pick up the vibration signals at the vocalization site, the bone conduction microphone can reduce the impact of ambient noise.
[0041] The air conduction microphone collects the air vibration signals caused by user 002's voice and converts them into electrical signals. The air conduction microphone can be a single microphone or a microphone array consisting of two or more air conduction microphones. The microphone array can be a beamforming microphone array or other similar microphone array. The microphone array can capture sounds from different directions or locations in space.
[0042] The first type of microphone 242 can output a first audio signal 243. The second type of microphone 244 can output a second audio signal 245. The microphone signal includes the first audio signal 243 and the second audio signal 245. In a low-noise scenario, the second audio signal 245 has better voice quality than the first audio signal 243. In a scenario with high ambient noise, the voice quality of the first audio signal 243 is higher in the low-frequency part, and the voice quality of the second audio signal 245 is higher in the high-frequency part. Therefore, in a scenario with high ambient noise, the audio signal obtained after feature fusion of the first audio signal 243 and the second audio signal 245 has good voice quality. In actual use, the ambient noise may change at any time, and the low-noise scenario and the high-noise scenario may be repeatedly switched.
[0043] The speaker 280 can convert an electrical signal into an audio signal. The speaker 280 can be configured to receive and output the speaker signal from the control device 400. For ease of description, the audio signal input to the speaker 280 is defined as a speaker input signal. In some embodiments, the speaker input signal can be the speaker signal. In some embodiments, the electronic device 200 can perform signal processing on the speaker signal and send the processed audio signal to the speaker 280 for output. In this case, the speaker input signal can be an audio signal obtained after the electronic device 200 performs signal processing on the speaker signal.
[0044] The sound of the speaker input signal output by speaker 280 can be transmitted to user 002 via air conduction or bone conduction. Speaker 280 can be a speaker that transmits sound by transmitting vibration signals to the human body, such as a bone conduction speaker, or a speaker that transmits vibration signals through the air, such as an air conduction speaker. A bone conduction speaker generates mechanical vibrations through a vibration module and transmits these mechanical vibrations to the ear through the bone. For example, speaker 280 can contact user 002's head directly or through a specific medium (e.g., one or more panels) and transmit the audio signal to the user's auditory nerve through skull vibrations. An air conduction speaker generates vibrations in the air through a vibration module and transmits these air vibrations to the ear through the air. Speaker 280 can also be a combination of a bone conduction speaker and an air conduction speaker. Speaker 280 can also be other types of speakers. The sound of the speaker input signal output by speaker 280 may be collected by microphone module 240, forming an echo. The greater the intensity of the speaker input signal, the greater the intensity of the sound output by speaker 280, and the stronger the echo signal.
[0045] It should be noted that the microphone module 240 and the speaker 280 can be integrated into the electronic device 200 or can be external devices of the electronic device 200 .
[0046] When the first microphone 242 and the second microphone 244 are in operation, they can not only capture the voice of the user 002, but also ambient noise and the sound emitted by the speaker 280. The electronic device 200 can collect audio signals and generate the microphone signal through the microphone module 240. The microphone signal may include a first audio signal 243 and a second audio signal 245. The voice quality of the first audio signal 243 and the second audio signal 245 may vary in different scenarios. To ensure the quality of voice communication, the electronic device 200 can select a target audio processing mode from multiple audio processing modes based on different application scenarios. This mode selects an audio signal with better voice quality from the microphone signal as the audio source signal, performs signal processing on the audio source signal using the target audio processing mode, and then outputs the signal to the control device 400. The audio source signal may be the input signal of the target audio processing mode. In some embodiments, the signal processing may include noise suppression to reduce noise signals. In some embodiments, the signal processing may include echo suppression to reduce echo signals. In some embodiments, the signal processing may include both noise suppression and echo suppression. In some embodiments, the signal processing may also directly output the audio source signal. For ease of illustration, the following description will be based on the signal processing including the echo suppression. Those skilled in the art should understand that other signal processing methods are within the scope of protection of this specification.
[0047] The selection of the target audio processing mode by the electronic device 200 is not only related to the ambient noise but also to the speaker signal. In some scenarios, for example, when the speaker signal is small and the sound output by the speaker 280 is also small, the voice quality of the audio signal after feature fusion of the first audio signal 243 output by the first type microphone 242 and the second audio signal 245 output by the second type microphone 244 is better than the voice quality of the second audio signal 245 output by the second type microphone 244.
[0048] However, in some special scenarios, such as when the speaker signal is large and the sound output by the speaker 280 is also large, the first audio signal 243 output by the first type microphone 242 is greatly affected, resulting in a large echo in the first audio signal 243. In some embodiments, the echo signal in the first audio signal 243 will exceed the voice signal of the user 002. In particular, when the speaker 280 is a bone conduction speaker, the echo signal in the first audio signal 243 is more obvious. Traditional echo cancellation algorithms have difficulty in eliminating the echo signal in the first audio signal 243 and cannot guarantee the effect of echo cancellation. At this time, the voice quality of the second audio signal 245 output by the second type microphone 244 is better than the voice quality of the audio signal after feature fusion of the first audio signal 243 output by the first type microphone 242 and the second audio signal 245 output by the second type microphone 244.
[0049] Therefore, the electronic device 200 may select the target audio processing mode from the plurality of audio processing modes based on the speaker signal to perform the signal processing on the microphone signal. The plurality of audio processing modes may include at least a first mode 1 and a second mode 2.
[0050] First mode 1 can perform signal processing on the first audio signal 243 and the second audio signal 245. As previously described, in some embodiments, the signal processing may include noise suppression to reduce noise signals. In some embodiments, the signal processing may include echo suppression to reduce echo signals. In some embodiments, the signal processing may include both the noise suppression and the echo suppression. For ease of presentation, the following description will be based on the assumption that the signal processing includes echo suppression. Those skilled in the art will appreciate that other signal processing methods are within the scope of this specification.
[0051] Second mode 2 can perform signal processing on the second audio signal 245. In some embodiments, the signal processing can include noise suppression to reduce noise signals. In some embodiments, the signal processing can include echo suppression to reduce echo signals. In some embodiments, the signal processing can include both noise suppression and echo suppression. For ease of presentation, the following description will assume that the signal processing includes echo suppression. Those skilled in the art will appreciate that other signal processing methods are within the scope of this specification.
[0052] The target audio processing mode is one of the first mode 1 and the second mode 2. The plurality of audio processing modes may further include other modes, for example, a processing mode for performing signal processing on the first audio signal 243 .
[0053] Therefore, when the speaker signal is relatively small, to ensure that the voice used for voice communication has a high quality, the electronic device 200 selects the first mode 1, uses the first audio signal 243 and the second audio signal 245 as the sound source signals, performs signal processing on the sound source signals, generates and outputs the first target audio 291 for use in voice communication. When the speaker signal is relatively large, to ensure that the voice used for voice communication has a high quality, the electronic device 200 selects the second mode 2, uses the second audio signal 245 as the sound source signal, performs signal processing on the sound source signal, generates and outputs the second target audio 292 for use in voice communication.
[0054] The electronic device 200 can execute the data or instructions of the method for audio signal processing for suppressing echo described in this specification to obtain the microphone signal and the speaker signal; the electronic device 200 can select the corresponding target audio processing mode to perform signal processing on the microphone signal based on the signal strength of the speaker signal. Specifically, the electronic device 200 can select a target audio processing mode corresponding to the strength of the speaker signal from a plurality of audio processing modes according to the strength of the speaker signal, select an audio signal with better voice quality or a combination thereof from the first audio signal 243 and the second audio signal 245 as the sound source signal, and use a corresponding signal processing algorithm to perform signal processing on the sound source signal (such as echo cancellation and noise reduction processing), generate a target audio and output it to reduce the echo in the target audio. The target audio may include one of the first target audio 291 and the second target audio 292. The electronic device 200 can output the target audio to the control device 400.
[0055] To sum up, in order to ensure the voice quality of communication, the electronic device 200 can control and select the target audio processing mode based on the strength of the speaker signal, thereby selecting an audio signal with better voice quality as the sound source signal of the electronic device 200, and performing signal processing on the sound source signal to obtain different target audios for different usage scenarios, thereby ensuring that the voice quality of the target audio is optimal in different usage scenarios.
[0056] Figure 2 FIG2 shows a schematic diagram of an electronic device 200. The electronic device 200 can execute the method for audio signal processing for suppressing echo described in this specification. The method for audio signal processing for suppressing echo is introduced in other parts of this specification. For example, Figures 5 to 7 The method for audio signal processing for echo suppression is introduced in the description.
[0057] like Figure 2As shown, the electronic device 200 may include a microphone module 240 and a speaker 280. In some embodiments, the electronic device 200 may further include at least one storage medium 230 and at least one processor 220.
[0058] The storage medium 230 may include a data storage device. The data storage device may be a non-transitory storage medium or a temporary storage medium. For example, the data storage device may include one or more of a disk, a read-only storage medium (ROM), or a random access storage medium (RAM). The storage medium 230 also includes at least one instruction set stored in the data storage device for audio signal processing for echo suppression. The instructions are computer program code, which may include programs, routines, objects, components, data structures, procedures, modules, etc. that execute the audio signal processing method for echo suppression provided in this specification.
[0059] like Figure 2 As shown, the at least one instruction set may include a control instruction, issued by the control module 231, configured to generate a control signal corresponding to the speaker signal based on the speaker signal or the speaker signal and the microphone signal. The control signal may include a first control signal or a second control signal. The first control signal corresponds to first mode 1. The second control signal corresponds to second mode 2. The control signal may be any signal; for example, the first control signal may be signal 1, the second control signal may be signal 2, and so on. The control instruction issued by the control module 231 may generate a corresponding control signal based on the signal strength of the speaker signal or the signal strength of the speaker signal and an evaluation parameter of the microphone signal. The correspondence between the control signal and the speaker signal or the speaker signal and the microphone signal will be described in detail below. The control module 231 may also select a target audio signal processing mode corresponding to the control signal based on the control signal. When the control signal is the first control signal, the control module 231 selects first mode 1; when the control signal is the second control signal, the control module 231 selects second mode 2.
[0060] In some embodiments, the at least one instruction set may further include an echo processing instruction, issued by the echo processing module 233, and configured to perform signal processing (such as echo suppression, noise reduction, etc.) on the microphone signal using the target audio processing mode of the electronic device 200 based on the control signal. When the control signal is the first control signal, the echo processing module 233 uses the first mode 1 to process the microphone signal. When the control signal is the second control signal, the echo processing module 233 uses the first mode 2 to process the microphone signal.
[0061] The echo processing module 233 may include a first algorithm 233-1 and a second algorithm 233-8. The first algorithm 233-1 corresponds to the first control signal and the first mode 1. The second algorithm 233-8 corresponds to the second control signal and the second mode 2.
[0062] In the first mode 1, the electronic device 200 uses the first algorithm 233 - 1 to perform signal processing on the first audio signal 243 and the second audio signal 245 respectively, and performs feature fusion on the first audio signal 243 and the second audio signal 245 after the signal processing, and outputs the first target audio 291.
[0063] Figure 3 FIG1 shows a working diagram of a first mode 1 provided according to an embodiment of this specification. Figure 3 As shown, in first mode 1, the first algorithm 233-1 can receive the first audio signal 243, the second audio signal 245, and the speaker input signal. The first algorithm 233-1 can use the first echo cancellation module 233-2 to perform echo cancellation on the first audio signal 243 based on the speaker input signal. The speaker input signal can be an audio signal that has undergone noise reduction processing. The first echo cancellation module 233-2 receives the first audio signal 243 and the speaker input signal and outputs the first audio signal 243 after echo cancellation. The first echo cancellation module 233-2 can be a single-microphone echo cancellation algorithm.
[0064] In some embodiments, the first algorithm 233-1 can use a second echo cancellation module 233-3 to perform echo cancellation on the second audio signal 245 based on the speaker input signal. The second echo cancellation module 233-3 receives the second audio signal 245 and the speaker input signal and outputs the echo-cancelled second audio signal 245. The second echo cancellation module 233-3 can be a single-microphone echo cancellation algorithm or a multi-microphone echo cancellation algorithm. The first echo cancellation module 233-2 and the second echo cancellation module 233-3 can be the same or different.
[0065] In some embodiments, the first algorithm 233-1 may use a first noise suppression module 233-4 to perform noise suppression on the first audio signal 243 and the second audio signal 245 after echo cancellation. The first noise suppression module 233-4 is configured to suppress noise signals in the first audio signal 243 and the second audio signal 245. The first noise suppression module 233-4 receives the first audio signal 243 and the second audio signal 245 after echo cancellation and outputs the first audio signal 243 and the second audio signal 245 after noise cancellation. The first noise suppression module 233-4 may perform noise reduction on the first audio signal 243 and the second audio signal 245 individually or simultaneously.
[0066] In some embodiments, the first algorithm 233-1 can use a feature fusion module 233-5 to perform feature fusion processing on the noise-suppressed first audio signal 243 and the second audio signal 245. The feature fusion module 233-5 receives the first audio signal 243 and the second audio signal 245 that have undergone noise reduction processing. The feature fusion module 233-5 can analyze the speech quality of the first audio signal 243 and the second audio signal 245. For example, the feature fusion module 233-5 can analyze the effective speech signal strength, noise signal strength, echo signal strength, and signal-to-noise ratio in the first audio signal 243 and the second audio signal 245 to determine the speech quality of the first audio signal 243 and the second audio signal 245, and then fuse the first audio signal 243 and the second audio signal 245 into the first target audio 291 and output it.
[0067] In some embodiments, the first algorithm 233-1 may further utilize a second noise suppression module 233-6 to perform noise suppression on the speaker signal. The second noise suppression module 233-6 is configured to suppress noise signals in the speaker signal. The second noise suppression module 233-6 receives the speaker signal sent by the control device 400, eliminates noise signals such as far-end noise, channel noise, and electronic noise in the electronic device 200 from the speaker signal, and outputs a processed speaker signal that has undergone noise reduction processing.
[0068] It should be noted that Figure 3This is merely an example. Those skilled in the art will appreciate that, in some embodiments, the first algorithm 233-1 may include a feature fusion module 233-5. In other embodiments, the first algorithm 233-2 may further include any one of a first echo cancellation module 233-2, a second echo cancellation module 233-3, a first noise suppression module 233-4, and a second noise suppression module 233-6, or any combination thereof.
[0069] In the second mode 2 , the electronic device 200 uses the second algorithm 233 - 8 to perform signal processing on the second audio signal 245 and outputs the second target audio 292 .
[0070] Figure 4 FIG2 shows a working diagram of a second mode 2 provided according to an embodiment of this specification. Figure 4 As shown, in second mode 2, the second algorithm 233-8 can receive the second audio signal 245 and the speaker input signal. The second algorithm 233-8 can use a third echo cancellation module 233-9 to perform echo cancellation on the second audio signal 245 based on the speaker input signal. The third echo cancellation module 233-9 receives the second audio signal 245 and the speaker input signal and outputs the echo-cancelled second audio signal 245. The third echo cancellation module 233-9 can be the same as or different from the second echo cancellation module 233-3.
[0071] In some embodiments, the second algorithm 233-8 may use a third noise suppression module 233-10 to perform noise suppression on the echo-cancelled second audio signal 245. The third noise suppression module 233-10 is configured to suppress noise signals in the second audio signal 245. The third noise suppression module 233-10 receives the echo-cancelled second audio signal 245 and outputs the noise-suppressed second audio signal 245 as the second target audio 292. The third noise suppression module 233-10 may be the same as or different from the first noise suppression module 233-4.
[0072] In some embodiments, the second algorithm 233-8 may also use a fourth noise suppression module 233-11 to perform noise suppression on the speaker signal. The fourth noise suppression module 233-11 is configured to suppress noise signals in the speaker signal. The fourth noise suppression module 233-11 receives the speaker signal sent by the control device 400, eliminates noise signals such as far-end noise, channel noise, and electronic noise in the electronic device 200, and outputs a processed speaker signal that has undergone noise reduction processing. The fourth noise suppression module 233-11 may be the same as or different from the second noise suppression module 233-6.
[0073] It should be noted that Figure 4 This is merely an example. Those skilled in the art will appreciate that, in some embodiments, the second algorithm 233-8 may include any one of the third echo cancellation module 233-9, the third noise suppression module 233-10, and the fourth noise suppression module 233-11, or any combination thereof. In other embodiments, the second algorithm 233-8 may not include any of the aforementioned signal processing modules and may directly output the second audio signal 245.
[0074] Only one of the first mode 1 and the second mode 2 can be run to save computing resources. When the first mode 1 is running, the second mode 2 can be turned off. When the second mode 2 is running, the first mode 1 can be turned off. The first mode 1 and the second mode 2 can also be run simultaneously. When one mode is running, the algorithm parameters of the other mode can be updated. When the electronic device 200 switches between the first mode 1 and the second mode 2, some parameters in the first mode 1 and the second mode 2 can be shared (such as noise parameters obtained by the noise estimation algorithm, human voice parameters obtained by the human voice estimation algorithm, signal-to-noise ratio parameters obtained by the signal-to-noise ratio algorithm, etc.), thereby saving computing resources and making the calculation results more accurate. The first algorithm 233-1 and the second algorithm 233-8 in the first mode 1 and the second mode 2 can also share some parameters in the control instructions issued by the control module 231, such as noise parameters obtained by the noise estimation algorithm, human voice parameters obtained by the human voice estimation algorithm, signal-to-noise ratio parameters obtained by the signal-to-noise ratio algorithm, etc., thereby saving computing resources and making the calculation results more accurate.
[0075] In some embodiments, the at least one instruction set may also include microphone control instructions, executed by the microphone control module 235, configured to smooth the target audio and output the smoothed target audio to the control device 400. The microphone control module 235 may receive the control signal generated by the control module 231 and the target audio, and perform the smoothing on the target audio based on the control signal. When the control signal is the first control signal, the first mode 1 is implemented, using the first target audio 291 output by the first algorithm 233-1 as the input signal. When the control signal is the second control signal, the second mode 2 is implemented, using the second target audio 292 output by the second algorithm 233-8 as the input signal. When the control signal switches between the first control signal and the second control signal, causing the target audio processing mode to switch between the first mode 1 and the second mode 2, the microphone control module 235 may smooth the target audio to avoid signal discontinuity caused by the switching between the first target audio 291 and the second target audio 292. Specifically, the microphone control module 235 can adjust the parameters of the first target audio 291 and the second target audio 291 to make the target audio continuous. The parameters can be pre-stored in the at least one storage medium 230. The parameters can include amplitude, phase, frequency response, etc. The adjustments can include adjusting the volume of the target audio, adjusting the EQ balance, adjusting the residual noise, etc. The microphone control module 235 can ensure that when the target audio processing mode switches between the first mode 1 and the second mode 2, the target audio is a continuous signal, making it difficult for the user 002 to perceive the switching between the two.
[0076] In some embodiments, the at least one instruction set may further include a speaker control instruction, which is executed by the speaker control module 237 and configured to adjust the speaker processing signal to obtain the speaker input signal, and output the speaker input signal to the speaker 280 to output sound. The speaker control module 237 can receive the speaker processing signals output by the first algorithm 233-1 and the second algorithm 233-8 and the control signal. When the control signal is the first control signal, the speaker control module 237 can control the speaker processing signal output by the first algorithm 233-1 to reduce or turn it off before outputting it to the speaker 280 for output, thereby reducing the sound output by the speaker 280, thereby reducing the echo and improving the echo cancellation effect of the first algorithm 233-1. When the control signal is the second control signal, the speaker control module 237 may not adjust the speaker processing signal output by the second algorithm 233-8. When the control signal switches between the first control signal and the second control signal, the speaker control module 237 can smooth the speaker processing signals output by the first algorithm 233-1 and the second algorithm 233-8 to avoid discontinuity in the sound output by the speaker 280. When switching between the first control signal and the second control signal, the speaker control module 237 strives to ensure continuity of the switching so that the user 002 does not easily perceive the switching.
[0077] In first mode 1, first algorithm 233-1 prioritizes the voice quality of user 002 picked up by near-end microphone module 240. When the speaker processing signal is too high, speaker control module 237 processes the speaker processing signal to reduce the speaker input signal, thereby reducing the sound output by speaker 280 and minimizing echo to ensure near-end voice quality. Second algorithm 233-8 prioritizes the speaker input signal from speaker 280 and does not utilize first audio signal 243 output by first-type microphone 242 to ensure voice quality and intelligibility of the speaker input signal from speaker 280.
[0078] At least one processor 220 can be communicatively connected to at least one storage medium 230, a microphone module 240, and a speaker 280. The communication connection refers to any form of connection capable of directly or indirectly receiving information. The at least one processor 220 is configured to execute the at least one instruction set described above. When the system 100 is running, the at least one processor 220 reads the at least one instruction set and, in accordance with the instructions of the at least one instruction set, obtains data from the microphone module 240 and the speaker 280, executing the audio signal processing method for echo suppression provided in this specification. The processor 220 can execute all steps included in the audio signal processing method for echo suppression. The processor 220 may be in the form of one or more processors. In some embodiments, the processor 220 may include one or more hardware processors, such as a microcontroller, a microprocessor, a reduced instruction set computer (RISC), an application-specific integrated circuit (ASIC), an application-specific instruction set processor (ASIP), a central processing unit (CPU), a graphics processing unit (GPU), a physical processing unit (PPU), a microcontroller unit, a digital signal processor (DSP), a field programmable gate array (FPGA), an advanced RISC machine (ARM), a programmable logic device (PLD), any circuit or processor capable of performing one or more functions, or any combination thereof. For illustrative purposes only, only one processor 220 is described in this specification in the electronic device 200. However, it should be noted that the electronic device 200 in this specification may also include multiple processors. Therefore, the operations and / or method steps disclosed in this specification may be performed by a single processor as described in this specification, or may be performed jointly by multiple processors. For example, if the processor 220 of the electronic device 200 performs step A and step B in this specification, it should be understood that step A and step B can also be performed jointly or separately by two different processors 220 (for example, the first processor performs step A, the second processor performs step B, or the first and second processors perform steps A and B together).
[0079] In some embodiments, the system 100 may select the target audio processing mode of the electronic device 200 according to the signal strength of the speaker signal. In some embodiments, the system 100 may select the target audio processing mode of the electronic device 200 according to the signal strength of the speaker signal and the microphone signal.
[0080] Figure 5 FIG1 shows a flow chart of an audio signal processing method P100 for suppressing echo according to an embodiment of the present specification. The method P100 is a flow chart of a method in which the system 100 selects the target audio processing mode of the electronic device 200 according to the signal strength of the speaker signal. Figure 5As shown, the method P100 may include executing, by at least one processor 220:
[0081] S120: Selecting a target audio processing mode of the electronic device 200 from the first mode 1 and the second mode 2 based at least on the speaker signal. As mentioned above, the target audio processing mode may include one of the first mode 1 and the second mode 2. Specifically, step S120 may include:
[0082] S121: Acquire the speaker signal.
[0083] S122: Generate a control signal corresponding to the speaker signal based at least on the strength of the speaker signal. The control signal includes a first control signal or a second control signal. Specifically, the electronic device 200 can receive the speaker signal sent by the control device 400, compare the strength of the speaker signal with a preset speaker threshold, and generate the control signal based on the comparison result. Step S122 may include one of the following situations:
[0084] S122-2: Determine that the intensity of the speaker signal is lower than the speaker threshold, and generate the first control signal; or
[0085] S122-4: Determine whether the intensity of the speaker signal is higher than a preset speaker threshold, and generate the second control signal.
[0086] Step S120 may further include:
[0087] S124: Based on the control signal, select the target audio processing mode corresponding to the control signal. The first control signal corresponds to first mode 1. The second control signal corresponds to second mode 2. When the control signal is the first control signal, first mode 1 is selected; when the control signal is the second control signal, second mode 2 is selected.
[0088] When the speaker signal strength is higher than the speaker threshold, the first algorithm 233-1 in first mode 1 is used to process the first audio signal 243 and the second audio signal 245. It is not possible to eliminate the echo signal while retaining the good human voice signal. As a result, the resulting first target audio 291 has poor voice quality. However, the second target audio 292 obtained by processing the second audio signal 245 using the second algorithm 233-8 in second mode 2 is of better quality. Therefore, when the speaker signal strength is higher than the speaker threshold, the electronic device 200 generates the second control signal corresponding to second mode 2.
[0089] When the speaker signal strength is lower than the speaker threshold, the first algorithm 233-1 in first mode 1 is used to process the first audio signal 243 and the second audio signal 245. This can eliminate echo signals while retaining a good human voice signal, resulting in better voice quality for the first target audio 291. Furthermore, the second algorithm 233-8 in second mode 2 is used to process the second audio signal 245, resulting in better quality for the second target audio 292. Therefore, when the speaker signal strength is lower than the speaker threshold, the electronic device 200 generates the first control signal corresponding to first mode 1 and may also generate the second control signal corresponding to second mode 2.
[0090] The control signal is generated by the control module 231. Specifically, the electronic device 200 can monitor the strength of the speaker signal in real time and compare it with the speaker threshold. The electronic device 200 can also periodically detect the strength of the speaker signal and compare it with the speaker threshold. The electronic device 200 can also compare the speaker signal with the speaker threshold when it detects a significant change in the strength of the speaker signal and the change exceeds a preset range.
[0091] When the intensity of the speaker signal is higher than the speaker threshold, the electronic device 200 generates the second control signal; when the speaker signal changes and the intensity of the speaker signal is lower than the speaker threshold, the electronic device generates the first control signal. When the intensity of the speaker signal is lower than the speaker threshold, the electronic device 200 generates the first control signal; when the speaker signal changes and the intensity of the speaker signal is higher than the speaker threshold, the electronic device generates the second control signal.
[0092] To ensure that the control signal is not perceived by user 002 when it is switched, the speaker threshold may be a range. The speaker threshold may be within a range between a first speaker critical value and a second speaker critical value. The first speaker critical value is less than the second speaker critical value. The intensity of the speaker signal being greater than the speaker threshold may include the intensity of the speaker signal being greater than the second speaker critical value. The intensity of the speaker signal being less than the speaker threshold may include the intensity of the speaker signal being less than the first speaker critical value.
[0093] When the speaker signal strength is equal to the speaker threshold, the electronic device 200 may generate the first control signal or the second control signal. When the speaker signal strength is higher than the second speaker critical value, the electronic device 200 may generate the second control signal; when the speaker signal strength decreases to between the first speaker critical value and the second speaker critical value, the electronic device 200 may generate the second control signal. When the speaker signal strength is lower than the first speaker critical value, the electronic device 200 may generate the first control signal; when the speaker signal strength increases to between the first speaker critical value and the second speaker critical value, the electronic device 200 may generate the first control signal.
[0094] The electronic device 200 may also obtain a control model through machine learning, input the speaker signal into the control model, and the control model outputs the control signal.
[0095] The method P100 may further include executing, by at least one processor 220:
[0096] S140: Processing the microphone signal in the target audio processing mode to generate the target audio, thereby at least reducing the echo in the microphone signal. Specifically, step S140 may include one of the following situations:
[0097] S142: Determine that the control signal is the first control signal. Based on the speaker input signal, perform signal processing and feature fusion on the first audio signal 243 and the second audio signal 245 using the first algorithm 233-1 in the first mode 1 corresponding to the first control signal to generate a first target audio signal 291. The specific process is as described above and will not be repeated here.
[0098] S144: Determine that the control signal is the second control signal, and perform signal processing on the second audio signal 245 based on the speaker input signal using the second algorithm 233-8 in the second mode 2 corresponding to the second control signal. The specific process is as described above and will not be repeated here.
[0099] S160: Output the target audio. The electronic device 200 may directly output the target audio. The electronic device 200 may also smooth the target audio so that the target audio is not perceived by the user 002 when switching between the first target audio 291 and the second target audio 292. Specifically, step S160 may include smoothing the target audio and outputting the smoothed target audio.
[0100] Specifically, the electronic device 200 can perform smoothing processing on the target audio through the microphone control module 235. When the target audio switches between the first target audio 291 and the second target audio 292, the microphone control module 235 can perform the smoothing processing on the connection between the first target audio 291 and the second target audio 292, that is, perform signal adjustment on the first target audio 291 and the second target audio 292 to ensure a smooth transition at the connection.
[0101] The method P100 may further include:
[0102] S180: Based on the control signal, control the strength of the speaker input signal of the speaker 280. Specifically, step S180 may be performed by the speaker control module 237. Step S180 may include the speaker control module 237 determining that the control signal is the first control signal; and processing the speaker processing signal to reduce the strength of the speaker input signal input to the speaker 280, thereby reducing the strength of the sound output by the speaker 280, thereby reducing the echo signal in the microphone signal, and improving the voice quality of the first target audio.
[0103] Table 1 shows Figure 5 Corresponding target audio processing mode result diagram. As shown in Table 1, for the convenience of comparison, we divide the scenarios into 4 scenarios, namely, the first scenario: the near-end sound signal is less than the threshold (for example, user 002 does not make a sound) and the speaker signal does not exceed the speaker threshold; the second scenario: the near-end sound signal is greater than the threshold (for example, user 002 makes a sound) and the speaker signal does not exceed the speaker threshold; the third scenario: the near-end sound signal is less than the threshold (for example, user 002 does not make a sound) and the speaker signal exceeds the speaker threshold; and the fourth scenario: the near-end sound signal is greater than the threshold (for example, user 002 makes a sound) and the speaker signal exceeds the speaker threshold. Among them, whether the near-end sound signal is greater than the threshold can be determined by the control module 231 based on the microphone signal. The near-end sound signal being greater than the threshold can mean that the intensity of the audio signal emitted by user 002 exceeds the preset threshold. The target audio processing modes corresponding to the 4 scenarios are: the first and second scenarios correspond to the first mode 1; the third and fourth scenarios correspond to the second mode 2.
[0104]
[0105] In the method P100, the electronic device 200 can select the target audio processing mode of the electronic device 200 according to the speaker signal to ensure that the voice quality processed by the target audio processing mode selected by the electronic device 200 in any scenario is optimal to ensure call quality.
[0106] In some embodiments, the selection of the target audio processing mode is not only related to the echo of the speaker signal, but also related to the ambient noise. The ambient noise can be evaluated by at least one of the ambient noise level and the signal-to-noise ratio in the microphone signal.
[0107] Figure 6 A flowchart of an audio signal processing method P200 for echo suppression provided according to an embodiment of this specification is shown. Method P200 is a flowchart of a method in which system 100 selects a target audio processing mode for electronic device 200 based on the signal strength of the speaker signal and the microphone signal. Specifically, method P200 is a flowchart of a method in which system 100 selects the target audio processing mode based on at least one of the ambient noise level and the signal-to-noise ratio in the speaker signal and the microphone signal. Method P200 may include executing, by at least one processor 220:
[0108] S220: Selecting a target audio processing mode of the electronic device 200 from the first mode 1 and the second mode 2 based at least on the speaker signal. Specifically, step S220 may include:
[0109] S222: Generate a control signal corresponding to the speaker signal based at least on the strength of the speaker signal. The control signal includes a first control signal or a second control signal. Specifically, step S222 may be that the electronic device 200 generates a corresponding control signal based on the strength of the speaker signal and the noise in the microphone signal. Step S222 may include:
[0110] S222-2: Obtain evaluation parameters of the speaker signal and the microphone signal. The evaluation parameter may be an environmental noise evaluation parameter in the microphone signal. The environmental noise evaluation parameter may include at least one of an environmental noise level and a signal-to-noise ratio. The electronic device 200 may obtain the environmental noise evaluation parameter in the microphone signal through the control module 231. Specifically, the electronic device 200 may obtain the environmental noise evaluation parameter based on at least one of the first audio signal 243 and the second audio signal 245. The electronic device 200 may obtain the environmental noise evaluation parameter through a noise estimation algorithm or the signal-to-noise ratio, which will not be described in detail in this specification.
[0111] S222-4: Generate the control signal based on the strength of the speaker signal and the environmental noise evaluation parameter. Specifically, the electronic device 200 may compare the strength of the speaker signal with a preset speaker threshold, and compare the environmental noise evaluation parameter with a preset noise evaluation range, and generate the control signal based on the comparison results. Step S222-4 may include one of the following situations:
[0112] S222-5: Determine whether the intensity of the speaker signal is higher than a preset speaker threshold, and generate a second control signal;
[0113] S222-6: Determine that the intensity of the speaker signal is lower than the speaker threshold and the environmental noise evaluation parameter is outside a preset noise evaluation range, and generate the first control signal;
[0114] S222-7: Determine that the intensity of the speaker signal is lower than the speaker threshold and the environmental noise evaluation parameter is within the noise evaluation range, and generate the first control signal or the second control signal.
[0115] Wherein, the environmental noise evaluation parameter being within the noise evaluation range may include at least one of the environmental noise level being lower than a preset environmental noise threshold, and the signal-to-noise ratio being higher than a preset signal-to-noise ratio threshold. At this time, the environmental noise is relatively small. The environmental noise evaluation parameter being outside the noise evaluation range may include at least one of the environmental noise level being higher than a preset environmental noise threshold, and the signal-to-noise ratio being lower than a preset signal-to-noise ratio threshold. At this time, the environmental noise is relatively large. Wherein, when the environmental noise evaluation parameter is outside the noise evaluation range, that is, in a high-noise environment, the voice quality of the first target audio 291 is better than the second target audio 292. When the environmental noise evaluation parameter is within the noise evaluation range, the voice quality of the first target audio 291 is not much different from the voice quality of the second target audio 292.
[0116] Step S220 may further include:
[0117] S224: Based on the control signal, select the target audio processing mode corresponding to the control signal. The first control signal corresponds to first mode 1. The second control signal corresponds to second mode 2. When the control signal is the first control signal, first mode 1 is selected; when the control signal is the second control signal, second mode 2 is selected.
[0118] When the speaker signal strength exceeds the speaker threshold, the first algorithm 233-1 in first mode 1 is unable to retain the good human voice signal while eliminating the echo signal in the first and second audio signals 243 and 245, resulting in poor voice quality in the first target audio 291. However, the second algorithm 233-8 in second mode 2 processes the second audio signal 245 to produce better quality second target audio 292. Therefore, when the speaker signal strength exceeds the speaker threshold, the electronic device 200 generates the second control signal corresponding to second mode 2, regardless of the ambient noise range.
[0119] When the speaker signal strength is lower than the speaker threshold, the first algorithm 233-1 in first mode 1 processes the first audio signal 243 and the second audio signal 245, preserving a good human voice signal while eliminating echo signals therein. As a result, the first target audio 291 obtained has good voice quality. Furthermore, the second algorithm 233-8 in second mode 2 processes the second audio signal 245 to obtain a good quality second target audio 292. Therefore, when the speaker signal strength is lower than the speaker threshold, the control signal generated by the electronic device 200 is related to the ambient noise.
[0120] When the ambient noise level is higher than the ambient noise threshold or the signal-to-noise ratio is lower than the signal-to-noise ratio threshold, it means that the ambient noise in the microphone signal is large. When the first algorithm 233-1 in the first mode 1 performs signal processing on the first audio signal 243 and the second audio signal 245, it can reduce the noise in the signal while retaining a good human voice signal, so the voice quality of the first target audio 291 obtained is better; while the voice quality of the second target audio 292 obtained by the second algorithm 233-8 in the second mode 2 performing signal processing on the second audio signal 245 is not as good as the voice quality of the first target audio 291. Therefore, when the intensity of the speaker signal is lower than the speaker threshold, and the ambient noise level is higher than the ambient noise threshold or the signal-to-noise ratio is lower than the signal-to-noise ratio threshold, the electronic device 200 generates the first control signal corresponding to the first mode 1.
[0121] It should be noted that when the ambient noise is low, that is, when the ambient noise evaluation parameter is within the noise evaluation range, the voice quality of the first target audio 291 and the voice quality of the second target audio 292 are similar. In this case, the electronic device 200 can always generate the second control signal to select the second algorithm 233-8 in the second mode 2 to process the second audio signal 245, thereby reducing the computational effort and conserving resources while ensuring the voice quality of the target audio.
[0122] When the ambient noise level is lower than the ambient noise threshold or the signal-to-noise ratio is higher than the signal-to-noise ratio threshold, it means that the ambient noise in the microphone signal is small. The first target audio 291 obtained when the first algorithm 233-1 in the first mode 1 performs signal processing on the first audio signal 243 and the second audio signal 245, and the second target audio 292 obtained when the second algorithm 233-8 in the second mode 2 performs signal processing on the second audio signal 245, both have good voice quality. Therefore, when the intensity of the loudspeaker signal is lower than the loudspeaker threshold, and the ambient noise level is lower than the ambient noise threshold or the signal-to-noise ratio is higher than the signal-to-noise ratio threshold, the electronic device 200 generates the first control signal or the second control signal. Specifically, the electronic device 200 can determine the control signal in the current scene based on the control signal of the previous scene. That is, when the electronic device generates the first control signal in the previous scene, when it is in the current scene, the electronic device also generates the first control signal, thereby ensuring signal continuity. And vice versa.
[0123] The control signal is generated by the control module 231. Specifically, the electronic device 200 can monitor the intensity of the speaker signal and the environmental noise evaluation parameter in real time, and compare them with the speaker threshold and the noise evaluation range. The electronic device 200 can also periodically detect the intensity of the speaker signal and the environmental noise evaluation parameter, and compare them with the speaker threshold and the noise evaluation range. The electronic device 200 can also compare the speaker signal and the environmental noise evaluation parameter with the speaker threshold and the noise evaluation range when it detects that the intensity of the speaker signal or the environmental noise evaluation parameter has changed significantly, and the change value exceeds a preset range.
[0124] To ensure that the control signal is not perceived by user 002 when switching, the speaker threshold, the ambient noise threshold, and the preset signal-to-noise ratio threshold may be within a range. The speaker threshold has been described above and will not be repeated here. The ambient noise threshold may be within the range of a first noise threshold and a second noise threshold. The first noise threshold is less than the second noise threshold. The ambient noise level being higher than the ambient noise threshold may include the ambient noise level being higher than the second noise threshold. The ambient noise level being lower than the ambient noise threshold may include the ambient noise level being lower than the first noise threshold. The signal-to-noise ratio threshold may be within the range of a first signal-to-noise ratio threshold and a second signal-to-noise ratio threshold. The first signal-to-noise ratio threshold is less than the second signal-to-noise ratio threshold. The signal-to-noise ratio being higher than the signal-to-noise ratio threshold may include the signal-to-noise ratio being higher than the second signal-to-noise ratio threshold. The signal-to-noise ratio being lower than the signal-to-noise ratio threshold may include the signal-to-noise ratio being lower than the first signal-to-noise ratio threshold.
[0125] The method P200 may include executing, by at least one processor 220:
[0126] S240: Processing the microphone signal in the target audio processing mode to generate target audio, thereby at least reducing the echo in the microphone signal. Specifically, step S240 may include one of the following situations:
[0127] S242: Determine that the control signal is the first control signal, select the first mode 1, and perform signal processing on the first audio signal 243 and the second audio signal 245 to generate the first target audio 291. Specifically, step S242 may be the same as step S142 and will not be repeated here.
[0128] S244: Determine that the control signal is the second control signal, select the second mode 2, perform echo suppression on the second audio signal 245, and generate the second target audio 292. Specifically, step S244 may be the same as step S144, and will not be repeated here.
[0129] S260: Output the target audio. Specifically, step S260 may be consistent with step S160, and will not be described in detail here.
[0130] The method P200 may further include:
[0131] S280: Based on the control signal, control the intensity of the speaker input signal of the speaker 280. Specifically, step S280 may be consistent with step S180, and will not be described in detail herein.
[0132] Table 2 shows Figure 6The corresponding target audio processing mode result diagram. As shown in Table 2, for the convenience of comparison, we divide the scenarios into 8 scenarios, namely: the first scenario: the near-end sound signal is less than the threshold (for example, user 002 does not make a sound), the speaker signal does not exceed the speaker threshold, and the environmental noise is small; the second scenario: the near-end sound signal is greater than the threshold (for example, user 002 makes a sound), the speaker signal does not exceed the speaker threshold, and the environmental noise is small; the third scenario: the near-end sound signal is less than the threshold (for example, user 002 does not make a sound), the speaker signal exceeds the speaker threshold, and the environmental noise is small; and the fourth scenario: the near-end sound signal is greater than the threshold (for example, user 002 makes a sound), the speaker signal exceeds the The speaker threshold is greater than the threshold, and the ambient noise is relatively low. The fifth scenario is that the near-end sound signal is less than the threshold (for example, user 002 does not make any sound), the speaker signal does not exceed the speaker threshold, and the ambient noise is relatively high. The sixth scenario is that the near-end sound signal is greater than the threshold (for example, user 002 makes a sound), the speaker signal does not exceed the speaker threshold, and the ambient noise is relatively high. The seventh scenario is that the near-end sound signal is less than the threshold (for example, user 002 does not make any sound), the speaker signal exceeds the speaker threshold, and the ambient noise is relatively high. The eighth scenario is that the near-end sound signal is greater than the threshold (for example, user 002 makes a sound), the speaker signal exceeds the speaker threshold, and the ambient noise is relatively high. Whether the near-end sound signal is greater than the threshold can be determined by the control module 231 based on the microphone signal. The near-end sound signal being greater than the threshold can mean that the strength of the audio signal emitted by user 002 exceeds a preset threshold. The target audio processing modes corresponding to the eight scenarios are: the fifth and sixth scenarios correspond to the first mode 1; the third, fourth, seventh, and eighth scenarios correspond to the second mode 2; and the remaining scenarios correspond to the first mode 1 or the second mode 2.
[0133]
[0134]
[0135] The method P200 can not only control the target audio processing mode of the electronic device 200 according to the speaker signal, but also control the target audio processing mode according to the near-end ambient noise signal, thereby ensuring that the voice quality of the voice signal output by the electronic device 200 is optimal in different scenarios to ensure call quality.
[0136] In some embodiments, the selection of the target audio processing mode is not only related to the echo and ambient noise of the speaker signal, but also to the voice signal of user 002 when speaking. The ambient noise signal can be evaluated by at least one of the ambient noise level and the signal-to-noise ratio in the microphone signal. The voice signal of user 002 when speaking can be evaluated by the strength of the human voice signal in the microphone signal. The human voice signal strength can be the strength of the human voice signal obtained by a noise estimation algorithm, or the strength of the audio signal obtained after noise reduction processing.
[0137] Figure 7 A flowchart of an audio signal processing method P300 for suppressing echo provided according to an embodiment of the present specification is shown. The method P300 is a flowchart of a method in which the system 100 selects a target audio processing mode for the electronic device 200 based on the signal strength of the speaker signal and the microphone signal. Specifically, the method P300 is a flowchart of a method in which the system 100 selects the target audio processing mode based on the speaker signal, the human voice signal strength in the microphone signal, and at least one of the ambient noise level and the signal-to-noise ratio. The method P300 may include executing, by at least one processor 220:
[0138] S320: Selecting a target audio processing mode of the electronic device 200 from the first mode 1 and the second mode 2 based at least on the speaker signal. Specifically, step S320 may include:
[0139] S322: Generate a control signal corresponding to the speaker signal based at least on the strength of the speaker signal. The control signal includes a first control signal or a second control signal. Step S320 may be that the electronic device 200 generates a corresponding control signal based on the strength of the speaker signal, the noise in the microphone signal, and the strength of the human voice signal in the microphone signal. Specifically, step S322 may include:
[0140] S322-2: Obtain evaluation parameters of the speaker signal and the microphone signal. The evaluation parameters may include an environmental noise evaluation parameter in the microphone signal, and may also include the human voice signal strength in the microphone signal. The environmental noise evaluation parameter may include at least one of the environmental noise level and the signal-to-noise ratio. The electronic device 200 may obtain the environmental noise evaluation parameter and the human voice signal strength in the microphone signal through the control module 231. Specifically, the electronic device 200 may obtain the evaluation parameters based on at least one of the first audio signal 243 and the second audio signal 245. The electronic device 200 may obtain the human voice signal, the environmental noise level and the signal-to-noise ratio through a noise estimation algorithm, which will not be described in detail in this specification.
[0141] S322-4: Generate the control signal based on the strength of the speaker signal and the evaluation parameter. Specifically, the electronic device 200 may compare the strength of the speaker signal with a preset speaker threshold, compare the environmental noise evaluation parameter with a preset noise evaluation range, and compare the strength of the human voice signal with a preset human voice threshold, and generate the control signal based on the comparison results. Step S322-4 may include one of the following situations:
[0142] S322-5: Determine that the intensity of the loudspeaker signal is higher than a preset loudspeaker threshold, the intensity of the human voice signal exceeds the human voice threshold, and the environmental noise evaluation parameter is outside a preset noise evaluation range, and generate the first control signal;
[0143] S322-6: Determine that the intensity of the loudspeaker signal is higher than the loudspeaker threshold, the intensity of the human voice signal exceeds the human voice threshold, and the environmental noise evaluation parameter is within the noise evaluation range, and generate the second control signal;
[0144] S322-7: Determine that the strength of the loudspeaker signal is higher than the loudspeaker threshold and the strength of the human voice signal is lower than the human voice threshold, and generate the second control signal;
[0145] S322-8: Determine that the intensity of the speaker signal is lower than the speaker threshold and the environmental noise evaluation parameter is outside the noise evaluation range, and generate the first control signal;
[0146] S322-9: Determine that the intensity of the speaker signal is lower than the speaker threshold and the environmental noise evaluation parameter is within the noise evaluation range, and generate the first control signal or the second control signal.
[0147] The environmental noise evaluation parameter being within the noise evaluation range may include at least one of the environmental noise level being lower than a preset environmental noise threshold and the signal-to-noise ratio being higher than a preset signal-to-noise ratio threshold. At this time, the environmental noise is relatively small. The environmental noise evaluation parameter being outside the noise evaluation range may include at least one of the environmental noise level being higher than a preset environmental noise threshold and the signal-to-noise ratio being lower than a preset signal-to-noise ratio threshold. At this time, the environmental noise is relatively large. When the environmental noise evaluation parameter is outside the noise evaluation range, that is, in a high-noise environment, the voice quality of the first target audio 291 is better than that of the second target audio 292. When the environmental noise evaluation parameter is within the noise evaluation range, the voice quality of the first target audio 291 is not much different from that of the second target audio 292. The speaker threshold, the environmental noise threshold, and the signal-to-noise ratio threshold have been described above and will not be repeated here.
[0148] The fact that the human voice signal strength exceeds the human voice threshold indicates that user 002 is speaking. In order to ensure the voice quality of user 002, electronic device 200 may generate the first control signal and reduce the speaker signal to ensure the voice quality of first target audio 291.
[0149] The speaker threshold, the ambient noise threshold, the signal-to-noise ratio threshold, and the human voice threshold may be pre-stored in the electronic device 200 .
[0150] Step S320 may further include:
[0151] S324: Based on the control signal, select the target audio processing mode corresponding to the control signal. The first control signal corresponds to first mode 1. The second control signal corresponds to second mode 2. When the control signal is the first control signal, first mode 1 is selected; when the control signal is the second control signal, second mode 2 is selected.
[0152] When the speaker signal strength exceeds the speaker threshold and the vocal signal strength exceeds the vocal threshold, and the environmental noise evaluation parameter is outside the preset noise evaluation range, this indicates that user 002 is speaking, and the echo and noise are high. To ensure the quality and intelligibility of user 002's voice, electronic device 200 can reduce or even shut down the speaker input signal to speaker 280 to reduce the echo in the microphone signal and ensure the voice quality of the target audio. In this case, the first target audio 291 obtained by signal processing the first audio signal 243 and the second audio signal 245 by the first algorithm 233-1 in first mode 1 has better voice quality than the second target audio 292 obtained by signal processing the second audio signal 245 by the second algorithm 233-8 in second mode 2. Therefore, when the speaker signal strength exceeds the speaker threshold and the vocal signal strength exceeds the vocal threshold, and the environmental noise evaluation parameter is outside the preset noise evaluation range, electronic device 200 generates the first control signal corresponding to first mode 1. In this case, the electronic device 200 can ensure the intelligibility of the voice quality of the near-end user 002. Although part of the speaker input signal is missing, the electronic device 200 can retain most of the voice quality and intelligibility of the speaker input signal, thereby improving the voice communication quality of both parties.
[0153] When the speaker signal strength is higher than the speaker threshold and the vocal signal strength is lower than the vocal threshold or exceeds the vocal threshold, and the environmental noise evaluation parameter is within the preset noise evaluation range, it indicates that user 002 is not speaking at this time, or that user 002 is speaking but the noise level is low. In this case, the voice quality of the first target audio 291 obtained by signal processing the first audio signal 243 and the second audio signal 245 by the first algorithm 233-1 in first mode 1 is worse than the second target audio 292 obtained by signal processing the second audio signal 245 by the second algorithm 233-8 in second mode 2. Therefore, when the speaker signal strength is higher than the speaker threshold and the vocal signal strength is lower than the vocal threshold or exceeds the vocal threshold, and the environmental noise evaluation parameter is within the preset noise evaluation range, the electronic device 200 generates the second control signal corresponding to second mode 2.
[0154] The other situations in step S322-4 are basically the same as those in step S222-4 and will not be repeated here.
[0155] The control signal is generated by the control module 231. Specifically, the electronic device 200 can monitor the intensity of the speaker signal and the evaluation parameter in real time, and compare them with the speaker threshold, the noise evaluation range, and the human voice threshold. The electronic device 200 can also periodically detect the intensity of the speaker signal and the evaluation parameter, and compare them with the speaker threshold, the noise evaluation range, and the human voice threshold. The electronic device 200 can also compare the speaker signal and the evaluation parameter with the speaker threshold, the noise evaluation range, and the human voice threshold when it detects that the intensity of the speaker signal or the evaluation parameter has changed significantly, and the change value exceeds a preset range.
[0156] The method P300 may include executing, by at least one processor 220:
[0157] S340: Processing the microphone signal in the target audio processing mode to generate the target audio, thereby at least reducing the echo in the microphone signal. Specifically, step S340 may include one of the following situations:
[0158] S342: Determine that the control signal is the first control signal, select the first mode 1, perform signal processing on the first audio signal 243 and the second audio signal 245, and generate the first target audio 291. Specifically, step S342 may be the same as step S142, and will not be repeated here.
[0159] S344: Determine that the control signal is the second control signal, select the second mode 2, perform signal processing on the second audio signal 245, and generate the second target audio 292. Specifically, step S344 may be the same as step S144, and will not be repeated here.
[0160] The method P300 may include executing, by at least one processor 220:
[0161] S360: Output the target audio. Specifically, step S360 may be consistent with step S160, and will not be described in detail here.
[0162] The method P300 may further include:
[0163] S380: Based on the control signal, control the intensity of the speaker input signal of the speaker 280. Specifically, step S380 may be consistent with step S180, and will not be described in detail herein.
[0164] Table 3 shows Figure 7The corresponding target audio processing mode result diagram. As shown in Table 3, for the convenience of comparison, we divide the scenarios into 8 scenarios, namely: the first scenario: the near-end sound signal is less than the threshold (for example, user 002 does not make a sound), the speaker signal does not exceed the speaker threshold, and the environmental noise is small; the second scenario: the near-end sound signal is greater than the threshold (for example, user 002 makes a sound), the speaker signal does not exceed the speaker threshold, and the environmental noise is small; the third scenario: the near-end sound signal is less than the threshold (for example, user 002 does not make a sound), the speaker signal exceeds the speaker threshold, and the environmental noise is small; and the fourth scenario: the near-end sound signal is greater than the threshold (for example, user 002 makes a sound), the speaker signal exceeds the The speaker threshold is greater than the threshold, and the ambient noise is relatively low. The fifth scenario is that the near-end sound signal is less than the threshold (for example, user 002 does not make any sound), the speaker signal does not exceed the speaker threshold, and the ambient noise is relatively high. The sixth scenario is that the near-end sound signal is greater than the threshold (for example, user 002 makes a sound), the speaker signal does not exceed the speaker threshold, and the ambient noise is relatively high. The seventh scenario is that the near-end sound signal is less than the threshold (for example, user 002 does not make any sound), the speaker signal exceeds the speaker threshold, and the ambient noise is relatively high. The eighth scenario is that the near-end sound signal is greater than the threshold (for example, user 002 makes a sound), the speaker signal exceeds the speaker threshold, and the ambient noise is relatively high. Whether the near-end sound signal is greater than the threshold can be determined by the control module 231 based on the microphone signal. The near-end sound signal being greater than the threshold can mean that the strength of the audio signal emitted by user 002 exceeds a preset threshold. The target audio processing modes corresponding to the eight scenarios are: the fifth, sixth, and eighth scenarios correspond to the first mode 1; the third, fourth, and seventh scenarios correspond to the second mode 2; and the remaining scenarios correspond to the first mode 1 or the second mode 2.
[0165]
[0166] It should be noted that Methods P200 and P300 are applicable to different application scenarios. When the loudspeaker signal is more important than the near-end speech quality, Method P200 can be selected to ensure the quality and intelligibility of the loudspeaker signal. When the near-end speech quality is more important than the loudspeaker signal, Method P300 can be selected to ensure both the quality and intelligibility of the near-end speech.
[0167] To sum up, the system 100, the method P100, the method P200 and the method P300 can control the target audio processing mode of the electronic device 200 according to the speaker signal for different scenarios, thereby controlling the sound source signal of the electronic device 200, so that the voice quality of the target audio is optimal in any scenario, thereby improving the quality of voice communication.
[0168] It should be noted that the signal strength of the ambient noise varies at different frequencies. The voice quality of the first target audio 291 and the second target audio 292 also varies at different frequencies. For example, at a first frequency, the voice quality of the first target audio 291 obtained after the first audio signal 243 and the second audio signal 245 are processed by the first algorithm 233-1 is better than the voice quality of the second target audio 292 obtained after the second audio signal 245 is processed by the second algorithm 233-8. At frequencies other than the first frequency, the voice quality of the first target audio 291 obtained after the first audio signal 243 and the second audio signal 245 are processed by the first algorithm 233-1 is similar to the voice quality of the second target audio 292 obtained after the second audio signal 245 is processed by the second algorithm 233-8. In this case, the electronic device 200 can further generate the control signal based on the frequency of the ambient noise, generating the first control signal at the first frequency and the second control signal at frequencies other than the first frequency.
[0169] When the environmental noise is low-frequency noise (such as in some situations such as subways and buses), it may happen that the first target audio 291 obtained by the first audio signal 243 and the second audio signal 245 under the signal processing of the first algorithm 233-1 has poor voice signal quality at low frequencies, that is, the first target audio 291 has poor voice intelligibility at low frequencies, but high voice intelligibility at high frequencies. At this time, the electronic device 200 can control the selection of the target audio processing mode according to the frequency of the environmental noise. For example, in the low-frequency range, the electronic device 200 can select the method P300 to control the target audio processing mode to ensure that the voice of the near-end user 002 is picked up, thereby ensuring the near-end voice quality; in the high-frequency range, the electronic device 200 can select the method P200 to control the target audio processing mode to ensure that the near-end user 002 can hear the speaker signal.
[0170] Another aspect of this specification provides a non-transitory storage medium storing at least one set of executable instructions for controlling an audio source signal based on a signal. When the executable instructions are executed by a processor, the executable instructions direct the processor to implement the steps of the audio signal processing method for echo suppression described in this specification. In some possible implementations, various aspects of this specification may also be implemented in the form of a program product comprising program code. When the program product is executed on an electronic device 200, the program code is configured to cause the electronic device 200 to perform the steps of controlling an audio source signal based on a signal described in this specification. The program product for implementing the above method may utilize a portable compact disc read-only memory (CD-ROM) comprising the program code and may be executed on the electronic device 200. However, the program product of this specification is not limited thereto. In this specification, a readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system (e.g., processor 220). The program product may utilize any combination of one or more readable media. The readable medium may be a readable signal medium or a readable storage medium. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination thereof. More specific examples of computer-readable storage media include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. The computer-readable storage medium may include a data signal propagated in baseband or as part of a carrier wave, which carries readable program code. This propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The computer-readable storage medium may also be any readable medium other than a computer-readable storage medium that can transmit, propagate, or transfer a program for use by or in conjunction with an instruction execution system, device, or component. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination thereof. Program code for performing the operations described herein may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the electronic device 200, partially on the electronic device 200, as a stand-alone software package, partially on the electronic device 200 and partially on a remote computing device, or entirely on the remote computing device.
[0171] The foregoing description of this specification describes specific embodiments. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the particular order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0172] In summary, after reading this detailed disclosure, those skilled in the art will appreciate that the foregoing detailed disclosure may be presented by way of example only and may not be limiting. Although not expressly stated herein, those skilled in the art will understand that this specification encompasses various reasonable changes, improvements, and modifications to the embodiments. Such changes, improvements, and modifications are intended to be suggested by this specification and are within the spirit and scope of the exemplary embodiments of this specification.
[0173] Furthermore, certain terms in this specification have been used to describe embodiments of this specification. For example, “one embodiment,” “an embodiment,” and / or “some embodiments” mean that a particular feature, structure, or characteristic described in connection with that embodiment can be included in at least one embodiment of this specification. Therefore, it is emphasized and should be understood that two or more references to “an embodiment,” “one embodiment,” or “an alternative embodiment” in various parts of this specification do not necessarily refer to the same embodiment. Furthermore, particular features, structures, or characteristics may be appropriately combined in one or more embodiments of this specification.
[0174] It should be understood that in the foregoing description of the embodiments of this specification, in order to help understand one feature, this specification combines various features in a single embodiment, drawing or description thereof for the purpose of simplifying this specification.
Claims
1. An audio signal processing method for suppressing echo, characterized in that: include: Acquire a speaker signal from a speaker, and select a target audio processing mode for the electronic device from a plurality of audio processing modes based on a magnitude relationship between a signal strength of the speaker signal and a preset speaker threshold, wherein the speaker signal is an audio signal sent by a control device to the electronic device; Processing a microphone signal in the target audio processing mode to generate a target audio, thereby at least reducing an echo in the target audio, wherein the microphone signal is at least two audio signals with different voice qualities output by a microphone module acquired by the electronic device; as well as The target audio signal is output.
2. The audio signal processing method according to claim 1, wherein: The microphone module includes at least one first type microphone and at least one second type microphone; The at least one first type microphone outputs a first audio signal; and The at least one second type microphone outputs a second audio signal, The microphone signal includes the first audio signal and the second audio signal.
3. The audio signal processing method according to claim 2, wherein: The at least one first type microphone is used to collect human body vibration signals; and The at least one second type microphone is used to collect air vibration signals.
4. The audio signal processing method according to claim 2, wherein: The multiple audio processing modes include at least: In a first mode, signal processing is performed on the first audio signal and the second audio signal; and In the second mode, signal processing is performed on the second audio signal.
5. The audio signal processing method according to claim 4, wherein: The selecting a target audio processing mode of the electronic device from a plurality of audio processing modes based on a magnitude relationship between the signal strength of the speaker signal and a preset speaker threshold comprises: generating a control signal corresponding to the speaker signal based at least on the strength of the speaker signal, the control signal comprising a first control signal or a second control signal; and Based on the control signal, a target audio processing mode corresponding to the control signal is selected, wherein the first mode corresponds to the first control signal and the second mode corresponds to the second control signal.
6. The audio signal processing method according to claim 5, wherein: The step of generating a control signal corresponding to the speaker signal based at least on the strength of the speaker signal comprises: determining that the intensity of the loudspeaker signal is lower than a preset loudspeaker threshold, and generating the first control signal; or It is determined that the strength of the loudspeaker signal is higher than the loudspeaker threshold, and the second control signal is generated.
7. The audio signal processing method according to claim 5, wherein: The step of generating a control signal corresponding to the speaker signal based at least on the strength of the speaker signal comprises: A corresponding control signal is generated based on the strength of the loudspeaker signal and the microphone signal.
8. The audio signal processing method according to claim 7, wherein: The generating a corresponding control signal based on the strength of the speaker signal and the microphone signal includes: Acquiring evaluation parameters of the microphone signal, the evaluation parameters including an environmental noise evaluation parameter, the environmental noise evaluation parameter including at least one of an environmental noise level and a signal-to-noise ratio; and The control signal is generated based on the strength of the speaker signal and the evaluation parameter.
9. The audio signal processing method according to claim 8, wherein: The generating of the control signal based on the intensity of the speaker signal and the evaluation parameter includes one of the following situations: determining that the intensity of the speaker signal is higher than a preset speaker threshold, and generating the second control signal; determining that the intensity of the speaker signal is lower than the speaker threshold and the environmental noise evaluation parameter is outside a preset noise evaluation range, and generating the first control signal; as well as It is determined that the intensity of the loudspeaker signal is lower than the loudspeaker threshold and the environmental noise evaluation parameter is within the noise evaluation range, and the first control signal or the second control signal is generated.
10. The audio signal processing method according to claim 9, wherein: The environmental noise evaluation parameter is within the noise evaluation range, including at least one of the following situations: The ambient noise level is lower than a preset ambient noise threshold; and The signal-to-noise ratio is higher than a preset signal-to-noise ratio threshold.
11. The audio signal processing method according to claim 8, wherein: The evaluation parameter further includes the strength of a human voice signal, and the generating of the control signal based on the strength of the loudspeaker signal and the evaluation parameter includes one of the following situations: determining that the intensity of the loudspeaker signal is higher than a preset loudspeaker threshold, the intensity of the human voice signal exceeds a preset human voice threshold, and the environmental noise evaluation parameter is outside a preset noise evaluation range, and generating the first control signal; determining that the intensity of the loudspeaker signal is higher than the loudspeaker threshold, the intensity of the human voice signal exceeds the human voice threshold, and the environmental noise evaluation parameter is within the noise evaluation range, and generating the second control signal; determining that the strength of the loudspeaker signal is higher than the loudspeaker threshold and the strength of the human voice signal is lower than the human voice threshold, and generating the second control signal; determining that the intensity of the speaker signal is lower than the speaker threshold and the environmental noise evaluation parameter is outside the noise evaluation range, and generating the first control signal; as well as It is determined that the intensity of the loudspeaker signal is lower than the loudspeaker threshold and the environmental noise evaluation parameter is within the noise evaluation range, and the first control signal or the second control signal is generated.
12. The audio signal processing method according to claim 11, wherein: The environmental noise evaluation parameter is within the noise evaluation range, including at least one of the following situations: The ambient noise level is lower than a preset ambient noise threshold; and The signal-to-noise ratio is higher than a preset signal-to-noise ratio threshold.
13. The audio signal processing method according to claim 5, wherein: Generating the target audio includes: performing signal processing on the first audio signal and the second audio signal using a first algorithm in the first mode to generate a first target audio signal; or performing signal processing on the second audio signal by a second algorithm in the second mode to generate a second target audio signal, The target audio includes one of the first target audio and the second target audio.
14. The audio signal processing method according to claim 13, wherein: The outputting the target audio includes: performing smoothing processing on the target audio, and performing the smoothing processing on a connection between the first target audio and the second target audio when the target audio switches between the first target audio and the second target audio; and The target audio that has undergone the smoothing process is output.
15. The audio signal processing method according to claim 5, wherein: The method further comprises: Based on the control signal, the strength of the speaker input signal of the speaker is controlled.
16. The audio signal processing method according to claim 15, wherein: The controlling the intensity of the speaker input signal of the speaker based on the control signal includes: The control signal is determined to be the first control signal, and the intensity of the speaker input signal input to the speaker is reduced, thereby reducing the intensity of the sound output by the speaker.
17. A system for audio signal processing for echo suppression, characterized in that: include: at least one storage medium storing at least one instruction set for audio signal processing for echo suppression; as well as at least one processor, in communication with the at least one storage medium; Wherein, when the system is running, the at least one processor reads the at least one instruction set and executes the audio signal processing method for suppressing echo according to any one of claims 1 to 16 according to instructions of the at least one instruction set.
Citation Information
Patent Citations
Active noise cancellation decisions in a portable audio device
CN102870154A