A method, apparatus, system and intelligent speaker device for processing an audio signal

By placing a vibration conduction microphone or sensor near the speaker of a smart device to collect the audio signal played by the speaker, and combining digital and analog signals for echo cancellation, the problem of low voice recognition rate caused by unsatisfactory echo cancellation in smart devices is solved, and the accuracy and reliability of voice wake-up are improved.

CN118972744BActive Publication Date: 2026-02-27ZHEJIANG FUTURE ELF ARTIFICIAL INTELLIGENCE TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202410877429.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-01
Publication Date
2026-02-27
Estimated Expiration
2044-07-01

AI Technical Summary

Technical Problem

In existing technologies, the echo cancellation of smart devices is not ideal, resulting in low voice recognition rates, especially when playing at high volumes, the signal-to-return ratio is very low, making it impossible to effectively achieve voice wake-up.

Method used

By placing a bone conduction microphone or vibration sensor near the speaker of a smart device to collect the audio signal played by the speaker as a reference signal, and combining the digital audio signal and the analog audio signal for echo cancellation processing, the target audio signal input by the user is preserved.

Benefits of technology

It improves speech recognition and wake-up rates, ensures more accurate echo cancellation in both high and low frequency bands, and enhances the reliability and effectiveness of voice wake-up.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118972744B_ABST
    Figure CN118972744B_ABST
Patent Text Reader

Abstract

The application discloses a processing method, device and system of an audio signal and an intelligent sound box device, wherein the processing method comprises the following steps: acquiring a mixed signal comprising a first audio signal and a second audio signal collected by a first input end of an intelligent device and a second audio collection signal collected by a second input end of the intelligent device, wherein the first audio signal is a target audio signal which needs to be further processed, the second audio signal is an audio signal output by an audio output end of the intelligent device, and the second audio collection signal is an audio signal obtained by collecting the second audio signal output by the audio output end of the intelligent device in a vibration conduction mode; taking the second audio collection signal as a reference signal, performing echo cancellation processing on the mixed signal, and determining the target audio signal after echo cancellation; and performing corresponding processing on the target audio signal according to a to-be-processed task, so that the accuracy of echo cancellation is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of speech signal processing, in particular to an audio signal processing method and device. The present application also relates to an audio signal processing system, an intelligent sound box device, an electronic device and a computer storage medium. BACKGROUND

[0002] Voice wake-up is a convenient and efficient human-computer interaction method, which can wake up a device by specifying a word and perform various operations, and has been widely applied to various intelligent devices, such as smart phones, intelligent sound boxes, smart home devices, etc. The so-called wake-up means to make the device enter the working state from the standby state, or to make the device enter another working state from the current working state, and to start listening, recognizing and responding to the user's speech, etc.

[0003] Voice wake-up technology has a wide range of application scenarios, for example, in the application scenario of smart home, users can activate and control intelligent home devices by specific wake-up words such as "turn on the TV" and "turn off the light", without using remote controls or mobile phone APPs and other intermediaries. In the application scenario of vehicle-mounted systems, voice wake-up technology enables drivers to activate the vehicle-mounted system through wake-up words, or directly issue commands such as making or receiving calls, navigation, controlling air conditioning temperature, etc., thereby ensuring driving safety. In the application scenario of wearable devices, voice wake-up technology is also applied to wearable devices such as smart watches and smart earphones, enabling users to interact with the devices through voice and complete various operations. In the application scenario of intelligent customer service and assistants, voice wake-up technology is used to provide voice assistant services in intelligent customer service and intelligent assistants, such as answering questions and providing information. In the application scenario of robots, voice wake-up technology can help robots better interact with humans, and control the behavior of robots through voice instructions. Of course, it can also provide more convenient and intelligent services for users in the application scenarios of medical devices, industrial automation, remote education, etc. SUMMARY

[0004] The present application provides an audio signal processing method to solve the problem of low voice recognition rate caused by unsatisfactory echo cancellation in intelligent devices in the prior art.

[0005] The present application provides an audio signal processing method, comprising:

[0006] Acquire a mixed signal including a first audio signal and a second audio signal collected by a first input terminal of the intelligent device, and a second audio collection signal collected by a second input terminal of the intelligent device, wherein the first audio signal is a target audio signal that needs to be further processed, the second audio signal is an audio signal output by an audio output terminal of the intelligent device, and the second audio collection signal is an audio signal obtained by collecting the second audio signal output by the audio output terminal of the intelligent device in a vibration conduction manner;

[0007] Perform echo cancellation processing on the mixed signal by taking the second audio collection signal as a reference signal to determine the target audio signal after echo cancellation;

[0008] According to the to-be-processed task, the target audio signal is processed accordingly.

[0009] Optionally, the method further comprises:

[0010] Acquire a digital audio signal of an audio signal of a sound source before digital-to-analog conversion processing, and / or acquire an analog audio signal of the audio signal of the sound source after digital-to-analog conversion processing and power amplification processing.

[0011] Optionally, the method further comprises:

[0012] The digital audio signal and the analog audio signal, and the second audio collection signal are determined as the reference signal; or the digital audio signal and the second audio collection signal are determined as the reference signal; or the analog audio signal and the second audio collection signal are determined as the reference signal.

[0013] According to the reference signal, the mixed signal is processed by echo cancellation to determine the target audio signal.

[0014] Optionally, the method further comprises:

[0015] The digital audio signal and the analog audio signal are fused with the second audio collection signal to determine the reference information; or the digital audio signal is fused with the second audio collection signal to determine the reference information; or the analog audio signal is fused with the second audio collection signal to determine the reference information.

[0016] Optionally, the processing of the target audio signal according to the to-be-processed task comprises:

[0017] According to the voice wake-up task, the target audio signal is identified to determine whether the target audio signal is a wake-up signal.

[0018] If yes, the smart device is woken up.

[0019] The application further provides an audio signal processing device, comprising:

[0020] An acquisition unit is configured to acquire a mixed signal comprising a first audio signal and a second audio signal collected by a first input end of a smart device, and the second audio signal collected by a second input end of the smart device, wherein the first audio signal is a target audio signal requiring further processing, the second audio signal is an audio signal output by an audio output end of the smart device, and the second audio signal collected is an audio signal obtained by collecting the second audio signal in a vibration conduction manner at the audio output end of the smart device.

[0021] A determination unit is configured to perform echo cancellation processing on the mixed signal by taking the second audio signal collected as a reference signal, and determine the target audio signal after echo cancellation.

[0022] A processing unit is configured to process the target audio signal according to a to-be-processed task.

[0023] The application further provides an audio signal processing system, comprising:

[0024] A first collection end provided on a smart device is configured to collect a mixed signal comprising a first audio signal and a second audio signal; wherein the first audio signal is a target audio signal requiring further processing, and the second audio signal is an audio signal played by a loudspeaker of the smart device.

[0025] A second collection end provided on the smart device is configured to collect a second audio signal collected by the loudspeaker of the smart device, which is an audio signal obtained by collecting the second audio signal played by the audio loudspeaker of the smart device in a vibration conduction manner, and the distance between the second collection end and the loudspeaker satisfies the requirement of collecting the second audio signal collected in a vibration conduction manner by the second collection end.

[0026] An echo cancellation processing end is configured to perform echo cancellation processing on the mixed signal by taking the second audio signal collected as a reference signal, and send the processed signal to a signal processing end as the target audio signal.

[0027] The signal processing end is configured to perform corresponding processing on the target audio signal according to the to-be-processed task.

[0028] The application further provides an intelligent sound box device, which comprises a first collecting end, a second collecting end, a loudspeaker, an echo cancellation processing end and a signal processing end.

[0029] The first collecting end is arranged on the intelligent sound box device and collects a mixed signal comprising a first audio signal and a second audio signal.

[0030] The second collecting end is arranged on the intelligent sound box device and collects a second audio collecting signal played by the loudspeaker of the intelligent device.

[0031] The echo cancellation processing end is configured to perform echo cancellation processing on the mixed signal by taking the second audio collecting signal as a reference signal, and determine the target audio signal after echo cancellation.

[0032] The signal processing end is configured to perform corresponding processing on the target audio signal according to the to-be-processed task.

[0033] Optionally, the signal processing end comprises: identifying the target audio signal according to a voice wake-up task, and determining whether the target audio signal is a wake-up signal.

[0034] The application further provides an electronic device, which comprises:

[0035] a processor;

[0036] a memory configured to store a program for processing data generated by the electronic device, wherein the program, when read and executed by the processor, performs the processing method of the audio signal as described above.

[0037] The application further provides a computer storage medium comprising a computer program, which, when running on an electronic device, causes the electronic device to perform the processing method of the audio signal as described above.

[0038] Compared with the prior art, the application has the following advantages:

[0039] The processing method of the audio signal provided by the application can collect the second audio collection signal played by the loudspeaker through the setting of the bone conduction microphone or vibration sensor and the like collection device with the vibration conduction principle in the vicinity of the loudspeaker of the smart device, and take the second audio collection signal as the reference signal. The echo cancellation processing is performed on the mixed signal of the target audio signal, i.e. the first audio signal collected by the conventional microphone in the smart device and the second audio signal played by the loudspeaker, the voice signal played by the loudspeaker in the mixed signal is removed, the target audio signal is retained, and the corresponding processing task is performed based on the target audio signal. The processing process can utilize the vibration conduction principle, on the one hand, the microphone as the second input end of the smart device has a more extensive setting position, which is not limited to the cavity of the voice signal transmission channel, and can meet the range that the loudspeaker can collect. On the other hand, the second input end of the smart device can collect the voice signal played by the loudspeaker of the smart device, and is not sensitive to other voice signals that are not in the collection range, so as to ensure the purity of the reference signal and improve the accuracy of the echo cancellation. Furthermore, in order to further improve the accuracy of the reference signal, the digital audio signal on the path between the sound source and the loudspeaker and / or the analog audio signal after the digital audio signal is converted from digital to analog and then amplified by power can be collected, one or both of the two and the second audio collection signal played by the loudspeaker collected by the second input end are taken as the reference signal for echo cancellation processing, so as to realize the complement of the voice signal in high and low frequencies and ensure the integrity and accuracy of the reference signal. BRIEF DESCRIPTION OF DRAWINGS

[0040] Figure 1 is a flowchart of the processing method of the audio signal provided by the application;

[0041] Figure 2 is a structural schematic diagram of the working principle embodiment of the processing method of the audio signal provided by the application;

[0042] Figure 3 is a structural schematic diagram of the working principle embodiment of the processing method of the audio signal provided by the application;

[0043] Figure 4 is a structural schematic diagram of the processing device of the audio signal provided by the application;

[0044] Figure 5 is a structural schematic diagram of the electronic device provided by the application. DETAILED DESCRIPTION

[0045] In the following description, numerous specific details are set forth in order to provide a thorough understanding of the present application. However, the present application can be practiced without the specific details. In other instances, well-known methods, procedures, components, and circuits have not been described in detail so as not to obscure the present application. The present application can also be practiced with different and / or additional components not depicted, or with only a subset of the components shown. Only the claims are active.

[0046] The terminology used in this application is for the purpose of describing particular embodiments only and is not intended to be limiting of this application. As used herein, the description using terms such as "a", "an" and "the" are not intended to refer to only a singular entity but include the general class of which a specific one can be an instance. Additionally, the use of "first", "second", "third" and the like does not indicate any ordering or sequence but is used for naming purposes only.

[0047] According to the above background art, the inventive concept of the present application is derived from the application scenario of voice wake-up. Specifically, in the scenario where voice wake-up is needed, voice wake-up is performed while the loudspeaker is playing audio, which results in low voice wake-up and / or voice recognition rate in the playing scenario, and ultimately leads to wake-up failure or low wake-up rate. For example, when the smart device is playing a song, the user needs to switch the song or switch to other functions, at which time the user needs to wake up the smart device. Therefore, the smart audio device with voice wake-up function provided in the prior art generally uses an acoustic echo canceller (AEC) for voice recognition. However, the conventional acoustic echo canceller (AEC) directly uses the signal 1 after the power amplifier (PA) processing or the signal 2 before the digital-to-analog conversion (DAC) as the reference signal, but there is a large difference between the reference signal and the signal affected by the non-linear distortion of the loudspeaker and the acoustic cavity structure, which makes the echo cancellation effect not good, and further leads to low voice recognition rate. Especially when playing at a large volume, the loudspeaker sound picked up by the microphone is much larger than the user input voice wake-up signal, which makes the signal echo ratio (SER) very low, thereby further reducing the voice recognition rate and failing to achieve voice wake-up of the smart device. Of course, based on improving the audio signal recognition rate and accuracy, those skilled in the art can apply it to any desired scenario other than voice wake-up. For example, in a noisy or complex scenario, the target audio signal needed is recognized to output other multimedia data or text data, i.e., the target audio signal is used as a trigger event for other services.

[0048] Therefore, the application provides a processing method of an audio signal, which can improve the recognition rate of the audio signal in a regular volume playing scene or a large volume playing scene. Before the processing method of the audio signal is described in detail, related technical terms or technical terms involved in the technical solution of the application are explained, so that the implementation principle of the processing method of the audio signal provided by the application can be better understood.

[0049] AEC: Acoustic Echo Canceller, AEC is based on the correlation between the loudspeaker signal and the multi-path echo generated by it, to establish a speech model of the far-end signal, to estimate the echo using it, and to continuously modify the coefficients of the filter to make the estimated value more close to the real echo. Then, the echo estimation value is subtracted from the input signal of the microphone (microphone), so as to achieve the purpose of eliminating echo.

[0050] DAC: Digital-to-Analog Converter, is an electronic component that converts digital signals (usually binary numbers) into analog signals (such as current, voltage or charge). In electronic systems, the input of DAC is digital signal (usually binary code), and the output is analog signal (such as voltage or current) proportional to the amplitude of the digital signal. DAC is a key component in many applications, such as audio and video systems, radar systems, wireless communication, control systems, test equipment, medical imaging, high-speed data acquisition, etc. In audio applications, DAC converts digital audio signals into analog audio signals for playback through speakers.

[0051] ADC: Analog-to-Digital Converter, is an electronic component that converts analog signals into digital signals. In audio applications, analog audio signals from various audio sources such as microphones, speakers (or players) are converted into digital signals to provide a basis for subsequent audio processing, storage and transmission, etc.

[0052] PA: Power Amplifier, used to amplify signals.

[0053] Bone conduction microphone: a microphone that transmits signals through physical conduction. Bone conduction microphone uses the principle of sound transmission through human bones. When sound is transmitted through bones, the bones will produce tiny mechanical vibrations. This vibration is detected and converted into an audio signal by a piezoelectric vibration sensor. Unlike conventional microphones, bone conduction microphones do not rely on air-borne sound, so they are not affected by external environmental noise.

[0054] Keyword Spotting (KWS): Real-time detection of a specific segment of a speaker within a continuous audio stream. In other words, it's a technology that identifies specific keywords or phrases within a continuous audio signal. This technology allows smart devices to enter working mode from standby mode by detecting a specific wake-up word uttered by the user, thereby responding to the user's voice commands.

[0055] Awakening rate: The probability that a wake word will be activated. The higher the awakening rate, the better the effect. It is usually expressed as a percentage.

[0056] Vibration sensor: A device that converts mechanical vibration into electrical signals. The vibration sensor uses an internal piezoelectric ceramic plate and spring-loaded weight structure to sense parameters of mechanical motion vibration (such as vibration velocity, frequency, acceleration, etc.) and convert them into usable output signals.

[0057] like Figure 1 As shown, Figure 1 This is a flowchart of an audio signal processing method provided in this application; the processing method includes:

[0058] Step S101: Acquire a mixed signal including a first audio signal and a second audio signal collected by the first input terminal of the smart device, and a second audio acquisition signal collected by the second input terminal of the smart device, wherein the first audio signal is the target audio signal that needs further processing, the second audio signal is the audio signal output by the audio output terminal of the smart device, and the second audio acquisition signal is the audio signal obtained by collecting the second audio signal output by the audio output terminal of the smart device using a vibration conduction method;

[0059] Step S102: Using the second audio acquisition signal as a reference signal, perform echo cancellation processing on the mixed signal to determine the target audio signal after echo cancellation;

[0060] Step S103: Process the target audio signal accordingly based on the task to be processed.

[0061] The steps S101 to S103 described above will be described in detail below with reference to specific embodiments.

[0062] First embodiment:

[0063] like Figure 2 As shown, Figure 2 This is a schematic diagram illustrating the working principle of an embodiment of the voice wake-up signal processing method provided in this application.

[0064] In step S101, a mixed signal including a first audio signal and a second audio signal collected by a first input of the intelligent device and a second audio collection signal collected by a second input of the intelligent device are obtained, wherein the first audio signal is a target audio signal that needs to be further processed, the second audio signal is an audio signal output by an audio output of the intelligent device, and the second audio collection signal is an audio signal collected by using a vibration conduction method on the second audio signal output by the audio output of the intelligent device.

[0065] As shown in Figure 2 The first input of the intelligent device can be an external conventional microphone of the intelligent device, that is, through air propagation, for example, a microphone of a mobile phone, a microphone of a smart sound box device, etc. The first input can include one or more, and the collected is a first audio signal (which can also be referred to as a first voice signal). The second input is a bone conduction microphone or a vibration sensor installed near the loudspeaker of the intelligent device, which can collect the voice signal played by the loudspeaker of the intelligent device, that is, a second audio signal (which can also be referred to as a second voice signal), that is, a collection device with a vibration conduction method. The installation distance satisfies the collection of the audio signal played by the loudspeaker of the second input.

[0066] The purpose of step S101 is to collect the first audio signal through the first input, that is, to wake up the audio signal of the intelligent device, and to collect the second audio signal from the loudspeaker of the intelligent device, and to take the two as the mixed signal collected by the first input. The second audio collection signal (which can be referred to as the second voice collection signal) collected by the second input from the loudspeaker of the intelligent device is taken as a reference signal.

[0067] The audio signal played by the loudspeaker collected by the second input through the vibration conduction method, because the principle of collecting the audio signal by the second input is different from that of the first input, therefore, the second input has no sensitivity to the first audio signal when collecting the second audio collection signal, thereby providing a more stable basis for subsequent echo cancellation processing of the mixed signal according to the reference signal.

[0068] In step S102, the second audio collection signal is taken as a reference signal to perform echo cancellation processing on the mixed signal, and the target audio signal after echo cancellation is determined.

[0069] Specifically, the reference signal is converted from an analog signal to a digital signal through analog-to-digital conversion, and is sent to an echo cancellation processing end. The echo cancellation processing end filters and cancels the second audio collection signal (REF0) in the mixed signal according to the received reference signal and mixed signal. For example, the coherence of the mixed signal is detected according to the reference signal to obtain the coherence value of the reference signal and the mixed signal in the corresponding frequency band. The signal component in the corresponding frequency band in the mixed signal is processed according to the coherence value to obtain the first voice signal, that is, the voice wake-up signal (which can also be understood as the audio signal of the voice wake-up signal) used to wake up the intelligent device. Therefore, the echo cancellation processing belongs to the prior art, and will not be described in detail here. It can be clearly understood that in the embodiment, the first and second audio signals collected by the first input end are used as the mixed signal (Mic signal), and the second audio collection signal (REF0) collected by the second input end is used as the reference signal. The second audio signal in the mixed signal is cancelled according to the reference signal, so that the first audio signal can be retained, which can be understood as a voice signal input by a user. Of course, the source of the first audio signal can be different in different scenarios. Therefore, after echo cancellation, a purer first audio signal can be obtained, and the problem of mistakenly deleting the first audio signal in the echo cancellation process can be avoided.

[0070] After step S102 is performed, step S103 is performed, that is, the target audio signal is processed according to the to-be-processed task. The target audio can be processed according to different to-be-processed tasks. In a voice wake-up scenario, for example, when an intelligent sound box is playing audio and wakes up the intelligent sound box to start other tasks, or when voice navigation is performed through a mobile phone and other tasks are performed on a wake-up voice assistant, etc., the target audio signal can be processed, which can include: identifying the target audio signal according to a voice wake-up task to determine whether the target audio signal is a wake-up signal; and if so, waking up the intelligent device. In an interactive scenario, the target audio signal can be output in text or other forms on an interactive interface. In an entertainment game scenario, the target audio signal can be processed to trigger a corresponding function, and the like. Therefore, the embodiments of the present application are described only in the voice wake-up manner, and are not used to limit the application scenarios of the target audio signal obtained by the audio signal processing method.

[0071] Second embodiment:

[0072] As Figure 3 shown, Figure 3 is a structural schematic diagram of another embodiment of the working principle of the audio signal processing method provided by the present application.

[0073] To further improve the accuracy of echo cancellation, the following can also be included:

[0074] Step S10a: obtaining the digital audio signal (REF2) of the audio signal of the sound source before digital-to-analog conversion processing, and / or, obtaining the analog audio signal (REF1) of the audio signal of the sound source after digital-to-analog conversion processing and power amplification processing.

[0075] The specific implementation process of the step S102 can include:

[0076] Step S102-1: determining the digital audio signal (REF2) and the analog audio signal (REF1), and the second audio acquisition signal (REF0) collected by the second input end as the reference signal; or, determining the digital audio signal (REF2) and the second audio acquisition signal (REF0) collected by the second input end as the reference signal; or, determining the analog audio signal (REF1) and the second audio acquisition signal (REF0) collected by the second input end as the reference signal.

[0077] Step S102-2: performing echo cancellation processing on the mixed signal according to the reference signal to determine the target audio signal. That is, the echo cancellation processing can include at least three implementation schemes:

[0078] The first is to include the digital audio signal (REF2), the analog audio signal (REF1) and the second audio acquisition signal (REF0) collected by the second input end as the reference signal, and to use the three audio signals as the reference signal to cancel the second audio signal in the mixed signal and retain the first audio signal in the mixed signal, i.e., the target audio signal; wherein the analog audio signal (REF1) and the second audio acquisition signal (REF0) are converted into digital signals after analog-to-digital conversion, and are sent to the echo cancellation processing system together with the digital audio signal (REF2) as reference information for echo cancellation processing. Specifically, the digital audio signal, the analog audio signal and the second audio acquisition signal collected by the second input end can be fused, and the fused audio signal can be used as the reference signal, for example: linear combination, linear transformation combination, frequency domain linear superposition, etc. of the digital audio signal, the analog audio signal and the second audio acquisition signal, that is, the fusion processing means can be used to integrate and / or enhance the audio signal.

[0079] The second kind is based on the reference signal including the digital audio signal and the second audio acquisition signal collected by the second input end, eliminating the second audio signal in the mixed signal, and retaining the first audio signal in the mixed signal, that is, the target audio signal. Specifically, the digital audio signal and the second audio acquisition signal collected by the second input end can be fused, and the fused audio signal can be used as the reference signal.

[0080] The third kind is based on the reference signal including the analog audio signal and the second audio acquisition signal collected by the second input end, eliminating the second audio signal in the mixed signal, and retaining the first audio signal in the mixed signal, that is, the target audio signal. Specifically, the analog audio signal and the second audio acquisition signal collected by the second input end can be fused, and the fused audio signal can be used as the reference signal.

[0081] In this embodiment, considering the strength of the audio signal, one or both of the digital audio signal and the analog audio signal and the second audio acquisition signal collected by the second input end are used as the reference signal, which can form a complement between low frequency and high frequency, so as to obtain better restoration in the full frequency band, improve the accuracy of elimination when the mixed signal is processed by echo elimination in the subsequent process, and further improve the processing efficiency in the corresponding task processing scene, such as voice wake-up rate, voice recognition rate, etc.

[0082] The above is a description of an embodiment of the audio signal processing method provided by the present application. As can be seen from the above, the audio signal processing method provided by the present application can set a collection device such as a bone conduction microphone or a vibration sensor with a vibration conduction principle near a loudspeaker of a smart device, collect a voice signal played through the loudspeaker, take the voice signal as a reference signal, perform echo cancellation processing on a mixed signal of a voice signal input by a user collected by a regular microphone and a voice signal played by the loudspeaker, remove the voice signal played by the loudspeaker from the mixed signal, retain the voice signal input by the user collected by the regular microphone, and send the voice signal as an audio signal to a signal processing wake-up end to wake up the smart device. The processing process can utilize the vibration conduction principle, on the one hand, so that the microphone as a second input end of the smart device has a more extensive setting position and does not need to be limited to a cavity of a voice signal transmission channel, but can meet a range that can be collected by the loudspeaker. On the other hand, the second input end of the smart device can collect the voice signal played by the loudspeaker of the smart device, and is not sensitive to the voice signal input by the user, so that the purity of the reference signal can be ensured, and the accuracy of echo cancellation is improved. Furthermore, to further improve the accuracy of the reference signal, a digital audio signal collected on a path between a sound source and the loudspeaker and / or an analog audio signal after digital-to-analog conversion and power amplification of the digital audio signal can be taken as the reference signal together with the voice signal played by the loudspeaker collected by the second input end to perform echo cancellation processing, so that the voice signal can be complemented in high and low frequencies, and the integrity and accuracy of the reference signal are ensured.

[0083] The above is a specific description of an embodiment of the audio signal processing method provided by the present application. Corresponding to the aforementioned embodiment of the audio signal processing method, the present application also discloses an embodiment of an audio signal processing device. Please refer to Figure 4 Since the device embodiment is basically similar to the method embodiment, the description is relatively simple, and the related parts can be referred to the part of the method embodiment. The device embodiment described below is only illustrative.

[0084] As Figure 4 shown, Figure 4 is a structural schematic diagram of an audio signal processing device provided by the present application. The device comprises:

[0085] The acquisition unit 401 is configured to acquire a mixed signal including a first audio signal and a second audio signal collected by a first input end of a smart device, and the second audio signal collected by a second input end of the smart device, wherein the first audio signal is a target audio signal that needs to be further processed, the second audio signal is an audio signal output by an audio output end of the smart device, and the second audio signal is an audio signal obtained by collecting the second audio signal in a vibration conduction manner at the audio output end of the smart device.

[0086] The acquisition unit can also be configured to acquire a digital audio signal of an audio signal of a sound source before digital-to-analog conversion processing, and / or acquire an analog audio signal of the audio signal of the sound source after digital-to-analog conversion processing and power amplification processing.

[0087] The determination unit 402 is configured to perform echo cancellation processing on the mixed signal by taking the second audio signal as a reference signal, and determine the target audio signal after echo cancellation.

[0088] The determination unit 402 can include a first determination subunit and a second determination subunit. The first determination subunit is configured to determine the digital audio signal and the analog audio signal, and the second audio signal as the reference signal, or determine the digital audio signal and the second audio signal as the reference signal, or determine the analog audio signal and the second audio signal as the reference signal. The second determination subunit is configured to perform echo cancellation processing on the mixed signal according to the reference signal, and determine the target audio signal.

[0089] The first determination subunit is specifically configured to fuse the digital audio signal and the analog audio signal with the second audio signal to determine the reference information, or fuse the digital audio signal with the second audio signal to determine the reference information, or fuse the analog audio signal with the second audio signal to determine the reference information.

[0090] The processing unit 403 is configured to perform corresponding processing on the target audio signal according to a to-be-processed task.

[0091] The description of the audio signal processing device can refer to the description of the audio signal processing method, which will not be described in detail here.

[0092] Based on the above, the present application further provides an audio signal processing system, which includes:

[0093] A first collection end arranged on the smart device collects a mixed signal including a first audio signal and a second audio signal; the first audio signal is a target audio signal that needs to be further processed, and the second audio signal is an audio signal played by a loudspeaker of the smart device;

[0094] A second collection end arranged on the smart device collects a second audio collection signal played by the loudspeaker of the smart device; the second audio collection signal is an audio signal obtained by collecting the second audio signal played by the audio loudspeaker of the smart device in a vibration conduction manner, and a distance between the second collection end and the loudspeaker meets a requirement that the second collection end collects the second audio collection signal in the vibration conduction manner;

[0095] An echo cancellation processing end, configured to take the second audio collection signal collected by the second collection end as a reference signal, perform echo cancellation processing on the mixed signal collected by the first collection end, and send a processed signal as a target audio signal to a signal processing end;

[0096] The signal processing end is configured to perform corresponding processing on the target audio signal according to a to-be-processed task.

[0097] The specific implementation process of the audio signal processing system can refer to the content of the audio signal processing method embodiments, which will not be described in detail here.

[0098] Based on the above, the application further provides a smart speaker device, including: a first collection end, a second collection end, a loudspeaker, an echo cancellation processing end, and a wake-up signal processing end;

[0099] The first collection end is arranged on the smart speaker device and collects a mixed signal including a first audio signal and a second audio signal; the first audio signal is a target audio signal that needs to be further processed, and the second audio signal is an audio signal played by the loudspeaker;

[0100] The second collection end is arranged on the smart speaker device and collects a second audio collection signal played by the loudspeaker of the smart device; the second audio collection signal is an audio signal obtained by collecting the second audio signal played by the audio loudspeaker of the smart device in a vibration conduction manner, and a distance between the second collection end and the loudspeaker meets a requirement that the second collection end collects the second audio collection signal in the vibration conduction manner;

[0101] The echo cancellation processing end takes the second audio collection signal as a reference signal, performs echo cancellation processing on the mixed signal, and determines a target audio signal after echo cancellation.

[0102] The signal processing end is configured to perform corresponding processing on the target audio signal according to a to-be-processed task. For example, according to a voice wake-up task, the target audio signal is identified to determine whether the target audio signal is a wake-up signal; if yes, the smart device is woken up.

[0103] Similarly, as to the first acquisition end, the second acquisition end, the loudspeaker, the echo cancellation processing end and the wake-up signal processing end in the smart speaker device, the above-mentioned processing method of the audio signal can be referred to, and details are not described herein.

[0104] It can be understood that the processing method of the audio signal provided by the present application can not only be used in the smart speaker device, but also be used in electronic devices such as mobile phones, tablets and the like having voice call function and device wake-up demand.

[0105] Based on the above, the present application further provides an electronic device, such as Figure 5 As shown in the figure, Figure 5 is a structural schematic diagram of an electronic device provided by the present application, which comprises:

[0106] a processor 501;

[0107] a memory 502 configured to store a program for processing data of the electronic device, wherein the program, when read and executed by the processor, performs the related steps of the above-mentioned processing method of the audio signal.

[0108] Based on the above, the present application further provides a computer storage medium, which comprises a computer program, and when the computer program runs on an electronic device, the electronic device performs the related steps of the above-mentioned processing method of the audio signal.

[0109] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or authorized by all parties, and the collection, use and processing of related data need to comply with relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation portal for user to choose authorization or refusal.

[0110] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.

[0111] The memory can include non-persistent memory in computer readable media, random access memory (RAM), and / or non-volatile memory, such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of computer readable media.

[0112] 1. Computer-readable media includes permanent and non-permanent, removable and non-removable media can be implemented by any method or technology to store information. Information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information that can be accessed by a computing device. According to the definition herein, computer-readable media does not include transitory media, such as modulated data signals and carriers.

[0113] 2. Those skilled in the art should understand that the embodiments of the present application can be provided as a method, system or computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0114] Although the present application is disclosed with reference to the preferred embodiments above, it is not intended to limit the present application, and any person skilled in the art can make possible changes and modifications without departing from the spirit and scope of the present application. Therefore, the scope of protection of the present application should be defined by the scope defined by the claims of the present application.

Claims

1. A method for processing audio signals, characterized in that, include: The system acquires a mixed signal including a first audio signal and a second audio signal collected by a first input terminal of the smart device, and a second audio acquisition signal collected by a second input terminal located near the speaker of the smart device. The first audio signal is a target audio signal that needs further processing, the second audio signal is an audio signal output by the audio output terminal of the smart device, and the second audio acquisition signal is an audio signal obtained by acquiring the second audio signal output by the audio output terminal of the smart device using a vibration conduction method. Using the second audio acquisition signal as a reference signal, perform echo cancellation processing on the mixed signal to determine the target audio signal after echo cancellation; including: eliminating the second audio signal in the mixed signal according to the reference signal, retaining the first audio signal; and using the first audio signal as the target audio signal; According to the task to be processed, the target audio signal is processed accordingly, including: according to the voice wake-up task, the target audio signal is identified to determine whether the target audio signal is a wake-up signal; if so, the smart device is woken up.

2. The audio signal processing method according to claim 1, characterized in that, Also includes: Acquire the digital audio signal of the audio source before digital-to-analog conversion processing, and / or acquire the analog audio signal of the audio source after digital-to-analog conversion processing and power amplification processing.

3. The audio signal processing method according to claim 2, characterized in that, The step of using the second audio acquisition signal as a reference signal to perform echo cancellation processing on the mixed signal and determining the target audio signal after echo cancellation further includes: The digital audio signal, the analog audio signal, and the second audio acquisition signal are determined as the reference signal; or, the digital audio signal and the second audio acquisition signal are determined as the reference signal; or, the analog audio signal and the second audio acquisition signal are determined as the reference signal. The mixed signal is subjected to echo cancellation processing based on the reference signal to determine the target audio signal.

4. The audio signal processing method according to claim 3, characterized in that, The digital audio signal, the analog audio signal, and the second audio acquisition signal are determined as the reference signal; or, the digital audio signal and the second audio acquisition signal are determined as the reference signal. Alternatively, the analog audio signal and the second audio acquisition signal can be determined as the reference signal, including: The digital audio signal and the analog audio signal are fused with the second audio acquisition signal to determine the reference signal; Alternatively, the digital audio signal can be fused with the second audio acquisition signal to determine the reference signal; or, the analog audio signal can be fused with the second audio acquisition signal to determine the reference signal.

5. An audio signal processing device, characterized in that, include: The acquisition unit is used to acquire a mixed signal including a first audio signal and a second audio signal collected by a first input terminal of the smart device, and a second audio acquisition signal collected by a second input terminal located near the speaker of the smart device. The first audio signal is a target audio signal that needs further processing, the second audio signal is an audio signal output by the audio output terminal of the smart device, and the second audio acquisition signal is an audio signal obtained by acquiring the second audio signal by vibration conduction at the audio output terminal of the smart device. A determining unit is configured to use the second audio acquisition signal as a reference signal to perform echo cancellation processing on the mixed signal and determine the target audio signal after echo cancellation; including: eliminating the second audio signal in the mixed signal according to the reference signal and retaining the first audio signal; and using the first audio signal as the target audio signal; The processing unit is configured to process the target audio signal according to the task to be processed, including: identifying the target audio signal according to the voice wake-up task, determining whether the target audio signal is a wake-up signal; if so, waking up the smart device.

6. An audio signal processing system, characterized in that, include: The first acquisition terminal, located on the smart device, acquires a mixed signal including a first audio signal and a second audio signal; wherein the first audio signal is the target audio signal that needs further processing, and the second audio signal is the audio signal played by the speaker of the smart device. A second acquisition terminal is disposed near the speaker of the smart device to acquire a second audio acquisition signal played by the speaker of the smart device. The second audio acquisition signal is an audio signal obtained by acquiring the second audio signal played by the audio speaker of the smart device through vibration conduction. The distance between the second acquisition terminal and the speaker meets the requirement that the second acquisition terminal acquires the second audio acquisition signal through vibration conduction. An echo cancellation processing unit is used to perform echo cancellation processing on the mixed signal using the second audio acquisition signal as a reference signal, and send the processed signal as the target audio signal to a signal processing unit; including: canceling the second audio signal in the mixed signal according to the reference signal, retaining the first audio signal; and using the first audio signal as the target audio signal; The signal processing terminal is used to process the target audio signal according to the task to be processed, including: identifying the target audio signal according to the voice wake-up task, and determining whether the target audio signal is a wake-up signal; if so, waking up the smart device.

7. A smart speaker device, characterized in that, include: The system includes a first acquisition terminal, a second acquisition terminal, a speaker, an echo cancellation processing terminal, and a signal processing terminal. The first acquisition terminal is installed on the smart speaker device to acquire a mixed signal including a first audio signal and a second audio signal; wherein the first audio signal is the target audio signal that needs further processing, and the second audio signal is the audio signal played by the speaker; The second acquisition end is located near the speaker of the smart speaker device to acquire the second audio acquisition signal played by the speaker of the smart speaker device. The second audio acquisition signal is an audio signal obtained by acquiring the second audio signal played by the speaker of the smart speaker device using a vibration conduction method. The distance between the second acquisition end and the speaker satisfies the requirement that the second acquisition end acquires the second audio acquisition signal by a vibration conduction method. An echo cancellation processing unit is used to perform echo cancellation processing on the mixed signal using the second audio acquisition signal as a reference signal, and to determine the target audio signal after echo cancellation; including: eliminating the second audio signal in the mixed signal according to the reference signal, retaining the first audio signal; and using the first audio signal as the target audio signal; The signal processing terminal is used to process the target audio signal according to the task to be processed, including: identifying the target audio signal according to the voice wake-up task, determining whether the target audio signal is a wake-up signal; if so, waking up the smart speaker device.

8. An electronic device, characterized in that, include: processor; A memory for storing a program for processing data generated by an electronic device, wherein when the program is read and executed by the processor, it performs the audio signal processing method as described in any one of claims 1 to 4.

9. A computer storage medium, characterized in that, The device includes a computer program that, when run on an electronic device, causes the electronic device to perform the audio signal processing method as described in any one of claims 1 to 4.

Citation Information

Patent Citations

  • Echo cancellation method and device, equipment and storage medium

    CN113345458A