Signal processing method and apparatus

By superimposing and modulating multiple audio channels in an audio playback device, and using transmission difference parameters to cancel out the differences between channels, the problems of speech recognition accuracy and computing power requirements of audio playback devices under multi-channel conditions are solved, achieving efficient speech recognition and low-complexity signal processing.

WO2026157621A1PCT designated stage Publication Date: 2026-07-30HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
HUAWEI TECH CO LTD
Filing Date
2025-12-10
Publication Date
2026-07-30

AI Technical Summary

Technical Problem

With multiple audio channels, the voice recognition device cannot accurately recognize the user's voice signal, while the computing power requirements of the audio playback device increase significantly.

Method used

By superimposing and modulating the signals of multiple audio channels in the audio playback device, the differences between channels are canceled out by the transmission difference parameter, and the modulated signal is output only on one audio channel, reducing the processing complexity of other channels. The speech recognition device performs filtering and demodulation processing to separate the user's speech signal.

Benefits of technology

It improves the accuracy of speech recognition, reduces the computing power requirements of audio playback devices, ensures the accuracy of speech recognition, and avoids increasing computational complexity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025141424_30072026_PF_FP_ABST
    Figure CN2025141424_30072026_PF_FP_ABST
Patent Text Reader

Abstract

A signal processing method and apparatus. The signal processing method is executed by an audio playback apparatus. The signal processing method comprises: on the basis of a first parameter, performing superposition processing on first audio signals to be outputted by a plurality of audio channels of the audio playback apparatus, so as to obtain a first mixed audio signal, wherein the first parameter is used for indicating transmission differences among the plurality of audio channels; and outputting a second mixed audio signal on a first audio channel, wherein the second mixed audio signal is a superposed signal of a modulation signal of the first mixed audio signal and the first audio signal of the first audio channel, the plurality of audio channels include the first audio channel and at least one second audio channel, and the second audio channel is used for outputting the first audio signal corresponding to the second audio channel. The signal processing method improves the recognition accuracy of a speech recognition apparatus for a user speech signal, and also reduces the computing power requirement of an audio playback apparatus.
Need to check novelty before this filing date? Find Prior Art

Description

Signal processing methods and devices

[0001] This application claims priority to Chinese Patent Application No. 202510127560.1, filed on January 27, 2025, entitled “Signal Processing Method and Apparatus”, the entire contents of which are incorporated herein by reference. Technical Field

[0002] This application relates to the field of communication technology, and in particular to a signal processing method and apparatus. Background Technology

[0003] During the process of a voice recognition device recognizing a user's voice signal, an audio playback device may be playing an audio signal. Therefore, the signal recognized by the voice recognition device may be a superposition of the user's voice signal and the audio signal. Thus, when performing voice recognition, the voice recognition device needs to perform echo cancellation on the recognized signal based on the audio signal to accurately determine the user's voice signal.

[0004] Currently, in scenarios where speech recognition devices cannot acquire audio signals, audio playback devices can modulate the low-frequency band (such as the human voice band) of the audio signal into a high-frequency band (such as the ultrasonic band) when they need to play the audio signal. Then, the resulting modulated signal is superimposed on the audio signal before playback. In this way, the speech recognition device can reconstruct the low-frequency band signal in the audio signal based on the modulated signal in the recognized signal, thereby accurately determining the user's voice signal based on the reconstructed signal.

[0005] However, when an audio playback device has multiple audio channels, if each audio channel uses the aforementioned signal processing methods (including modulation and superposition processing), the computing power requirements of the audio playback device will increase significantly. Therefore, how to improve the accuracy of speech recognition devices in recognizing user voice signals while reducing the computing power requirements of audio playback devices has become an urgent technical problem to be solved. Summary of the Invention

[0006] This application provides a signal processing method and apparatus that can improve the accuracy of speech recognition devices in recognizing user speech signals while reducing the computing power requirements of audio playback devices.

[0007] Firstly, this application provides a signal processing method, which can be executed by an audio playback device or by other entities, and this application does not limit the scope of the method. For ease of description, the following explanation will take the execution of the method by an audio playback device as an example.

[0008] The method includes: superimposing first audio signals to be output from multiple audio channels of an audio playback device based on a first parameter to obtain a first mixed audio signal; and outputting a second mixed audio signal from the first audio channel.

[0009] The first parameter is used to indicate the transmission differences among multiple audio channels. For example, the first parameter is used to indicate the transmission differences when the same audio signal (which can be any audio signal) is transmitted to the speech recognition device from multiple audio channels.

[0010] The first audio signal to be output from multiple audio channels includes the first audio signal to be output from each of the multiple audio channels. The first audio signals to be output from each of the multiple audio channels may be the same or different. The first audio signal to be output from an audio channel may be a signal from a portion of the original audio signal of that audio channel (e.g., non-ultrasonic frequency bands) or the entire frequency band.

[0011] In the process of superimposing multiple first audio signals based on a first parameter, the audio playback device can superimpose signals from the low-frequency band (e.g., the human voice frequency band) or all frequency bands of the multiple first audio signals based on the first parameter. That is, the first mixed audio signal can be a signal obtained by superimposing signals from the low-frequency band or all frequency bands of the multiple first audio signals.

[0012] The second mixed audio signal is the superposition of the modulation signal of the first mixed audio signal and the first audio signal of the first audio channel.

[0013] The modulation signal of the first mixed audio signal is a high-frequency signal obtained by the audio playback device modulating the signal in the low-frequency band or all frequency bands of the first mixed audio signal. The minimum frequency of this high-frequency band is greater than the maximum frequency of the first audio signal (i.e., the frequency band of the modulation signal does not overlap with the frequency band of the first audio signal), for example, this high-frequency band is the ultrasonic frequency band.

[0014] In some examples, if the first mixed audio signal is obtained by superimposing signals from the low-frequency bands of multiple first audio signals, then the modulation signal of the first mixed audio signal can be obtained by modulating signals from all frequency bands of the first mixed audio signal. In other examples, if the first mixed audio signal is obtained by superimposing signals from all frequency bands of multiple first audio signals, then the modulation signal of the first mixed audio signal can be obtained by modulating signals from the low-frequency bands of the first mixed audio signal.

[0015] The plurality of audio channels includes a first audio channel and at least one second audio channel. The first audio channel is used to output a second mixed audio signal, and the second audio channel is used to output a first audio signal corresponding to the second audio channel. In some examples, the first audio channel is any one of the plurality of audio channels, and the at least one second audio channel is any other channel among the plurality of audio channels besides the first audio channel.

[0016] In the technical solution provided in this application, when the speech recognition device cannot obtain the original audio signal from the audio playback device, and the audio playback device has multiple audio channels, when it is necessary to output audio signals from multiple audio channels, the audio playback device can first superimpose the first audio signals to be output from the multiple audio channels to obtain a first mixed audio signal. Then, the audio playback device can modulate the first mixed audio signal to obtain a high-frequency modulated signal. Afterwards, the audio playback device can superimpose the modulated signal with the first audio signal of the first audio channel and output it on the first audio channel, and output the first audio signal corresponding to the second audio channel on the second audio channel. Therefore, the signal recognized by the speech recognition device includes the user's speech signal, the first audio signal output from each of the multiple audio channels, and the modulated signal output from the first audio channel. Since the modulated signal is a high-frequency signal (the minimum frequency of this high-frequency band is greater than the maximum frequency of the first audio signal), and the user's speech signal is a low-frequency signal, the speech recognition device can separate the modulated signal from the recognized signal by filtering it. Furthermore, since the modulation signal is obtained by modulating the superposition of multiple first audio signals, the speech recognition device can obtain the re-sampled signals of the multiple first audio signals by demodulating the modulation signal. Therefore, the speech recognition device can accurately separate the user's speech signal from the recognized signal based on the re-sampled signal. In addition, different audio channels have transmission differences. If the audio playback device directly superimposes the multiple first audio signals and then modulates them, the re-sampled signal (i.e., demodulated signal) of the multiple first audio signals determined by the speech recognition device from the recognized signal will differ significantly from the actual first audio signals received by the speech recognition device from the multiple audio channels. Based on this, this application also pre-determines a first parameter to indicate the transmission differences of the multiple audio channels. When the audio playback device superimposes the first audio signals from the multiple audio channels, it refers to this first parameter to cancel out the transmission differences of the multiple audio channels. This improves the accuracy of the re-sampled signal determined by the speech recognition device, thereby improving the accuracy of the speech recognition device in recognizing the user's speech signal.

[0017] As can be seen, in the technical solution provided in this application, the signal processing of only one audio channel (i.e., the first audio channel) in the audio playback device involves complex processing such as modulation and superposition. Other audio channels besides the first audio channel do not involve such complex processing. Therefore, this application will not significantly increase the computing power requirements of the audio playback device. Furthermore, when the audio playback device outputs the modulated signal resulting from the superposition of multiple first audio signals through the first audio channel, the transmission differences between the multiple audio channels are canceled out. Therefore, this application can ensure the accuracy of the speech recognition device in recognizing the user's speech signal. Thus, this application can improve the accuracy of the speech recognition device in recognizing the user's speech signal while reducing the computing power requirements of the audio playback device.

[0018] In one possible implementation, the first parameter includes second parameters of multiple audio channels, which are used to indicate the signal gain of the audio signal transmitted from the audio channel to the speech recognition device; based on the first parameter, the first audio signals to be output from the multiple audio channels of the audio playback device are superimposed to obtain a first mixed audio signal, including: determining a third parameter of a second audio channel based on the second parameters of the multiple audio channels; the third parameter is used to indicate the deviation between the second parameters of the second audio channel and the second parameters of the first audio channel; based on the third parameter, the first audio signals of the multiple audio channels are superimposed to obtain the first mixed audio signal.

[0019] The transmission differences in audio channels include differences in signal gain. Therefore, when the audio playback device outputs the modulation signal of the first mixed audio signal through the first audio channel, it cancels out the difference in signal gain between the second audio channel and the first audio channel. This reduces the impact of the difference in signal gain of the audio channels on the user's speech signal recognition process, thereby ensuring the accuracy of the speech recognition device in recognizing the user's speech signal.

[0020] In another possible implementation, the first parameter includes a third parameter of the second audio channel, which indicates the amount of deviation between the second parameter of the second audio channel and the second parameter of the first audio channel. The second parameter of the audio channel indicates the signal gain of the audio signal transmitted from the audio channel to the speech recognition device.

[0021] Measuring relative signal gain (the third parameter) is easier than measuring absolute signal gain (the second parameter). Furthermore, in the process of superimposing the first audio signals from multiple audio channels to obtain the first mixed audio signal, the relative signal gain of the audio channels can be directly used as the superposition coefficient to cancel out the difference between the signal gains of the second and first audio channels. Therefore, when the audio playback device superimposes multiple first audio signals based on the first parameter, directly using the relative signal gain as the first parameter reduces the computational complexity of the signal processing, thereby further reducing the computing power requirements of the audio playback device.

[0022] In another possible implementation, multiple audio channels are determined from candidate audio channels based on a second parameter of the candidate audio channels of the audio playback device.

[0023] When the signal gain of an audio channel (i.e., the second parameter of the audio channel) is very small, the audio signal output by the audio channel has minimal interference with the user's speech signal, meaning its impact on the recognition accuracy of the speech recognition device is minimal. Therefore, in this application, when the audio playback device superimposes the first audio signals of multiple audio channels with very small signal gains, it can only superimpose the first audio signals of multiple audio channels. This further reduces the computational complexity of the signal processing, thereby further reducing the computing power requirements of the audio playback device.

[0024] In another possible implementation, the first parameter includes a fourth parameter of multiple audio channels, which indicates the signal delay of the audio signal transmitted from the audio channel to the speech recognition device. Based on the first parameter, the first audio signals to be output from the multiple audio channels of the audio playback device are superimposed to obtain a first mixed audio signal, including: determining a fifth parameter of a second audio channel based on the fourth parameter of the multiple audio channels; the fifth parameter indicates the deviation between the fourth parameter of the second audio channel and the fourth parameter of the first audio channel; and superimposing the first audio signals of the multiple audio channels based on the fifth parameter to obtain the first mixed audio signal.

[0025] The transmission differences in audio channels include differences in signal delay. Therefore, when the audio playback device outputs the modulation signal of the first mixed audio signal through the first audio channel, it cancels out the difference in signal delay between the second audio channel and the first audio channel. This reduces the impact of the difference in signal delay between the audio channels on the recognition process of the user's voice signal, thereby ensuring the accuracy of the voice recognition device in recognizing the user's voice signal.

[0026] In another possible implementation, the first parameter includes a fifth parameter of the second audio channel, which indicates the amount of deviation between the fourth parameter of the second audio channel and the fourth parameter of the first audio channel. The fourth parameter of the audio channel indicates the signal delay of the audio signal transmitted from the audio channel to the speech recognition device.

[0027] Measuring the relative signal delay (i.e., the fifth parameter of the second audio channel) is easier than measuring the absolute signal delay (i.e., the fourth parameter of the second audio channel). Furthermore, in the process of superimposing the first audio signals from multiple audio channels to obtain the first mixed audio signal, the relative signal delay of the audio channels can be directly used as the superposition coefficient of the first audio signals, thus canceling out the difference in signal delay between the second and first audio channels. Therefore, when the audio playback device superimposes multiple first audio signals based on the first parameter, directly using the relative signal delay as the first parameter reduces the computational complexity of the signal processing, thereby further reducing the computing power requirements of the audio playback device.

[0028] In another possible implementation, the first audio channel is determined from multiple audio channels based on a first parameter.

[0029] In this application, determining the first audio channel from multiple audio channels based on the first parameter can further reduce the computing power requirements of the audio playback device or further improve the accuracy of the speech recognition device in recognizing user voice signals.

[0030] Secondly, this application provides a signal processing method, which can be executed by a speech recognition device or by other entities, and this application does not limit the scope of the method. For ease of description, the following explanation will take the execution of the method by a speech recognition device as an example.

[0031] The method includes: acquiring a third mixed audio signal; filtering and demodulating the third mixed audio signal to obtain a demodulated signal of the modulated signal; and performing echo cancellation processing on the third mixed audio signal based on the demodulated signal to obtain a user speech signal.

[0032] The third mixed audio signal is a superposition of the user's voice signal, the second mixed audio signal output from the first audio channel, and the first audio signal output from at least one second audio channel. The second mixed audio signal is a superposition of the modulation signal of the first mixed audio signal and the first audio signal from the first audio channel. The first mixed audio signal is obtained by superimposing the first audio signals from multiple audio channels of the audio playback device based on a first parameter. The first parameter is used to indicate the transmission differences among the multiple audio channels. The multiple audio channels include the first audio channel and the second audio channel.

[0033] Filtering the third mixed audio signal yields a high-frequency modulated signal. Demodulating the modulated signal yields a demodulated signal. This demodulated signal can be a low-frequency signal or a signal from the entire frequency range of the first mixed audio signal.

[0034] In some examples, if the modulation signal is a signal obtained by modulating signals across all frequency bands of the first mixed audio signal, then the demodulated signal can be a signal across all frequency bands of the first mixed audio signal. In other examples, if the modulation signal is a signal obtained by modulating signals in the low-frequency band of the first mixed audio signal, then the demodulated signal can be a signal in the low-frequency band of the first mixed audio signal.

[0035] In one possible implementation, before acquiring the third mixed audio signal, the signal processing method provided in this application further includes: acquiring a first parameter; sending first information; the first information being used to instruct the audio playback device to determine the first mixed audio signal based on the first parameter.

[0036] In this application, the voice recognition device can send a first message to cause the audio playback device to switch playback modes to adapt to the needs of different application scenarios.

[0037] In another possible implementation, the first parameter includes a third parameter of the second audio channel, which indicates the deviation between the second parameter of the second audio channel and the second parameter of the first audio channel. The second parameter of the audio channel indicates the signal gain of the audio signal transmitted from the audio channel to the speech recognition device. Obtaining the first parameter includes: sending second information; the second information instructing multiple audio channels to output preset audio signals sequentially; and determining the third parameter based on the signal amplitude of the multiple preset audio signals received from the multiple audio channels.

[0038] In another possible implementation, the first parameter includes a fifth parameter of the second audio channel, which indicates the deviation between the fourth parameter of the second audio channel and the fourth parameter of the first audio channel. The fourth parameter of the audio channel indicates the signal delay of the audio signal transmitted from the audio channel to the speech recognition device. Obtaining the first parameter includes: sending third information; the third information instructing multiple audio channels to sequentially output preset audio signals based on a preset output interval; and determining the fifth parameter based on the preset output interval and the reception time of the multiple preset audio signals received from the multiple audio channels.

[0039] In another possible implementation, the preset frequency band of the audio signal is the same as the frequency band of the demodulated signal.

[0040] In this application, selecting an audio signal with the same frequency band as the demodulated signal as the preset audio signal allows for the determination of a more accurate first parameter.

[0041] In another possible implementation, before acquiring the third mixed audio signal, the signal processing method provided in this application further includes: determining a first audio channel from multiple audio channels according to a first parameter.

[0042] Thirdly, this application provides a signal processing apparatus comprising at least one module or at least one unit, wherein the at least one module or at least one unit is configured to perform the method provided by the first aspect or any implementation thereof. The at least one module or at least one unit may be implemented in hardware, in software, or in a combination of hardware and software.

[0043] In one possible implementation, the signal processing device includes a processing module and an output module. The processing module is used to superimpose first audio signals to be output from multiple audio channels of the audio playback device based on a first parameter to obtain a first mixed audio signal; the first parameter is used to indicate the transmission differences among the multiple audio channels; the output module is used to output a second mixed audio signal from the first audio channel; the second mixed audio signal is a superposition signal of the modulation signal of the first mixed audio signal and the first audio signal from the first audio channel, the multiple audio channels include a first audio channel and at least one second audio channel, and the second audio channel is used to output the first audio signal corresponding to the second audio channel.

[0044] In another possible implementation, the first parameter includes second parameters of multiple audio channels, which indicate the signal gain of the audio signal transmitted from the audio channel to the speech recognition device. The processing module is specifically configured to: determine a third parameter of the second audio channel based on the second parameters of the multiple audio channels; and, based on the third parameter, superimpose the first audio signals from the multiple audio channels to obtain a first mixed audio signal. The third parameter indicates the deviation between the second parameters of the second audio channel and the second parameters of the first audio channel.

[0045] In another possible implementation, the first parameter includes a third parameter of the second audio channel, which indicates the amount of deviation between the second parameter of the second audio channel and the second parameter of the first audio channel. The second parameter of the audio channel indicates the signal gain of the audio signal transmitted from the audio channel to the speech recognition device.

[0046] In another possible implementation, multiple audio channels are determined from candidate audio channels based on a second parameter of the candidate audio channels of the audio playback device.

[0047] In another possible implementation, the first parameter includes a fourth parameter of multiple audio channels, which indicates the signal delay of the audio signal transmitted from the audio channel to the speech recognition device. The processing module is specifically configured to: determine a fifth parameter of the second audio channel based on the fourth parameters of the multiple audio channels; and, based on the fifth parameter, superimpose the first audio signals from the multiple audio channels to obtain a first mixed audio signal. The fifth parameter indicates the deviation between the fourth parameters of the second audio channel and the fourth parameters of the first audio channel.

[0048] In another possible implementation, the first parameter includes a fifth parameter of the second audio channel, which indicates the amount of deviation between the fourth parameter of the second audio channel and the fourth parameter of the first audio channel. The fourth parameter of the audio channel indicates the signal delay of the audio signal transmitted from the audio channel to the speech recognition device.

[0049] In another possible implementation, the first audio channel is determined from multiple audio channels based on a first parameter.

[0050] Fourthly, this application provides a signal processing apparatus comprising at least one module or at least one unit, wherein the at least one module or at least one unit is used to perform the method provided by the second aspect or any implementation thereof. The at least one module or at least one unit may be implemented in hardware, in software, or in a combination of hardware and software.

[0051] In one possible implementation, the signal processing device includes an acquisition module and a processing module. The acquisition module acquires a third mixed audio signal; the third mixed audio signal is a superposition of a user's voice signal, a second mixed audio signal output from a first audio channel, and a first audio signal output from at least one second audio channel. The second mixed audio signal is a superposition of the modulation signal of the first mixed audio signal and the first audio signal from the first audio channel. The first mixed audio signal is obtained by superimposing the first audio signals from multiple audio channels of the audio playback device based on a first parameter, which indicates the transmission differences among the multiple audio channels. The multiple audio channels include a first audio channel and a second audio channel. The processing module performs filtering and demodulation processing on the third mixed audio signal to obtain a demodulated signal of the modulation signal. The processing module is further configured to perform echo cancellation processing on the third mixed audio signal based on the demodulated signal to obtain the user's voice signal.

[0052] In another possible implementation, the signal processing device further includes a transmitting module. The processing module is also configured to acquire a first parameter before the acquisition module acquires the third mixed audio signal; the transmitting module is configured to transmit first information; the first information is used to instruct the audio playback device to determine the first mixed audio signal based on the first parameter.

[0053] In another possible implementation, the first parameter includes a third parameter of the second audio channel. This third parameter indicates the deviation between the second parameter of the second audio channel and the second parameter of the first audio channel. The second parameter of the audio channel indicates the signal gain of the audio signal transmitted from the audio channel to the speech recognition device. Specifically, the acquisition module is used to: send second information via the sending module; and determine the third parameter based on the signal amplitudes of multiple preset audio signals received from multiple audio channels. The second information is used to instruct the multiple audio channels to sequentially output preset audio signals.

[0054] In another possible implementation, the first parameter includes a fifth parameter of the second audio channel. This fifth parameter indicates the deviation between the fourth parameter of the second audio channel and the fourth parameter of the first audio channel. The fourth parameter of the audio channel indicates the signal delay of the audio signal transmitted from the audio channel to the speech recognition device. The acquisition module is specifically used to: send third information through the sending module; and determine the fifth parameter based on a preset output interval and the reception time of multiple preset audio signals received from multiple audio channels. The third information instructs the multiple audio channels to sequentially output preset audio signals based on the preset output interval.

[0055] In another possible implementation, the preset frequency band of the audio signal is the same as the frequency band of the demodulated signal.

[0056] In another possible implementation, the processing module is also used to determine the first audio channel from multiple audio channels based on the first parameter before the acquisition module acquires the third mixed audio signal.

[0057] Fifthly, this application provides a signal processing apparatus, which includes a memory and at least one processor; the memory is coupled to the processor; wherein the memory stores computer program code, which includes computer instructions, and when the computer instructions are executed by the processor, the signal processing apparatus performs the method provided by the first aspect or any implementation thereof, or performs the method provided by the second aspect or any implementation thereof.

[0058] Sixthly, this application provides a computer-readable storage medium including computer instructions that, when executed on a computer, cause the computer to perform the method provided by the first aspect or any implementation thereof, or to perform the method provided by the second aspect or any implementation thereof.

[0059] In a seventh aspect, this application provides a computer program product that, when run on a computer, causes the computer to execute the method provided by the first aspect or any implementation thereof, or to execute the method provided by the second aspect or any implementation thereof.

[0060] Eighthly, this application provides a chip or chip system. Exemplarily, the chip or chip system may be a signal processing apparatus provided in the third or fourth aspect, or a signal processing apparatus provided in the fifth aspect. The chip or chip system includes: at least one processor, configured to execute a computer program or instructions to cause the method provided in the first aspect or any implementation thereof to be executed, or to cause the method provided in the second aspect or any implementation thereof to be executed.

[0061] Ninthly, this application provides a signal processing system. The signal processing system includes an audio playback device and a speech recognition device. The audio playback device is used to execute the method provided in the first aspect or any implementation thereof, and the speech recognition device is used to execute the method provided in the second aspect or any implementation thereof. Exemplarily, the audio playback device may be a signal processing device provided in the third aspect, and the speech recognition device may be a signal processing device provided in the fourth aspect.

[0062] It is understood that the technical effects of the technical solutions provided in the second to ninth aspects of this application can be referred to the technical effects corresponding to the first aspect or any implementation of the first aspect, and will not be repeated here.

[0063] Based on the implementation methods provided in the above aspects, this application can be further combined to provide more implementation methods. Attached Figure Description

[0064] Figure 1 is a schematic diagram of the structure of an electronic device provided in an embodiment of this application;

[0065] Figure 2 is a schematic diagram of a scenario in which a voice recognition device interacts with an audio playback device according to an embodiment of this application;

[0066] Figure 3 is a schematic flowchart of a signal processing method provided in an embodiment of this application;

[0067] Figure 4 is a simplified flowchart illustrating the process of determining a second mixed audio signal according to an embodiment of this application;

[0068] Figure 5 is a simplified flowchart of a signal processing method provided in an embodiment of this application;

[0069] Figure 6 is a flowchart illustrating another signal processing method provided in an embodiment of this application;

[0070] Figure 7 is a schematic diagram of a signal processing device provided in an embodiment of this application;

[0071] Figure 8 is a schematic diagram of another signal processing device provided in an embodiment of this application;

[0072] Figure 9 is a schematic diagram of another signal processing device provided in an embodiment of this application. Detailed Implementation

[0073] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the embodiments described herein are some, but not all, of the embodiments of this application. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the protection scope of this application.

[0074] Before describing the technical solution provided in this application in detail, some terms involved in this application will be explained.

[0075] The terms "first," "second," etc., used in the embodiments and accompanying drawings of this application are for distinguishing purposes only and should not be construed as indicating or implying relative importance or order.

[0076] The terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion, such as including a series of steps or units. A method, apparatus, or product is not limited to the steps or units listed, but may optionally include other steps or units not listed, or may optionally include other steps or units inherent to such methods, apparatus, or products.

[0077] The term "at least one" means one or more, and "more than one" means two or more. "And / or" is used to describe the relationship between related objects, indicating that there can be three relationships. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. A and B can be singular or plural.

[0078] The terms “exemplary,” “for example,” and “e.g.,” are used to indicate examples, illustrations, or descriptions. Any embodiment or implementation described herein as “exemplary,” “for example,” and “e.g.,” should not be construed as being more preferred or advantageous than other embodiments or implementations. Rather, the use of terms such as “exemplary” is intended to present the relevant concepts in a specific manner.

[0079] The term "instruction" or "for instruction" can include both direct and indirect instruction. When describing an instruction as being used to instruct A, it can include whether the instruction directly or indirectly instructs A, but does not necessarily mean that the instruction necessarily carries A. The instruction methods involved in the embodiments of this application can be understood to cover various methods that enable the party to be instructed to know the instruction information. The instruction information can be sent as a whole or divided into multiple sub-information messages sent separately, and the sending period and / or timing of these sub-information messages can be the same or different.

[0080] The term "protocol" can refer to a standard protocol in the field of communications. For example, it can refer to fifth-generation (5G) protocols, new radio (NR) protocols, and related protocols applied to future communication systems, which are not limited in this application.

[0081] Furthermore, in this application, "sending information to XX (device or equipment)" can be understood as the destination of the information being the device, and may include sending information directly or indirectly to the device. "Receiving or obtaining information from XX (device or equipment)" or "receiving or obtaining information from XX (device or equipment)" can be understood as the source of the information being the device, and may include receiving information directly or indirectly from the device. The information may undergo necessary processing between the source of the information transmission and the destination of the information reception, such as format conversion.

[0082] Furthermore, the electronic device in this application can be a portable electronic device capable of processing audio signals. For example, it can be a mobile phone, wearable device (e.g., a smart bracelet), tablet computer, laptop computer, ultra-mobile personal computer (UMPC), netbook, or personal digital assistant (PDA). Alternatively, the electronic device in this application can also be a non-portable electronic device capable of processing audio signals. For example, it can be a desktop computer, set-top box, speaker, etc. In other examples, the electronic device in this application can also be a smart home device capable of processing audio signals, such as a smart TV, smart refrigerator, smart air conditioner, or a smart speaker, smart humidifier, etc., used to provide smart living services.

[0083] For example, referring to FIG1, a schematic diagram of the structure of an electronic device provided in an embodiment of this application is shown. As shown in FIG1, the electronic device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone jack 170D, a sensor module 180, buttons 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc.

[0084] For example, the sensor module 180 may include at least one of the following: a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an accelerometer sensor 180E, a distance sensor 180F, a proximity sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, and a bone conduction sensor 180M.

[0085] Processor 110 may include one or more processing units. In one possible implementation, processor 110 may include a central processing unit (CPU), a graphics processing unit (GPU), a modem processor, an image signal processor (ISP), a controller, memory, a video codec, a digital signal processor (DSP), a baseband processor, an application processor (AP), and a neural network processing unit (NPU). These different processing units may be independent devices or integrated into one or more processors.

[0086] The controller can be the nerve center and command center of the electronic device 100. The controller can generate operation control signals according to the instruction opcode and timing signals to complete the control of fetching and executing instructions.

[0087] The processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory can store instructions or data that the processor 110 has just used or is cyclically used in a short period of time. If the processor 110 needs to use the instruction or data again, it can directly retrieve it from the aforementioned memory. This avoids repeated access to instructions or data and reduces the waiting time of the processor 110.

[0088] In one possible implementation, processor 110 includes one or more interfaces. These interfaces may include at least one of the following: an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), and a general-purpose input / output (GPIO) interface. It should be understood that in practical applications, processor 110 may also include other interfaces, and this application does not limit this.

[0089] An I2C interface (also known as an I2C bus interface, I2C bus, etc.) can be a bidirectional synchronous serial bus. An I2C interface may include a serial data line (SDA) and a serial clock line (SCL). In some examples, processor 110 may include multiple I2C interfaces. Processor 110 can couple touch sensor 180K, camera 193, etc., through different I2C interfaces. For example, processor 110 can couple touch sensor 180K through an I2C interface, enabling processor 110 to communicate with touch sensor 180K via the I2C interface, thereby realizing the touch function of electronic device 100.

[0090] The I2S interface (also known as the I2S bus interface, I2S bus, etc.) can be used for audio communication. In some examples, processor 110 may include multiple I2S interfaces. Processor 110 can couple with audio module 170 through the I2S interface, thereby enabling communication between processor 110 and audio module 170. In some examples, audio module 170 can transmit audio signals to wireless communication module 160 through the I2S interface, thereby enabling the function of answering phone calls through Bluetooth headsets.

[0091] The PCM interface can also be used for audio communication. For example, the PCM interface can be used to sample, quantize, or encode analog audio signals. In some examples, the audio module 170 and the wireless communication module 160 can be coupled via the PCM bus interface. For instance, the audio module 170 can also transmit audio signals to the wireless communication module 160 via the PCM interface to enable the function of answering phone calls through a Bluetooth headset.

[0092] The UART interface can be a universal serial data bus for asynchronous communication. It can also be a bidirectional communication bus. Furthermore, the UART interface can convert data between serial and parallel communication. In some examples, the UART interface can be used to connect processor 110 and wireless communication module 160. For instance, processor 110 can communicate with the Bluetooth module in wireless communication module 160 via the UART interface to enable Bluetooth functionality. In some examples, audio module 170 can transmit audio signals to wireless communication module 160 via the UART interface to enable music playback through Bluetooth headphones.

[0093] The MIPI interface can be used to connect the processor 110 to the display screen 194, or to connect peripheral devices such as the camera 193. The MIPI interface may include a camera serial interface (CSI), a display serial interface (DSI), etc. In some examples, the processor 110 and the camera 193 can communicate via the CSI interface to enable the electronic device 100 to capture images; the processor 110 and the display screen 194 can communicate via the DSI interface to enable the electronic device 100 to display images.

[0094] The GPIO interface can be configured via software. It can be configured to handle control signals or data signals. In some embodiments, the GPIO interface can be used to connect the processor 110 to a camera 193, a display screen 194, a wireless communication module 160, an audio module 170, or a sensor module 180, etc. The GPIO interface can also be configured as an I2C interface, an I2S interface, a UART interface, or a MIPI interface, etc.

[0095] USB interface 130 can be an interface compliant with the USB standard specification. USB interface 130 can be used to connect a charger to charge electronic device 100, and can also be used for data transfer between electronic device 100 and peripheral devices. For example, USB interface 130 can be used to connect headphones to play audio. USB interface 130 can also be used to connect other electronic devices.

[0096] The charging management module 140 can be used to receive charging input from a charger. The charger can be a wireless charger or a wired charger, and this application is not limited to either. In some examples of wired charging, the charging management module 140 can receive charging input from the wired charger via the USB interface 130. In some examples of wireless charging, the charging management module 140 can receive wireless charging input via the wireless charging coil of the electronic device 100. While charging the battery 142, the charging management module 140 can also supply power to the electronic device via the power management module 141.

[0097] The power management module 141 can be used to connect to the battery 142, the charging management module 140, and the processor 110. The power management module 141 can receive input from the battery 142 and / or the charging management module 140, thereby powering the processor 110, internal memory 121, external memory, display 194, camera 193, or wireless communication module 160, etc. The power management module 141 can also be used to monitor parameters such as battery capacity, battery cycle count, and battery health status (leakage current, impedance). In some examples, the power management module 141 can be located within the processor 110. In other examples, the power management module 141 and the charging management module 140 can be located in the same device.

[0098] Antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, modem processor, and baseband processor, etc., can work together to realize the wireless communication function of electronic device 100.

[0099] Antennas 1 and 2 are used to transmit and receive electromagnetic wave signals. Each antenna in electronic device 100 can be used to cover one or more communication frequency bands. Different antennas can also be reused to improve antenna utilization. For example, antenna 1 can be reused as a diversity antenna for a wireless local area network. In one possible implementation, the antennas can also be used in conjunction with a tuning switch.

[0100] The mobile communication module 150 can provide wireless communication solutions, including 2G, 3G, 4G, 5G, or future communication technologies, for use on the electronic device 100. The mobile communication module 150 may include at least one filter, switch, power amplifier, low noise amplifier (LNA), etc. The mobile communication module 150 can receive electromagnetic waves via antenna 1, and perform filtering, amplification, and other processing on the received electromagnetic waves before transmitting them to a modem processor for demodulation. The mobile communication module 150 can also amplify the signal modulated by the modem processor and convert it into electromagnetic waves for radiation via antenna 1. In one possible implementation, at least some functional modules of the mobile communication module 150 may be housed in the processor 110. In another possible implementation, at least some functional modules of the mobile communication module 150 and at least some modules of the processor 110 may be housed in the same device.

[0101] A modem processor may include a modulator and a demodulator. The modulator can be used to modulate a low-frequency baseband signal to be transmitted into a mid-to-high frequency signal. The demodulator can be used to demodulate a received electromagnetic wave signal into a low-frequency baseband signal. The demodulator can also be used to transmit the demodulated low-frequency baseband signal to a baseband processor for processing. After processing by the baseband processor, the low-frequency baseband signal can be transmitted to an application processor. The application processor can output audio signals through audio devices (e.g., speaker 170A, receiver 170B, or other devices) or display images or videos through a display screen 194. In some examples, the modem processor may be a separate device. In other examples, the modem processor may be set up independently of the processor 110, but within the same device as the mobile communication module 150 or other functional modules.

[0102] The wireless communication module 160 can provide solutions for wireless communication applications on the electronic device 100, including wireless local area networks (WLAN), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), and infrared (IR). The wireless communication module 160 can be one or more devices integrating at least one communication processing module. The wireless communication module 160 receives electromagnetic waves via antenna 2, performs frequency modulation and filtering of the electromagnetic wave signals, and sends the processed signal to processor 110. The wireless communication module 160 can also receive signals to be transmitted from processor 110, perform frequency modulation and amplification, and convert them into electromagnetic waves for radiation via antenna 2.

[0103] In one possible implementation, antenna 1 of electronic device 100 is coupled to mobile communication module 150, and antenna 2 is coupled to wireless communication module 160, enabling electronic device 100 to communicate with other devices via wireless communication technology. The wireless communication technology may include at least one of the following: Global System for Mobile Communications (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Time-Division Code Division Multiple Access (TD-SCDMA), Long Term Evolution (LTE), BT, GNSS, WLAN, NFC, FM, and IR technology.

[0104] The CPU, GPU, and display screen 194 work together to realize the display function of the electronic device 100. The GPU is a microprocessor for image processing and can be connected to the display screen 194 and the CPU. The GPU can handle a large number of mathematical and geometric operations and can be used for rendering and drawing images. The processor 110 may include one or more GPUs, which execute program instructions to generate or modify display information.

[0105] The display screen 194 may include a display panel for displaying images, videos, etc. The display panel may be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a quantum dot light-emitting diode (QLED), etc. For example, the electronic device 100 may include one or N display screens 194, where N is a positive integer greater than 1.

[0106] The ISP, camera 193, video codec, CPU, GPU, display 194, and CPU, etc., can work together to realize the shooting function of electronic device 100.

[0107] The Information Service Provider (ISP) processes data fed back from the camera 193. For example, when a user takes a photo, the shutter is opened, and light is transmitted through the lens to the camera's image sensor. The light signal is converted into an electrical signal, which is then transmitted to the ISP for processing, transforming it into a visible image. The ISP can also perform algorithmic optimizations on image noise, brightness, and skin tone. Additionally, the ISP can optimize parameters such as exposure and color temperature for the shooting scene. In one possible implementation, the ISP could be integrated into the camera 193 itself.

[0108] Camera 193 is used to capture still images or videos. An object is projected onto a photosensitive element by generating an optical image through a lens. The photosensitive element can be a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. For example, the photosensitive element can convert light signals into electrical signals and transmit them to an ISP; the ISP can convert the electrical signals into digital image signals and output them to a DSP for processing; the DSP converts the digital image signals into image signals in a standard format. In one possible implementation, the electronic device 100 may include one or N cameras 193, where N is a positive integer greater than 1.

[0109] A digital signal processor (DSP) can be used to process digital signals, such as digital image signals or other digital signals. For example, when the electronic device 100 is selecting a frequency point, the DSP can be used to perform a Fourier transform on the frequency energy, etc.

[0110] Video codecs can be used to compress or decompress digital video. Electronic device 100 can support one or more video codecs to enable playback or recording of videos in various encoding formats. For example, the encoding format could be Moving Picture Experts Group (MPEG) 1, MPEG 2, MPEG 3, MPEG 4, etc.

[0111] An NPU can be a neural network (NN) computing processor. By borrowing the structure of biological neural networks (such as the transmission patterns between neurons in the human brain), an NPU can rapidly process input information and continuously learn itself. NPUs can also enable intelligent cognitive functions in electronic devices, such as image recognition, facial recognition, speech recognition, or text understanding.

[0112] The external memory interface 120 can connect to an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 100. The external memory card communicates with the processor 110 through the external memory interface 120 to perform data storage.

[0113] Internal memory 121 can be used to store computer program code, which includes computer instructions. Processor 110 can execute these computer instructions to implement the signal processing method provided in this application. Internal memory 121 may include a program storage area and a data storage area. For example, the program storage area may store the operating system of electronic device 100, applications required for at least one function, etc. The data storage area may store data generated during the use of electronic device 100 (e.g., audio data), etc. In addition, internal memory 121 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, universal flash storage (UFS), etc.

[0114] The audio module 170, speaker 170A, receiver 170B, microphone 170C, headphone jack 170D, and CPU work together to realize the audio functions of electronic device 100, such as music playback and recording.

[0115] The audio module 170 is used to convert digital audio signals into analog audio signals for output, and can also be used to convert analog audio signals into digital audio signals, as well as to encode and decode audio signals. In some examples, the audio module 170 can be located in the processor 110, or some functional modules of the audio module 170 can be located in the processor 110.

[0116] The speaker 170A, also known as a "loudspeaker," is used to convert audio electrical signals into sound signals. The electronic device 100 can listen to music or make hands-free calls through the speaker 170A.

[0117] The receiver 170B, also known as the "earpiece," is used to convert audio electrical signals into sound signals. When the electronic device 100 answers a phone call or receives a voice signal, the user can listen to the voice by bringing the receiver 170B close to their ear.

[0118] Microphone 170C, also known as a "microphone" or "voice transducer," is used to convert sound signals into electrical signals. When making a phone call or sending a voice signal, the user can bring microphone 170C close to their mouth to speak, thus inputting the sound signal into microphone 170C. Electronic device 100 can be equipped with at least one microphone 170C. Multiple microphones 170C, in addition to acquiring sound signals, can also perform noise reduction, sound source identification, and directional recording functions.

[0119] The 170D headphone jack can be used to connect wired headphones. The 170D headphone jack can also be a USB interface or a 3.5mm Open Mobile Terminal Platform (OMTP) standard interface, etc.

[0120] Pressure sensor 180A can be used to sense pressure signals and convert them into electrical signals. In some examples, pressure sensor 180A can be located in display screen 194. Pressure sensor 180A can be a resistive pressure sensor, an inductive pressure sensor, or a capacitive pressure sensor, etc., and this application is not limited thereto. A capacitive pressure sensor can include at least two parallel plates with conductive material. When a force is applied to pressure sensor 180A, the capacitance between the electrodes changes. Electronic device 100 can determine the intensity of the pressure based on this change in capacitance. For example, when a touch operation is applied to display screen 194, electronic device 100 can detect the intensity of the touch operation based on pressure sensor 180A. Electronic device 100 can also calculate the touch position based on the detection signal from pressure sensor 180A. In some examples, touch operations applied to the same touch position but with different touch operation intensities can trigger different operation commands. For example, when a touch operation with a touch operation intensity less than a predetermined pressure threshold is applied to the SMS application icon, a command to view the SMS message can be triggered. When a touch operation with a force greater than or equal to the pressure threshold is applied to the SMS application icon, a command to create a new SMS message can be triggered.

[0121] The gyroscope sensor 180B can be used to determine the motion attitude of the electronic device 100. In some examples, the gyroscope sensor 180B can determine the angular velocity of the electronic device 100 around three axes (e.g., the three axes of a spatial Cartesian coordinate system). The gyroscope sensor 180B can be used to implement image stabilization. For example, when a user uses the camera function and presses the shutter, the gyroscope sensor 180B can detect the angle of the electronic device 100's shake, calculate the distance that the lens module needs to compensate based on the angle, and allow the lens to counteract the shake of the electronic device 100 by moving in the opposite direction, thus achieving image stabilization.

[0122] The barometric pressure sensor 180C can be used to measure air pressure. In some examples, the electronic device 100 can calculate altitude using the air pressure value measured by the barometric pressure sensor 180C, to assist in positioning and navigation.

[0123] The magnetic sensor 180D may include a Hall sensor. The electronic device 100 can use the magnetic sensor 180D to detect the opening and closing of the flip cover of the electronic device 100. In some examples, when the electronic device 100 is a flip phone, the electronic device 100 can detect the opening and closing of the flip cover using the magnetic sensor 180D, and then set functions such as automatic flip unlocking based on the detected opening and closing state of the flip cover.

[0124] The accelerometer 180E can detect the magnitude of the acceleration of the electronic device 100 in multiple directions (e.g., the three directions corresponding to the three axes of a Cartesian coordinate system). For example, the posture of the electronic device 100 can be identified based on the detected acceleration, enabling functions such as landscape / portrait screen switching or a pedometer.

[0125] The distance sensor 180F is used to measure distance. For example, in a shooting scenario, the electronic device 100 can use the distance sensor 180F to measure distance for fast focusing.

[0126] The proximity sensor 180G may include a light-emitting diode (LED) and a light detector (e.g., a photodiode). The LED may be an infrared LED or other light-emitting diode. The electronic device 100 may emit infrared light outward through the LED. The electronic device 100 may use the photodiode to detect infrared reflected light from nearby objects. When sufficient reflected light is detected, it can be determined that there is an object near the electronic device 100. When insufficient reflected light is detected, the electronic device 100 can determine that there is no object near the electronic device 100. In addition, the electronic device 100 may also use the proximity sensor 180G to detect whether the user is holding the electronic device 100 close to their ear to make a call, so as to realize an automatic screen-off function to save power.

[0127] The ambient light sensor 180L can be used to sense ambient light intensity. The electronic device 100 can adaptively adjust the brightness of its display screen 194 based on the sensed ambient light intensity. The ambient light sensor 180L can also be used to automatically adjust the white balance when taking photos. The ambient light sensor 180L can also work in conjunction with the proximity sensor 180G to detect whether the electronic device 100 is in a pocket, preventing accidental touches.

[0128] The fingerprint sensor 180H is used to collect fingerprints. The electronic device 100 can use the fingerprint sensor to collect fingerprint data. If the collected fingerprint data matches the stored template image, fingerprint unlocking can be achieved, thereby enabling access to application locks, fingerprint photography, fingerprint answering of calls, etc.

[0129] Temperature sensor 180J can be used to detect temperature. In some embodiments, electronic device 100 can use the temperature detected by temperature sensor 180J to execute a temperature handling strategy. For example, when the temperature reported by temperature sensor 180J exceeds a preset threshold, electronic device 100 performs actions to reduce the temperature of the processor located near temperature sensor 180J in order to reduce power consumption and implement thermal protection. In other examples, when the temperature is below another preset threshold, electronic device 100 can heat battery 142 to prevent abnormal shutdown of electronic device 100 due to low temperature. In still other examples, when the temperature is below yet another preset threshold, electronic device 100 boosts the output voltage of battery 142 to prevent abnormal shutdown due to low temperature.

[0130] Touch sensor 180K, also known as a "touch panel," can be located on display screen 194. The touch sensor 180K and display screen 194 together form a touchscreen, also known as a "touch screen." Touch sensor 180K detects touch operations applied to or near it. It transmits the detected touch operation to the CPU to determine the type of touch event. Additionally, display screen 194 can provide visual output related to the touch operation. Alternatively, touch sensor 180K can also be located on the surface of electronic device 100, in a different position than display screen 194.

[0131] The bone conduction sensor 180M can acquire vibration signals. In some examples, the bone conduction sensor 180M can acquire vibration signals from the vibrating bone segments of the human vocal cords. The bone conduction sensor 180M can also contact the human pulse to receive blood pressure signals. In some examples, the bone conduction sensor 180M can also be integrated into headphones to form bone conduction headphones. The audio module 170 can parse voice signals from the vibration signals acquired by the bone conduction sensor 180M to realize voice functionality. The application processor can parse heart rate information based on the blood pressure signals acquired by the bone conduction sensor 180M to realize heart rate detection functionality.

[0132] Buttons 190 include a power button, volume buttons, etc. Buttons 190 can be mechanical buttons or touch buttons. Electronic device 100 can receive button input and generate key signal inputs related to user settings and function control of electronic device 100.

[0133] Motor 191 can generate vibration alerts. Motor 191 can be used for incoming call vibration alerts or for touch vibration feedback. For example, different touch operations applied to different applications can correspond to different vibration feedback effects. Touch operations applied to different areas of the display screen 194 can also correspond to different vibration feedback effects from motor 191. In different application scenarios (e.g., time reminders, receiving messages, alarm clocks, games, etc.), motor 191 can also achieve different vibration feedback effects.

[0134] Indicator 192 can be an indicator light used to indicate the charging status and power changes of electronic device 100, or it can be used to indicate messages, missed calls, notifications, etc.

[0135] The SIM card interface 195 is used to connect a SIM card. The SIM card can be inserted into or removed from the SIM card interface 195 to make contact with and separate from the electronic device 100. The electronic device 100 can support one or N SIM card interfaces, where N is a positive integer greater than 1. The SIM card interface 195 can support Micro SIM cards, SIM cards, etc. Multiple cards can be inserted into the same SIM card interface 195 simultaneously; the types of cards can be the same or different. The SIM card interface 195 is also compatible with different types of SIM cards and external memory cards. The electronic device 100 interacts with the network through the SIM card to achieve functions such as making calls and data communication. In one possible implementation, the electronic device 100 can use an embedded SIM card (eSIM), which can be embedded in the electronic device 100 and cannot be separated from it.

[0136] It should be noted that the structure shown in Figure 1 does not constitute a limitation on the electronic device. In addition to the components shown in Figure 1, the electronic device may include more or fewer components, or combine certain components, or have different component arrangements. The components shown can be implemented in hardware, software, or a combination of both. Furthermore, the interface connections between the modules shown in Figure 1 are merely examples and do not constitute a structural limitation on the electronic device. In practical applications, the electronic device may also employ different interface connection methods than those shown in Figure 1, or a combination of multiple interface connection methods.

[0137] The following provides an illustrative description of the application scenarios of this application. The technical solution provided in this application can be applied to scenarios where a voice recognition device interacts with an audio playback device. Specifically, it can be applied to scenarios where the voice recognition device cannot obtain the original audio signal from the audio playback device, and the audio playback device has multiple audio channels.

[0138] An audio playback device can be the audio playback device itself, or a chip or circuit within it. For example, it can be a modem chip (also known as a baseband chip), a system-on-a-chip (SoC) chip containing a modem core, or a system-in-package (SIP) chip. Alternatively, an audio playback device can also be a functional module within an audio playback device that can call and execute programs. The following explanation uses the example of an audio playback device as an example.

[0139] The audio playback device can be an electronic device with audio playback functionality. For example, it can be a mobile phone, tablet computer, or other electronic device. The specific structure of the audio playback device can be referred to the aforementioned electronic device 100.

[0140] In one possible implementation, the audio playback function of the audio playback device can be implemented through built-in devices. For example, the audio playback device can be a speaker enclosure (e.g., a soundbar, stereo speaker, etc.) comprising multiple speakers (also called "loudspeakers"), with different speakers corresponding to different audio channels (also called "channels"). Taking a soundbar as an example, it can include a left channel speaker, a right channel speaker, and a sky speaker, etc. The audio playback device can achieve audio playback by outputting the corresponding audio signal from the audio channel of each of the built-in speakers.

[0141] In other examples, the audio playback device may include multiple speakers, each speaker including at least one loudspeaker, with one loudspeaker corresponding to one audio channel. For example, the multiple speakers may include a soundbar (which may include a speaker for the left channel, a speaker for the right channel, and a speaker for the overhead sound, etc.), a speaker for the left surround channel, and a speaker for the right surround channel, etc. The audio playback device can achieve audio playback functionality by outputting the corresponding audio signal from the audio channel of each speaker in each speaker through the multiple speakers.

[0142] In another possible implementation, the audio playback function of the audio playback device can also be implemented through a peripheral device. For example, the audio playback device can be connected to a speaker (which includes multiple loudspeakers) or multiple speakers, and the audio playback device can output audio signals through the one or more speakers, or it can output audio signals through an internal device.

[0143] In other examples, the audio playback device has no built-in speakers and cannot play audio itself, but can achieve audio playback through external devices. For example, the audio playback device can be a set-top box, which can process the audio signal (including decoding the audio encoding) and transmit the processed audio signal to the speaker for output, thereby achieving audio playback.

[0144] The following description of this application will use an audio playback device implementing audio playback function through its built-in devices as an example.

[0145] A voice recognition device can be the voice recognition device itself, or a chip or circuit within it. For example, it can be a modem chip, a SoC chip containing a modem core, or a SIP chip. Alternatively, a voice recognition device can also be a functional module within a voice recognition device that can call and execute programs. The following explanation uses the example of a voice recognition device.

[0146] Voice recognition devices can be electronic devices with voice recognition capabilities. For example, they can be televisions (also known as "smart TVs," "large screens," "intelligent screens," "smart large screens," etc.), mobile phones, tablets, and other electronic devices. The specific structure of a voice recognition device can be referenced in the aforementioned electronic device 100.

[0147] It should be understood that the aforementioned audio playback device may also be called an audio processing device or other names, the aforementioned speech recognition device may also be called a speech receiving device or other names, the aforementioned audio playback device may also be called an audio processing device or other names, and the aforementioned speech recognition device may also be called a speech receiving device or other names, and this application does not limit it in this regard.

[0148] Taking a television as the voice recognition device and a speaker as the audio playback device as an example, Figure 2 illustrates a scenario where a voice recognition device and an audio playback device interact, according to an embodiment of this application. As shown in Figure 2, the television 210 can interact with a speaker (such as the soundbar 220 in Figure 2), for example, the television 210 can acquire the audio signal output by that speaker. Alternatively, the television 210 can also interact with multiple speakers (such as speakers 232, 234, 236, and 238 in Figure 2), for example, the television 210 can acquire the audio signals output by multiple speakers.

[0149] It is understood that the application scenarios described in the embodiments of this application are for the purpose of more clearly illustrating the technical solutions of this application, and do not constitute a limitation on the technical solutions of this application. Those skilled in the art will recognize that with the emergence of new application scenarios, the technical solutions provided in this application are also applicable to technical problems similar to those in this application.

[0150] Taking the interactive scenario shown in Figure 2 as an example, when a user interacts with the TV via voice commands, the TV needs to recognize the user's voice signal. During this recognition process, the TV may be playing audio signals through speakers. Both the audio signal from the speakers and the user's voice signal will reach the TV's microphone. Therefore, the signal recognized by the TV may be a superposition of the user's voice signal and the audio signal from the speakers. To improve the microphone's pickup accuracy and thus the TV's accuracy in recognizing the user's voice signal, the TV needs to perform echo cancellation on the recognized signal based on the original audio signal from the speakers.

[0151] Currently, when the television cannot obtain the original audio signal from the speaker, it cannot use the original audio signal as a re-sampled signal, and therefore cannot perform echo cancellation on the identified signal. For example, in the interactive scenario shown in Figure 2, the television may not decode the audio data stream, but directly pass the encoded audio data stream through to the speaker, which then decodes the audio data stream. In this case, the television only passes through the audio data stream and cannot obtain the audio signal. Alternatively, if the speaker is connected to a set-top box, the set-top box may directly transmit the audio data stream to the speaker, only transmitting the picture data stream to the television, in which case the television also cannot obtain the audio signal. To address this, related technologies can adjust the signal processing method of the speaker. Specifically, when audio signals need to be played through multiple audio channels (for example, the multiple audio channels of the soundbar 220 in Figure 2, or the audio channels corresponding to speakers 232, 234, 236, and 238 in Figure 2), for each audio channel, the low-frequency band (e.g., the human voice band) of the audio signal to be played in that channel can first be modulated to a high-frequency band (e.g., the ultrasonic band). Then, the resulting modulated signal is superimposed on the audio signal of that audio channel before playback. In this way, the television can reconstruct the low-frequency band signal from the audio signals of multiple audio channels based on the modulated signal in the identified signal, thereby accurately determining the user's voice signal based on the reconstructed signal.

[0152] However, if the audio signal output from each of the multiple audio channels has to go through the above signal processing process (including modulation processing and superposition processing), the computing power requirement of the speaker will increase significantly.

[0153] As can be seen, in the technical solutions provided by related technologies, when the audio playback device has multiple audio channels, the computational power requirement of the audio playback device increases significantly in order to achieve the echo cancellation function of the speech recognition device. Based on this, this application provides a signal processing method that can improve the accuracy of the speech recognition device in recognizing user speech signals (i.e., achieve the echo cancellation function of the speech recognition device) while reducing the computational power requirement of the audio playback device.

[0154] The signal processing method provided in this application will be described below by way of example. Referring to FIG3, a flowchart illustrating a signal processing method provided in an embodiment of this application is shown. Exemplarily, this method can be applied to the interactive scenario shown in FIG2, or to other scenarios where a speech recognition device interacts with an audio playback device. As shown in FIG3, the method includes the following steps:

[0155] S310. Based on the first parameter, the audio playback device performs superposition processing on the first audio signals to be output from multiple audio channels of the audio playback device to obtain a first mixed audio signal.

[0156] The first parameter is used to indicate the transmission differences among multiple audio channels. For example, the first parameter is used to indicate the transmission differences when the same audio signal (which can be any audio signal) is transmitted to the speech recognition device from multiple audio channels. For instance, the first parameter may include transmission parameters for the same audio signal transmitted to the speech recognition device from multiple audio channels.

[0157] The first audio signal to be output from multiple audio channels includes the first audio signal to be output from each of the multiple audio channels. The first audio signals to be output from each of the multiple audio channels may be the same or different. The first audio signal to be output from one audio channel may be a signal from a portion of the original audio signal of that audio channel (e.g., non-ultrasonic frequency bands) or the entire frequency band. For example, the original audio signal of an audio channel may be music, film / TV audio, etc., to be output from that audio channel.

[0158] The raw audio signal of the audio channel can be generated (or provided) by the audio playback device, or it can be obtained by the audio playback device from other devices. For example, when the audio playback device is an audio playback device, it can obtain the raw audio signal from internal or external memory, or it can decode the audio data stream stored in internal or external memory to generate the raw audio signal, or it can obtain the raw audio signal from an external device (such as a set-top box, game console, etc.).

[0159] In some examples, if the original audio signal of the audio channel does not include the ultrasonic frequency band, then the first audio signal of the audio channel can be a signal across all frequency bands of the original audio signal of the audio channel. In other examples, if the original audio signal of the audio channel includes the ultrasonic frequency band, then the first audio signal of the audio channel can be a signal across the non-ultrasonic frequency bands of the original audio signal of the audio channel. For example, an audio playback device can filter the original audio signal of the audio channel to remove signals in the ultrasonic frequency band, thus obtaining the first audio signal of the audio channel. That is to say, the frequency band of the first audio signal in this application does not overlap with the ultrasonic frequency band.

[0160] During the process of superimposing multiple first audio signals based on the first parameter, the audio playback device can superimpose signals from the low-frequency band or all frequency bands of the multiple first audio signals based on the first parameter. That is, the first mixed audio signal can be a signal obtained by superimposing signals from the low-frequency band or all frequency bands of the multiple first audio signals by the audio playback device.

[0161] In some examples, the audio playback device may first process each of the multiple first audio signals based on a first parameter to obtain a second audio signal for each first audio signal; then, the audio playback device may superimpose the multiple second audio signals obtained, and the superposition result is the first mixed audio signal.

[0162] In other examples, the audio playback device may first process each of the plurality of first audio signals based on a first parameter to obtain a second audio signal for each first audio signal; then, the plurality of second audio signals may be superimposed to obtain a fourth mixed audio signal; then, the fourth mixed audio signal may be filtered to retain the low-frequency signals of the fourth mixed audio signal to obtain a first mixed audio signal. It should be understood that in practical applications, the execution order of the above three processing procedures (the signal processing based on the first parameter, the signal superposition processing, and the signal filtering processing) may vary, and this application does not limit this.

[0163] The low-frequency band in this application can be the frequency band of the user's voice signal, or it can include the frequency band of the user's voice signal. That is, the minimum frequency of the low-frequency band can be less than or equal to the minimum frequency of the user's voice signal frequency band, and the maximum frequency of the low-frequency band can be greater than or equal to the maximum frequency of the user's voice signal frequency band. The frequency band of the user's voice signal can be the human voice frequency band, or it can be a frequency band with strong human voice power. The frequency band of the user's voice signal does not overlap with the ultrasonic frequency band. For example, the frequency band of the user's voice signal can be 0-4 kHz, 0-2 kHz, or 0-24 kHz, etc.

[0164] S320, the audio playback device outputs a second mixed audio signal through the first audio channel.

[0165] The multiple audio channels of an audio playback device can be all or some of the audio channels of the audio playback device. The audio channels of the audio playback device can be its own audio channels or the audio channels of external devices connected to the audio playback device. The multiple audio channels include a first audio channel and at least one second audio channel. The first audio channel is used to output a second mixed audio signal, and the second audio channel is used to output a first audio signal corresponding to the second audio channel. In some examples, the first audio channel is any one of the multiple audio channels, and the at least one second audio channel is any other channel among the multiple audio channels besides the first audio channel.

[0166] The second mixed audio signal is the superposition of the modulation signal of the first mixed audio signal and the first audio signal of the first audio channel.

[0167] The modulation signal of the first mixed audio signal is a high-frequency signal obtained by the audio playback device modulating the signal in the low-frequency band or all frequency bands of the first mixed audio signal. In this application, the minimum frequency of the high-frequency band is greater than the maximum frequency of the low-frequency band (i.e., the frequency band of the modulation signal does not overlap with the frequency band of the first audio signal). For example, the high-frequency band in this application can be the ultrasonic frequency band.

[0168] In some examples, if the first mixed audio signal is a signal obtained by superimposing signals in the low-frequency bands of multiple first audio signals by the audio playback device, then the modulation signal of the first mixed audio signal can be a signal obtained by the audio playback device modulating signals in all frequency bands of the first mixed audio signal.

[0169] In other examples, if the first mixed audio signal is a signal obtained by superimposing signals across all frequency bands of multiple first audio signals by the audio playback device, then the modulation signal of the first mixed audio signal can be a signal obtained by modulating the low-frequency band of the first mixed audio signal by the audio playback device. For example, the audio playback device can first filter the first mixed audio signal, retaining the low-frequency band signal to obtain a fifth mixed audio signal; then, it can modulate the fifth mixed audio signal to obtain the modulation signal of the first mixed audio signal. Alternatively, the audio playback device can first modulate the first mixed audio signal, and then filter the modulation result to obtain the modulation signal of the first mixed audio signal.

[0170] For example, the audio playback device modulates the signal in the low-frequency band or all frequency bands of the first mixed audio signal. This can be done by up-converting the signal in the low-frequency band or all frequency bands of the first mixed audio signal; or, it can be done by first compressing the signal in the low-frequency band or all frequency bands of the first mixed audio signal, and then up-converting the compressed result. It is understood that in practical applications, the audio playback device can also determine the modulated signal of the first mixed audio signal through other modulation processing methods, and this application does not limit this. Taking the low-frequency band or all frequency bands of the first mixed audio signal as 0-4kHz as an example, if the sampling rate of the original audio signal is 48kHz, then the signal in the low-frequency band or all frequency bands of the first mixed audio signal can be up-converted to obtain a modulated signal with a frequency band of 20-24kHz.

[0171] The second mixed audio signal can be a signal obtained by directly superimposing the modulation signal of the first mixed audio signal with the first audio signal of the first audio channel. The frequency band of the second mixed audio signal includes the frequency band of the modulation signal and the frequency band of the first audio signal of the first audio channel.

[0172] For example, referring to FIG4, a simplified flowchart of determining a second mixed audio signal is provided in an embodiment of this application. As shown in FIG4, (a) represents the first audio signal of the first audio channel, (b) represents the modulation signal of the first mixed audio signal, and (c) represents the second mixed audio signal. In FIG4, the horizontal axis represents the frequency of the signal, and the vertical axis represents the amplitude of the signal. Taking FIG4(c) as an example, the horizontal axis f represents the frequency of the second mixed audio signal, and the vertical axis Z(f) represents the amplitude of the second mixed audio signal.

[0173] It should be understood that the waveforms of the signals in Figure 4 are for illustrative purposes only and do not constitute a limitation on the waveforms of the modulated signal, the second mixed audio signal, and the first audio signal of the first audio channel.

[0174] The audio playback device outputs a second mixed audio signal through its first audio channel. This can be done either by the audio playback device outputting the second mixed audio signal through its own first audio channel or by the audio playback device outputting the second mixed audio signal through the first audio channel of another device. For example, when the audio playback device is an audio playback device, it can output the second mixed audio signal through the first audio channel of a built-in speaker, or it can output the second mixed audio signal through the first audio channel of an external speaker.

[0175] It should be noted that although the first audio signal in the second mixed audio signal output from the first audio channel is not the original audio signal of the first audio channel, it is a filtered signal in the non-ultrasonic frequency band. Since the human ear cannot hear ultrasonic signals, the perception of the original audio signal from the first audio channel is the same as that of the original audio signal from the first audio channel. Therefore, the output of the first audio signal instead of the original audio signal from the first audio channel does not affect human perception. Similarly, the output of the first audio signal instead of the original audio signal from the second audio channel also does not affect human perception.

[0176] S330, the voice recognition device acquires the third mixed audio signal.

[0177] The third mixed audio signal is a superposition signal of the user's voice signal, the second mixed audio signal output from the first audio channel, and the first audio signal output from at least one second audio channel. The second mixed audio signal is a superposition signal of the modulation signal of the first mixed audio signal and the first audio signal of the first audio channel. The first mixed audio signal is obtained by superimposing the first audio signals of multiple audio channels of the audio playback device based on a first parameter. The first parameter is used to indicate the transmission differences of the multiple audio channels. The multiple audio channels include the first audio channel and the second audio channel. Therefore, the third mixed audio signal is also a superposition signal of the user's voice signal, the modulation signal of the first mixed audio signal, and the first audio signal of each of the multiple audio channels.

[0178] It should be understood that in practical applications, the third mixed audio signal acquired by the speech recognition device may also include other signals, such as environmental noise signals. The process of removing other signals from the third mixed audio signal can be referred to the description in related technologies, and will not be repeated here. For example, the speech recognition device can remove environmental noise from the third mixed audio signal using a noise reduction algorithm.

[0179] In some examples, where the speech recognition device is a speech recognition apparatus, it may acquire a third mixed audio signal via a built-in microphone. Where the speech recognition device is a component (e.g., a chip or circuit) within a speech recognition apparatus, it may acquire the third mixed audio signal from the component's external memory or interface.

[0180] S340, the voice recognition device performs filtering and demodulation processing on the third mixed audio signal to obtain the demodulated signal of the modulated signal.

[0181] Since the frequency band of the modulation signal of the first mixed audio signal (e.g., the ultrasonic frequency band) does not overlap with the frequency band of the user's voice signal and does not overlap with the frequency band of the first audio signal, the voice recognition device can filter the third mixed audio signal to separate the modulation signal of the first mixed audio signal from the third mixed audio signal.

[0182] Furthermore, the speech recognition device demodulates the modulated signal of the first mixed audio signal to obtain a demodulated signal. This demodulated signal can be a signal from the low-frequency band or all frequency bands of the first mixed audio signal. In some examples, if the modulated signal is obtained by modulating a signal from all frequency bands of the first mixed audio signal, then the demodulated signal can be a signal from all frequency bands of the first mixed audio signal. In other examples, if the modulated signal is obtained by modulating a signal from the low-frequency band of the first mixed audio signal, then the demodulated signal can be a signal from the low-frequency band of the first mixed audio signal.

[0183] For example, the demodulation processing of the modulated signal by the speech recognition device can be achieved by down-converting the signal across all frequency bands of the modulated signal. It is understood that in practical applications, the speech recognition device can also determine the demodulated signal through other demodulation processing methods, and this application does not limit this. Taking a modulated signal with a frequency band of 20-24kHz as an example, if the sampling rate of the original audio signal is 48kHz, then down-converting the modulated signal can yield a demodulated signal with a frequency band of 0-4kHz.

[0184] S350: The voice recognition device performs echo cancellation processing on the third mixed audio signal based on the demodulated signal to obtain the user's voice signal.

[0185] For example, the speech recognition device can filter out the modulation signal from the third mixed audio signal to obtain a sixth mixed audio signal. This sixth mixed audio signal is a superposition of the user's speech signal and the first audio signal of each of the multiple audio channels. Then, the speech recognition device can use the demodulated signal as a re-sampling signal to perform echo cancellation processing on the sixth mixed audio signal to obtain the user's speech signal.

[0186] In some examples, the speech recognition device can use the demodulated signal as the re-acquisition signal and perform echo cancellation processing on the sixth mixed audio signal using a pre-trained echo cancellation model to obtain the user's speech signal. The echo cancellation model uses the sample demodulated signal and the sample sixth mixed audio signal (which can be a superposition of the sample user's speech signal and the sample demodulated signal) as model inputs and the sample user's speech signal as model output, obtained through a deep neural network algorithm. In other examples, the speech recognition device can also use the demodulated signal as the re-acquisition signal and perform echo cancellation processing on the sixth mixed audio signal using an echo cancellation algorithm (e.g., an adaptive filtering echo cancellation algorithm) to obtain the user's speech signal. The specific process of performing echo cancellation processing on the sixth mixed audio signal can be found in the descriptions in related technologies, and will not be repeated here.

[0187] In this application, the audio playback device superimposes the modulation signal of the first mixed audio signal with the first audio signal of the first audio channel and outputs it in the first audio channel, and outputs the first audio signal of the second audio channel in the second audio channel. Therefore, the signal recognized by the speech recognition device can include the user's voice signal, the modulation signal of the first mixed audio signal, and multiple first audio signals (including the first audio signal of the first audio channel and the first audio signal of each second audio channel). Since the modulation signal is a high-frequency signal (e.g., a signal in the ultrasonic frequency band), and the first audio signal is a filtered non-high-frequency signal (e.g., a signal in a non-ultrasonic frequency band), and the minimum frequency of the high-frequency band is greater than the maximum frequency of the first audio signal, the separation of the modulation signal and the first audio signal can be achieved through filtering. Furthermore, since the user's voice signal is a low-frequency signal, and the maximum frequency of the low-frequency band is less than the minimum frequency of the high-frequency band, the separation of the modulation signal and the user's voice signal can also be achieved through filtering. Thus, the speech recognition device can separate the modulation signal from the recognized signal by filtering it.

[0188] Furthermore, in this application, the modulation signal of the first mixed audio signal can be obtained by modulating a low-frequency signal. For example, it can be obtained by modulating the low-frequency band of the first mixed audio signal, or the first mixed audio signal itself is a low-frequency signal (low-pass filtering has already been performed when multiple first audio signals are superimposed). And, the first mixed audio signal is obtained by superimposing multiple first audio signals. Therefore, demodulating the modulation signal of the first mixed audio signal yields a demodulated signal that is the low-frequency portion of the multiple first audio signals. Since the user's voice signal is a low-frequency signal, the part of the signal recognized by the voice recognition device that interferes with the user's voice signal is also the low-frequency portion of the multiple first audio signals. Therefore, by using the demodulated signal as the retrieval signal, the voice recognition device can accurately determine the user's voice signal from the recognized signal.

[0189] In practical applications, the modulation signal of the first mixed audio signal can also be obtained by modulating all frequency bands of the first mixed audio signal, and the first mixed audio signal is obtained by superimposing signals from all frequency bands of multiple first audio signals (the frequency bands of the first audio signals do not overlap with the ultrasonic frequency band). Taking the entire frequency band of the first mixed audio signal as 0-24kHz as an example, if the sampling rate of the original audio signal is 96kHz, then up-conversion processing of the signals from all frequency bands of the first mixed audio signal can obtain a modulation signal with a frequency band of 24-48kHz. Then, demodulating the modulation signal of the first mixed audio signal yields a demodulated signal, which is also a plurality of first audio signals. In this way, the speech recognition device can accurately determine the user's speech signal by filtering out the modulation signal portion from the recognized signal and removing the portion related to the plurality of first audio signals from the remaining signal based on the demodulated signal.

[0190] Furthermore, in related technologies, each audio channel outputs a modulated signal of its first audio signal. In this application, however, the first audio channel replaces the second audio channel to output the modulated signal corresponding to the second audio channel; that is, the first audio channel outputs a mixed modulated signal of multiple audio channels (i.e., the modulated signal of the first mixed audio signal). However, different audio channels exhibit transmission differences. If the first audio channel directly replaces the second audio channel, the transmission differences between the first and second audio channels will affect the accuracy of the ultimately determined user voice signal. Therefore, this application can pre-determine a first parameter to indicate the transmission differences of multiple audio channels. When the audio playback device performs superposition processing on the first audio signals of multiple audio channels, it can refer to this first parameter to cancel out the transmission differences of the multiple audio channels. This reduces the impact of audio channel transmission differences on the user voice signal recognition process, thereby ensuring the accuracy of the recognized user voice signal.

[0191] In summary, the signal processing method provided in this application embodiment, when the speech recognition device cannot obtain the original audio signal from the audio playback device, and the audio playback device has multiple audio channels, can first superimpose the first audio signals to be output from multiple audio channels to obtain a first mixed audio signal when it is necessary to output audio signals from multiple audio channels. Then, the audio playback device can modulate the first mixed audio signal to obtain a high-frequency modulated signal. Afterwards, the audio playback device can superimpose the modulated signal with the first audio signal from the first audio channel and output it from the first audio channel, and output the first audio signal corresponding to the second audio channel from the second audio channel. Therefore, the signal recognized by the speech recognition device includes the user's speech signal, the first audio signal output from each of the multiple audio channels, and the modulated signal output from the first audio channel. Since the modulated signal is a high-frequency signal (the minimum frequency of this high-frequency band is greater than the maximum frequency of the first audio signal), and the user's speech signal is a low-frequency signal, the speech recognition device can separate the modulated signal from the recognized signal by performing filtering processing. Furthermore, since the modulation signal is obtained by modulating the superposition of multiple first audio signals, the speech recognition device can obtain the re-sampled signals of multiple first audio signals by demodulating the modulation signal. Therefore, the speech recognition device can accurately separate the user's speech signal from the recognized signal based on the re-sampled signal. In addition, different audio channels have transmission differences. If the audio playback device directly superimposes multiple first audio signals and then modulates them, the re-sampled signal (i.e., demodulated signal) of the multiple first audio signals determined by the speech recognition device from the recognized signal will differ significantly from the first audio signals actually received by the speech recognition device from the multiple audio channels. Based on this, this embodiment of the application also pre-determines a first parameter to indicate the transmission differences of multiple audio channels. When the audio playback device superimposes the first audio signals of multiple audio channels, it refers to this first parameter to cancel out the transmission differences of the multiple audio channels. This improves the accuracy of the re-sampled signal determined by the speech recognition device, thereby improving the accuracy of the speech recognition device in recognizing the user's speech signal.

[0192] As can be seen, the signal processing method provided in this application embodiment involves complex processing procedures such as modulation and superposition for only one audio channel (i.e., the first audio channel) in the audio playback device. Other audio channels besides the first audio channel do not involve such complex processing procedures. Therefore, this application embodiment does not significantly increase the computing power requirements of the audio playback device. Furthermore, when the audio playback device outputs the modulated signal resulting from the superposition of multiple first audio signals through the first audio channel, the transmission differences between the multiple audio channels are canceled out. Therefore, the signal processing method provided in this application embodiment can ensure the accuracy of the speech recognition device in recognizing the user's speech signal. Thus, this application embodiment can improve the accuracy of the speech recognition device in recognizing the user's speech signal while reducing the computing power requirements of the audio playback device.

[0193] It should be noted that in this application, the actual waveform of the same audio signal may differ to some extent in different signal processing stages. For example, the actual waveform of the second mixed audio signal in the third mixed audio signal acquired in S330 may differ to some extent from that of the second mixed audio signal output in S320. Even though the signals referred to by the same term may have different actual forms (e.g., different waveforms) in different signal processing stages, those skilled in the art will understand that these signals with different actual forms are still the signals referred to by the same term mentioned above.

[0194] In one possible implementation, the first parameter includes second parameters for multiple audio channels. S310 described above may include: the audio playback device determining a third parameter for a second audio channel based on the second parameters of the multiple audio channels; and the audio playback device performing superposition processing on the first audio signals from the multiple audio channels based on the third parameter to obtain a first mixed audio signal.

[0195] The second parameter of the audio channel is used to indicate the signal gain of the audio signal transmitted from the audio channel to the speech recognition device (or, the second parameter of the audio channel can be the signal gain of the audio channel), and the third parameter is used to indicate the deviation between the second parameter of the second audio channel and the second parameter of the first audio channel (or, the third parameter of the second audio channel is used to indicate the difference in signal gain between the second audio channel and the first audio channel).

[0196] For example, if the original signal amplitude of an audio signal (e.g., any audio signal) is P u The amplitude of the audio signal output from a certain audio channel is P. v Then the second parameter α of the audio channel is P v / P u (representing P) vWith P u The ratio of P to P', that is, the signal gain of the audio signal transmitted from the audio channel to the speech recognition device is P. v / P u If the second parameter of the first audio channel is α x The second parameter of the second audio channel is α. y Then the third parameter of the second audio channel is α. y / α x The signal amplitude can also be replaced by the signal playback volume, signal playback strength, signal playback power, etc.

[0197] Taking the first mixed audio signal as the result of superimposing signals from all frequency bands of the first audio signals from multiple audio channels as an example, if the multiple audio channels include N channels, and the third parameter of the k-th channel among the N channels is α k The first audio signal of the k-th channel out of N channels is S. k (t), then the first mixed audio signal It should be understood that the third parameter of the first audio channel out of N channels is 1, indicating that the signal gain of the first audio channel is no different from that of the first audio channel itself. For example, the second parameter of the first audio channel is α. x The second parameter of the k-th channel (any second audio channel) is α. y (Then the third parameter α of the k-th channel) k yes In the case where the first audio signal S is output by the k-th channel itself... k (t), then the first audio signal recognized by the speech recognition device is α. y S k (t); If the first audio signal S is output from the first audio channel k (t), then the first audio signal recognized by the speech recognition device is α. x S k (t); if the third parameter is output from the first audio channel after passing through the k-th channel. Processed audio signal The first audio signal recognized by the voice recognition device is

[0198] It should be understood that in practical applications, the method of determining the third parameter based on the second parameter can also be other methods, and correspondingly, the method of determining the first mixed audio signal based on the third parameter can also be other methods, which are not limited in this application.

[0199] As can be seen from the above example, when superimposing the first audio signals of multiple audio channels, this application processes the first audio signal of the second audio channel using the third parameter of the second audio channel, which can offset the difference between the signal gain of the second audio channel and the signal gain of the first audio channel.

[0200] Differences in audio channel transmission can include differences in signal gain between the audio channels. Therefore, when the first audio channel replaces the second audio channel to output the modulation signal corresponding to the second audio channel—that is, when the first audio channel outputs a mixed modulation signal of multiple audio channels (i.e., the modulation signal of the first mixed audio signal)—canceling the difference in signal gain between the second and first audio channels can reduce the impact of this difference on the user's speech signal recognition process, thereby ensuring the accuracy of the speech recognition device in recognizing the user's speech signal. Therefore, this embodiment of the application, by canceling the differences in signal gain between multiple audio channels, can improve the accuracy of the speech recognition device in recognizing the user's speech signal while reducing the computational power requirements of the audio playback device.

[0201] In another possible implementation, the first parameter includes a third parameter of the second audio channel, which indicates the amount of deviation between the second parameter of the second audio channel and the second parameter of the first audio channel. The second parameter of the audio channel indicates the signal gain of the audio signal transmitted from the audio channel to the speech recognition device.

[0202] Because measuring relative signal gain (which is the signal gain of the second audio channel relative to the first audio channel, i.e., the third parameter of the second audio channel) is easier than measuring absolute signal gain (which is the signal gain of the audio channel, i.e., the second parameter of the audio channel), and because in the process of superimposing the first audio signals from multiple audio channels to obtain the first mixed audio signal, the difference in signal gain between the second and first audio channels can be offset by directly using the relative signal gain as the superposition coefficient of the first audio signals of the audio channels. Therefore, when the audio playback device superimposes multiple first audio signals based on the first parameter, directly using the relative signal gain as the first parameter can reduce the computational complexity of the signal processing process, thereby further reducing the computing power requirements of the audio playback device.

[0203] Optionally, multiple audio channels are determined from candidate audio channels based on a second parameter of the candidate audio channels of the audio playback device.

[0204] The candidate audio channels of an audio playback device are all the audio channels of the audio playback device, including all the audio channels of the built-in speakers of the audio playback device and all the audio channels of the external speakers of the audio playback device.

[0205] In some examples, multiple audio channels can be the M audio channels with the largest second parameter among the candidate audio channels (i.e., the M audio channels with the largest signal gain). M can be a predetermined positive integer. When the number of candidate audio channels is less than or equal to M, the candidate audio channels can be determined as multiple audio channels. In other examples, multiple audio channels can be the audio channels among the candidate audio channels whose second parameter is greater than a predetermined signal gain threshold. It should be understood that in practical applications, the method of determining multiple audio channels based on the second parameter can also be other, and this application does not limit this.

[0206] When the signal gain of an audio channel is very small, the audio signal output from the audio channel has little impact on the accuracy of the speech recognition device in recognizing the user's speech signal. Based on this, in this application, when the audio playback device superimposes the first audio signals from multiple audio channels with very small signal gains, it can only superimpose the first audio signals from these channels. This further reduces the computational complexity of the signal processing, thereby further reducing the computing power requirements of the audio playback device.

[0207] In another possible implementation, the first parameter includes a fourth parameter of multiple audio channels. S310 described above may include: the audio playback device determining a fifth parameter of the second audio channel based on the fourth parameter of the multiple audio channels; and the audio playback device performing superposition processing on the first audio signals of the multiple audio channels based on the fifth parameter to obtain a first mixed audio signal.

[0208] The fourth parameter of the audio channel is used to indicate the signal delay of the audio signal from the audio channel to the speech recognition device (or, the fourth parameter of the audio channel can be the signal delay of the audio channel), and the fifth parameter is used to indicate the deviation between the fourth parameter of the second audio channel and the fourth parameter of the first audio channel (or, the fifth parameter of the second audio channel is used to indicate the difference in signal delay between the second audio channel and the first audio channel).

[0209] For example, if the output time of an audio playback device outputting an audio signal (e.g., any audio signal) through an audio channel is t... u The reception time of the audio signal received by the voice recognition device is t. v Then the fourth parameter τ of this audio channel is t v -t u That is, the signal delay of the audio signal from the audio channel to the speech recognition device is t. v -t u If the fourth parameter of the first audio channel is τ xThe fourth parameter of the second audio channel is τ. y Then the fifth parameter of the second audio channel is τ. y -τ x The signal delay may include the transmission delay of the audio playback device outputting the audio signal through the audio channel, and the transmission delay of the audio signal from the audio channel to the speech recognition device.

[0210] Taking the example that the first mixed audio signal is the result of superimposing signals from all frequency bands of the first audio signals from multiple audio channels, if the multiple audio channels include N channels, and the fifth parameter of the k-th channel among the N channels is τ k The first audio signal of the k-th channel out of N channels is S. k (t), then the first mixed audio signal It should be understood that the fifth parameter of the first audio channel out of N channels is 0, indicating that the signal delay of the first audio channel is no different from that of the first audio channel itself. For example, the fourth parameter of the first audio channel is τ. x The fourth parameter of the k-th channel (any second audio channel) is τ. y (Then the fifth parameter τ of the k-th channel) k It is τ y -τ x In the case where the first audio signal S is output by the k-th channel itself... k (t), then the first audio signal recognized by the speech recognition device is S. k (t+τ y If the first audio signal S is output from the first audio channel; k (t), then the first audio signal recognized by the speech recognition device is S. k (t+τ x If the fifth parameter τ is output from the first audio channel and passes through the k-th channel... y -τ x Processed audio signal S k (t+τ y -τ x If the first audio signal recognized by the speech recognition device is S, then the first audio signal recognized by the speech recognition device is S. k (t+τ y -τ x +τ x ) = S k (t+τ y ).

[0211] It should be understood that in practical applications, the method of determining the fifth parameter based on the fourth parameter can also be other methods, and correspondingly, the method of determining the first mixed audio signal based on the fifth parameter can also be other methods, which are not limited in this application.

[0212] As can be seen from the above example, when superimposing the first audio signals of multiple audio channels, this application processes the first audio signal of the second audio channel through the fifth parameter of the second audio channel, which can offset the difference between the signal delay of the second audio channel and the signal delay of the first audio channel.

[0213] Differences in audio channel transmission can include differences in signal delay. Therefore, when the first audio channel replaces the second audio channel to output the modulation signal corresponding to the second audio channel—that is, when the first audio channel outputs a mixed modulation signal of multiple audio channels (i.e., the modulation signal of the first mixed audio signal)—canceling the difference in signal delay between the second and first audio channels can reduce the impact of this difference on the user's speech signal recognition process, thereby ensuring the accuracy of the speech recognition device in recognizing the user's speech signal. Therefore, this embodiment of the application, by canceling the differences in signal delay between multiple audio channels, can improve the accuracy of the speech recognition device in recognizing the user's speech signal while reducing the computational power requirements of the audio playback device.

[0214] In another possible implementation, the first parameter includes a fifth parameter of the second audio channel, which indicates the amount of deviation between the fourth parameter of the second audio channel and the fourth parameter of the first audio channel. The fourth parameter of the audio channel indicates the signal delay of the audio signal transmitted from the audio channel to the speech recognition device.

[0215] Because measuring relative signal delay (which is the signal delay of the second audio channel relative to the first audio channel, i.e., the fifth parameter of the second audio channel) is easier than measuring absolute signal delay (which is the signal delay of the audio channel, i.e., the fourth parameter of the audio channel), and because in the process of superimposing the first audio signals from multiple audio channels to obtain the first mixed audio signal, the difference in signal delay between the second and first audio channels can be offset by directly using the relative signal delay as the superposition coefficient of the first audio signals of the audio channels. Therefore, when the audio playback device superimposes multiple first audio signals based on the first parameter, directly using the relative signal delay as the first parameter can reduce the computational complexity of the signal processing process, thereby further reducing the computing power requirements of the audio playback device.

[0216] Optionally, the first audio channel is determined from multiple audio channels based on a first parameter.

[0217] In some examples, if the first parameter includes the fourth parameter of multiple audio channels, then the first audio channel can be the audio channel with the smallest fourth parameter (i.e., the smallest signal delay) among the multiple audio channels. In this way, the fifth parameter of the second audio channel is always a positive value. When the audio playback device superimposes the first audio signals of multiple audio channels, the computational complexity can be reduced, thereby further reducing the computing power requirements of the audio playback device.

[0218] In other examples, if the first parameter includes second parameters for multiple audio channels, then the first audio channel can be the audio channel with the largest second parameter (i.e., the largest signal gain) among the multiple audio channels. This can improve the pickup accuracy of the speech recognition device for the modulated signal of the first mixed audio signal, thereby further improving the accuracy of the speech recognition device in recognizing the user's speech signal.

[0219] In another possible implementation, before S330 (in practical applications, it may be before S310), the signal processing method provided in this application embodiment further includes: the voice recognition device acquiring a first parameter; and the voice recognition device sending first information.

[0220] The first information is used to instruct the audio playback device to determine the first mixed audio signal based on the first parameter.

[0221] In some examples, the audio playback device includes two playback modes: a normal playback mode and a re-acquisition playback mode. The normal playback mode is suitable for scenarios where the speech recognition device does not need to acquire re-acquisition signals (e.g., when the speech recognition device is not powered on, or when the speech recognition device can acquire the original audio signal from the audio playback device). The re-acquisition playback mode is suitable for scenarios where the speech recognition device needs to acquire re-acquisition signals (e.g., when the speech recognition device cannot acquire the original audio signal from the audio playback device). When the audio playback device operates in normal playback mode, each audio channel outputs either its first audio signal or the original audio signal. When the audio playback device operates in re-acquisition playback mode, the audio playback device needs to determine a first mixed audio signal based on a first parameter, and outputs a superposition signal of the modulation signal of the first mixed audio signal and the first audio signal of the first audio channel in the first audio channel, and outputs the first audio signal of the second audio channel in the second audio channel.

[0222] As can be seen, in this application, the voice recognition device can send first information to cause the audio playback device to switch playback modes to adapt to the needs of different application scenarios.

[0223] In another possible implementation, the first parameter includes a third parameter of the second audio channel. This third parameter indicates the deviation between the second parameter of the second audio channel and the second parameter of the first audio channel. The second parameter of the audio channel indicates the signal gain of the audio signal transmitted from the audio channel to the speech recognition device. The aforementioned step of "the speech recognition device acquiring the first parameter" may include: the speech recognition device sending second information; and the speech recognition device determining the third parameter based on the signal amplitudes of multiple preset audio signals received from multiple audio channels.

[0224] The second piece of information is used to instruct multiple audio channels to output preset audio signals sequentially. For example, the preset audio signal can be any pre-acquired audio signal. The output interval of the preset audio signals from the multiple audio channels can be a predetermined interval, which must ensure that the multiple audio channels do not output the preset audio signals at the same time to avoid mutual interference between the preset audio signals output by the multiple audio channels.

[0225] In some examples, after receiving multiple preset audio signals from multiple audio channels, the speech recognition device can identify the audio channel corresponding to the preset audio signal with the largest signal amplitude among the multiple preset audio signals as the first audio channel, and identify the audio channels other than the first audio channel as the second audio channels. If the signal amplitude of the preset audio signal received by the speech recognition device from the first audio channel is P1, and the signal amplitude of the preset audio signal received by the speech recognition device from the second audio channel is P2, then P2 / P1 (i.e., the ratio of P2 to P1) can be determined as the third parameter of the second audio channel.

[0226] In this application, the speech recognition device can send a second message to cause multiple audio channels of the audio playback device to sequentially output preset audio signals. In this way, the speech recognition device can quickly determine the relative signal gain (i.e., the third parameter of the second audio channel) based on the signal amplitudes of the received multiple preset audio signals, without needing to determine the absolute signal gain (i.e., the second parameter), thereby reducing the computational complexity of the speech recognition device.

[0227] In another possible implementation, the first parameter includes a fifth parameter of the second audio channel, which indicates the deviation between the fourth parameter of the second audio channel and the fourth parameter of the first audio channel. The fourth parameter of the audio channel indicates the signal delay of the audio signal transmitted from the audio channel to the speech recognition device. The above-described step of "the speech recognition device acquiring the first parameter" may include: the speech recognition device sending third information; and the speech recognition device determining the fifth parameter based on a preset output interval and the reception time of multiple preset audio signals received from multiple audio channels.

[0228] The third piece of information is used to instruct multiple audio channels to output preset audio signals sequentially based on a preset output interval. The preset output interval can be a predetermined output interval that ensures that multiple audio channels do not output preset audio signals at the same time, so as to avoid mutual interference between the preset audio signals output by multiple audio channels.

[0229] In some examples, after a speech recognition device receives multiple preset audio signals from multiple audio channels, it can identify the audio channel corresponding to the first received preset audio signal as the first audio channel, and identify the audio channels other than the first audio channel as the second audio channels. If the reception time of the first received preset audio signal is t1, the reception time of the kth received preset audio signal is tk, and the preset output interval is t, then the fifth parameter of the audio channel corresponding to the kth received preset audio signal is tk-t1-(k-1)t. For example, if the reception time of the second received preset audio signal is t2, then the fifth parameter of the audio channel corresponding to the second received preset audio signal is t2-t1-t.

[0230] In this application, the speech recognition device can send a third message to cause multiple audio channels of the audio playback device to sequentially output preset audio signals based on a preset output interval. In this way, the speech recognition device can quickly determine the relative signal delay (i.e., the fifth parameter of the second audio channel) based on the preset output interval and the reception time of the multiple preset audio signals, without having to determine the absolute signal delay (i.e., the fourth parameter), thereby reducing the computational complexity of the speech recognition device.

[0231] Optionally, the preset frequency band of the audio signal is the same as the frequency band of the demodulated signal.

[0232] To more accurately assess the transmission differences among multiple audio channels and thus determine a more precise first parameter, in this application, the preset audio signal can be an audio signal with the same frequency band as the demodulated signal; that is, the frequency band of the preset audio signal can be the low-frequency band in this application. For example, the frequency band of the preset audio signal can be 0-4kHz. In other examples, if the frequency response curve (which characterizes the gain ratio of a signal in the 20-24kHz band to a signal in the 0-4kHz band) is known, then the frequency band of the preset audio signal can also be 20-24kHz. When determining the signal gain, the speech recognition device can calculate the signal gain of the low-frequency signal by combining the frequency response curve.

[0233] In another possible implementation, before S330 (in practical applications, it could be before S310), the signal processing method provided in this application embodiment further includes: the speech recognition device determining a first audio channel from multiple audio channels based on the first parameter. The specific process by which the speech recognition device determines the first audio channel can be referred to the relevant descriptions above, and will not be repeated here.

[0234] To more clearly illustrate the technical solution provided in this application, a simplified flowchart of a signal processing method provided in an embodiment of this application is shown in Figure 5. This method can be applied to an audio playback device. As shown in Figure 5, the audio playback device includes k audio channels, each of which outputs its corresponding original audio signal. Audio channel 1 is the first audio channel, and the other audio channels are second audio channels. The audio playback device can first perform low-pass filtering on the original audio signal of each of the k audio channels to filter out the ultrasonic frequency band portion of the original audio signal of each audio channel, obtaining the first audio signal of each audio channel. Then, the audio playback device can perform superposition processing on the first audio signal of each audio channel based on a first parameter to obtain a first mixed audio signal; and, it can perform modulation processing on the first mixed audio signal to obtain a modulated signal of the ultrasonic frequency band. Then, the modulated signal and the first audio signal of audio channel 1 can be superimposed, and the superposition processing result is output on audio channel 1; simultaneously, the first audio signal corresponding to each channel is output on the other channels.

[0235] It should be understood that, in practical applications, various possible implementations of the above signal processing methods can be combined. Below, one such combination is illustrated by example. Referring to Figure 6, a flowchart of another signal processing method provided by an embodiment of this application is shown. For example, this method can be applied to the interactive scenario shown in Figure 2, or to other scenarios where a speech recognition device interacts with an audio playback device. As shown in Figure 6, the method includes the following steps:

[0236] S610, the voice recognition device sends third information to the audio playback device.

[0237] The third information is used to instruct multiple audio channels to output preset audio signals sequentially based on preset output intervals.

[0238] In another possible implementation, prior to S610, the signal processing method provided in this application embodiment may further include: establishing a connection between the voice recognition device and the audio playback device. Specific connection establishment methods can be found in the descriptions in related technologies, and will not be repeated here.

[0239] S620, the audio playback device outputs preset audio signals sequentially through multiple audio channels based on preset output intervals.

[0240] The S630 voice recognition device receives multiple preset audio signals through multiple audio channels.

[0241] S640, the voice recognition device determines a third parameter based on the signal amplitude of multiple preset audio signals, and determines a fifth parameter based on the preset output interval and the reception time of multiple preset audio signals.

[0242] S650, the voice recognition device sends the first information to the audio playback device.

[0243] The first information is specifically used to instruct the audio playback device to determine the first mixed audio signal based on the fifth parameter and the third parameter.

[0244] Since the fifth and third parameters of the audio channel may change (for example, the third parameter may change when the playback volume of the audio channel changes; the fifth parameter may change when the deployment distance between the speech recognition device and the audio playback device changes), in another possible implementation, the speech recognition device may also periodically update the fifth and third parameters based on a predetermined time interval, and send the first information to the audio playback device after the update. This can further improve the accuracy of the speech recognition device in recognizing user voice signals.

[0245] In another possible implementation, prior to S650, the signal processing method provided in this application embodiment may further include: the voice recognition device detecting whether it can acquire the original audio signal of the audio playback device; if the voice recognition device cannot acquire the original audio signal of the audio playback device, the voice recognition device sending first information; if the voice recognition device can acquire the original audio signal of the audio playback device, the voice recognition device not sending first information.

[0246] S660, the audio playback device, based on the fifth parameter and the third parameter, performs superposition processing on the first audio signals to be output from multiple audio channels to obtain the first mixed audio signal.

[0247] Taking the first mixed audio signal as the result of superimposing signals from all frequency bands of the first audio signals from multiple audio channels as an example, if the multiple audio channels include N channels, and the third parameter of the k-th channel among the N channels is α k The fifth parameter of the k-th channel is τ. k The first audio signal of the k-th channel is S. k (t), then the first mixed audio signal In this configuration, the third parameter of the first audio channel out of the N channels is 1, and the fifth parameter of the first audio channel out of the N channels is 0. For example, the second parameter of the first audio channel is α. x The second parameter of the k-th channel (any second audio channel) is α. y (Then the third parameter α of the k-th channel) k yes The fourth parameter of the first audio channel is τ. x The fourth parameter of the k-th channel (any second audio channel) is τ. y (Then the fifth parameter τ of the k-th channel) k It is τ y -τ x In the case where the first audio signal S is output by the k-th channel itself... k (t), then the first audio signal recognized by the speech recognition device is α. y S k (t+τ y If the first audio signal S is output from the first audio channel; k (t), then the first audio signal recognized by the speech recognition device is α. x S k (t+τ x If the third parameter is output from the first audio channel after passing through the k-th channel; And the fifth parameter τ of the k-th channel y -τ x Processed audio signal The first audio signal recognized by the voice recognition device is

[0248] As can be seen from the above examples, when superimposing the first audio signals from multiple audio channels, this application, by using the third and fifth parameters of the second audio channel to process the first audio signal of the second audio channel, can offset the difference in signal gain between the second and first audio channels, and also offset the difference in signal delay between the second and first audio channels. This not only reduces the impact of differences in signal gain between audio channels on the user's speech signal recognition process, but also reduces the impact of differences in signal delay between audio channels on the user's speech signal recognition process, thereby further improving the accuracy of the speech recognition device in recognizing user's speech signals.

[0249] S670, The audio playback device modulates the first mixed audio signal to obtain the modulated signal of the first mixed audio signal.

[0250] S680, the audio playback device outputs a second mixed audio signal in the first audio channel and outputs the first audio signal corresponding to the second audio channel in the second audio channel.

[0251] S690, the voice recognition device acquires the third mixed audio signal.

[0252] S6100, the voice recognition device filters the third mixed audio signal to obtain the modulated signal of the first mixed audio signal.

[0253] S6110, The voice recognition device demodulates the modulation signal of the first mixed audio signal to obtain a demodulated signal.

[0254] S6120, the voice recognition device performs echo cancellation processing on the third mixed audio signal based on the demodulated signal to obtain the user's voice signal.

[0255] The implementation process of some steps in the signal processing method shown in Figure 6 can be referred to the relevant descriptions above, and will not be repeated here.

[0256] It should be noted that the order of the methods provided in the embodiments of this application can be appropriately adjusted, and the process can also be added or removed as appropriate. Any variations that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the protection scope of this application, and the embodiments of this application do not limit this.

[0257] The signal processing apparatus that performs the signal processing method provided in this application will now be described. Details of the signal processing apparatus performing the signal processing method can be found in the foregoing description of the method embodiments, and will not be repeated here.

[0258] Referring to Figure 7, which is a schematic diagram of a signal processing device provided in an embodiment of this application, the signal processing device 700 includes a processor 710 and a communication interface 720, which can be interconnected via a bus 730. The signal processing device 700 can be an audio playback device or a speech recognition device.

[0259] Optionally, as shown in FIG7, the signal processing device 700 may further include a memory 740. The memory 740 includes, but is not limited to, random access memory (RAM), read-only memory (ROM), cache, erasable programmable read-only memory (EPROM), synchronous dynamic random access memory (SDRAM), hard disk drive (HDD), solid-state drive (SSD), or compact disc read-only memory (CD-ROM). The memory 740 is used to store program instructions or information accessed by application processes. The processor 710 can implement the methods provided in this application by executing the program instructions in the memory 740. The memory 740 may be integrated with the processor 710 or disposed separately.

[0260] Processor 710 can be one or more of the following: CPU, application-specific integrated circuit (ASIC), DSP, microprocessor unit (MPU), microcontroller unit (MCU), GPU, field-programmable gate array (FPGA), and NPU, or a combination thereof. If processor 710 is a CPU, the CPU can be a single-core CPU or a multi-core CPU. Processor 710 can be a signal processor, chip, or other integrated circuit capable of implementing the methods provided in this application, or it can be a portion of the circuitry within the aforementioned processor, chip, or integrated circuit used for signal or data processing. Additionally, communication interface 720 can be an input / output interface, used for inputting or outputting signals or data, or it can be an input / output circuit.

[0261] The above description of memory 740 and processor 710 also applies to other memories and processors in the embodiments of this application, such as internal memory 121 and processor 110 of FIG1.

[0262] Taking an audio playback device as an example, the processor 710 can perform the following operations: based on the first parameter, superimpose the first audio signals to be output from multiple audio channels of the audio playback device to obtain a first mixed audio signal; and output a second mixed audio signal from the first audio channel.

[0263] Taking a speech recognition device as an example, the processor 710 can perform the following operations: acquire a third mixed audio signal; perform filtering and demodulation processing on the third mixed audio signal to obtain a demodulated signal of the modulated signal; and perform echo cancellation processing on the third mixed audio signal based on the demodulated signal to obtain a user speech signal.

[0264] It is understood that, in order to achieve the above-mentioned functions (i.e., to execute the methods provided in this application), the signal processing device includes hardware and / or software modules corresponding to each function. Referring to FIG8, a schematic diagram of another signal processing device provided in an embodiment of this application is shown. This signal processing device can be an audio playback device. As shown in FIG8, when each functional module is divided according to its corresponding function, the signal processing device 800 may include a processing module 810 and an output module 820.

[0265] The processing module 810 is used to perform superposition processing on the first audio signals to be output from multiple audio channels of the audio playback device based on the first parameter to obtain a first mixed audio signal; the first parameter is used to indicate the transmission differences of multiple audio channels; the output module 820 is used to output a second mixed audio signal from the first audio channel; the second mixed audio signal is the superposition signal of the modulation signal of the first mixed audio signal and the first audio signal of the first audio channel, the multiple audio channels include the first audio channel and at least one second audio channel, and the second audio channel is used to output the first audio signal corresponding to the second audio channel.

[0266] In another possible implementation, the first parameter includes second parameters of multiple audio channels, which indicate the signal gain of the audio signal transmitted from the audio channel to the speech recognition device. The processing module 810 is specifically configured to: determine a third parameter of the second audio channel based on the second parameters of the multiple audio channels; and perform superposition processing on the first audio signals of the multiple audio channels based on the third parameter to obtain a first mixed audio signal. The third parameter indicates the deviation between the second parameters of the second audio channel and the second parameters of the first audio channel.

[0267] In another possible implementation, the first parameter includes a third parameter of the second audio channel, which indicates the amount of deviation between the second parameter of the second audio channel and the second parameter of the first audio channel. The second parameter of the audio channel indicates the signal gain of the audio signal transmitted from the audio channel to the speech recognition device.

[0268] In another possible implementation, multiple audio channels are determined from candidate audio channels based on a second parameter of the candidate audio channels of the audio playback device.

[0269] In another possible implementation, the first parameter includes a fourth parameter of multiple audio channels, which indicates the signal delay of the audio signal transmitted from the audio channel to the speech recognition device. The processing module 810 is specifically configured to: determine a fifth parameter of the second audio channel based on the fourth parameters of the multiple audio channels; and, based on the fifth parameter, perform superposition processing on the first audio signals of the multiple audio channels to obtain a first mixed audio signal. The fifth parameter indicates the deviation between the fourth parameters of the second audio channel and the fourth parameters of the first audio channel.

[0270] In another possible implementation, the first parameter includes a fifth parameter of the second audio channel, which indicates the amount of deviation between the fourth parameter of the second audio channel and the fourth parameter of the first audio channel. The fourth parameter of the audio channel indicates the signal delay of the audio signal transmitted from the audio channel to the speech recognition device.

[0271] In another possible implementation, the first audio channel is determined from multiple audio channels based on a first parameter.

[0272] The signal processing apparatus 800 provided in this application embodiment can execute some steps of the aforementioned signal processing method. Taking the signal processing method shown in FIG3 as an example, the processing module 810 can be used to execute S310 in FIG3, and the output module 820 can be used to execute S320 in FIG3. The specific implementation process of the signal processing apparatus 800 and the corresponding beneficial effects can be referred to the relevant description of the aforementioned method embodiments, and will not be repeated here.

[0273] For example, the functions implemented by the processing module 810 and the output module 820 can both be implemented by the processor 710 in FIG7 executing the program instructions in the memory 740 in FIG7.

[0274] Referring to Figure 9, a schematic diagram of another signal processing device provided in an embodiment of this application is shown. This signal processing device can be a speech recognition device. As shown in Figure 9, when each functional module is divided according to its corresponding function, the signal processing device 900 may include an acquisition module 910 and a processing module 920.

[0275] The acquisition module 910 is used to acquire a third mixed audio signal; the third mixed audio signal is a superposition signal of the user's voice signal, a second mixed audio signal output from the first audio channel, and a first audio signal output from at least one second audio channel. The second mixed audio signal is a superposition signal of the modulation signal of the first mixed audio signal and the first audio signal of the first audio channel. The first mixed audio signal is obtained by superimposing the first audio signals of multiple audio channels of the audio playback device based on a first parameter, which is used to indicate the transmission differences of multiple audio channels. The multiple audio channels include the first audio channel and the second audio channel. The processing module 920 is used to perform filtering and demodulation processing on the third mixed audio signal to obtain a demodulated signal of the modulation signal. The processing module 920 is also used to perform echo cancellation processing on the third mixed audio signal based on the demodulated signal to obtain the user's voice signal.

[0276] In another possible implementation, the signal processing device 900 further includes a transmitting module. The processing module 920 is also configured to acquire a first parameter before the acquisition module 910 acquires the third mixed audio signal; the transmitting module is configured to transmit first information; the first information is configured to instruct the audio playback device to determine the first mixed audio signal based on the first parameter.

[0277] In another possible implementation, the first parameter includes a third parameter of the second audio channel. This third parameter indicates the deviation between the second parameter of the second audio channel and the second parameter of the first audio channel. The second parameter of the audio channel indicates the signal gain of the audio signal transmitted from the audio channel to the speech recognition device. The acquisition module 910 is specifically used to: send second information via the sending module; and determine the third parameter based on the signal amplitudes of multiple preset audio signals received from multiple audio channels. The second information is used to instruct the multiple audio channels to sequentially output preset audio signals.

[0278] In another possible implementation, the first parameter includes a fifth parameter of the second audio channel. This fifth parameter indicates the deviation between the fourth parameter of the second audio channel and the fourth parameter of the first audio channel. The fourth parameter of the audio channel indicates the signal delay of the audio signal transmitted from the audio channel to the speech recognition device. The acquisition module 910 is specifically used to: send third information through the sending module; and determine the fifth parameter based on a preset output interval and the reception time of multiple preset audio signals received from multiple audio channels. The third information instructs the multiple audio channels to sequentially output preset audio signals based on the preset output interval.

[0279] In another possible implementation, the preset frequency band of the audio signal is the same as the frequency band of the demodulated signal.

[0280] In another possible implementation, the processing module 920 is further configured to determine a first audio channel from multiple audio channels based on a first parameter before the acquisition module 910 acquires the third mixed audio signal.

[0281] The signal processing apparatus 900 provided in this application embodiment can execute some steps of the aforementioned signal processing method. Taking the signal processing method shown in FIG3 as an example, the acquisition module 910 can be used to execute S330 in FIG3, and the processing module 920 can be used to execute S340 and S350 in FIG3. The specific implementation process of the signal processing apparatus 900 and the corresponding beneficial effects can be referred to the relevant description of the aforementioned method embodiments, and will not be repeated here.

[0282] For example, the functions implemented by the acquisition module 910 and the processing module 920 can both be implemented by the processor 710 in FIG7 executing the program instructions in the memory 740 in FIG7.

[0283] It should be noted that the module division in Figures 8 and 9 is exemplary and only represents one logical functional division. In actual implementation, other division methods are possible. For example, two or more functions can be integrated into one module. The integrated module described above can be implemented in hardware or as a software functional module.

[0284] This application also provides a computer-readable storage medium including computer instructions that, when executed on a computer, cause the computer to perform the signal processing method provided in this application.

[0285] This application also provides a computer program product that, when run on a computer, causes the computer to execute the signal processing method provided in this application.

[0286] This application also provides a chip or chip system. Exemplarily, the chip or chip system may be the aforementioned signal processing apparatus. The chip or chip system includes: at least one processor, configured to execute a computer program or instructions to cause the signal processing method provided in this application to be performed.

[0287] This application also provides a signal processing system. The signal processing system includes an audio playback device and a speech recognition device. The audio playback device and the speech recognition device can cooperate to execute the signal processing method provided in this application.

[0288] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A signal processing method, characterized in that, The method includes: Based on the first parameter, the first audio signals to be output from multiple audio channels of the audio playback device are superimposed to obtain a first mixed audio signal; the first parameter is used to indicate the transmission differences of the multiple audio channels. A second mixed audio signal is output from a first audio channel; the second mixed audio signal is a superposition signal of the modulation signal of the first mixed audio signal and the first audio signal of the first audio channel, the plurality of audio channels include the first audio channel and at least one second audio channel, the second audio channel is used to output the first audio signal corresponding to the second audio channel.

2. The method according to claim 1, characterized in that, The first parameter includes the second parameters of the plurality of audio channels, wherein the second parameters of the audio channels are used to indicate the signal gain of the audio signal transmitted from the audio channel to the speech recognition device; the step of superimposing the first audio signals to be output from the plurality of audio channels of the audio playback device based on the first parameter to obtain the first mixed audio signal includes: Based on the second parameters of the plurality of audio channels, a third parameter of the second audio channel is determined; the third parameter is used to indicate the amount of deviation between the second parameter of the second audio channel and the second parameter of the first audio channel. Based on the third parameter, the first audio signals of the multiple audio channels are superimposed to obtain the first mixed audio signal.

3. The method according to claim 1, characterized in that, The first parameter includes a third parameter of the second audio channel, which is used to indicate the deviation between the second parameter of the second audio channel and the second parameter of the first audio channel. The second parameter of the audio channel is used to indicate the signal gain of the audio signal transmitted from the audio channel to the speech recognition device.

4. The method according to claim 2 or 3, characterized in that, The plurality of audio channels are determined from the candidate audio channels based on a second parameter of the candidate audio channels of the audio playback device.

5. The method according to any one of claims 1-4, characterized in that, The first parameter includes a fourth parameter of the plurality of audio channels, the fourth parameter of the audio channel being used to indicate the signal delay of the audio signal transmitted from the audio channel to the speech recognition device; the step of superimposing the first audio signals to be output from the plurality of audio channels of the audio playback device based on the first parameter to obtain a first mixed audio signal includes: Based on the fourth parameters of the plurality of audio channels, a fifth parameter of the second audio channel is determined; the fifth parameter is used to indicate the amount of deviation between the fourth parameter of the second audio channel and the fourth parameter of the first audio channel. Based on the fifth parameter, the first audio signals of the multiple audio channels are superimposed to obtain the first mixed audio signal.

6. The method according to any one of claims 1-4, characterized in that, The first parameter includes a fifth parameter of the second audio channel, which is used to indicate the deviation between the fourth parameter of the second audio channel and the fourth parameter of the first audio channel. The fourth parameter of the audio channel is used to indicate the signal delay of the audio signal transmitted from the audio channel to the speech recognition device.

7. The method according to any one of claims 1-6, characterized in that, The first audio channel is determined from the plurality of audio channels based on the first parameter.

8. A signal processing method, characterized in that, The method includes: A third mixed audio signal is acquired; the third mixed audio signal is a superposition signal of the user's voice signal, a second mixed audio signal output from the first audio channel, and a first audio signal output from at least one second audio channel; the second mixed audio signal is a superposition signal of the modulation signal of the first mixed audio signal and the first audio signal of the first audio channel; the first mixed audio signal is obtained by superimposing the first audio signals of multiple audio channels of the audio playback device based on a first parameter, the first parameter being used to indicate the transmission differences of the multiple audio channels; the multiple audio channels include the first audio channel and the second audio channel; The third mixed audio signal is filtered and demodulated to obtain the demodulated signal of the modulated signal; The user's voice signal is obtained by performing echo cancellation processing on the third mixed audio signal based on the demodulated signal.

9. The method according to claim 8, characterized in that, Prior to acquiring the third mixed audio signal, the method further includes: Obtain the first parameter; Send a first message; the first message is used to instruct the audio playback device to determine the first mixed audio signal based on the first parameter.

10. The method according to claim 9, characterized in that, The first parameter includes a third parameter of the second audio channel, the third parameter indicating the deviation between the second parameter of the second audio channel and the second parameter of the first audio channel, the second parameter of the audio channel indicating the signal gain of the audio signal transmitted from the audio channel to the speech recognition device; obtaining the first parameter includes: Send a second message; the second message is used to instruct the plurality of audio channels to output preset audio signals sequentially. The third parameter is determined based on the signal amplitude of the multiple preset audio signals received from the multiple audio channels.

11. The method according to claim 9 or 10, characterized in that, The first parameter includes a fifth parameter of the second audio channel, which is used to indicate the deviation between the fourth parameter of the second audio channel and the fourth parameter of the first audio channel. The fourth parameter of the audio channel is used to indicate the signal delay of the audio signal transmitted from the audio channel to the speech recognition device. Obtaining the first parameter includes: Send a third message; the third message is used to instruct the plurality of audio channels to output a preset audio signal sequentially based on a preset output interval; The fifth parameter is determined based on the preset output interval and the reception time of the multiple preset audio signals received from the multiple audio channels.

12. The method according to claim 10 or 11, characterized in that, The frequency band of the preset audio signal is the same as the frequency band of the demodulated signal.

13. The method according to any one of claims 8-12, characterized in that, Prior to acquiring the third mixed audio signal, the method further includes: The first audio channel is determined from the plurality of audio channels based on the first parameter.

14. A signal processing apparatus, characterized in that, It includes at least one module or at least one unit, said at least one module or said at least one unit being used to perform the method as described in any one of claims 1-13.

15. A signal processing apparatus, characterized in that, include: Memory and at least one processor; The memory is coupled to the processor; wherein the memory stores computer program code, the computer program code including computer instructions, and when the computer instructions are executed by the processor, the signal processing device performs the method as described in any one of claims 1-13.

16. A computer-readable storage medium comprising computer instructions, characterized in that, When the computer instructions are executed on the computer, the computer causes the computer to perform the method as described in any one of claims 1-13.

17. A computer program product, characterized in that, When the computer program product is run on a computer, it causes the computer to perform the method as described in any one of claims 1-13.

18. A chip or chip system, characterized in that, include: At least one processor, the at least one processor being configured to execute a computer program or instructions to cause the method of any one of claims 1-13 to be performed.