Audio processing method and electronic device
By generating an inaudible audio signal as a reference signal, modulating it onto the ultrasonic frequency band, and transmitting it through an air path, the echo noise problem between external playback devices and smart devices is solved, achieving efficient audio return and echo suppression, and improving the accuracy and efficiency of user audio processing.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HUAWEI TECH CO LTD
- Filing Date
- 2024-11-30
- Publication Date
- 2026-06-02
AI Technical Summary
In existing technologies, when an external playback device is connected to a smart device via an HDMI interface, the audio received by the smart device from the external playback device becomes echo noise, affecting the accuracy and efficiency of the user's audio recognition and response. Furthermore, the cost and difficulty of adding hardware interfaces or wireless devices are relatively high.
By generating an inaudible audio signal as a reference signal, modulating it onto the ultrasonic band or the near-ultrasonic part of the audible sound band, and transmitting it through an air path, the smart device identifies and removes echo noise to avoid affecting the user's audio listening experience.
Without altering the hardware structure, audio return and echo cancellation are implemented simply and efficiently, improving the accuracy and efficiency of smart devices in recognizing and responding to user audio, while reducing costs.
Smart Images

Figure CN122138116A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of electronic equipment technology, and in particular to an audio processing method and an electronic device. Background Technology
[0002] External playback devices (such as home theater systems and external speakers) can serve as audio playback devices for smart devices (such as smart screens, computers, and projectors) and play the audio required by the smart devices. Currently, most external playback devices and smart devices connect via a high-definition multimedia interface (HDMI), and audio transmission is achieved through the audio return channel (ARC) or enhanced audio return channel (eARC) defined by the HDMI protocol.
[0003] Most smart devices possess features such as voice wake-up and call functionality. To achieve these features when connected to an external playback device, the smart device needs to simultaneously receive and process user audio while the external playback device plays audio. In this scenario, because the external playback device is typically close to the smart device, the audio played by the external playback device will act as an echo, which will be received by the smart device. Therefore, in scenarios where the smart device needs to receive user audio, the echo from the external playback device, received simultaneously with the user audio, becomes noise in the user audio, affecting the accuracy and efficiency of the smart device's recognition and response to the user audio. Therefore, it is necessary to use the audio played by the external playback device as a reference signal to send back to the smart device, so that the smart device can recognize the echo based on the audio returned by the external playback device and perform echo suppression (i.e., removing the audio from the external playback device contained in the audio received by the smart device). However, the HDMI interface between the external playback device and the smart device cannot achieve audio backhaul from the external playback device to the smart device. Therefore, it is necessary to add a new hardware interface or wireless device between the external playback device and the smart device for audio backhaul. This method is costly and difficult to implement. Summary of the Invention
[0004] This application provides an audio processing method and an electronic device for easily and efficiently transmitting a reference signal for audio identification, while reducing costs.
[0005] In a first aspect, embodiments of this application provide an audio processing method applied to a first electronic device. The method includes: acquiring a first audio to be played; wherein the frequency of the first audio belongs to a first frequency band that is audible to the human ear; generating a second audio based on the first audio; wherein the frequency of the second audio belongs to a second frequency band that is not audible to the human ear, and the second audio is used to identify the first audio; and playing the first audio and the second audio.
[0006] In this method, a first electronic device, while playing a first audio, generates and plays a second audio signal, based on the first audio, used to identify the first audio. The first audio is audible to the human ear, while the second audio is inaudible. Therefore, the first electronic device can play the second audio signal, i.e., the reference signal, used to identify the first audio, without interfering with the user's listening to the first audio. On one hand, this method, without changing the hardware structure of the first electronic device, improves the software processing method by using inaudible audio as the reference signal and an air passage as the transmission path for the reference signal, thus easily and efficiently transmitting the reference signal, i.e., the second audio, used to identify the first audio. On the other hand, this method enables other electronic devices in or near the first electronic device to identify and remove the first audio based on the second audio when receiving audio including both the audio played by the first electronic device and the audio they need to receive, thereby preventing the first audio from becoming noise when other electronic devices receive and identify the audio they need. For example, in scenarios where other electronic devices in or around a first electronic device need to receive, respond to, or process user audio in the environment, even if the audio received by the other electronic devices includes the audio played by the first electronic device, the other electronic devices can extract a second audio from the audio and identify the first audio based on the second audio, thereby extracting user audio with no or very low noise interference from the received audio, thus improving the accuracy and efficiency of responding to and processing user audio.
[0007] In one possible design, the upper limit of the first frequency band is lower than or equal to a set frequency, and the lower limit of the second frequency band is higher than the set frequency. Optionally, the set frequency is greater than or equal to 16 kHz and less than or equal to 20 kHz.
[0008] In this method, the first frequency band audible to the human ear and the second frequency band inaudible to the human ear can be preset frequency bands, with the lower limit of the second frequency band higher than the upper limit of the first frequency band. This allows for relatively flexible setting of the frequency band for audio playback by the first electronic device, while ensuring that the playback of the second audio does not affect the playback and listening experience of the first audio. Furthermore, the first frequency band can be an audible sound band or a lower frequency band within the audible sound band, and the second frequency band can be the near-ultrasound portion and / or the ultrasonic band within the audible sound band. This ensures that the reference signal used to identify audible sounds is played through an inaudible audio frequency band, thereby guaranteeing that the user's listening experience to audible sounds is not affected.
[0009] In one possible design, the first audio frequency belongs to the audible frequency band, and the second audio frequency belongs to the ultrasonic frequency band.
[0010] In one possible design, when one of the receivers of the first audio and the second audio is a second electronic device located in the same spatial area as the first electronic device, the second audio is used by the second electronic device to identify and remove the first audio included in the audio collected by the second electronic device. Optionally, the distance between the first electronic device and the second electronic device is less than or equal to a set distance. Optionally, the second electronic device can be a smart device, and the first electronic device can be an external playback device connected to the second electronic device, serving as the audio playback device for the second electronic device. In this scenario, based on the above method, the second electronic device can perform echo suppression based on the second audio in the audio played by the first electronic device, thereby preventing the first audio in the audio played by the first electronic device from becoming an echo of the second electronic device and interfering with the second electronic device's reception, identification, and processing of other audio (e.g., user audio).
[0011] In one possible design, acquiring the first audio to be played includes: receiving the first audio transmitted by a second electronic device via a high-definition multimedia interface (HDMI); or receiving the first audio transmitted by a third electronic device via HDMI; or acquiring the first audio from an application in the first electronic device.
[0012] In this method, the audible audio played by the first electronic device can come from multiple sources. Based on the above method, reference signal return and noise suppression can be easily and efficiently achieved in various audio playback scenarios, thus demonstrating high practicality.
[0013] In one possible design, generating the second audio based on the first audio includes: modulating the first audio onto the second frequency band to obtain the second audio; or, performing a first preprocessing and / or a second preprocessing on the first audio and modulating the obtained audio onto the second frequency band to obtain the second audio; wherein the first preprocessing is used to filter the first audio, and the second preprocessing is used to simulate the changes in the first audio during its transmission from the first electronic device to the second electronic device, thereby processing the first audio.
[0014] Based on this method, the first electronic device can generate the second audio in different ways, offering high flexibility and practicality. Specifically, in the process of generating the second audio from the first audio, performing a first preprocessing step on the first audio can reduce audio bandwidth and data volume while preserving the main and key information, thereby improving playback efficiency. Performing a second preprocessing step on the first audio allows the audio characteristics of the second audio played by the first electronic device to more closely resemble the audio characteristics of the first audio reaching the second electronic device, thus improving the accuracy of the second electronic device in recognizing and removing the first audio based on the second audio.
[0015] In one possible design, the first preprocessing of the first audio includes: filtering the first audio based on a set filtering bandwidth; or filtering the first audio based on a target filtering bandwidth; wherein the target filtering bandwidth is a filtering bandwidth corresponding to a target sampling rate among a plurality of filtering bandwidths, the different filtering bandwidths among the plurality of filtering bandwidths correspond to different sampling rates, and the target sampling rate is the smallest audio sampling rate among the audio sampling rates of the first electronic device and the second electronic device. Optionally, the set filtering bandwidth or the target filtering bandwidth meets the following condition: less than the difference between half of the target sampling rate and the lower limit of the second frequency band. Optionally, before generating the second audio based on the first audio, the method further includes: obtaining the audio sampling rate of the second electronic device from the second electronic device.
[0016] In this method, the first electronic device can obtain the filtering bandwidth in different ways, offering high flexibility. Specifically, by selecting an appropriate filtering bandwidth based on the hardware capabilities of both the first and second electronic devices (i.e., the audio sampling rate), the success rate of transmitting the obtained second audio to the second electronic device can be further improved, as can the accuracy of the second electronic device in restoring or recognizing the first audio based on the second audio, thereby enhancing the accuracy of noise reduction based on the second audio.
[0017] In one possible design, the first audio includes audio from at least one channel; the second preprocessing of the first audio includes: processing the audio of each channel based on the audio transfer function corresponding to each of the at least one channel; wherein the audio transfer function corresponding to each channel is used to indicate the correspondence between the audio of each channel played by the first electronic device and the audio of each channel received by the second electronic device, or it can be understood that the audio transfer function corresponding to each channel is used to indicate the changes in the audio of each channel played by the first electronic device during the transmission to the second electronic device.
[0018] This method processes the audio of each channel based on its corresponding transfer function, simulating the changes that occur in the audio of each channel from playback to arrival at the second electronic device. This ensures that the audio portion corresponding to each channel in the second audio obtained in this way is as close as possible to the audio of that channel received by the second electronic device. This method helps improve the accuracy of the second electronic device in reconstructing or recognizing the first audio based on the second audio, thereby improving the accuracy of noise reduction based on the second audio.
[0019] In one possible design, the method further includes: receiving correction information from the second electronic device; wherein the correction information is used to correct the audio transfer function corresponding to each channel; and correcting the audio transfer function corresponding to each channel according to the correction information.
[0020] Based on this method, the first electronic device can correct the transfer function corresponding to each channel on the first electronic device side based on feedback from the other end, i.e., the second electronic device side, which helps to improve the accuracy of the subsequent generation of reference signals for identifying audible audio.
[0021] In one possible design, playing the first audio and the second audio includes: when the first audio includes multiple audio channels, playing the first audio channel and the second audio channel through a first audio output device among multiple audio output devices; and playing the second audio channel through a second audio output device among multiple audio output devices. Optionally, the audio of each channel among the multiple audio channels is played through one audio output device among the multiple audio output devices, and the audio of different channels among the multiple audio channels is played through different audio output devices among the multiple audio output devices. Optionally, playing the first audio channel and the second audio channel through the first audio output device among the multiple audio output devices includes: mixing the first audio channel and the second audio to obtain a third audio, and playing the third audio through the first audio output device.
[0022] Based on this method, the first electronic device can play the second audio generated by the first electronic device based on the first audio through an existing audio output device in the hardware. Therefore, the playback of the second audio can be achieved easily and efficiently through simple mixing without increasing the additional hardware cost, and the second audio can be transmitted to the second electronic device based on the air channel. Moreover, this solution has high practicality.
[0023] In one possible design, the first audio output device is the audio output device that is closest to the second electronic device among the plurality of audio output devices; or, the first audio output device is positioned facing the second electronic device.
[0024] In the above method, the second audio played by the first electronic device is used by the second electronic device to identify the audible audio played by the first electronic device, i.e., the first audio. Therefore, by using the audio output device that is closest to or facing the second electronic device to play the second audio, the first electronic device can further improve the accuracy of the second audio received by the second electronic device, which helps to improve the accuracy of the second electronic device in subsequent processing based on the second audio.
[0025] Secondly, embodiments of this application provide an audio processing method applied to a second electronic device. The method includes: receiving a first audio; wherein the first audio includes audio from a user and audio played by the first electronic device; dividing the first audio into a second audio and a third audio; wherein the frequency of the second audio belongs to a first frequency band audible to the human ear, and the frequency of the third audio belongs to a second frequency band inaudible to the human ear; identifying and removing a fourth audio included in the second audio based on the third audio to obtain a fifth audio; wherein the third audio and the fourth audio originate from the first electronic device, and the third audio is used to identify the fourth audio. Optionally, the third audio is an audio generated by the first electronic device based on the fourth audio for identifying the fourth audio.
[0026] In this method, when the audio received by the second electronic device includes audio played by the first electronic device, the second electronic device can identify, based on the third audio (a second frequency band inaudible to the human ear) in the audio played by the first electronic device, the fourth audio (a first frequency band audible to the human ear) in the audio played by the first electronic device. This allows the second electronic device to remove the fourth audio from the audible audio received by the second electronic device. Based on this method, interference from the audio played by the first electronic device on the second electronic device's reception, recognition, response, or processing of user audio can be eliminated, thereby improving the accuracy and efficiency of the second electronic device's response and processing of user audio.
[0027] Optionally, the second electronic device and the first electronic device can be located in the same spatial area. Optionally, the distance between the first electronic device and the second electronic device is less than or equal to a set distance. Optionally, the second electronic device can be a smart device, and the first electronic device can be an external playback device connected to the second electronic device, or the first electronic device can serve as an audio playback device for the second electronic device.
[0028] In one possible design, identifying and removing the fourth audio included in the second audio based on the third audio includes: demodulating the third audio, and identifying and removing the fourth audio in the second audio based on the demodulated third audio.
[0029] In this method, the second electronic device demodulates the third audio, which belongs to the second frequency band, and converts it to the first frequency band. Then, based on the converted third audio, it can identify the fourth audio in the second audio that has similar features to the converted third audio, and remove the fourth audio in the second audio to achieve noise reduction.
[0030] In one possible design, the method further includes: analyzing the error between the second audio and the fourth audio based on the second audio and the third audio, and generating correction information based on the error; wherein the correction information is used to correct the audio transfer function corresponding to each channel of the first electronic device, the audio transfer function corresponding to each channel is used to indicate the correspondence between the audio of each channel played by the first electronic device and the audio of each channel received by the second electronic device, and the transfer function corresponding to each channel is used to process the audio of each channel in the fourth audio during the process of the first electronic device generating the second audio based on the fourth audio; and sending the correction information to the first electronic device via a high-definition multimedia interface (HDMI). Optionally, the second electronic device can use an artificial intelligence model or algorithm to extract the fourth audio in the third audio that has similar features to the second audio, analyze the error between the second audio and the fourth audio, and then generate correction information based on the error.
[0031] In the above method, the error between the second and fourth audio frequencies can, to some extent, reflect the deviation between the audio processed by the first electronic device based on the audio transfer function and the audio actually received by the second electronic device. Therefore, the audio transfer function of the first electronic device can be adjusted based on this deviation. Thus, in this method, the second electronic device generates correction information based on this error and sends it to the first electronic device, which helps the first electronic device correct the audio transfer function used in generating the second audio, thereby improving the accuracy of related processing during subsequent audio transmission.
[0032] Thirdly, embodiments of this application provide an audio processing method applied to a system composed of a first electronic device and a second electronic device. The method includes: the first electronic device acquiring a first audio to be played; wherein the frequency of the first audio belongs to a first frequency band audible to the human ear; the first electronic device generating a second audio based on the first audio; wherein the frequency of the second audio belongs to a second frequency band inaudible to the human ear, and the second audio is used to identify the first audio; the first electronic device playing the first audio and the second audio; the second electronic device receiving a third audio; wherein the third audio includes the first audio, the second audio, and user audio from a user; the second electronic device dividing the third audio into the second audio and a fourth audio including the first audio and the user audio; the second electronic device identifying and removing the first audio from the fourth audio based on the second audio to obtain the user audio.
[0033] In one possible design, the upper limit of the first frequency band is lower than or equal to a set frequency, and the lower limit of the second frequency band is higher than the set frequency.
[0034] In one possible design, the set frequency is greater than or equal to 16 kHz and less than or equal to 20 kHz.
[0035] In one possible design, the second electronic device identifies and removes the first audio from the fourth audio based on the second audio, including: the second electronic device demodulates the second audio, and identifies and removes the first audio from the fourth audio based on the demodulated second audio.
[0036] In one possible design, the method further includes: the second electronic device analyzing the error between the second audio and the first audio based on the second audio and the fourth audio, and generating correction information based on the error; wherein the correction information is used to correct the audio transfer function corresponding to each channel of the first electronic device, and the audio transfer function corresponding to each channel is used to indicate the correspondence between the audio played by the first electronic device for each channel and the audio received by the second electronic device for each channel; the second electronic device sends the correction information to the first electronic device via a high-definition multimedia interface (HDMI).
[0037] Other related methods on the first electronic device side in the above system can be referred to the description in the first aspect above, and will not be repeated here.
[0038] Fourthly, this application provides an audio processing system, which includes the first electronic device and the second electronic device described in the third aspect.
[0039] Fifthly, this application provides an electronic device, the electronic device including a memory and one or more processors; wherein the memory is used to store computer program code, the computer program code including computer instructions; when the computer instructions are executed by the one or more processors, the electronic device causes the electronic device to perform the method described in the first aspect or any possible design of the first aspect, or to perform the method described in the second aspect or any possible design of the second aspect.
[0040] In a sixth aspect, this application provides a computer-readable storage medium storing a computer program that, when run on an electronic device, causes the electronic device to perform the method described in the first aspect or any possible design of the first aspect, or to perform the method described in the second aspect or any possible design of the second aspect.
[0041] In a seventh aspect, this application provides a computer program product comprising a computer program or instructions that, when executed on an electronic device, cause the electronic device to perform the method described in the first aspect or any possible design of the first aspect, or to perform the method described in the second aspect or any possible design of the second aspect.
[0042] Eighthly, this application provides a chip system including a processor and a memory, wherein the memory stores instructions; when the instructions are executed by the processor, they implement the methods described in the first aspect or any possible design of the first aspect, or implement the methods described in the second aspect or any possible design of the second aspect. The chip system may be composed of chips or may include chips and other discrete devices.
[0043] The beneficial effects of the third to eighth aspects mentioned above can be found in the beneficial effects of the first or second aspects mentioned above, and will not be repeated here. Attached Figure Description
[0044] Figure 1 This is a schematic diagram of the architecture of an audio transmission system.
[0045] Figure 2 This is a schematic diagram of the architecture of an audio transmission system.
[0046] Figure 3 A schematic diagram of the hardware architecture of an electronic device provided in an embodiment of this application;
[0047] Figure 4 A schematic diagram of the software architecture of an electronic device provided in an embodiment of this application;
[0048] Figure 5 A schematic diagram illustrating an audio processing method provided in an embodiment of this application;
[0049] Figure 6 This application provides a schematic diagram of the architecture of an audio processing system.
[0050] Figure 7 This application provides a schematic diagram of the architecture of an audio processing system.
[0051] Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0052] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the embodiments of this application will be further described in detail below with reference to the accompanying drawings.
[0053] In the description of the embodiments of this application, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Therefore, a feature defined with "first" and "second" may explicitly or implicitly include one or more of that feature.
[0054] It should be understood that in the embodiments of this application, "at least one" means one or more, and "more than one" means two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can represent: a, b, c, a and b, a and c, b and c, or a, b, and c, where a, b, and c can be single or multiple.
[0055] Figure 1 This is a schematic diagram of the architecture of an audio transmission system. Figure 1 As shown, the system includes a smart device and an external playback device, wherein the smart device can play audio through the external playback device. The smart device includes at least an HDMI interface, a microphone, and an analog-to-digital converter module. The external playback device may include an HDMI interface, an audio decoding module, an audio effects processing module, a digital-to-analog converter module, a power amplifier (hereinafter referred to as a power amplifier), and at least one physical channel (e.g., [missing information]). Figure 1 The physical channels 1 to n mentioned above (where n is a positive integer) can specifically be audio playback devices (such as speakers) connected to external playback equipment. Figure 1 The system architecture shown illustrates an audio processing method that may include the following stages:
[0056] Phase 1: The smart device transmits audio data to the external playback device.
[0057] During this stage, the smart device can generate digital audio data based on the audio to be played. This digital audio data can be encoded audio data. The smart device can transmit the digital audio data to an external playback device via an HDMI interface. The external playback device can receive the digital audio data from the smart device via an HDMI interface. The audio corresponding to this digital audio data can be played through at least one channel.
[0058] Phase 2: External playback devices play audio based on digital audio data from smart devices.
[0059] In this stage, digital audio data received via the HDMI interface in the external playback device can be transmitted to the audio decoding module. The audio decoding module decodes the received digital audio data to obtain at least one audio data channel, where each audio data channel corresponds to a physical channel, and each audio data channel can be played through its corresponding physical channel. The at least one audio data channel obtained by the audio decoding module can be transmitted to the sound effects processing module. The sound effects processing module performs sound effects processing on each audio data channel and then transmits the at least one audio data channel to the digital-to-analog converter (DAC) module. The DAC module converts the at least one audio data channel from digital format to analog format and then transmits the at least one audio data channel to the power amplifier. The power amplifier amplifies the at least one audio data channel and then transmits each audio data channel to its corresponding physical channel. Each physical channel can play the received audio data channel. Based on the above method, the external playback device can play audio from smart devices.
[0060] Phase 3: Smart devices receive and process user audio.
[0061] In this stage, the smart device receives user audio (also known as user voice, such as voice wake-up commands or user call voice) and processes it. Specifically, the microphone in the smart device receives ambient audio, which is analog audio. The microphone transmits the received audio to an analog-to-digital converter (ADC). The ADC converts the received audio from analog to digital format. The smart device then performs subsequent processing based on the converted digital audio, such as responding or transmitting data.
[0062] During this stage, since the smart device receives ambient audio, this typically includes audio played by an external playback device near the smart device (essentially an echo of the audio played by the smart device). Because the smart device needs to receive and process user audio, the audio played by the external playback device is considered noise, which can affect the smart device's ability to recognize and process user audio. This reduces the accuracy and efficiency of the smart device's audio recognition and processing, ultimately impacting the user experience.
[0063] To avoid the above problems, the current solution can refer to Figure 2 .like Figure 2As shown, an integrated circuit (IC) with a built-in audio bus (inter-IC sound, I2S) interface for audio return can be added to the hardware structure of both the smart device and the external playback device. The I2S interface in the external playback device is used to transmit audio to the smart device, and the I2S interface in the smart device is used to receive audio from the external playback device. Figure 1 Based on the aforementioned scheme, the external playback device may further include a mixing module. The mixing module can combine at least one audio data stream processed by the audio effects processing module into a single audio data stream, and transmit this audio data stream to the smart device via an I2S interface. This audio data stream can serve as a reference signal (or reference data) for echo removal by the smart device. Figure 1 Based on the illustrated scheme, the smart device can further include an echo suppression module. This module receives audio data from the analog-to-digital converter (ADC) and the I2S interface, and based on the audio data from the I2S interface, removes noise from the audio data from the ADC (i.e., the audio received by the smart device from an external playback device), thereby achieving echo suppression. The smart device can then perform further processing based on the audio processed by the echo suppression module.
[0064] The above Figure 2 The scheme shown is the same as Figure 1 Compared to the proposed solution, while it can improve the accuracy and efficiency of user audio processing, it requires additional hardware interfaces for audio transmission in both smart devices and external playback devices, increasing the complexity and cost of inter-device connections. Furthermore, due to the need to modify the hardware architecture, it may not be feasible to implement in mature hardware architectures.
[0065] To address the above issues and achieve simple and efficient audio feedback without altering the device's hardware structure, this application provides an audio processing method and electronic device. In this method, when playing audio, the first electronic device can modulate the audio to the ultrasonic frequency band or the near-ultrasonic portion of the audible sound frequency band before playback. The audio played by the first electronic device from the second electronic device is low-frequency audio audible to the human ear, while the modulated audio played by the first electronic device is high-frequency audio inaudible to the human ear. Based on this method, the first electronic device can play the audio as high-frequency audio without affecting the user's audio listening experience. In scenarios where the audio played by the first electronic device causes echo interference to the second electronic device, the ambient audio received by the second electronic device includes a mixture of user audio (low-frequency audio), low-frequency audio played by the first electronic device, and high-frequency audio. The second electronic device can separate the high-frequency and low-frequency audio from the mixture, demodulate the high-frequency audio, and then remove the audio played by the first electronic device from the separated low-frequency audio to obtain the user audio. Based on the above method, audio transmission from the first electronic device to the second electronic device can be achieved easily and efficiently without increasing the hardware structure of the first and second electronic devices, and echo suppression can be achieved easily and efficiently on the second electronic device side, thereby better processing user voice and improving user experience.
[0066] The technical solutions provided in this application can be executed by multiple electronic devices with audio transmission, reception, and processing capabilities. In some embodiments of this application, the electronic devices can be portable devices, such as mobile phones, tablets, wearable devices with wireless communication functions (e.g., watches, bracelets, etc.), in-vehicle terminal devices, augmented reality (AR) / virtual reality (VR) devices, laptops, ultra-mobile personal computers (UMPCs), netbooks, personal digital assistants (PDAs), smart home devices (e.g., smart TVs, smart speakers, etc.), smart robots, workshop equipment, wireless terminals in self-driving vehicles, wireless terminals in remote medical surgery, wireless terminals in smart grids, wireless terminals in transportation safety, wireless terminals in smart cities, or wireless terminals in smart homes, and flying devices (e.g., smart robots, drones, airplanes), etc. Among them, wearable devices are portable devices that users can wear directly on their bodies or integrate into their clothing or accessories.
[0067] In some embodiments of this application, the electronic device may also be a portable terminal device that includes other functions such as a personal digital assistant and / or a music player. Exemplary embodiments of the portable terminal device include, but are not limited to, devices equipped with... Alternatively, it could be a portable terminal device with another operating system. The aforementioned portable terminal device could also be other portable terminal devices, such as a laptop computer with a touch-sensitive surface (e.g., a touch panel). It should also be understood that in some other embodiments of this application, the aforementioned electronic device may not be a portable terminal device, but rather a desktop computer with a touch-sensitive surface (e.g., a touch panel).
[0068] For a description of the performance of electronic devices, please refer to the relevant descriptions below.
[0069] See below. Figure 3 The structure of the electronic device to which the method provided in the embodiments of this application is applicable will be described.
[0070] like Figure 3As shown, the electronic device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a USB interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone jack 170D, a sensor module 180, buttons 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a SIM card interface 195, etc.
[0071] The sensor module 180 may include a gyroscope sensor, an accelerometer, a proximity sensor, a fingerprint sensor, a touch sensor, a temperature sensor, a pressure sensor, a distance sensor, a magnetic sensor, an ambient light sensor, a barometric pressure sensor, a bone conduction sensor, etc.
[0072] Understandable, Figure 3 The electronic device 100 shown is merely an example and does not constitute a limitation on the electronic device. The electronic device may have more or fewer components than those shown in the figure, may combine two or more components, or may have different component configurations. Figure 3 The various components shown can be implemented in hardware, software, or a combination of hardware and software, including one or more signal processing and / or application-specific integrated circuits.
[0073] Processor 110 may include one or more processing units, such as application processors (APs), modem processors, graphics processing units (GPUs), image signal processors (ISPs), controllers, memory, video codecs, digital signal processors (DSPs), baseband processors, and / or neural network processing units (NPUs). Different processing units may be independent devices or integrated into one or more processors. The controller may serve as the central nervous system and command center of the electronic device 100. The controller can generate operation control signals based on instruction opcodes and timing signals to control instruction fetching and execution.
[0074] The processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory can store instructions or data that the processor 110 has just used or that are used repeatedly. If the processor 110 needs to use the instruction or data again, it can retrieve it directly from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.
[0075] The audio processing method provided in this application embodiment can be executed by the processor 110 controlling or calling other components. For example, it can call the processing program of this application embodiment stored in the internal memory 121, or call the processing program of this application embodiment stored in a third-party device through the external memory interface 120 to control the wireless communication module 160 to perform data communication with other devices, thereby improving the intelligence and convenience of the electronic device 100 and enhancing the user experience. The processor 110 may include different devices. For example, when integrating a CPU and a GPU, the CPU and GPU can cooperate to execute the audio processing method provided in this application embodiment. For example, some algorithms in the audio processing method are executed by the CPU, and other algorithms are executed by the GPU to obtain faster processing efficiency.
[0076] Display screen 194 is used to display images, videos, etc. Display screen 194 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a Miniled LED, a MicroLED, a Micro-OLED, a quantum dot light-emitting diode (QLED), etc. In some embodiments, electronic device 100 may include one or N displays screens 194, where N is a positive integer greater than 1. Display screen 194 can be used to display information input by the user or information provided to the user, as well as various graphical user interfaces (GUIs). For example, display screen 194 can display photos, videos, web pages, or documents, etc.
[0077] In this embodiment of the application, the display screen 194 can be a single flexible display screen, or it can be a splicing display screen composed of two rigid screens and a flexible screen located between the two rigid screens.
[0078] Camera 193 (a front-facing camera or a rear-facing camera, or a single camera that can function as both) is used to capture still images or videos. Typically, camera 193 may include a photosensitive element such as a lens assembly and an image sensor. The lens assembly includes multiple lenses (convex or concave lenses) for collecting light signals reflected from the object being photographed and transmitting these signals to the image sensor. The image sensor generates a raw image of the object being photographed based on the light signals.
[0079] The internal memory 121 can be used to store computer executable program code, which includes instructions. The processor 110 executes various functional applications and data processing of the electronic device 100 by running the instructions stored in the internal memory 121. The internal memory 121 may include a program storage area and a data storage area. The program storage area can store the operating system, application program code (such as the functions corresponding to the scheme in this application), etc. The data storage area can store data created during the use of the electronic device 100.
[0080] The internal memory 121 may also store one or more computer programs corresponding to the algorithm of this application. The one or more computer programs are stored in the internal memory 121 and configured to be executed by one or more processors 110. The one or more computer programs include instructions that can be used to perform the various steps in the following embodiments.
[0081] In addition, the internal memory 121 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, universal flash storage (UFS), etc.
[0082] Of course, the algorithm code of the embodiment of this application can also be stored in external memory. In this case, the processor 110 can run the algorithm code of the embodiment of this application stored in external memory through the external memory interface 120.
[0083] A touch sensor, also known as a "touch panel," can be located on the display screen 194. The touch sensor and display screen 194 together form a touch display screen, also called a "touch screen." The touch sensor detects touch operations applied to or near it. The touch sensor can transmit the detected touch operation to the application processor to determine the type of touch event. Visual output related to the touch operation can be provided through the display screen 194. In other embodiments, the touch sensor may also be located on the surface of the electronic device 100, in a different position than the display screen 194.
[0084] For example, the display screen 194 of the electronic device 100 can display a main interface, which includes icons for multiple applications (such as a camera application, a fitness and health application, etc.). For instance, a user can tap the camera application icon on the main interface using a touch sensor, triggering the processor 110 to launch the camera application and open the camera 193. The display screen 194 then displays the camera application's interface, such as the viewfinder.
[0085] The wireless communication function of electronic device 100 can be realized through antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, modem processor and baseband processor, etc.
[0086] Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in electronic device 100 can be used to cover one or more communication frequency bands. Different antennas can also be multiplexed to improve antenna utilization. For example, antenna 1 can be multiplexed as a diversity antenna for a wireless local area network. In some other embodiments, the antennas can be used in conjunction with tuning switches.
[0087] The mobile communication module 150 can provide solutions for wireless communication, including 2G / 3G / 4G / 5G, applied to the electronic device 100. The mobile communication module 150 may include at least one filter, switch, power amplifier, low noise amplifier (LNA), etc. The mobile communication module 150 can receive electromagnetic waves via antenna 1, and perform filtering, amplification, and other processing on the received electromagnetic waves before transmitting them to a modem processor for demodulation. The mobile communication module 150 can also amplify the signal modulated by the modem processor and convert it into electromagnetic waves for radiation via antenna 1. In some embodiments, at least some functional modules of the mobile communication module 150 may be housed in the processor 110. In some embodiments, at least some functional modules of the mobile communication module 150 and at least some modules of the processor 110 may be housed in the same device. In this embodiment, the mobile communication module 150 can also be used for information interaction with other devices.
[0088] The modem processor may include a modulator and a demodulator. The modulator modulates the low-frequency baseband signal to be transmitted into a mid-to-high frequency signal. The demodulator demodulates the received electromagnetic wave signal into a low-frequency baseband signal. The demodulator then transmits the demodulated low-frequency baseband signal to the baseband processor for processing. After processing by the baseband processor, the low-frequency baseband signal is transmitted to the application processor. The application processor outputs sound signals through an audio device (not limited to speaker 170A, receiver 170B, etc.) or displays images or videos through the display screen 194. In some embodiments, the modem processor may be a separate device. In other embodiments, the modem processor may be independent of the processor 110 and may be housed in the same device as the mobile communication module 150 or other functional modules.
[0089] The wireless communication module 160 can provide solutions for wireless communication applications on the electronic device 100, including wireless local area networks (WLANs) (such as wireless fidelity (WiFi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), and infrared (IR) technologies. The wireless communication module 160 can be one or more devices integrating at least one communication processing module. The wireless communication module 160 receives electromagnetic waves via antenna 2, performs frequency modulation and filtering of the electromagnetic wave signals, and sends the processed signal to processor 110. The wireless communication module 160 can also receive signals to be transmitted from processor 110, perform frequency modulation and amplification, and convert them into electromagnetic waves for radiation via antenna 2. In this embodiment, the wireless communication module 160 can be used to establish connections with other electronic devices for data interaction. Alternatively, the wireless communication module 160 can be used to access access point devices, send control commands to other electronic devices, or receive data from other electronic devices.
[0090] In addition, the electronic device 100 can implement audio functions through an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone jack 170D, and an application processor, such as music playback and recording. The electronic device 100 can receive input from buttons 190, generating key signal inputs related to user settings and function control. The electronic device 100 can use a motor 191 to generate vibration alerts (such as vibration alerts for incoming calls). The indicator 192 in the electronic device 100 can be an indicator light, used to indicate charging status, battery level changes, messages, missed calls, notifications, etc. The SIM card interface 195 in the electronic device 100 is used to connect a SIM card. The SIM card can be inserted into or removed from the SIM card interface 195 to achieve contact and separation with the electronic device 100.
[0091] It should be understood that, in practical applications, electronic device 100 may include more than Figure 3 The number of more or fewer components shown is not limited in the embodiments of this application. The illustrated electronic device 100 is merely an example, and the electronic device 100 may have more or fewer components than shown in the figure, may combine two or more components, or may have different component configurations. The various components shown in the figure may be implemented in hardware, software, or a combination of hardware and software, including one or more signal processing and / or application-specific integrated circuits.
[0092] The software system of electronic device 100 can adopt a layered architecture, event-driven architecture, microkernel architecture, microservice architecture, or cloud architecture. A layered architecture divides the software into several layers, each with a clear role and function. Layers communicate with each other through software interfaces. For example, ... Figure 4 As shown, the software architecture can be divided into four layers, from top to bottom: the application layer, the application framework layer (framework, FWK), the runtime and system libraries, and the kernel layer.
[0093] The application layer is the top layer of the operating system and includes native operating system applications such as camera, gallery, calendar, Bluetooth, music, video, and messaging applications, as well as third-party applications. The applications discussed in this application embodiment are referred to as applications (APPs), which are software programs capable of performing one or more specific functions. Typically, multiple applications can be installed on an electronic device, such as camera applications and email applications. The applications mentioned below can be system applications pre-installed on the electronic device at the factory, or third-party applications downloaded by the user from the network or obtained from other electronic devices during the use of the electronic device.
[0094] Of course, for developers, they can write applications and install them into this layer. In one possible implementation, the application can be developed using the Java language, by calling the application programming interface (API) provided by the application framework layer. Developers can then interact with the underlying operating system (such as the kernel layer) through the application framework to develop their own applications.
[0095] The application framework layer provides the API and programming framework for the application layer. It can include predefined functions. The application framework layer may include a window manager, content provider, view system, phone manager, resource manager, notification manager, etc.
[0096] The window manager is used to manage window programs. The window manager can obtain the screen size, determine if a status bar is present, lock the screen (or display), and capture the screen, among other things.
[0097] Content providers are used to store and retrieve data, and make that data accessible to applications. This data may include files (e.g., documents, videos, images, audio), text, and other information.
[0098] A view system includes visual controls, such as controls that display text, images, documents, and other content. View systems can be used to build applications. An interface in a display window can consist of one or more views. For example, a display interface including a text message notification icon could include a view that displays text and a view that displays images.
[0099] The phone manager provides communication functionality for electronic devices. The notification manager allows applications to display notification information in the status bar; it can be used to convey informative messages and can disappear automatically after a short pause without user interaction.
[0100] The runtime includes the core libraries and the virtual machine. The runtime is responsible for system scheduling and management.
[0101] The system's core library consists of two parts: one part contains the functionalities that the Java language needs to call, and the other part is the system's core library. The application layer and application framework layer run in a virtual machine. Taking Java as an example, the virtual machine executes the Java files in the application layer and application framework layer as binary files. The virtual machine is used to perform functions such as object lifecycle management, stack management, thread management, security and exception management, and garbage collection.
[0102] The system library can include multiple functional modules. Examples include: a surface manager, a media library, a 3D graphics processing library (e.g., OpenGL ES), a 2D graphics engine (e.g., SGL), and an image processing library. The surface manager manages the display subsystem and provides fusion of 2D and 3D layers for multiple applications. The media library supports playback and recording of various common audio and video formats, as well as still image files. It supports multiple audio and video encoding formats, such as MPEG4, H.564, MP3, AAC, and AMR. The 3D graphics processing library implements 3D graphics drawing, image rendering, compositing, and layer processing. The 2D graphics engine is the drawing engine for 2D graphics.
[0103] The kernel layer provides core system services for the operating system, such as security, memory management, process management, network protocol stack, and driver models, all of which are implemented based on the kernel layer. The kernel layer also serves as an abstraction layer between the hardware and software stacks. This layer contains many drivers related to electronic devices, including: display drivers; keyboard drivers as input devices; Flash drivers for memory-based devices; camera drivers; audio drivers; Bluetooth drivers; and WiFi drivers.
[0104] It is important to understand that the functional services described above are just an example. In practical applications, electronic devices may be divided into more or fewer functional services based on other factors, or the functions of each service may be divided in other ways, or they may not be divided into functional services but work as a whole.
[0105] The solution provided in this application can be applied to a system composed of a first electronic device and a second electronic device. The first and second electronic devices can be located in the same spatial environment, and / or the distance between them can be less than or equal to a set distance. The first electronic device can play audible audio (low-frequency audio) and audio modulated to a high-frequency band (inaudible to the human ear). The second electronic device can receive ambient audio and identify and respond to user audio contained within it. The ambient audio received by the second electronic device can include both types of audio played by the first electronic device. The second electronic device can divide the ambient audio into low-frequency and high-frequency audio, identify the audible audio played by the first electronic device based on the high-frequency audio, remove the audible audio played by the second electronic device from the low-frequency audio, and retain the user audio. This results in noise-removed user audio, which is then further processed.
[0106] In the above scheme, the first electronic device is a device for playing audio, and the second electronic device is a device for receiving audio. Therefore, the first electronic device can also be called a transmitting device, and the second electronic device can also be called a receiving device.
[0107] The above solution will be described in detail below with reference to specific embodiments.
[0108] Example 1: Audio Return and Echo Suppression
[0109] Reference Figure 5 An audio processing method provided in this application embodiment may include:
[0110] S501: The first electronic device generates a second audio based on the first audio to be played; wherein the frequency of the first audio belongs to a first frequency band that is audible to the human ear, the frequency of the second audio belongs to a second frequency band that is not audible to the human ear, and the second audio is used to identify the first audio.
[0111] The audio described in this embodiment can also be understood as audio signal or audio data. The frequency band described in this embodiment can be used to represent the frequency distribution range of audio or audio signal. Frequency refers to the number of vibrations per second of a sound wave, and the unit of frequency is Hertz (Hz). Therefore, a frequency band can also be understood as a frequency range or frequency interval. Based on the frequency, sound can be divided into different frequency bands, such as the audible sound band and the ultrasonic sound band. Audible sound is defined as sound that can be heard by the human ear. The audible sound band range (i.e., frequency range) is 20–20 kHz, and the ultrasonic sound band range (i.e., frequency range) is above 20 kHz.
[0112] In some embodiments of this application, the first frequency band can be an audible sound band, that is, the first audio frequency can belong to the 20-20kHz frequency band. The second frequency band can be an ultrasonic frequency band, that is, the second audio frequency can belong to the frequency band above 20kHz.
[0113] In some embodiments of this application, in practical applications, there are also some frequency bands within the audible sound bands that are practically inaudible or very difficult to hear by the human ear. These frequency bands can also be used as the frequency band to which the second audio signal belongs. Therefore, the first frequency band can also be a frequency band within the audible sound band whose upper limit frequency is lower than or equal to a set frequency, and the second frequency band can be a frequency band whose lower limit frequency is higher than a set frequency. Optionally, the set frequency can be greater than or equal to 16kHz and less than or equal to 20kHz.
[0114] Based on this method, in one example, the first audio frequency may belong to a third frequency band defined within the audible sound frequency band, and the second audio frequency may be the audio obtained by modulating the first audio frequency to a fourth frequency band and / or an ultrasonic frequency band defined within the audible sound frequency band, wherein the lower limit of the fourth frequency band is higher than or equal to the upper limit of the third frequency band. Modulating the first audio frequency to the fourth frequency band and ultrasonic frequency band defined within the audible sound frequency band refers to modulating the first audio frequency to the frequency band composed of the fourth frequency band and the ultrasonic frequency band. For example, the third frequency band may be 20–12 kHz, and the fourth frequency band may be 12 kHz–20 kHz. Another example is that the third frequency band may be 20–16 kHz, and the fourth frequency band may be 16 kHz–20 kHz. Yet another example is that the third frequency band may be 20–15 kHz, and the fourth frequency band may be 16 kHz–20 kHz.
[0115] Optionally, the third frequency band set in the audible sound band can also be called the low frequency band set in the audible sound band. The fourth frequency band set in the audible sound band can also be called the high frequency band set in the audible sound band, or the near-ultrasound band, or the near-ultrasound portion of the audible sound band.
[0116] In this method, the first audio is audible to the human ear, and the second audio is inaudible to the human ear. This method allows a first electronic device to play the second audio used to identify the first audio without interfering with the user's listening to the first audio. This enables other electronic devices, while receiving the user's audio, to simultaneously receive the first and second audio, and to identify and remove the first audio from the audible audio based on the second audio, thus preserving the user's audio and preventing the first audio from becoming noise when other electronic devices receive and identify the user's audio.
[0117] The following is a detailed explanation of the methods for the first electronic device to acquire the first audio.
[0118] In some embodiments of this application, the source of the first audio can be any of the following:
[0119] 1) The first electronic device itself.
[0120] In this approach, the first audio file can be audio generated by the first electronic device itself and intended for playback. For example, the first audio file can be audio generated by a first application within the first electronic device that needs to be played. For instance, the first electronic device can be a smart device such as a smart screen.
[0121] 2) Second electronic device.
[0122] In this method, the first audio source can be a second electronic device, which can act as an audio playback device for the second electronic device and play the first audio from the second electronic device. The first and second electronic devices can be connected via HDMI (or other similar methods), and the second electronic device can send the first audio through the HDMI interface, while the first electronic device can receive the first audio through the HDMI interface.
[0123] For example, in this scenario, the first electronic device can be an external playback device such as a home theater playback system, speakers, soundbars, or smart speakers, and the second electronic device can be a smart device such as a mobile phone, tablet, computer, smart screen, smart TV, or projector.
[0124] In one example scenario, the second electronic device can play audio through the first electronic device and also interact with the user via voice (e.g., voice wake-up, voice command response). In another example scenario, the second electronic device can play audio through the first electronic device and also conduct voice or video calls with other electronic devices. For example, audio provided by an application window (or application interface) of the second electronic device can be played through the first electronic device, and the user's voice received by the second electronic device can be used to participate in a call corresponding to another application window (or call interface) used for communication.
[0125] 3) Third electronic devices.
[0126] In other embodiments of this application, the first audio source may be a third electronic device, which may be used as an audio playback device for the third electronic device to play the first audio source. Optionally, the first electronic device and the second electronic device may be connected via HDMI, near-field communication (e.g., Bluetooth), wireless connection, or other means.
[0127] For example, in this scenario, the first electronic device can be an external playback device such as a smart speaker or Bluetooth speaker, and the second and third electronic devices can be smart devices such as mobile phones, tablets, computers, smart screens, smart TVs, and projectors.
[0128] In the above method, when the first audio comes from other electronic devices (such as the second or third electronic device mentioned above), the first electronic device can obtain audio in an encoded format from the other electronic device. The first electronic device can obtain the first audio by decoding the audio from the other electronic device. For example, the encoded audio format could be...
[0129] The following provides a detailed description of the method by which the first electronic device generates a second audio based on a first audio and performs playback-related processing on the first and second audio.
[0130] In this embodiment of the application, the first electronic device can generate the second audio based on the first audio in any of the following ways:
[0131] 1) The first electronic device can modulate the first audio to the second frequency band to obtain the second audio.
[0132] 2) The first electronic device can preprocess the first audio and then modulate the resulting audio onto the second frequency band to obtain the second audio. The preprocessing may include decoding, sound effects processing, etc.
[0133] In this embodiment, the first audio may include audio from one or more channels. Each channel of audio can be played through an audio output device. In scenarios where the first audio includes one channel, the first electronic device can play both the first and second audio through the same audio output device, or through different audio output devices. In scenarios where the first audio includes multiple channels, each channel can be played through one of the multiple audio output devices of the first electronic device, and different channels can be played through different audio output devices. The second audio can be played through one of the multiple audio output devices. For example, the first electronic device can play both the first and second channels of audio from multiple channels through the first audio output device, and play the second channel audio from multiple channels through the second audio output device. The playback method for other channels is similar to that for the second channel. The first electronic device can mix the first and second channels of audio to obtain a third audio, and play the third audio through the first audio output device.
[0134] The audio output device (also referred to as an audio playback device or audio playback device, such as a speaker) mentioned in the above method can be an audio output device built into the first electronic device, or it can be an audio output device connected to the first electronic device in a wired or wireless manner.
[0135] In this embodiment, each audio output device can be used to play audio from one channel of the first audio; therefore, each audio output device can also be understood as a physical channel. For ease of description, the audio output device can be simply referred to as a physical channel in the following embodiments.
[0136] The above methods will be explained in detail below with specific application scenarios.
[0137] In some embodiments of this application, initially, the first audio generated by the first electronic device itself or obtained from other electronic devices can be digital audio (also referred to as digital audio data or digital audio signal). The first audio can include at least one audio channel, and each audio channel can be played through a physical channel. When the first audio comes from other electronic devices, the first electronic device can obtain at least one audio channel included in the first audio by performing audio decoding. Specifically, the physical channel of the first electronic device can be an audio playback device (such as a speaker) or an audio playback apparatus of the first electronic device.
[0138] In the first possible method, the first electronic device can modulate the acquired digital-formatted first audio signal onto the aforementioned ultrasonic frequency band using digital modulation methods such as amplitude modulation (AM) and frequency modulation (FM) in the digital domain, thereby obtaining ultrasonic audio, i.e., the second audio signal. When playing the audio, for each of at least one audio signal, the first electronic device can perform analog-to-digital conversion, power amplification, and other processing on that audio signal before playing it through the corresponding physical channel. For the second audio signal, the first electronic device can perform analog-to-digital conversion, power amplification, and other processing on the second audio signal before playing it through a separate physical channel.
[0139] In the second possible method, the first electronic device can modulate the acquired digital-formatted first audio signal onto the aforementioned ultrasonic frequency band using digital modulation methods such as amplitude modulation (AM) and frequency modulation (FM) in the digital domain, thereby obtaining ultrasonic audio, i.e., the second audio signal. When playing the audio, for the second audio signal, the first electronic device can mix the second audio signal with at least one other audio signal, perform analog-to-digital conversion and power amplification on the resulting mixed audio signal, and then play the mixed audio signal through the physical channel corresponding to that audio signal. For each of the at least one audio signal other than the one mixed with the second audio signal, the first electronic device can perform analog-to-digital conversion and power amplification on that audio signal before playing it through the corresponding physical channel.
[0140] The first electronic device can select any one audio source from at least one audio source and mix it with the second audio source to obtain a mixed audio, or it can select one audio source from at least one audio source and mix it with the second audio source in any of the following ways to obtain a mixed audio:
[0141] 1) Select at least one physical channel that is closest to the second electronic device and mix it with the second audio.
[0142] In this embodiment of the application, the audio channel corresponding to the physical channel refers to the audio channel played through the physical channel.
[0143] Optionally, the physical channel closest to the second electronic device among at least one physical channel can be pre-configured in the first electronic device. For example, a user can pre-store indication information for indicating the physical channel closest to the second electronic device among at least one physical channel in the first electronic device, and the first electronic device can determine the physical channel closest to the second electronic device among at least one physical channel based on the indication information.
[0144] Optionally, the physical channel closest to the second electronic device among at least one physical channel can also be determined by the first electronic device. For example, the first electronic device can use ultrasonic detection technology to detect the relative positional relationship between the first electronic device and the second electronic device, and combine this with the distribution position of at least one physical channel to determine the physical channel closest to the second electronic device among at least one physical channel.
[0145] 2) Select at least one physical channel corresponding to the physical channel facing the second electronic device and mix it with the second audio.
[0146] Optionally, at least one physical channel facing the second electronic device may be pre-configured in the first electronic device.
[0147] In both methods described above, the power amplification processing performed by the first electronic device on different audio paths in at least one audio path can be the same, and the power amplification processing performed by the first electronic device on at least one audio path and the second audio path can be the same. Using these methods, the first electronic device can generate a second audio path based on the acquired first audio path, and after processing the first and second audio paths, play the first and second audio paths through at least one physical channel.
[0148] Optionally, in either of the above two methods, the first electronic device may further perform specific sound effect processing on the first audio, then generate a second audio based on the first audio, and play the first and second audio. For example, before performing analog-to-digital conversion or other processing on the acquired first audio, and before generating the second audio based on the first audio, the first electronic device may further perform specific sound effect processing on each of at least one audio path, mix the at least one audio path after sound effect processing, and then modulate the mixed audio onto the aforementioned ultrasonic frequency band to obtain the second audio.
[0149] Optionally, the preprocessing of the first audio by the first electronic device may further include a first preprocessing and / or a second preprocessing, wherein the first preprocessing is used to filter the first audio, and the second preprocessing is used to simulate the changes in the first audio during its transmission from the first electronic device to the second electronic device, and to process the first audio. Regarding the method for the first preprocessing of the first audio by the first electronic device, refer to the method described in Embodiment 2 below, which will not be detailed here. Regarding the method for the second preprocessing of the first audio by the first electronic device, refer to the method described in Embodiment 4 below, which will not be detailed here.
[0150] S502: The first electronic device plays the first audio and the second audio.
[0151] S503: The second electronic device receives a third audio; wherein the third audio includes a first audio, a second audio, and a user audio.
[0152] The third audio can be audio received by the second electronic device from the environment, such as audio received (or acquired) by the audio output device of the second electronic device.
[0153] In some embodiments of this application, since the first electronic device and the second electronic device are in the same spatial environment or are close to each other, in scenarios where the second electronic device needs to receive, process, or respond to user audio, the audio played by the first electronic device becomes ambient noise. When the audio played by the first electronic device originates from the second electronic device, the audio played by the first electronic device can also be understood as an echo. Therefore, when the second electronic device needs to collect user audio, it will actually collect a mixed audio (i.e., ambient audio) that includes both the user audio and the audio played by the first electronic device (i.e., the first and second audio mentioned above). The audible sound played by the first electronic device in this mixed audio will affect the accuracy and efficiency of the second electronic device in recognizing user audio. Therefore, noise reduction (or denoising) is performed to obtain the user audio, thereby ensuring more accurate and efficient recognition of user audio.
[0154] Optionally, the second electronic device may receive third audio via an audio receiving device such as a microphone.
[0155] S504: The second electronic device divides the third audio into a second audio and a fourth audio that includes the first audio and the user audio.
[0156] The third audio received by the second electronic device can be analog audio (or analog audio data or analog audio signal). The second electronic device can first perform analog-to-digital conversion on the third audio, and then filter the converted third audio to filter out the second audio in the ultrasonic frequency band and the fourth audio in the audible frequency band. At this time, the second and fourth audio are digital audio.
[0157] Optionally, in the above method, the second electronic device can filter the third audio signal using a bandpass filter to obtain the second audio signal. The bandpass filter allows the signal to pass through the ultrasonic frequency band. The second electronic device can also filter the third audio signal using a low-pass filter to obtain the fourth audio signal. The low-pass filter allows the signal to pass through the audible frequency band.
[0158] S505: The second electronic device demodulates the second audio and, based on the demodulated second audio, performs noise reduction processing on the fourth audio to obtain the fifth audio.
[0159] In this method, the second electronic device can demodulate the second audio in the digital domain. The demodulated second audio can be used as audio in the audible frequency band played by the first electronic device, or it can be understood as a reference signal for echo suppression. Noise reduction processing is used to identify and remove the first audio included in the fourth audio. The second electronic device can identify the first audio in the fourth audio based on the demodulated second audio, and remove the first audio from the fourth audio to obtain the fifth audio. The second electronic device can use the fifth audio as the received user audio and perform subsequent processing based on the fifth audio.
[0160] In the above method, the noise reduction process is used to remove the audio in the fourth audio that is played by the first electronic device and is an echo for the second electronic device. Therefore, the noise reduction process can also be understood as echo suppression.
[0161] In the above method, when the first electronic device plays audible audio, it can simultaneously play ultrasonic audio used to identify the audio without interfering with the user's normal listening to the audible audio played by the first electronic device. The second electronic device can identify the audible audio played by the first electronic device based on the ultrasonic audio played by the first electronic device, and then remove the audible audio played by the first electronic device from the received ambient audio, retaining the user's audio. Therefore, the above method can minimize the interference of the audio played by the first electronic device on the second electronic device's reception and processing of the user's audio, ensuring that the second electronic device accurately and efficiently identifies the user's audio. In this method, the processes of audio processing and transmission by the first electronic device and audio reception and processing by the second electronic device can both be completed using the hardware components of the devices themselves, requiring only limited software modifications and optimizations. Therefore, the implementation cost is very low, the efficiency is very high, and it can support the implementation of the above functions in devices with already defined hardware. Therefore, the above method has high feasibility and practicality.
[0162] It should be understood that the above implementation process is merely an illustrative example of the method flow applicable to the embodiments of this application. The execution order of each step can be adjusted according to actual needs, and other steps can be added or some steps can be removed. The execution order between steps that are not temporally related can be arbitrary, and this embodiment of the application does not impose any restrictions on this.
[0163] Example 2: Audio Filtering Processing
[0164] In one possible solution, in the solution provided in Embodiment 1 above, the first electronic device and the second electronic device can be devices whose hardware capabilities meet the following conditions: the supported sampling rate is greater than or equal to twice the upper limit of the frequency band to which the first audio belongs.
[0165] In this embodiment, the sampling rate can also be called the sampling frequency. The sampling rate defines the sampling format for extracting and assembling a discrete signal from a continuous signal per unit time. The sampling rate can be represented in Hertz (Hz). In this embodiment, the sampling rate refers to the sampling rate used to sample an audio signal.
[0166] In another possible solution, as provided in Embodiment 1 above, during the process of generating the second audio based on the first audio, the first electronic device may perform filtering processing (i.e., the first preprocessing described in Embodiment 1) on the audio obtained in the previous processing (such as the first audio described in Embodiment 1, or the audio obtained by mixing at least one audio channel after sound effect processing as described in Embodiment 1) before modulating it to the ultrasonic frequency band. This reduces the audio bandwidth and improves the success rate and efficiency of subsequent playback and reception of the second audio. The method is described in detail below.
[0167] In this scheme, the filtering parameters when the first electronic device filters the audio obtained from the previous processing step may include a filtering bandwidth. The first electronic device can perform the filtering process based on the filtering bandwidth. This filtering bandwidth can also be understood as the bandwidth of the signal obtained after filtering the audio obtained from the previous processing step by the first electronic device. Bandwidth refers to the frequency range occupied by the audio signal, i.e., the width of the frequency range occupied by the audio signal. In some embodiments of this application, the filtering bandwidth can be a preset value. Alternatively, the first electronic device can determine the filtering bandwidth based on the hardware capabilities of both the first and second electronic devices. The hardware capability of any electronic device refers to the sampling rate it supports when sampling the audio signal, i.e., the audio sampling rate. In some embodiments of this application, the size of the filtering bandwidth is positively correlated with the size of the reference sampling rate, which is the minimum sampling rate supported by the first and second electronic devices. For example, when the sampling rate of the first electronic device is 48kHz and the sampling rate of the second electronic device is 96kHz, the reference sampling rate can be 48kHz. When the sampling rates of both the first and second electronic devices are 48kHz, the reference sampling rate can also be 48kHz. The first electronic device can obtain the sampling rate supported by the second electronic device.
[0168] As an optional implementation, after determining the reference sampling rate based on the hardware capabilities of the first electronic device and the second electronic device, the first electronic device can select the filtering bandwidth corresponding to the reference sampling rate as the filtering bandwidth used (i.e., the filtering bandwidth used when performing the above filtering process) from a set correspondence between sampling rates and filtering bandwidths. Different filtering bandwidths correspond to different sampling rates. In one example, the set correspondence between sampling rates and filtering bandwidths can have any of the following conditions:
[0169] 1) A sampling rate value can correspond to a filter bandwidth value.
[0170] In this case, the first electronic device can select a filter bandwidth corresponding to the reference sampling rate from a plurality of candidate filter bandwidth values.
[0171] For example, the correspondence between the set sampling rate and the filtering bandwidth can include: a sampling rate of 48kHz corresponds to a filtering bandwidth of 4kHz, a sampling rate of 72kHz corresponds to a filtering bandwidth of 16kHz, etc. In this scenario, when the reference sampling rate is 48kHz, the filtering bandwidth selected by the first electronic device is 4kHz.
[0172] 2) A sampling rate range (or sampling rate value interval) can correspond to a filter bandwidth value.
[0173] In this case, the first electronic device can select a filter bandwidth corresponding to the sampling rate range to which the reference sampling rate belongs from a plurality of candidate filter bandwidth values.
[0174] 3) A sampling rate value can correspond to a filtering bandwidth range (or a range of filtering bandwidth values).
[0175] In this case, the first electronic device can select a filter bandwidth value within a corresponding filter bandwidth range based on the value of the reference sampling rate.
[0176] As another optional implementation, after determining a reference sampling rate based on the hardware capabilities of the first electronic device and the second electronic device, the first electronic device may select one of a plurality of candidate filter bandwidths that meets set conditions as the filter bandwidth to be used, based on the reference sampling rate; or, based on the reference sampling rate, it may set one filter bandwidth that meets set conditions as the filter bandwidth to be used. The set conditions may include any of the following:
[0177] 1) The filter bandwidth is greater than zero or the set value, and less than half of the reference sampling rate and the upper limit of the frequency band to which the first audio belongs.
[0178] 2) The filter bandwidth is greater than zero or the set value, and less than half of the reference sampling rate and the lower limit of the frequency band to which the second audio belongs.
[0179] Optionally, since a higher filtering bandwidth results in a higher degree of signal restoration, the accuracy of noise reduction or echo suppression based on the restored signal is improved. Therefore, when there are multiple filtering bandwidths that meet the set conditions, the first electronic device can use the largest filtering bandwidth that meets the set conditions as the filtering bandwidth to be used.
[0180] For example, in a scenario where the reference sampling rate is 48kHz and the first audio frequency belongs to the audible frequency band (with an upper limit of 20kHz), the difference between half the reference sampling rate and the upper limit of the first audio frequency band is 4kHz. Therefore, the filtering bandwidth used by the first electronic device is greater than 0 and less than or equal to 4kHz. For example, the first electronic device can choose 4kHz as the filtering bandwidth. Optionally, in this scenario, the first electronic device can filter out audio frequencies in the 0–4kHz range from the first audio and modulate these frequencies to above 20kHz (i.e., modulate to the ultrasonic frequency band) to obtain the second audio.
[0181] For example, in a scenario where the reference sampling rate is 48kHz, the first audio frequency band is the first frequency band described in Embodiment 1, and the first frequency band is 0–16kHz (the upper limit of this frequency band is 16kHz), the difference between half of the reference sampling rate and the upper limit of the first audio frequency band is 8kHz. Therefore, the range of the filtering bandwidth used by the first electronic device is greater than 0 and less than or equal to 8kHz. For example, the first electronic device can select 8kHz as the filtering bandwidth. Optionally, in this scenario, the first electronic device can filter out audio in the 0–8kHz range from the first audio and modulate that audio to above 16kHz to obtain the second audio.
[0182] Example 3
[0183] Based on the method provided in Embodiment 1 above, or based on the methods provided in Embodiments 1 and 2 above, a possible system architecture applicable to the solution provided in this application embodiment can be referred to. Figure 6 .like Figure 6 As shown, the system may include a first electronic device and a second electronic device.
[0184] like Figure 6 As shown, the second electronic device may include an HDMI interface. The second electronic device can transmit the audio to be played to the first electronic device via the HDMI interface.
[0185] like Figure 6 As shown, the first electronic device may include an HDMI interface, an audio processing module, and at least one physical channel (e.g., Figure 6 The diagram shows physical channels 1 to n, where n is a positive integer. A physical channel can specifically be an audio playback device such as a speaker. The HDMI interface can be used to receive first audio from a second electronic device. The audio processing module can generate second audio based on the first audio. The audio processing module can also be used for preprocessing the first and second audio before playback (e.g., sound effects processing, digital-to-analog conversion, power amplification, etc.). At least one physical channel can be used to play the first and second audio.
[0186] In one example, such as Figure 6 As shown, the audio processing module in the first electronic device may include an audio decoding module, a sound effects processing module, a mixing and filtering module, an ultrasonic modulation module, a mixing module, a digital-to-analog converter module, and a power amplifier. Based on this structure, the process of processing audio received by the first electronic device through the HDMI interface within the audio processing module may include:
[0187] 1) The audio decoding module can receive audio data in encoded format from the HDMI interface, decode the encoded audio to obtain at least one audio channel, and transmit at least one audio channel to the audio effects processing module.
[0188] 2) The audio processing module can perform specific audio processing on each audio channel in the received at least one audio channel, and send one of the processed audio channels to the mixing module, send the other audio channels to the digital-to-analog conversion module, and send at least one audio channel to the mixing and filtering module.
[0189] 3) The mixing and filtering module can mix at least one received audio signal to obtain another audio signal and send the audio signal to the ultrasonic modulation module; or, the mixing and filtering module can mix at least one received audio signal to obtain another audio signal, filter the audio signal based on the method provided in Embodiment 2, and send the obtained audio signal to the ultrasonic modulation module.
[0190] 4) The ultrasonic modulation module can modulate the received audio to the ultrasonic frequency band and send the resulting audio to the mixing module.
[0191] 5) The mixing module can mix the two received audio streams to obtain one audio stream, and then send the audio stream to the digital-to-analog converter module.
[0192] 6) The digital-to-analog conversion module can perform digital-to-analog conversion on each received audio channel separately, and send each converted audio channel to the power amplifier separately.
[0193] 7) The power amplifier can amplify the power of each received audio channel and send it to the corresponding physical channel for playback.
[0194] like Figure 6 As shown, based on the above method, the audio played by the first electronic device includes audible sound (i.e., the audio to be played from the second electronic device) and ultrasound. In scenarios where user audio needs to be received and processed, the audio received by the second electronic device includes the audible sound played by the first electronic device, ultrasound, and user speech.
[0195] like Figure 6 As shown, the second electronic device may further include an audio receiving device (e.g., a microphone) and an audio processing module. The audio receiving device can receive ambient audio and send it to the audio processing module. This audio may include audible sounds, ultrasound, and user voice (i.e., user audio) from the first electronic device. The audio processing module can be used to acquire the user audio based on the received audio.
[0196] In one example, such as Figure 6As shown, the audio processing module in the second electronic device may include an analog-to-digital conversion module, an audio separation module, an ultrasonic demodulation module, and an echo suppression module. Based on this structure, the process by which the audio received by the second electronic device through the audio receiving device is processed within the audio processing module may include:
[0197] 1) The analog-to-digital conversion module can perform analog-to-digital conversion on the audio from the audio receiving device and send the converted audio to the audio separation module.
[0198] 2) The audio separation module can divide the received audio into audible audio and ultrasonic audio, and send the audible audio to the echo suppression module and the ultrasonic audio to the ultrasonic demodulation module.
[0199] Optional, such as Figure 6 As shown, the audio separation module may include a bandpass filter and a low-pass filter. The bandpass filter can extract ultrasonic audio from the audio received by the audio separation module and send the ultrasonic audio to the ultrasonic demodulation module. The low-pass filter can extract audible audio from the audio received by the audio separation module and send the audible audio to the echo suppression module.
[0200] 3) The ultrasonic demodulation module can perform ultrasonic demodulation on the received audio and send the resulting audio to the echo suppression module.
[0201] 4) The echo suppression module can remove audio from the audible audio from the audio separation module that belongs to the first electronic device, based on the audio from the ultrasonic demodulation module, and retain the user's audio.
[0202] The second electronic device can perform subsequent processing or response based on the user audio obtained by the echo suppression module.
[0203] Figure 6 In the system shown, based on the existing hardware of the first and second electronic devices, the audio return transmission and echo suppression processes provided in the above embodiments can be easily and efficiently implemented through improvements to software modules such as the mixing and filtering module, ultrasonic modulation module, mixing module, audio separation module, and ultrasonic demodulation module.
[0204] It is necessary to understand that Figure 6The system architecture shown is merely an example. In practical applications, modules can be added or removed, or modules can be split or merged. This application does not impose specific limitations. The division of functional modules in this system architecture is illustrative and represents only a logical functional division. In actual implementation, there may be other division methods, and the specific functions performed by each functional module may also be divided in other ways. For example, each electronic device may be divided into more or fewer functional services according to other factors, or the functions of each service may be divided in other ways, or the functions may not be divided into services, but rather work as a whole. The device names or functional module names mentioned above are merely examples. For instance, the first electronic device mentioned above may also be called a transmitting device or a sending device, and the second electronic device mentioned above may also be called a receiving device. Furthermore, the audio processing module in the first electronic device mentioned above may also be called a first audio processing module, and the audio processing module in the second electronic device mentioned above may also be called a second audio processing module. The names of the functional modules in this application embodiment are not limited.
[0205] Example 4: Transfer Function Processing
[0206] In some embodiments of this application, in the solutions provided in the above embodiments, each physical channel of the first electronic device may have a transfer function. The transfer function of each physical channel is used to represent the relationship between an audio signal played by that physical channel and the corresponding audio signal received by the second electronic device. After each physical channel plays an audio signal, there may be loss during the process of that audio signal reaching the audio receiving device (e.g., a microphone) of the second electronic device. Therefore, there may be a deviation between the audio signal actually received by the second electronic device and the audio signal actually played by the physical channel. The transfer function of the physical channel is used to represent the relationship between the audio signal actually played by the physical channel and the audio signal actually received by the second electronic device when this deviation exists. Optionally, the transfer function can be a function that calculates the corresponding audio signal received by the second electronic device based on the audio signal played by the physical channel. In this method, the audio played by the physical channel can be understood as the input audio of the transfer function, and the audio signal received by the audio receiving device of the second electronic device from that physical channel can be understood as the output audio of the transfer function. The transfer function can be understood as a function used to represent the relationship between the input audio and the output audio.
[0207] Initially, as an optional implementation, the transfer function of the physical channel can be configured based on the acoustic characteristics of the physical channel. These acoustic characteristics may include, for example, the electroacoustic response of the physical channel and its deployment location. The electroacoustic response of the physical channel represents the relationship between the input signal and the output sound of the physical channel. As another optional implementation, while maintaining a quiet external environment, a second electronic device can send audio to a first electronic device, which then plays the audio. In this scenario, the audio received by the second electronic device from the first electronic device does not contain user voice or other interference signals. Therefore, the transfer function of the physical channel of the first electronic device can be determined more accurately based on the relationship between the audio signal played by the physical channel of the first electronic device and the corresponding audio received by the second electronic device.
[0208] In subsequent processes, the second electronic device can utilize a pre-defined AI model or AI algorithm to determine correction parameters for the transfer functions of the physical channels of the first electronic device by comparing the fourth audio obtained according to the method described in Embodiment 1 with the demodulated second audio, and then instruct the first electronic device to provide these correction parameters. The first electronic device can then correct or adjust the transfer functions of each physical channel based on these correction parameters. For example, the second electronic device can first use the pre-defined AI model or AI algorithm to identify and remove user audio from the fourth audio to obtain a sixth audio (which is the audio received by the second electronic device from the first electronic device), and then determine the correction parameters for the transfer functions of the physical channels of the first electronic device by comparing the sixth audio with the demodulated second audio and based on the error between the sixth audio and the demodulated second audio. Of course, the second electronic device can also determine these correction parameters in other specific ways, and this embodiment does not impose specific limitations on this.
[0209] In some embodiments of this application, based on the above method, in the solutions provided in the above embodiments, before the first electronic device performs analog-to-digital conversion on each audio channel (for example, before modulating the first audio to the ultrasonic band as described in the first possible method in Embodiment 1, or before mixing at least one audio channel as described in the second possible method in Embodiment 1), it can also perform transformation processing on the audio channel according to the transfer function of the physical channel corresponding to each audio channel (i.e., the second preprocessing described in Embodiment 1), so that the transformed audio is closer to the audio channel actually received by the second electronic device. Based on this method, after the second audio generated based on the transformed audio is transmitted to the second electronic device, the audio obtained by the second electronic device after demodulating the second audio can be closer to the audio played by the first electronic device and received by the second electronic device, thereby further improving the accuracy of subsequent noise reduction processing.
[0210] Example 5
[0211] Based on the methods provided in Embodiments 1 and 4 above, or based on the methods provided in Embodiments 1, 2, and 4 above, a possible system architecture applicable to the solutions provided in this application can be referred to. Figure 7 .like Figure 7 As shown, the system may include a first electronic device and a second electronic device.
[0212] like Figure 7 As shown, the first electronic device may include an HDMI interface, an audio processing module, and at least one physical channel (e.g., Figure 7 The physical audio channels shown are 1 to n, where n is a positive integer. The second electronic device may include an HDMI interface, an audio receiving device (such as a microphone), and an audio processing module.
[0213] In one example, such as Figure 7 As shown, the audio processing module in the first electronic device may include an audio decoding module, a sound effects processing module, a transfer function processing module, a mixing and filtering module, an ultrasonic modulation module, a mixing module, a digital-to-analog conversion module, and a power amplifier. Based on this structure, the process of processing audio received by the first electronic device through the HDMI interface within the audio processing module may include:
[0214] 1) The audio decoding module can receive audio data in encoded format from the HDMI interface, decode the encoded audio to obtain at least one audio channel, and transmit at least one audio channel to the audio effects processing module.
[0215] 2) The audio effects processing module can perform specific audio effects processing on each of the received at least one audio channel, and send one of the processed audio channels to the mixing module, send the other audio channels to the digital-to-analog conversion module, and send at least one audio channel to the transfer function processing module.
[0216] 3) The transfer function processing module can process each audio channel in at least one received audio channel through the transfer function of the corresponding physical channel and send the resulting audio to the mixing and filtering module.
[0217] In this process, one audio channel is passed through the transfer function of the corresponding physical channel, which means that the audio channel is used as the input of the transfer function, and the output audio corresponding to the audio channel is calculated according to the relationship between the input audio and the output audio indicated by the transfer function.
[0218] 4) The mixing and filtering module, ultrasonic modulation module, mixing module, digital-to-analog conversion module, and power amplifier process the received audio according to the method described in the aforementioned embodiment 3, which will not be repeated here.
[0219] In one example, such as Figure 7 As shown, the audio processing module in the second electronic device may include an analog-to-digital conversion module, an audio separation module, an ultrasonic demodulation module, an AI correction module, and an echo suppression module. Based on this structure, the process by which the audio received by the second electronic device through the audio receiving device is processed within the audio processing module may include:
[0220] 1) The analog-to-digital conversion module can perform analog-to-digital conversion on the audio from the audio receiving device and send the converted audio to the audio separation module.
[0221] 2) The audio separation module can divide the received audio into audible audio and ultrasonic audio, and send the audible audio to the AI correction module and the echo suppression module respectively, and send the ultrasonic audio to the ultrasonic demodulation module.
[0222] Optional, such as Figure 7 As shown, the audio separation module may include a bandpass filter and a low-pass filter. The bandpass filter extracts ultrasonic audio from the audio received by the audio separation module and sends the ultrasonic audio to the ultrasonic demodulation module. The low-pass filter extracts audible audio from the audio received by the audio separation module and sends the audible audio to the AI correction module and the echo suppression module, respectively.
[0223] 3) The ultrasonic demodulation module can perform ultrasonic demodulation on the received audio and send the resulting audio to the AI correction module and the echo suppression module respectively.
[0224] 4) The echo suppression module can remove audio from the audible audio from the audio separation module that belongs to the first electronic device, based on the audio from the ultrasonic demodulation module, and retain the user's audio.
[0225] 5) The AI correction module can determine the correction parameters for the transfer function of the physical channel of the first electronic device based on the audible audio from the audio separation module and the audio from the ultrasonic demodulation module, and send the correction parameters to the first electronic device via the HDMI interface.
[0226] There is no sequential relationship between this step and step 4), and they can be executed independently.
[0227] Based on the above method, the transfer function processing module in the first electronic device is also used to perform the following method: receive correction parameters from the second electronic device through the HDMI interface, and correct the transfer function of each physical channel based on the correction parameters.
[0228] In this embodiment, Figure 7 The system architecture shown is Figure 6Compared to the system architecture shown, except for the addition of a transfer function processing module and related processing methods in the first electronic device and an AI correction module and related processing methods in the second electronic device, the other modules and processing methods in the first and second electronic devices can be referred to [the original text is missing here, likely due to a formatting error]. Figure 6 The corresponding descriptions in the corresponding methods will not be repeated in this embodiment.
[0229] Figure 7 In the system shown, based on the existing hardware of the first electronic device and the second electronic device, the audio return transmission and echo suppression processes provided in the above embodiments can be easily and efficiently implemented by improving software modules such as the transfer function processing module, the mixing and filtering module, the ultrasonic modulation module, the mixing module, the audio separation module, the ultrasonic demodulation module, and the AI correction module.
[0230] It is necessary to understand that Figure 7 The system architecture shown is merely an example. In practical applications, modules can be added or removed, or modules can be split or merged. This application does not impose specific limitations. The division of functional modules in this system architecture is illustrative and represents only a logical functional division. In actual implementation, there may be other division methods, and the specific functions performed by each functional module may also be divided in other ways. For example, the first electronic device or the second electronic device may be divided into more or fewer functional services according to other factors, or the functions of each service may be divided in other ways, or the functional services may not be divided and the device may work as a whole. The names of the functional modules mentioned above are merely examples, and the names of the functional modules in this application embodiment are not limited.
[0231] It should be noted that the application scenarios provided in the above embodiments are merely illustrative examples of the applicable scenarios for the embodiments of this application, and do not limit the applicable scenarios of the solutions in this application. Some methods or the same technical concepts provided in any of the above embodiments can also be applied in other embodiments or other scenarios, or can be combined with methods provided in other embodiments. For example, based on the solution provided in Embodiment 1, the audio filtering processing method provided in Embodiment 2 and / or the transfer function processing method provided in Embodiment 4 can be combined. The methods provided in the above embodiments can also be applied in combination with specific embodiments or specific scenarios, and will not be listed and described one by one in this application.
[0232] Based on the above embodiments and the same technical concept, this application also provides an electronic device for implementing the audio processing method provided in this application for a first electronic device or a second electronic device. Figure 8As shown, the electronic device 800 may include: a memory 801, one or more processors 802, and one or more computer programs (not shown). These devices may be coupled via one or more communication buses 803. Optionally, the electronic device 800 may also include a display screen 804.
[0233] The memory 801 stores one or more computer programs (code), and the one or more computer programs include computer instructions; one or more processors 802 call the computer instructions stored in the memory 801, causing the electronic device 800 to execute the audio processing method applied to the first electronic device or the second electronic device provided in the embodiments of this application.
[0234] In a specific implementation, memory 801 may include high-speed random access memory, and may also include non-volatile memory, such as one or more disk storage devices, flash memory devices, or other non-volatile solid-state storage devices. Memory 801 may store an operating system (hereinafter referred to as the system), such as embedded operating systems like Android, iOS, Windows, or Linux. Memory 801 can be used to store implementation programs of the embodiments of this application. Memory 801 may also store network communication programs, which can be used to communicate with one or more additional devices, one or more user devices, or one or more network devices.
[0235] One or more processors 802 may be a general-purpose central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits used to control the execution of programs according to the present application.
[0236] Display screen 804 is used to display application interfaces and other related user interfaces.
[0237] It should be noted that, Figure 8 This is merely one implementation of the electronic device 800 provided in this application embodiment. In practical applications, the electronic device 800 may include more or fewer components, as detailed in the following references. Figure 3 The specific structure and description shown are not limited here.
[0238] Based on the above embodiments and the same technical concept, this application also provides a computer-readable storage medium storing a computer program that, when run on a computer, causes the computer to perform the method provided in the above embodiments.
[0239] Based on the above embodiments and the same technical concept, this application also provides a computer program product, which includes a computer program or instructions. When the computer program or instructions are run on a computer, the computer performs the method provided in the above embodiments.
[0240] Based on the above embodiments and the same technical concept, this application also provides a computer program product, which includes a computer program or instructions. When the computer program or instructions are run on a computer, the computer performs the method provided in the above embodiments.
[0241] The methods provided in this application can be implemented, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, they can be implemented, in whole or in part, in the form of a computer program product. A computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the flow or function according to the embodiments of the present invention is generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, a network device, a user equipment, or other programmable device. Computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, computer instructions can be transmitted from one website, computer, server, or data center to another via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. A computer-readable storage medium can be any available medium that a computer can access, or a data storage device such as a server or data center that integrates one or more available media. Available media can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., digital video discs (DVDs), or semiconductor media (e.g., SSDs), etc.
[0242] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.
Claims
1. An audio processing method, applied to a first electronic device, characterized in that, The method includes: Obtain the first audio to be played; wherein the frequency of the first audio belongs to the first frequency band that is audible to the human ear; A second audio is generated based on the first audio; wherein the frequency of the second audio belongs to a second frequency band that is inaudible to the human ear, and the second audio is used to identify the first audio; Play the first audio and the second audio.
2. The method as described in claim 1, characterized in that, The upper limit of the first frequency band is lower than or equal to the set frequency, and the lower limit of the second frequency band is higher than the set frequency.
3. The method as described in claim 2, characterized in that, The set frequency is greater than or equal to 16 kHz and less than or equal to 20 kHz.
4. The method according to any one of claims 1 to 3, characterized in that, When one of the receiving ends of the first audio and the second audio is a second electronic device located in the same spatial area as the first electronic device, the second audio is used by the second electronic device to identify and remove the first audio included in the audio collected by the second electronic device.
5. The method as described in claim 4, characterized in that, The step of generating the second audio based on the first audio includes: Modulate the first audio signal to the second frequency band to obtain the second audio signal; or The first audio is subjected to a first preprocessing and / or a second preprocessing, and the resulting audio is modulated to the second frequency band to obtain the second audio. The first preprocessing is used to filter the first audio, and the second preprocessing is used to simulate the changes in the first audio during the process of being transmitted from the first electronic device to the second electronic device, and to process the first audio.
6. The method as described in claim 5, characterized in that, The first preprocessing of the first audio includes: The first audio signal is filtered based on the set filtering bandwidth; or The first audio is filtered based on the target filtering bandwidth; wherein, the target filtering bandwidth is the filtering bandwidth corresponding to the target sampling rate among a plurality of filtering bandwidths, and the sampling rates corresponding to different filtering bandwidths among the plurality of filtering bandwidths are different, and the target sampling rate is the smallest audio sampling rate among the audio sampling rates of the first electronic device and the audio sampling rate of the second electronic device.
7. The method as described in claim 6, characterized in that, The set filtering bandwidth or the target filtering bandwidth meets the following condition: it is less than the difference between half of the target sampling rate and the lower limit of the second frequency band.
8. The method according to any one of claims 5 to 7, characterized in that, The first audio includes audio from at least one channel; The second preprocessing of the first audio includes: Based on the audio transfer function corresponding to each of the at least one audio channel, the audio of each channel is processed; wherein, the audio transfer function corresponding to each channel is used to indicate the correspondence between the audio of each channel played by the first electronic device and the audio of each channel received by the second electronic device.
9. The method as described in claim 8, characterized in that, The method further includes: The system receives correction information from the second electronic device; wherein the correction information is used to correct the audio transfer function corresponding to each channel. Based on the correction information, the audio transfer function corresponding to each channel is corrected.
10. The method according to any one of claims 4 to 9, characterized in that, Playing the first audio and the second audio includes: When the first audio includes audio from multiple channels, the audio from the first channel and the second audio are played through the first audio output device among the multiple audio output devices, and the audio from the second channel among the multiple audio output devices is played through the second audio output device among the multiple audio output devices.
11. The method as described in claim 10, characterized in that, Playing the audio of the first channel and the second audio of the multiple channels through the first audio output device of the multiple audio output devices includes: The audio from the first channel and the second audio are mixed to obtain a third audio, which is then played through the first audio output device.
12. The method as described in claim 10 or 11, characterized in that, The first audio output device is the audio output device that is closest to the second electronic device among the plurality of audio output devices; or The first audio output device is positioned facing the second electronic device.
13. An audio processing method, applied to a system consisting of a first electronic device and a second electronic device, characterized in that, The method includes: The first electronic device acquires a first audio to be played; wherein the frequency of the first audio belongs to a first frequency band that is audible to the human ear; The first electronic device generates a second audio based on the first audio; wherein the frequency of the second audio belongs to a second frequency band that is inaudible to the human ear, and the second audio is used to identify the first audio; The first electronic device plays the first audio and the second audio; The second electronic device receives a third audio; wherein the third audio includes the first audio, the second audio, and user audio from the user; The second electronic device divides the third audio into the second audio and a fourth audio that includes the first audio and the user audio; The second electronic device obtains the user audio by recognizing the first audio in the fourth audio and removing the second audio.
14. The method as described in claim 13, characterized in that, The upper limit of the first frequency band is lower than or equal to the set frequency, and the lower limit of the second frequency band is higher than the set frequency.
15. The method as described in claim 14, characterized in that, The set frequency is greater than or equal to 16 kHz and less than or equal to 20 kHz.
16. The method according to any one of claims 13 to 15, characterized in that, The second electronic device identifies and removes the first audio from the fourth audio based on the second audio, including: The second electronic device demodulates the second audio and identifies and removes the first audio from the fourth audio based on the demodulated second audio.
17. The method according to any one of claims 13 to 16, characterized in that, The method further includes: The second electronic device analyzes the error between the second audio and the first audio based on the second audio and the fourth audio, and generates correction information based on the error; wherein, the correction information is used to correct the audio transfer function corresponding to each channel of the first electronic device, and the audio transfer function corresponding to each channel is used to indicate the correspondence between the audio of each channel played by the first electronic device and the audio of each channel received by the second electronic device; The second electronic device sends the correction information to the first electronic device via a high-definition multimedia interface (HDMI).
18. An electronic device, characterized in that, The electronic device includes a memory and one or more processors; The memory is used to store computer program code, which includes computer instructions; when the computer instructions are executed by the one or more processors, the electronic device performs the method as described in any one of claims 1 to 12.
19. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when run on an electronic device, causes the electronic device to perform the method as described in any one of claims 1 to 12.
20. A computer program product, characterized in that, The computer program product includes a computer program or instructions that, when executed on an electronic device, cause the electronic device to perform the method as described in any one of claims 1 to 12.