Audio signal processing method and related devices

The method addresses transmission delays in spatial audio systems by determining binaural signals and encoded audio without head tracking, enhancing user experience through improved accuracy and efficiency.

WO2026156875A1PCT designated stage Publication Date: 2026-07-30HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
HUAWEI TECH CO LTD
Filing Date
2025-01-27
Publication Date
2026-07-30

AI Technical Summary

Technical Problem

Existing spatial audio systems experience transmission delays due to the need for head tracking information exchange between host and client devices, which degrades user experience.

Method used

An audio signal processing method that determines binaural signals and encoded audio signals without requiring head tracking information transmission, using decorrelation and compression techniques to improve accuracy and reduce delays.

Benefits of technology

This approach reduces transmission delays and enhances user experience by eliminating the need for real-time head tracking information exchange, thereby improving the accuracy and efficiency of spatial audio rendering.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025075469_30072026_PF_FP_ABST
    Figure CN2025075469_30072026_PF_FP_ABST
Patent Text Reader

Abstract

Provide an audio signal processing method and related device. The audio signal processing method includes the following steps: determining a first audio signal according to an audio signal resource; performing binauralization on the first audio signal to obtain N pair (s) of binaural signals, where the N pair (s) of binaural signals is in one-to-one correspondence with N location (s); determining M set (s) of binaural signals according to the N location (s), where each of the M set (s) of binaural signals includes at least one pair of binaural signals; determining an encoded audio signal based on the M set (s) of binaural signals, where the encoded audio signal is a K-channel encoded signal, where K is a positive integer and greater than or equal to two; transmitting the encoded audio signal to a client device.
Need to check novelty before this filing date? Find Prior Art

Description

AUDIO SIGNAL PROCESSING METHOD AND RELATED DEVICESTECHNICAL FIELD

[0001] Embodiments of the present application relate to the field of signal processing, more specifically, related to an audio signal processing method and related devices.BACKGROUND

[0002] Spatial audio is a technology that simulates the natural auditory perception mechanism of the human ear to create three-dimensional (3D) sound effects. Unlike traditional stereo or surround sound systems, spatial audio not only localizes sound in the horizontal plane but also extends to the vertical plane, providing a more immersive and directional auditory experience. It achieves this by simulating the time and level differences of sound as it travels and by using head-related transfer functions (HRTF) to recreate the perceived location of sound sources. For example, when sound comes from one side, the ear detects slight differences in timing and loudness. Spatial audio technology mimics these natural auditory features, making the sound seem to come from different directions in the real world. This not only improves the accuracy of sound localization but also enhances the sense of depth and immersion.

[0003] In practice, spatial audio often combines with dynamic head tracking technology to adjust the audio signal in real-time, ensuring that the position of the sound source changes as the user moves their head. For instance, when using headphones that support spatial audio, the direction of sound changes as the user turns their head, further enhancing immersion and realism. This process mimics how humans use the relative position of their ears and environmental cues to perceive the location of sounds, delivering a more natural and precise auditory experience. The implementation of spatial audio can be applied across various devices such as headphones, smart speakers, and cinema systems. By leveraging innovative acoustic technologies, users can enjoy a higher level of spatial and dynamic audio, as though they are right in the middle of the sound's origin.SUMMARY

[0004] Embodiments of the present application provide an audio signal processing method and related device, which can reduce a transmission delay between the host device and the client device and improve user experience

[0005] According to a first aspect, an embodiment of the present application provides an audio signal processing method. The audio signal processing method includes the following steps: determining a first audio signal according to an audio signal resource; performing binauralization on the first audio signal to obtain N pair (s) of binaural signals, where the N pair (s) of binaural signals is in one-to-one correspondence with N location (s) , N is a positive integer greater than or equal to one; determining M set (s) of binaural signals according to the N location (s) , where each of the M set (s) of binaural signals includes at least one pair of binaural signals, M is a positive integer and less than or equal to N; determining an encoded audio signal based on the M set (s) of binaural signals, where the encoded audio signal is a K-channel encoded signal, where K is a positive integer and greater than or equal to two; transmitting the encoded audio signal to a client device.

[0006] According to the aforementioned technical solution, tracking information obtained by the client device do not need to transmit to a host device. The tracking information includes a motion posture of a user’s head who wears the client device. In this case, the host device does not need to use the tracking information to perform binauralization to obtain binaural signals. Therefore, a transmission delay caused by the tracking information may be avoided. This can reduce a transmission delay between the host device and the client device and improve user experience. Further, the binaural signals are divided into different sets according to their locations. This may improve accuracy for a subsequent encoding procedure.

[0007] In a possible design of the first aspect, the determining an encoded audio signal based on the M set (s) of binaural signals, includes: performing a decorrelation operation on the M set (s) of binaural signals to obtain L processed audio signals, where L is a positive integer greater than or equal to K; compressing the L processed audio signals to obtain the encoded audio signal.

[0008] The decorrelation operation may further improve the accuracy for the subsequent encoding procedure.

[0009] In a possible design of the first aspect, the decorrelation operation is configured to delay and / or to do a phase shifting the M set (s) of binaural signals.

[0010] In a possible design of the first aspect, the compressing the L processed binaural signals to obtain the encoded audio signal, includes: compressing the L processed audio signals to obtain the encoded audio signal by using a matrix encoder, where the encoded audio signal is a 2-channel encoded signal, the 2-channel encoded signal includes a left-channel encoded signal and a right-channel encoded signal.

[0011] In a possible design of the first aspect, each location among the N location (s) includes a first component and a second component, where the first component is front or rear, and the second component is left, center, or right; when a set of binaural signals among the M set (s) of binaural signals includes at least two pairs of binaural signals, two first components of two locations are the same, where the two locations corresponds to two pairs of binaural signals among the at least two pairs of binaural signals respectively.

[0012] In a possible design of the first aspect, when M = 1, the first component of each binaural signal among the M set (s) of binaural signals is the front; when M is a positive integer greater than one, the M set (s) of binaural signals includes M1 first set (s) of binaural signals and M2 second set (s) of binaural signals, where M1 and M2 are positive integer, and M1 + M2 = M, the first component of each binaural signal among the first set of binaural signals is the front, while the first component of each binaural signal among the second set of binaural signals is the rear.

[0013] In a possible design of the first aspect, the determining a first audio signal according to an audio signal resource, includes: decoding the audio signal resource to obtain a second audio signal; rendering the second audio signal to obtain the first audio signal.

[0014] According to a second aspect, an embodiment of the present application provides an audio signal processing method. The audio signal processing method includes the following steps: obtaining an encoded audio signal; determining M set(s) of reconstructed binaural signals according to the encoded audio; determining a pair of binaural signals based on tracking information obtained from a sensor unit in the client device and the M set (s) of reconstructed binaural signals; determining an output audio signal according to the pair of binaural signals.

[0015] The tracking information includes a motion posture of a user’s head who wears the client device. According to the aforementioned technical solution, the client device does not need to transmit the tracking information to a host device. The host device does not need to use the tracking information to obtain binaural signals. Therefore, a transmission delay caused by the tracking information may be avoided. This can reduce a transmission delay between the host device and the client device and improve user experience.

[0016] In a possible design of the second aspect, the determining M set (s) of reconstructed binaural signals according to the encoded audio, includes: determining L reconstructed audio signals according the encoded audio; performing an inverse decorrelation operation on the L reconstructed audio signals to obtain the M set (s) of reconstructed binaural signals.

[0017] In a possible design of the second aspect, the determining L reconstructed audio signals according the encoded audio, includes: determining three sets of audio signals according to the encoded audio signal, where the three sets of audio signals includes a first audio signal set, a second audio signal set and a third audio signal set, the first audio signal set includes two positively correlated audio signals, the second audio signal set includes two low correlated audio signals, and the third audio signal set includes two negatively correlated audio signals; determining the L reconstructed audio signals according to the three sets of audio signals.

[0018] In a possible design of the second aspect, the determining the L reconstructed audio signals according to the three sets of audio signals, includes: decoding, by using a first matrix decoder, audio signals of the first audio signal set to obtain a plurality of first target audio signals; decoding, by using a second matrix decoder, audio signals of the second audio signal set to obtain a plurality of second audio binaural signals; decoding, by using a third matrix decoder, audio signals of the third audio signal set to obtain a plurality of second target audio signals; determining the L reconstructed audio signals according to the plurality of the first target audio signals, the plurality of the second target audio signals, and the plurality of the third target audio signals.

[0019] In a possible design of the second aspect, the determining a pair of binaural signals based on tracking information obtained from a sensor unit in the client device and the M set (s) of reconstructed binaural signals, includes: filtering, by using a spectral cue filter, the M set (s) of reconstructed binaural signals to obtain M set (s) of filtered binaural signals based on the tracking information; delaying the M set (s) of filtered binaural signals to obtain M set (s) of delayed binaural signals; combining left-channel audio signals of the M set (s) of delayed binaural signals to obtain a left-channel binaural signal of the pair of binaural signals; combining right-channel audio signals of the M set (s) of delayed binaural signals to obtain a right-channel binaural signal of the pair of binaural signals.

[0020] According to a third aspect, an embodiment of the present application provides an electronic device, and the electronic device has a function of implementing the method in the first aspect. The function may be implemented by hardware, or may be implemented by hardware executing corresponding software. The hardware of the software includes one or more modules corresponding to the function.

[0021] According to a fourth aspect, an embodiment of the present application provides an electronic device, and the electronic device has a function of implementing the method in the second aspect. The function may be implemented by hardware, or may be implemented by hardware executing corresponding software. The hardware of the software includes one or more modules corresponding to the function.

[0022] According to a fifth aspect, an embodiment of the present application provides an electronic device, including a processor and a memory. The processor is connected to the memory. The memory is configured to store instructions, and the processor is configured to execute the instructions. When the processor executes the instructions stored in the memory, the processor is enabled to perform the method in the first aspect or any possible design of the first aspect.

[0023] According to a sixth aspect, an embodiment of the present application provides an electronic device, including a processor and a memory. The processor is connected to the memory. The memory is configured to store instructions, and the processor is configured to execute the instructions. When the processor executes the instructions stored in the memory, the processor is enabled to perform the method in the second aspect or any possible design of the second aspect.

[0024] According to a seventh aspect, an embodiment of the present application provides a computer readable storage medium, including instructions. When the instructions run on an electronic device, the electronic device is enabled to perform the method in the first aspect or any possible design of the first aspect.

[0025] According to an eighth aspect, an embodiment of the present application provides a computer readable storage medium, including instructions. When the instructions run on an electronic device, the electronic device is enabled to perform the method in the second aspect or any possible design of the second aspect.

[0026] According to a ninth aspect, an embodiment of the present application provides a chip system, where the chip system includes a communication interface and a processing circuit, the communication interface is configured to obtain to-be-processed data, and the processing circuit is configured to process the to-be-processed data according to the method in the first aspect or any possible design of the first aspect.

[0027] According to a tenth aspect, an embodiment of the present application provides a chip system, where the chip system includes a communication interface and a processing circuit, the communication interface is configured to obtain to-be-processed data, and the processing circuit is configured to process the to-be-processed data according to the method in the second aspect or any possible design of the second aspect.

[0028] According to an eleventh aspect, an embodiment of the present application provides a computer program product, where when the computer program product runs on an electronic device, the electronic device is enabled to perform the method in the first aspect or any possible design of the first aspect.

[0029] According to a twelfth aspect, an embodiment of the present application provides a computer program product, where when the computer program product runs on an electronic device, the electronic device is enabled to perform the method in the second aspect or any possible design of the second aspect.

[0030] According to a thirteenth aspect, an embodiment of the present application provides a computer readable storage medium, including a data structure, the data structure includes the encoded audio signal obtained according to the method in the first aspect or any possible design of the first aspect.DESCRIPTION OF DRAWINGS

[0031] FIG. 1 illustrates an audio system according to some embodiments of the present application.

[0032] FIG. 2 is a schematic structural diagram of a host device 100.

[0033] FIG. 3 is a schematic structural diagram of the client device 200.

[0034] FIG. 4 illustrates an audio signal processing method provided by some embodiments of the present application.

[0035] FIG. 5 illustrates an example of a relationship between the N pair (s) of binaural signals and the N location (s) .

[0036] FIG. 6 illustrates another example of a relationship between the N pair (s) of binaural signals and the N location (s) .

[0037] FIG. 7 illustrates another example of a relationship between the N pair (s) of binaural signals and the N location (s) .

[0038] FIG. 8 illustrates an audio signal processing procedure according to some embodiment of the present application.

[0039] FIG. 9 depicts a procedure for obtaining the 2-channel encoded audio signal.

[0040] FIG. 10 depicts a procedure for determining the binaural signal.

[0041] FIG. 11 is a schematic block diagram of an electronic device 1100 according to some embodiments of the present application.

[0042] FIG. 12 is a schematic block diagram of an electronic device 1200 according to some embodiments of the present application.DESCRIPTION OF EMBODIMENTS

[0043] The following describes the technical solutions in the present application with reference to the accompanying drawings.

[0044] The terms such as "first" and "second" below are merely for a descriptive purpose, and cannot be understood as indicating or implying relative importance, or implicitly indicating a quantity of indicated technical features. Therefore, the features defined by "first" and "second" can explicitly or implicitly include one or more features.

[0045] As used herein, "at least one" means one or more, and "a plurality of" means two or more. "and / or" describes an association relationship of associated objects, and indicates that there may be three relationships. For example, A and / or B may indicate cases includes “only A” , “both A and B” , and “only B” , where A and B may be singular or plural. The character " / " generally indicates that the associated objects are in an OR relationship. "At least one of the following items" or a similar expression thereof refers to any combination of these items, including any combination of a single item or a plurality of items. For example, “at least one of a, b, or c” may represent a, b, c, “a and b” , “a and c” , “b and c” , or “a, b and c” , where a, b, and c may be a single or multiple form.

[0046] For a better understanding of embodiments provided by the present application, the following briefly describes some related basic concepts.

[0047] Binaural cue refers to various audio signal characteristics that a human ear uses to locate a position of sound sources, judge sound direction, and perceive spatial awareness. A human brain deduces the position of the sound sources by comparing differences in sound signals received by left and right ears.

[0048] The binaural cues may include interaural time difference (ITD) , interaural level difference (ILD) , and head-related transfer function (HRTF) .

[0049] ITD refers to a time difference between a sound reaching left and right ears. Typically, when a sound is located to one side, the sound will reach an ear closer to the source first and then an ear farther from the source.

[0050] ILD refers to a difference in an intensity of sound received by the left and right ears. Due to a head's blocking effect, when the sound comes from one direction, one ear receives a stronger signal than the other.

[0051] HRTF is a comprehensive mathematical model that describes the complex changes in frequency response, time difference, and level difference when sound reaches both ears from any direction in space. It takes into account not only ITD and ILD but also the changes in the sound's frequency components caused by filtering effects from the head, ear shape, and other factors during propagation.

[0052] FIG. 1 illustrates an audio system according to some embodiments of the present application. Referring to FIG. 1, an audio system 10 includes a host device 100 and a client device 200.

[0053] The host device 100 may be a mobile phone, a laptop, a tablet, an augmented reality (AR) device, a virtual reality (VR) device, or the like. The host device may obtain an audio signal resource from a multimedia file or an application (APP) , determine an encoded audio signal corresponding to the audio signal resource, and transmit the encoded audio signal to the client device 200. The audio signal resource may be a music file, a sound track obtained from a video file, a sound signal obtained from an application.

[0054] The client device 200 may be earphones or headphones. The client device 200 may be a wired device or a wireless device. For example, in some embodiments, the client device 200 connects with the host device via a connector (e.g., a 3.5mm audio connector, a universal serial bus (USB) connector or the like) . In some other embodiments, the client device 200 may wirelessly connect to the host device 100. For example, the client device 200 may connect to the host device 100 through a short-range wireless technology (e.g., Bluetooth, NearLink, or the like) . The client device 200 may determine a sound signal from the obtained encoded audio signal and supply the sound signal to right and left acoustic transducers of the client device 200.

[0055] For example, FIG. 2 is a schematic structural diagram of a host device 100. The host device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communications module 150, a wireless communications module 160, an audio module 170, a loudspeaker 170A, a telephone receiver 170B, a microphone 170C, a headset jack 170D, a sensor module 180, and the like. The sensor module 180 may include a pressure sensor, a gyroscope sensor, an acceleration sensor, a distance sensor, an optical proximity sensor, a fingerprint sensor, a touch sensor, and the like.

[0056] It may be understood that the schematic structure in this embodiment of this application constitutes no specific limitation on the host device 100. In some other embodiments of this application, the host device 100 may include more or fewer components than those shown in the figure, or some components may be combined, or some components may be split, or components are arranged in different manners. The components shown in the figure may be implemented by using hardware, software, or a combination of software and hardware.

[0057] The processor 110 may include one or more processing units. For example, the processor 110 may include an application processor (AP) , a modem processor, a graphics processing unit (GPU) , an image signal processor (ISP) , a controller, a memory, a video codec, a digital signal processor (DSP) , a baseband processor, and / or a neural-network processing unit (NPU) , and the like. Different processing units may be separate components, or may be integrated into one or more processors.

[0058] The controller may be a nerve center and a command center of the host device 100. The controller may generate an operation control signal based on an instruction operation code and a time sequence signal, to complete control of instruction reading and instruction execution.

[0059] The memory may be further disposed in the processor 110, to store an instruction and data. In some embodiments, the memory in the processor 110 is a cache. The memory may store an instruction or data that is used or cyclically used by the processor 110. If the processor 110 needs to use the instruction or the data again, the processor 110 may directly invoke the instruction or the data from the memory, so as to avoid repeated access, and reduce a waiting time of the processor 110, thereby improving system efficiency.

[0060] In some embodiments, the processor 110 may include one or more interfaces. The interface may be an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI) , a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, a pulse density modulation (PDM) interface, an USB interface, and / or the like.

[0061] The I2S interface may be configured to perform audio communication. In some embodiments, the processor 110 may include a plurality of groups of I2S buses. The processor 110 may be coupled to the audio module 170 over an I2S bus, to implement communication between the processor 110 and the audio module 170. In some embodiments, the audio module 170 may transmit an audio signal to the wireless communications module 160 over an I2S interface, to implement a function of answering a call over a wireless headset.

[0062] The PCM interface may be also configured to perform audio communication, to perform sampling, quantization, and encoding on an analog signal. In some embodiments, the audio module 170 may be coupled to the wireless communications module 160 over a PCM bus interface. In some embodiments, the audio module 170 may also transmit an audio signal to the wireless communications module 160 over a PCM interface, to implement a function of answering a call over a wireless headset. Both the I2S interface and the PCM interface may be configured to perform audio communication.

[0063] The UART interface is a universal serial data line, and is configured to perform asynchronous communication. The bus may be a two-way communications bus. The UART interface switches to-be-transmitted data between serial communication and parallel communication. In some embodiments, the UART interface is usually configured to connect the processor 110 to the wireless communications module 160. For example, the processor 110 communicates with a short-range wireless communication module in the wireless communications module 160 over the UART interface, to implement a short-range wireless communication function. In some embodiments, the audio module 170 may transmit an audio signal to the wireless communications module 160 over the UART interface, to implement a function of playing music over a wireless headset.

[0064] The GPIO interface may be configured by using software. The GPIO interface may be configured as a control signal, or may be configured as a data signal. In some embodiments, the GPIO interface may be configured to connect the processor 110 to the wireless communications module 160, the audio module 170, the sensor module 180, and the like. The GPIO interface may be further configured as an I2C interface, an I2S interface, a UART interface, an MIPI interface, or the like.

[0065] The USB interface 130 is an interface that meets a USB standard specification, and may be specifically a Mini USB interface, a Micro USB interface, a USB Type C interface, or the like. The USB interface 130 may be configured to connect to the charger to charge the electronic device, or may be configured to transmit data between the host device 100 and a peripheral device, or may be configured to connect to a headset, to play audio over the headset. The interface may be further configured to connect to another electronic device such as an AR device.

[0066] It may be understood that a schematic interface connection relationship between the modules in this embodiment of this application is merely an example for description, and constitutes no limitation on the structure of the host device 100. In some other embodiments of this application, the host device 100 may alternatively use an interface connection manner different from that in the foregoing embodiment, or use a combination of a plurality of interface connection manners.

[0067] The charging management module 140 is configured to receive a charging input from the charger. The charger may be a wireless charger, or may be a wired charger. When charging the battery 142, the charging management module 140 may further supply power to the electronic device over the power management module 141.

[0068] The power management module 141 is configured to connect to the battery 142, the charging management module 140, and the processor 110. The power management module 141 receives an input of the battery 142 and / or the charging management module 140, to supply power to the processor 110, the internal memory 121, an external memory, the wireless communications module 160, and the like. The power management module 141 may be further configured to monitor parameters such as a battery capacity, a battery cycle count, and a battery state of health (electric leakage and impedance) . In some other embodiments, the power management module 141 may be alternatively disposed in the processor 110. In some other embodiments, the power management module 141 and the charging management module 140 may be alternatively disposed in a same component.

[0069] A wireless communication function of the host device 100 may be implemented by using the antenna 1, the antenna 2, the mobile communications module 150, the wireless communications module 160, the modem processor, the baseband processor, and the like.

[0070] The antenna 1 and the antenna 2 are configured to transmit and receive an electromagnetic wave signal. Each antenna of the host device 100 may be configured to cover one or more communication frequency bands. Different antennas may be multiplexed to improve utilization of the antennas. For example, the antenna 1 may be multiplexed as a diversity antenna of a wireless local area network. In some other embodiments, the antenna may be used in combination with a tuning switch.

[0071] The mobile communications module 150 may provide a solution to wireless communication such as 2G / 3G / 4G / 5G applied to the host device 100. The mobile communications module 150 may include at least one filter, a switch, a power amplifier, a low noise amplifier (LNA) , and the like. The mobile communications module 150 may receive an electromagnetic wave over the antenna 1, perform processing such as filtering and amplification on the received electromagnetic wave, and transmit a processed electromagnetic wave to the modem processor for demodulation. The mobile communications module 150 may further amplify a signal modulated by the modem processor, and convert the signal into an electromagnetic wave for radiation over the antenna 1. In some embodiments, at least some function modules of the mobile communications module 150 may be disposed in the processor 110. In some embodiments, at least some function modules of the mobile communications module 150 and at least some modules of the processor 110 may be disposed in a same component.

[0072] The modem processor may include a modulator and a demodulator. The modulator is configured to modulate a to-be-sent low-frequency baseband signal into an intermediate-and-high frequency signal. The demodulator is configured to demodulate a received electromagnetic wave signal into a low-frequency baseband signal. Then, the demodulator transmits the low-frequency baseband signal obtained through demodulation to the baseband processor for processing. After being processed by the baseband processor, the low-frequency baseband signal is transmitted to the application processor. The application processor outputs a sound signal over an audio device (which is not limited to the loudspeaker 170A, the telephone receiver 170B, and the like) , or displays an image or a video over a display screen. In some embodiments, the modem processor may be an independent component. In some other embodiments, the modem processor may be separate from the processor 110, and the modem processor and the mobile communications module 150 or another function module may be disposed in a same component.

[0073] The wireless communications module 160 may provide a solution to wireless communication applied to the host device 100, for example, a wireless local area network (WLAN) (for example, a Wi-Fi network) , Bluetooth (BT) , a global navigation satellite system (GNSS) , frequency modulation (FM) , near field communication (NFC) technology, Nearlink and an infrared (infrared, IR) technology. The wireless communications module 160 may be one or more components into which at least one communication processing module is integrated. The wireless communications module 160 receives an electromagnetic wave over the antenna 2, performs frequency modulation and filtering processing on an electromagnetic wave signal, and sends a processed signal to the processor 110. The wireless communications module 160 may further receive a to-be-sent signal from the processor 110, perform frequency modulation and amplification on the signal, and convert the signal into an electromagnetic wave for radiation over the antenna 2.

[0074] In some embodiments, the antenna 1 and the mobile communications module 150 of the host device 100 are coupled, and the antenna 2 and the wireless communications module 160 of the host device 100 are coupled, so that the host device 100 can communicate with a network and another device by using a wireless communications technology. The wireless communications technology may include a global system for mobile communications (GSM) , a general packet radio service (GPRS) , code division multiple access (CDMA) , wideband code division multiple access (WCDMA) , time-division code division multiple access (TD-CDMA) , long term evolution (LTE) , BT, a GNSS, a WLAN, NFC, FM, an IR technology, and / or the like. The GNSS may include a global positioning system (GPS) , a global navigation satellite system (GLONASS) , a beidou navigation satellite system (BDS) , a quasi-zenith satellite system (QZSS) , and / or a satellite based augmentation system (SBAS) .

[0075] The host device 100 implements a display function over the GPU, the display screen, the application processor, and the like. The display screen is configured to display an image, a video, and the like. The display screen includes a display panel. The display panel may use a liquid crystal display (LCD) , an organic light-emitting diode (OLED) , an active-matrix organic light emitting diode (AMOLED) , a flexible light-emitting diode (FLED) , a MiniLed, a MicroLED, a Micro-OLED, a quantum dot light emitting diode (QLED) , and the like.

[0076] The host device 100 may implement a photographing function over the ISP, a camera lens, the video codec, the GPU, the display screen, the application processor, and the like.

[0077] The digital signal processor is configured to process a digital signal, and in addition to a digital image signal, may further process another digital signal. For example, when the host device 100 performs frequency selection, the digital signal processor is configured to perform Fourier transform and the like on frequency energy.

[0078] The external memory interface 120 may be configured to connect to an external storage card such as a micro SD card, to extend a storage capability of the host device 100. The external storage card communicates with the processor 110 over the external memory interface 120, to implement a data storage function, for example, to store files such as music and videos in the external storage card.

[0079] The internal memory 121 may be configured to store computer executable program code, and the executable program code includes an instruction. The processor 110 runs the instruction stored in the internal memory 121, to perform various function applications and data processing of the host device 100. The internal memory 121 may include a program storage region and a data storage region. The program storage region may store an operating system, an application required by at least one function (for example, a voice playing function or an image playing function) , and the like. The data storage region may store data (for example, audio data and an address book) and the like created when the host device 100 is used. In addition, the internal memory 121 may include a high-speed random access memory, or may include a non-volatile memory such as at least one magnetic disk memory, a flash memory, or a universal flash storage (UFS) .

[0080] The host device 100 may implement an audio function such as music playing or recording over the audio module 170, the loudspeaker 170A, the telephone receiver 170B, the microphone 170C, the headset jack 170D, the application processor, and the like.

[0081] The audio module 170 is configured to convert digital audio information into an analog audio signal output, and is further configured to convert an analog audio input into a digital audio signal. The audio module 170 may be further configured to encode and decode an audio signal. In some embodiments, the audio module 170 may be disposed in the processor 110, or some function modules of the audio module 170 are disposed in the processor 110.

[0082] The loudspeaker 170A is configured to convert an audio electrical signal into a sound signal. The host device 100 may be used to listen to music or answer a call in a hands-free mode over the loudspeaker 170A.

[0083] The telephone receiver 170B is configured to convert an audio electrical signal into a voice signal. When the host device 100 is used to answer a call or receive voice information, the telephone receiver 170B may be put close to a human ear, to receive the voice information.

[0084] The microphone 170C is configured to convert a sound signal into an electrical signal. When making a call or sending voice information, a user may speak with the mouth approaching the microphone 170C, to input a sound signal to the microphone 170C. At least one microphone 170C may be disposed in the host device 100. In some other embodiments, two microphones 170C may be disposed in the host device 100, to collect a sound signal and implement a noise reduction function. In some other embodiments, three, four, or more microphones 170C may be alternatively disposed in the host device 100, to collect a sound signal, implement noise reduction, recognize a sound source, implement a directional recording function, and the like.

[0085] The headset jack 170D is configured to connect to a wired headset. The headset jack 170D may be a USB interface 130, or may be a 3.5 mm open mobile terminal platform (OMTP) standard interface or cellular telecommunications industry association of the USA (CTIA) standard interface.

[0086] FIG. 3 is a schematic structural diagram of the client device 200. The client device 200 includes a left acoustic transducer 201, a right acoustic transducer 202, a processing module 210, a memory 220, a communication module 230, a sensor module 240, and the like.

[0087] In some embodiments, the communication module 230 may be a wired communication module. For example, the communication module 230 may include a connector and a cable. The connector is configured to connect to the host device (e.g., headset jack 170D) and obtain signals from the host device. The cable is configured to transmit the signals obtained by the connector to the processing module 210.

[0088] In some other embodiments, the communication module 230 may be wireless communication module. For example, the communication module 230 may be a Bluetooth module, a NearLink module, or the like. In theses embodiments, the host device 100 may compress the encoded audio signal into an audio encoding format that is efficiently sent via the wireless communication technology. For example, the audio encoding format may be subband coding (SBC) , advanced audio coding (AAC) , or the like. Once the encoded audio signal is compressed into a wireless-compatible format, it is encapsulated into wireless data packets. The host device 100 may transmit the wireless data packet via the wireless communications module 160. Correspondingly, the client device 200 may receive the wireless data packets via the communication module 230. After receiving the wireless data packets, the communication module 230 may unpack the wireless data packets and decode data carried by the wireless data packets to obtain the encoded audio signal. Then, the commination module 230 may send the encoded audio signal to the processing module 210.

[0089] The processing module 210 may include one or more processing units. For example, the processing module 210 may include a controller, a DSP, and the like.

[0090] The processing module 210 is configured to process the encoded audio signal obtained from the communication module 230 to obtain an output audio signal. The output audio signal may be transmitted to the left acoustic transducer 201 and the right acoustic transducer 202. The left acoustic transducer 201 and the right acoustic transducer 220 are configured to convert the output audio signal into sound.

[0091] The memory 220 is configured to store instructions, data and the like. For example, the wireless data packets, the encoded audio signal and / or the output audio signal may be temporally stored in the memory 220. The processing module 210 runs the instructions stored in the memory 220 to processing the encoded audio signal obtained from the communication module 230. It may be understood that, in some embodiments, the processing module 210 may include one or more logic circuits which are configured to perform one or more steps during a procedure for processing the encoded audio signal, the logic circuits may perform these steps by specialized hardware and not need to read instructions from the memory 200.

[0092] The sensor module 240, also referred to as an "inertial measurement unit (IMU) " , is configured to measure a motion posture of a user’s head who wears the client device 200. For example, the sensor module 240 may include one or more gyroscope sensors 240A, one or more acceleration sensors 240B, and so one. For example, the gyroscope sensor 240A may be used to determine angular velocities of the user’s head around three axes (namely, axes x, y, and z) . The acceleration sensor 240B may be used to detect magnitude of acceleration of the user’s head in various directions (usually on three axes) . When the user’s head is static, the acceleration sensor 240B may detect magnitude and a direction of gravity. Information collected by the sensor module 240 may be referred to as head tracking information.

[0093] It may be understood that the schematic structure in this embodiment of this application constitutes no specific limitation on the client device 200. In some other embodiments of this application, the client device 200 may include more or fewer components than those shown in the figure, or some components may be combined, or some components may be split, or components are arranged in different manners. The components shown in the figure may be implemented by using hardware, software, or a combination of software and hardware. For example, in some embodiments, the client device 200 may be true wireless stereo (TWS) earbuds. In these embodiments, the client device may include a left TWS earbud and a right TWS earbud. Each of the left TWS earbud and the right TWS earbud may include an acoustic transducer, the processing module 210, the memory 220, the communication module 230 and the sensor module 240.

[0094] FIG. 4 illustrates an audio signal processing method provided by some embodiments of the present application. The method illustrated in FIG. 4 may be performed by a host device and a client device, or a component of the host device and a component of the client device (e.g., a chip of the host device / client device, a processing circuit of the host device / client device, or the like) . For convenience, it is assumed that the method illustrated in FIG. 4 is performed by the host device and the client device. The host device may be the host device 100 illustrated in FIG. 1 and FIG. 2, while the client device may be the client device 200 illustrated in FIG. 1 and FIG. 3.

[0095] 401, The host device determines a first audio signal according to an audio signal resource.

[0096] The audio signal resource may be a music file, a sound track obtained from a video file, a sound signal obtained from an application (e.g., a video meeting application, a game, or the like) .

[0097] In some embodiments, the first audio signal may be a mono-channel audio signal. In some other embodiments, the first audio signal may be a multi-channel audio signal. For example, the first audio signal may be a stereo 2-channel audio signal, a Dolby 5.1.4 audio signal, a Dolby 7.1.4 audio signal, a Huawei audio vivid audio signal or the like. The first audio signal may be a pulse code modulation signal (PCM) , a differential pulse code modulation (DPCM) signal, an adaptive differential pulse code modulation (ADPCM) signal, or the like, and the present application is not limited thereto.

[0098] In some embodiments, an original audio signal may be compressed into the audio signal resource (e.g., an AAC file, a windows media audio (WMA) file, or the like) . In these cases, the host device may decode the audio signal resource to obtain a second audio signal.

[0099] In some embodiments, the audio signal resource may be a lossless compression file (e.g., a waveform audio file format (WAV) file, a free lossless audio codec (FLAC) file, or the like) . In these cases, the host device may obtain the second audio signal directly from the audio signal resource without decoding the audio signal resource to reconstructed the second audio signal.

[0100] In some embodiments, after obtaining the second audio signal, the host device may render the second audio signal, and the rendered second audio signal is the first audio signal.

[0101] In some other embodiments, the host device does not need to render the second audio signal. For these cases, the second audio signal is the first audio signal.

[0102] In some embodiments, the first audio signal may be a channel based audio signal. In these embodiments, the first audio signal is organized into fixed channels (like stereo or multi-channel setups) , with each channel carrying independent audio data.

[0103] In some other embodiments, the audio signal resource may carry spatial location information. In these embodiments, the first audio signal may be object based audio signal. The first audio signal is treated as independent objects with additional data about position, movement, and sound properties.

[0104] In some other embodiments, the first audio signal may include both the channel based audio signal and the object based audio signal.

[0105] In some other embodiments, the first audio signal may be an audio signal obtained by using scene-based audio with Ambisonics.

[0106] 402, The host device may perform binauralization on the first audio signal to obtain N pair (s) of binaural signals.

[0107] The N pair (s) of binaural signals are in one-to-one correspondence with N location (s) . In other words, each of the N pair (s) of binaural signals has a corresponding location.

[0108] FIG. 5 illustrates an example of a relationship between the N pair (s) of binaural signals and the N location (s) .

[0109] For the example illustrated in (a) of FIG. 5, N is equal to 5. Referring to (a) of FIG. 5, the 5 locations are: front center, front left, front right, rear left, and rear right. Corresponding to the 5 locations, there are 5 pairs of binaural signals. For convenience, the 5 pairs of binaural signals may be referred to as a binaural signal pair 1, a binaural signal pair 2, a binaural signal pair 3, a binaural signal pair 4, and a binaural signal pair 5. The 5 pairs of binaural signals are in one-to-one correspondence with the 5 locations. For example, the binaural signal pair 1 corresponds to the front center, the binaural signal pair 2 corresponds to the front left, the binaural signal pair 3 corresponds to the front right, the binaural signal pair 4 corresponds to the rear left, and the binaural signal pair 5 corresponds to the rear right.

[0110] FIG. 6 illustrates another example of a relationship between the N pair (s) of binaural signals and the N location (s) .

[0111] For the example illustrated in FIG. 6, N is equal to 13. Table 1 illustrates 13 locations and the corresponding binaural signal pairs. Table 1

[0112] FIG. 7 illustrates another example of a relationship between the N pair (s) of binaural signals and the N location (s) .

[0113] For the example illustrated in FIG. 7, N is equal to 2.2 locations include front and rear. Corresponding to the 2 locations, there are 2 pairs of binaural signals. A binaural signal pair 1 corresponds to front, while a binaural signal pair 2 corresponds to rear.

[0114] The host device may determine the N pair (s) of binaural signals according to a reference dataset. The reference dataset may include a head related impulse response (HRIR) dataset, a binaural room impulse response (BRIR) dataset, or the like. The number of the HRIR / BRIR dataset may be equal to the number of the binaural signal pair. Each of the N pair (s) of binaural signals has a corresponding HRIR / BRIR dataset, and each of the N pair (s) of binaural signals is determined according to the corresponding HRIR / BRIR dataset. In other words, each HRIR / BRIR dataset corresponds to a location, and a binaural signal pair corresponding to the location is determined based on the first audio signal and the HRIR / BRIR dataset corresponding to the location. Taking the example illustrated in FIG. 5 as an example, there are 5 HRIR / BRIR datasets. For convenience, the 5 HRIR / BRIR datasets may be referred to as a first HRIR / BRIR dataset, a second HRIR / BRIR dataset, a third HRIR / BRIR dataset, a fourth HRIR / BRIR dataset, and a fifth HRIR / BRIR dataset. The first HRIR / BRIR dataset corresponds to the front center. The binaural signal pair 1 corresponding to the front center is determined according to the first audio signal and the first HRIR / BRIR dataset. The second HRIR / BRIR dataset corresponds to the front left, and the binaural signal pair 2 corresponding to the front left is determined according to the first audio signal and the second HRIR / BRIR dataset. Similarly, the third HRIR / BRIR dataset corresponds to the front right, the binaural signal pair 3 is determined according to the first audio signal and the third HRIR / BRIR dataset; the fourth HRIR / BRIR dataset corresponds to the rear left, and the binaural signal pair 4 is determined based on the first audio signal and the fourth HRIR / BRIR dataset; the fifth HRIR / BRIR dataset corresponds to the rear right, and the binaural signal pair 5 is determined based on the built-channel audio signal and the fifth HRIR / BRIR dataset. The host device convolves the first audio signal with the corresponding HRIR / BRIR dataset, applying the acoustic characteristics (such as spatial location of the sound, room reflections, etc. ) from the HRIR / BRIR to the first audio signal, thus achieving binaural processing and obtaining the pair of binaural signals.

[0115] 403, The host device determines M set (s) of binaural signals according to the N locations.

[0116] Each of the M set (s) of binaural signals comprises at least one pair of binaural signals. M is a positive integer and less than or equal to N.

[0117] Referring to the abovementioned examples, each location may include at least one component (also referred to as “characteristic” ) . The at least one component includes a first component that is used to indicate the location is front or rear. In some embodiments, in addition to the first component, the at least one component may further include a second component that is used to indicate the location is right, center, or right. In some other embodiments, the at least one component may further include a third component that is used to indicate the location is upper, lower, or middle. In some embodiments, a feature “middle” may be omitted.

[0118] When a binaural signal set includes at least two pairs of binaural signals, the first components of any two of the at least two pairs of binaural signals are the same. For the example illustrated in FIG. 5 as an example, the 5 pairs of binaural signals may be divided into 2 sets of binaural signals. For convenience, the 2 sets of binaural signals may be referred to as a first binaural signal set and a second binaural signal set. The first binaural signal set includes the first binaural signal pair, the second binaural signal pair, and the third binaural signal pair, while the second binaural signal set includes the fourth binaural signal pair and the fifth binaural signal pair. For the example illustrated in FIG. 6 as an example, the 13 pairs of binaural signals may be divided into 2 sets of binaural signals. Similarly, the 2 sets of binaural signals may be referred to as a first binaural signal set and a second binaural signal set. The first binaural signal set includes the biannual signal pairs 1 to 9, while the second binaural signal set includes the binaural signal pairs 10 to 13.

[0119] In some embodiments, when the location includes both the first component and the third component, the binaural signal pairs may be divided based on the first component and the third component. For example, when the binaural signal set includes two or more binaural signal pairs, the first components of any two of the binaural signal pairs are the same, and the second components of any two of the binaural signal pairs are the same. Taking the example illustrated in FIG. 6 as an example, the 13 binaural signal pairs may be divided into 4 binaural signal sets. For convenience, the 4 binaural signal sets may be referred to as a first binaural signal set, a second binaural signal set, a third binaural signal set, and a fourth binaural signal set. The first binaural set includes the binaural signal pair 1, the binaural signal pair 2 and the binaural signal pair 3, the second binaural set includes the binaural signal pair 4, the binaural signal pair 5, and the binaural signal pair 6, the third binaural signal set includes the binaural signal pair 7, the binaural signal pair 8, and the binaural signal pair 9, the fourth binaural signal set includes the binaural signal pair 10 and the binaural signal pair 11, and the fifth binaural signal set includes the binaural signal pair 12 and the binaural signal pair 13.

[0120] In some embodiments, the host device may determine one set of binaural signals. Under this case, the first component of the binaural signal pair (s) in the binaural signal set is front. For example, the host device may determine one binaural signal set according to the two pairs of binaural signals illustrated in FIG. 7, and the binaural signal set includes the binaural signal pair 1.

[0121] 404, the host device determines an encoded audio signal based on the M set (s) of binaural signals.

[0122] The encoded audio signal includes K-channel audio signals. Therefore, the encoded audio signal may be referred to as a K-channel encoded signal. K is a positive integer greater than or equal to two.

[0123] In some embodiments, the host device may perform a decorrelation operation on the M set (s) of binaural signals to obtain L processed audio signals. L is a positive integer greater than or equal to K. Then, the host device may compress the L processed audio signal to obtain the encoded audio signal.

[0124] For the decorrelation operation, the host device may use any possible method to decorrelate the M set (s) of binaural signals to obtain the L processed audio signals. For example, the possible method may be a delay, a phase shift, or the like.

[0125] In some embodiments, the host device may use a matrix encoder to compress the L processed audio signals to obtain the encoded audio signal.

[0126] In some other embodiments, the host device may use other compressing method to compress the L processed audio signal. For example, the host device may use an artificial intelligence (AI) module, or the like to compress the L processed audio signal to obtain the encoded audio signal.

[0127] In some embodiments, the encoded audio signal may be a 2-channel audio signal.

[0128] 405, The host device transmits the encoded audio signal to the client device. Correspondingly, the client device may receive the encoded audio signal from the host device.

[0129] As previously mentioned, the encoded audio signal may be transmitted via a wired link or a wireless link.

[0130] 406, The client device determines M set (s) of reconstructed binaural signals according to the encoded audio signal.

[0131] In some embodiments, the client device may determine L reconstructed audio signal and determine the M set (s) of reconstructed binaural signals according to the L reconstructed audio signal.

[0132] In some embodiments, the client device may determine three sets of audio signals according to the encoded audio signal. For convenience, the three sets of audio signals may be referred to as a first audio signal set, a second audio signal set, and a third audio signal set. The first audio signal set includes two positively correlated audio signals with positive correlation, the second audio signal set includes two low correlated audio signals, and the third audio signal set includes two negatively correlated audio signals. Then, the client may determine the L reconstructed audio signal. An operation for determining the L reconstructed audio signal is an inverse operation for obtaining the encoded audio signal. For example, when the encoded audio signal is obtained by using the matrix encoder, the L reconstructed audio signal may be obtained by using a matrix decoder. For example, a first matrix decoder may be used to decode audio signals of the first audio signal set to obtain a plurality of first target audio signals, a second matrix decoder may be used to decode audio signals of the second audio signal set to obtain a plurality of second target audio signals, and a third matrix decoder may be used to decode audio signals of the third audio signal set to obtain a plurality of third target audio signals. Then, the client device may determine the L reconstructed audio signals according to the first target audio signals, the second target audio signals and the third audio signals. Each of the reconstructed audio signal is obtained by combining two target audio signals (e.g., one first target audio signal and one second target audio signal, one second target audio and one third target audio signal, or one firs target audio signal and one third target audio signal) and performing an inverse time-frequency transform on a combination of the two target audio signals. After obtaining the L reconstructed audio signals, the client device may perform an inverse decorrelation operation on the L reconstructed audio signals to obtain the M set (s) of reconstructed binaural signals.

[0133] Step 406 is an inverse of the step 404. According to the previous embodiments, when the host device compresses the L processed audio signals by using the matrix encoder, the client device may use the matrix decoder to obtain the M set (s) of reconstructed binaural signals; when the host device performs the decorrelation operation to obtain the L processed audio, the client device may perform the inverse decorrelation to obtain the M set (s) of reconstructed binaural signals. Correspondingly, when the L processed audio signals are compressed by using an AI model, the client device may use a corresponding AI model to decompress the encoded audio signals to obtain the L reconstructed audio signals.

[0134] 407, The client device determines a pair of binaural signals based on tracking information obtained from a sensor unit in the client device and the M set (s) of reconstructed binaural signals.

[0135] In some embodiments, a spectral cue filter may be configured to filter the M set (s) of reconstructed binaural signals to obtain M set (s) of filtered binaural signals based on the tracking information. In some embodiments, reconstructed binaural signals belonging a same set may be added, and the added binaural signals may be filtered by the spectral cue filter. For example, a set of reconstructed binaural signals may include a left-channel front binaural signal and a right-channel front binaural signal. The left-channel front binaural signal and the right-channel front binaural signal may be added. For example, the left-channel front binaural signal added to the right-channel front binaural signal obtains a first added front binaural signal, the right-channel front binaural signal added to the left-channel front binaural signal obtains a second added front binaural signal. The spectral cue filter filters the first added front binaural signal and the second added front binaural signal to obtain a set of filtered binaural signal based on the tracking information where the set of filtered binaural signal corresponds to the set of reconstructed binaural signals including the left-channel front binaural signal and the right-channel front binaural signal.

[0136] The spectral cue filter holds the difference spectrum information of the binaural signal, HRIR or BRIR which is standardized at a certain angle for each group. For example, when M = 2, the standardized the reference angles for the spectral cue filter may be the centers of the front and the rear. A spectral cue filter database may consist data for all horizontal, vertical, and all position angles for the head movement. That is, the spectral cue filter G_m with respect to mth set of reconstructed binaural signals are designed so that

[0137]

[0138] where H denotes a minimum-phase filter that approximates HRTF corresponding to a source angle  and  is the reference source angle of the mth set of reconstructed binaural signals. m = 1, .., M.

[0139] The minimum-phase approximation aims to model the spectral cue filter G with a less degree. Furthermore, the minimum-phase approximation also makes it easier to interpolate the filter owing to decreasing the filter changes with respect to angular changes. It is known that phase information except for interaural time difference (ITD) has little effect on the localization. The spectral cue filter G with a less degree may decrease the cost of computing resources. In some embodiments, the degree of the spectral cue filter may be about 20 (e.g., 15, 17, 19, 21, 23, or the like) .

[0140] In some other embodiments, the minimum-phase filter may be replaced by a finite impulse response (FIR) filter or an infinite impulse response (IIR) filter that approximates HRTF corresponding to a source angle

[0141] In some embodiments, the ITD module is configured to delay the M set (s) of filtered binaural signals according to the source angle to obtain M set (s) of delayed binaural signals so that the ITD of the delayed binaural signals corresponds to the source angle. The delay applied to each set of binaural signals is calculated based on the following ITD model:

[0142]

[0143]

[0144] where d, c, θ, and  are the diameter of the head, the sound speed, the elevation, and azimuth angle of the source, respectively.

[0145] The source angle indicates the position information expressed in elevation (θ) and azimuth of the sound source from the center of the user's head to perform binauralization. For example,  of Front Left in (a) of FIG. 5 is (0, 330) degree and Front Right in (a) of FIG. 5 is (0, 30) degree, when the Front Center is defined (0, 0) degree with clock wise.

[0146] is “the reference source angle” of the mth set of reconstructed binaural signals. For example, in the case of 2 sets of binaural signals with the 5 pairs of binaural signals in (a) of FIG. 5, the reference source angle in the first binaural set may be (0, 0) and the second one may be (0, 180) .

[0147] The spectral cue filter and the ITD are determined by a sound source angle with the user’s head angle  as a center position when the sensor unit detects the head movement of the user. The input angle for the spectral cue filter and ITD may be  (b) of FIG. 5 and (c) of FIG. 5 illustrate an example of the user’s head angle

[0148] In some embodiments, when M is a positive integer greater than one, each of the M sets of delayed binaural signals may include a left-channel audio signal and a right-channel audio signal. In these embodiments, left-channel audio signals from the M sets of delayed binaural signals may be combined to obtain a left-channel binaural signal, while right-channel audio signals from the M sets of delayed binaural signal may be combined to obtain a right-channel binaural signal. The pair of binaural signals includes the left-channel binaural signal and the right-channel binaural signal.

[0149] In some embodiments, when M is equal to one, the left-channel audio signal of the M set (s) of delayed binaural signals is the left-channel binaural signals of the pair of binaural signals, while the right-channel audio signal of the M set (s) of delayed binaural signal is the right-channel binaural signal of the pair of binaural signals.

[0150] In some embodiments, when the sensor unit detects that the head of the user keeps unmoving, the pair of binaural signals may be directly determined according to the M set (s) of reconstructed binaural signals. In other words, the M set (s) of reconstructed binaural signals does not require spectral cues filtering and delay processing.

[0151] In some embodiments, when the sensor unit detects that the head of the user keeps unmoving, the spectral cues filter and delay processing may be performed by using predefined parameters.

[0152] 408, The client device determines an output signal according to the pair of binaural signals.

[0153] For example, a digital-analog-converter may be configured to convert the pair of binaural signals into analog signals, and the analog signals may be played through the left acoustic transducer and the right acoustic transducer of the client device.

[0154] In some embodiments, an amplifier may be configured to amplify the analog signals, and the amplified analog signals may be played through the left acoustic transducer and the right acoustic transducer of the client device.

[0155] FIG. 8 illustrates an audio signal processing procedure according to some embodiment of the present application.

[0156] Referring to FIG. 8, the host device may obtain an audio signal from an audio signal resource. A plurality of HRIR datasets is used to convolve a plurality of channels in the audio signal respectively. In some embodiments, the HRIR data sets may be replaced by BRIR datasets, HRTF datasets, or, BRTF datasets. Each of the HRIR datasets can be used to obtain a two-channel audio signal for binaural left and right channel. For example, the audio signal includes four channels, channel 1 to channel 4. Four HRIR datasets are used to convolve the four channels. The channel 1 is convolved by an HRIR dataset to obtain a binaural signal pair 1, the channel 2 is convolved by another HRIR dataset to obtain a binaural signal pair 2, and the like. In some other embodiments, BRIR datasets may be divided between HRIR and later BRIR to process separately.

[0157]

[0158] Each of the binaural signal pair corresponds to a location. Based on locations corresponding to the plurality of pairs of binaural signals, the plurality of pairs of binaural signals may be grouped into different sets. For example, the plurality of pairs of binaural signals may be divided into two sets, locations corresponding to binaural signal pairs in one of the two sets is front, locations corresponding to binaural signal pairs in another set is rear. Referring to FIG. 8, it is assumed that a location corresponding to the binaural signal pair 1 is front left, a location corresponding to the binaural signal pair 2 is front right, a location corresponding to the binaural signal pair 3 is rear left, and a location corresponding to the binaural signal pair 4 is rear right. The binaural signal set 1 includes the binaural signal pair 1 and the binaural signal pair 2, while the binaural signal set 2 includes the binaural signal pair 3 and the binaural signal pair 4. The binaural signal sets may be encoded to obtain a 2-channel encoded audio signal via an encoder. The 2-channel encoded audio signal may be transmitted to the client device via a wired link or a wireless link.

[0159] The client device obtains the 2-channel encoded audio signal from the host device. The 2-channel encoded audio signal may be decoded to obtain a plurality of sets of reconstructed binaural signals according to the encoder for encoding the binaural signal sets in host device. For example, the plurality of sets of reconstructed binaural signals may be decoded to the sets in front and the sets in rear from the 2-channel encoded audio signal. A spectral cue filter and an ITD module are configured to process the reconstructed binaural signals to obtain a 2-channel binaural signals. Details about the spectral cue filter and the ITD module may refer to aforementioned embodiments, and we will not elaborate further for the sake of conciseness.

[0160] FIG. 9 depicts a procedure for obtaining the 2-channel encoded audio signal.

[0161] Referring to FIG. 9, the host device performs binauralization on an audio signal to obtain two pairs of binaural signals. One of the two pairs of binaural signals include a left-channel front binaural signal (Front Binaural Lch) and a right-channel front binaural signal (Front Binaural Rch) , while another one of the two pairs of binaural signals includes a left-channel rear binaural signal (Rear Binaural Lch) and a right-channel rear binaural signal (Rear Binaural Rch) . For this embodiment, each pair of binaural signals may be regarded as a set of binaural signals, since the first component “front” only has one corresponding binaural signal pair and the first component “rear” only has one corresponding binaural signal pair. In other words, the two pairs of binaural signals are two sets of binaural signals.

[0162] A first decorrelator is configured to delay Front Binaural Lch and the Front Binaural Rch, a second decorrelator is configured to delay Rear Binaural Lch and Rear Binaural Rch. A matrix decoder obtains the delayed signals and compresses the delayed signals into a left-channel encoded signal (Encoded Signal Lch) . and a right-channel encoded signal (Encoded Signal Rch) .

[0163] FIG. 10 depicts a procedure for determining the binaural signal.

[0164] The client device obtains Encoded Signal Lch and Encoded Signal Rch and performs a time-frequency transform operation on Encoded Signal Lch and Encoded Signal Rch. A separator is configured to separate the transformed signals based on cross-correlation and outputs three audio signal sets. Each of the three audio signal sets includes two audio signals. The three audio signal sets include a first audio signal set, a second audio signal set and a third audio signal set. There is a positive correlation between the audio signals of the first audio signal set. There is a low correlation between the audio signals of the second audio signal set. There is a negative correlation between the audio signals of the third audio signal set. A first matrix decoder is configured to process the first audio signal set. The first matrix decoder is an inverse matrix for front channels. A second matrix decoder is configured to process the second audio signal set. The second matrix decoder is a pseudoinverse of an encode matrix. A third matrix encoder is configured to process the third audio signal set. The third matrix encoder is an inverse matrix for rear channels. As illustrate in FIG. 10, the first matrix decoder determines a target audio signal T_f1 and a target audio signal T_f2 based on the first audio signal set, the second matrix decoder determines a target audio signal T_s1, a target audio signal T_s2, a target audio signal T_s3 and a target audio signal T_s4 based on the second audio signal set, and the third matrix decoder determines a target audio signal T_t1 and a target audio signal T_t2 based on the third audio signal set.

[0165] A first reconstructed audio signal R1 may be obtained by combining the target audio signal T_f1 and the target audio signal T_s1, a second reconstructed audio signal R2 may be obtained by combining the target audio signal T_f2 and the target audio signal T_s2, a third reconstructed audio signal R3 may be obtained by combining the target audio signal T_s3 and the target audio signal T_t1, and a fourth reconstructed audio signal R4 may be obtained by combining the target audio signal T_s4 and the target audio signal T_t2.

[0166] An inverse time-frequency transform and an inverse transform of decorrelation may be performed on the first reconstructed audio signal R1 to the fourth reconstructed audio signal R4 in sequence. Then, the client device may obtain a left-channel front binaural signal (Front Binaural Lch’ ) , a right-channel front binaural signal (Front Binaural Rch’ ) , a left-channel rear binaural signal (Rear Binaural Lch’ ) , and a right-channel rear binaural signal (Rear Binaural Rch’ ) . The four binaural signals may be combined, and the combined binaural signals may be input to a spectral cues (SC) filter and a SC filter. A spectral cue filter module and an ITD module may obtain the filtered binaural signals and delay the obtained binaural signals. The delayed binaural signals may be combined to obtain a pair of binaural signals, and the pair of binaural signals includes a left-channel binaural signal (Binaural Lch) and a right-channel binaural signal (Binaural Rch) .

[0167] FIG. 11 is a schematic block diagram of an electronic device 1100 according to some embodiments of the present application. Referring to FIG. 11, the electronic device 1100 includes a processing module 1101 and a transmitting module 1102. The electronic device 1100 may be aforementioned host device.

[0168] The processing module 1101 is configured to determine a first audio signal according to an audio signal resource.

[0169] The processing module 1101 is further configured to perform binauralization on the first audio signal to obtain N pair (s) of binaural signals, wherein the N pair (s) of binaural signals is in one-to-one correspondence with N location (s) , N is a positive integer greater than or equal to one.

[0170] The processing module 1101 is further configured to determine M set (s) of binaural signals according to the N location (s) , wherein each of the M set (s) of binaural signals comprises at least one pair of binaural signals, M is a positive integer and less than or equal to N.

[0171] The processing module 1101 is further configured to to determine an encoded audio signal based on the M set (s) of binaural signals, wherein the encoded audio signal is a K-channel encoded signal, wherein K is a positive integer and greater than or equal to two.

[0172] The transmitting module 1102 is configured to transmit the encoded audio signal to a client device.

[0173] Details on how to process the first audio signal may refer to the above-mentioned embodiments and will not be described here.

[0174] FIG. 12 is a schematic block diagram of an electronic device 1200 according to some embodiments of the present application. Referring to FIG. 12, the electronic device 1200 includes an obtaining module 1201 and a processing module 1202. The electronic device 1200 may be aforementioned client device.

[0175] The obtaining module 1201 is configured to obtain an encoded audio signal.

[0176] The processing module 1202 is configured to determine M set (s) of reconstructed binaural signals according to the encoded audio.

[0177] The processing module 1202 is further configured to determine a pair of binaural signals based on tracking information obtained from a sensor unit in the client device and the M set (s) of reconstructed binaural signals;

[0178] The processing module 1202 is further configured to determine an output audio signal according to the pair of binaural signals.

[0179] Details on how to process the obtained encoded audio signal may refer to the above-mentioned embodiments and will not be described here.

[0180] The present application provides a computer readable storage medium including instructions. When the instructions run on an electronic device, the electronic device is enabled to perform the aforementioned method.

[0181] The present application provides a computer readable storage medium including a data structure. The data structure includes the encoded audio signal obtained according to the aforementioned method. The data structure may be a file, a bitstream, or the like.

[0182] The present application provides a chip system. The chip system includes a communication interface and a processing circuit, and the communication interface is configured to obtain to-be-processed data, and the processing circuit is configured to process the to-be-processed data according to the aforementioned method.

[0183] The present application provides a computer program product. When the computer program product runs on an electronic device, the electronic device is enabled to perform the aforementioned method.

[0184] A person of ordinary skill in the art may be aware that, in combination with the examples described in the embodiments disclosed in this specification, units and algorithm steps can be implemented by electronic hardware or a combination of computer software and electronic hardware. Whether the functions are performed by hardware or software depends on particular applications and design constraints of the technical solutions. A person skilled in the art may use different methods to implement the described functions for each particular application, but it should not be considered that the implementation goes beyond the scope of this application.

[0185] It may be clearly understood by a person skilled in the art that, for the purpose of convenient and brief description, for a detailed working process of the foregoing system, apparatus, and unit, refer to a corresponding process in the foregoing method embodiment. Details are not described herein again.

[0186] In the several embodiments provided in this application, it should be understood that the disclosed system, apparatus, and method may be implemented in other manners. For example, the described apparatus embodiment is merely an example. For example, the unit division is merely logical function division and may be other division in actual implementation. For example, a plurality of units or components may be combined or integrated into another system, or some features may be ignored or not performed. In addition, the displayed or discussed mutual couplings or direct couplings or communication connections may be implemented through some interfaces. The indirect couplings or communication connections between the apparatuses or units may be implemented in electronic, mechanical, or other forms.

[0187] The units described as separate parts may be or may not be physically separate, and parts displayed as units may be or may not be physical units, may be located in one position, or may be distributed on a plurality of network units. Some or all of the units may be selected based on actual requirements to achieve the objectives of the solutions of the embodiments.

[0188] In addition, functional units in the embodiments of this application may be integrated into one processing unit, or each of the units may exist alone physically, or two or more units are integrated into one unit.

[0189] When the functions are implemented in a form of a software functional unit and sold or used as an independent product, the functions may be stored in a computer readable storage medium. Based on such an understanding, the technical solutions in this application essentially, or the part contributing to the prior art, or some of the technical solutions may be implemented in a form of a software product. The computer software product is stored in a storage medium, and includes several instructions for instructing a computer device (which may be a personal computer, a server, a network device, or the like) to perform all or some of the steps of the methods described in the embodiments of this application. The foregoing storage medium includes: any medium that can store program code, such as a USB flash drive, a removable hard disk, a read-only memory (Read-Only Memory, ROM) , a random access memory (Random Access Memory, RAM) , a magnetic disk, or an optical disc.

[0190] The foregoing descriptions are merely specific implementations of this application, but are not intended to limit the protection scope of this application. Any variation or replacement readily figured out by a person skilled in the art within the technical scope disclosed in this application shall fall within the protection scope of this application. Therefore, the protection scope of this application shall be subject to the protection scope of the claims.

Claims

1.An audio signal processing method, wherein comprising:determining a first audio signal according to an audio signal resource;performing binauralization on the first audio signal to obtain N pair (s) of binaural signals, wherein the N pair (s) of binaural signals is in one-to-one correspondence with N location (s) , N is a positive integer greater than or equal to one;determining M set (s) of binaural signals according to the N location (s) , wherein each of the M set (s) of binaural signals comprises at least one pair of binaural signals, M is a positive integer and less than or equal to N;determining an encoded audio signal based on the M set (s) of binaural signals, wherein the encoded audio signal is a K-channel encoded signal, wherein K is a positive integer and greater than or equal to two;transmitting the encoded audio signal to a client device.2.The method according to claim 1, wherein the determining an encoded audio signal based on the M set (s) of binaural signals, comprises:performing a decorrelation operation on the M set (s) of binaural signals to obtain L processed audio signals, wherein L is a positive integer greater than or equal to K;compressing the L processed audio signals to obtain the encoded audio signal.3.The method according to claim 2, wherein the decorrelation operation is configured to delay the M set (s) of binaural signals.4.The method according to claim 2 or 3, wherein the compressing the L processed binaural signals to obtain the encoded audio signal, comprises:compressing the L processed audio signals to obtain the encoded audio signal by using a matrix encoder, wherein the encoded audio signal is a 2-channel encoded signal, the 2-channel encoded signal comprises a left-channel encoded signal and a right-channel encoded signal.5.The method according to any one of claims 1 to 4, wherein each location among the N location (s) comprises a first component and a second component, wherein the first component is front or rear, and the second component is left, center, or right;when a set of binaural signals among the M set (s) of binaural signals comprises at least two pairs of binaural signals, two first components of two locations are the same, wherein the two locations corresponds to two pairs of binaural signals among the at least two pairs of binaural signals respectively.6.The method according to 5, wherein when M = 1, the first component of each binaural signal among the M set (s) of binaural signals is the front;when M is a positive integer greater than one, the M set (s) of binaural signals comprises M1 first set (s) of binaural signals and M2 second set (s) of binaural signals, wherein M1 and M2 are positive integer, and M1 + M2 = M, the first component of each binaural signal among the first set of binaural signals is the front, while the first component of each binaural signal among the second set of binaural signals is the rear.7.The method according to any one of claims 1 to 6, wherein the determining a first audio signal according to an audio signal resource, comprises:decoding the audio signal resource to obtain a second audio signal;rendering the second audio signal to obtain the first audio signal.8.An audio signal processing method, wherein comprising:obtaining an encoded audio signal;determining M set (s) of reconstructed binaural signals according to the encoded audio;determining a pair of binaural signals based on tracking information obtained from a sensor unit in the client device and the M set (s) of reconstructed binaural signals;determining an output audio signal according to the pair of binaural signals.9.The method according to claim 8, wherein the determining M set (s) of reconstructed binaural signals according to the encoded audio, comprises:determining L reconstructed audio signals according the encoded audio;performing an inverse decorrelation operation on the L reconstructed audio signals to obtain the M set (s) of reconstructed binaural signals.10.The method according to claim 9, wherein the determining L reconstructed audio signals according the encoded audio, comprises:determining three sets of audio signals according to the encoded audio signal, wherein the three sets of audio signals comprises a first audio signal set, a second audio signal set and a third audio signal set, the first audio signal set comprises two positively correlated audio signals, the second audio signal set comprises two low correlated audio signals, and the third audio signal set comprises two negatively correlated audio signals;determining the L reconstructed audio signals according to the three sets of audio signals.11.The method according to claim 10, wherein the determining the L reconstructed audio signals according to the three sets of audio signals, comprises:decoding, by using a first matrix decoder, audio signals of the first audio signal set to obtain a plurality of first target audio signals;decoding, by using a second matrix decoder, audio signals of the second audio signal set to obtain a plurality of second audio binaural signals;decoding, by using a third matrix decoder, audio signals of the third audio signal set to obtain a plurality of second target audio signals;determining the L reconstructed audio signals according to the plurality of the first target audio signals, the plurality of the second target audio signals, and the plurality of the third target audio signals.12.The method according to any one of claims 8 to 11, wherein the determining a pair of binaural signals based on tracking information obtained from a sensor unit in the client device and the M set (s) of reconstructed binaural signals, comprises:filtering, by using a spectral cue filter, the M set (s) of reconstructed binaural signals to obtain M set (s) of filtered binaural signals based on the tracking information;delaying the M set (s) of filtered binaural signals to obtain M set (s) of delayed binaural signals;combining left-channel audio signals of the M set (s) of delayed binaural signals to obtain a left-channel binaural signal of the pair of binaural signals;combining right-channel audio signals of the M set (s) of delayed binaural signals to obtain a right-channel binaural signal of the pair of binaural signals.13.An electronic device, wherein comprising,a processing module, configured to determine a first audio signal according to an audio signal resource;the processing module, further configured to perform binauralization on the first audio signal to obtain N pair (s) of binaural signals, wherein the N pair (s) of binaural signals is in one-to-one correspondence with N location (s) , N is a positive integer greater than or equal to one;the processing module, further configured to determine M set (s) of binaural signals according to the N location (s) , wherein each of the M set (s) of binaural signals comprises at least one pair of binaural signals, M is a positive integer and less than or equal to N;the processing module, further configured to determine an encoded audio signal based on the M set (s) of binaural signals, wherein the encoded audio signal is a K-channel encoded signal, wherein K is a positive integer and greater than or equal to two;a transmitting module, configured to transmit the encoded audio signal to a client device.14.The electronic device according to claim 12, wherein the processing module is specifically configured to perform a decorrelation operation on the M set (s) of binaural signals to obtain L processed audio signals, wherein L is a positive integer greater than or equal to K;compress the L processed audio signals to obtain the encoded audio signal.15.The electronic device according to claim 14, wherein the decorrelation operation is configured to delay the M set (s) of binaural signals.16.The electronic device according to claim 14 or 15, wherein the processing module is specifically configured to compress the L processed audio signals to obtain the encoded audio signal by using a matrix encoder, wherein the encoded audio signal is a 2-channel encoded signal, the 2-channel encoded signal comprises a left-channel encoded signal and a right-channel encoded signal.17.The electronic device according to any one of claims 13 to 16, wherein each location among the N location (s) comprises a first component and a second component, wherein the first component is front or rear, and the second component is left, center, or right;when a set of binaural signals among the M set (s) of binaural signals comprises at least two pairs of binaural signals, two first components of two locations are the same, wherein the two locations corresponds to two pairs of binaural signals among the at least two pairs of binaural signals respectively.18.The electronic device according to claim 17, wherein when M = 1, the first component of each binaural signal among the M set (s) of binaural signals is the front;when M is a positive integer greater than one, the M set (s) of binaural signals comprises M1 first set (s) of binaural signals and M2 second set (s) of binaural signals, wherein M1 and M2 are positive integer, and M1 + M2 = M, the first component of each binaural signal among the first set of binaural signals is the front, while the first component of each binaural signal among the second set of binaural signals is the rear.19.The electronic device according to any one of claims 13 to 18, wherein the processing module is specifically configured to decode the audio signal resource to obtain a second audio signal;render the second audio signal to obtain the first audio signal.20.An electronic device, wherein comprising:an obtaining module, configured to obtain an encoded audio signal;a processing module, configured to determine M set (s) of reconstructed binaural signals according to the encoded audio;the processing module, further configured to determine a pair of binaural signals based on tracking information obtained from a sensor unit in the client device and the M set (s) of reconstructed binaural signals;the processing module, further configured to determine an output audio signal according to the pair of binaural signals.21.The electronic device according to claim 20, wherein the processing module is specifically configured to determine L reconstructed audio signals according the encoded audio;perform inverse decorrelation operation on the L reconstructed audio signals to obtain the M set (s) of reconstructed binaural signals.22.The electronic device according to claim 21, wherein the processing module is specifically configured to determine three sets of audio signals according to the encoded audio signal, wherein the three sets of audio signals comprises a first audio signal set, a second audio signal set and a third audio signal set, the first audio signal set comprises two positively correlated audio signals, the second audio signal set comprises two low correlated audio signals, and the third audio signal set comprises two negatively correlated audio signals;determine the L reconstructed audio signals according to the three sets of audio signals.23.The electronic device according to claim 22, wherein the processing module is specifically configured to decode, by using a first matrix decoder, audio signals of the first audio signal set to obtain a plurality of first target audio signals;decode, by using a second matrix decoder, audio signals of the second audio signal set to obtain a plurality of second audio binaural signals;decode, by using a third matrix decoder, audio signals of the third audio signal set to obtain a plurality of second target audio signals;determine the L reconstructed audio signals according to the plurality of the first target audio signals, the plurality of the second target audio signals, and the plurality of the third target audio signals.24.The electronic device according to any one of claims 20 to 23, wherein the processing module is specifically configured to filter, by using a spectral cue filter, the M set (s) of reconstructed binaural signals to obtain M set (s) of filtered binaural signals based on the tracking information;delay the M set (s) of filtered binaural signals to obtain M set (s) of delayed binaural signals;combine left-channel audio signals of the M set (s) of delayed binaural signals to obtain a left-channel binaural signal of the pair of binaural signals;combine right-channel audio signals of the M set (s) of delayed binaural signals to obtain a right-channel binaural signal of the pair of binaural signals.25.An electronic device, comprising a memory and a processor, wherein the memory is configured to store instructions, and the processor is configured to invoke the instructions from the memory and run the instructions, so that the electronic device performs the method according to any one of claims 1 to 7, or, any one of claims 8 to 12.26.A chip system, comprising a communication interface and a processing circuit, wherein the communication interface is configured to obtain to-be-processed data, and the processing circuit is configured to process the to-be-processed data according to the method according to any one of claims 1 to 7, or, any one of claims 8 to 12.27.A computer readable storage medium, wherein the computer readable storage medium stores instructions, and when the instructions run on an electronic device, the electronic device is enabled to perform the method according to any one of claims 1 to 7, or, any one of claims 8 to 12.28.A computer program product, wherein when the computer program product runs on an electronic device, the electronic device is enabled to perform the method according to any one of claims 1 to 7, or, any one of claims 8 to 12.29.A computer readable storage medium, wherein the computer readable storage medium stores a data structure comprising the encoded audio signal obtained according to any one of claims 1 to 7.