Audio Processing Method, Apparatus, Wireless Headphone, and Storage Medium

By independently receiving and rendering audio signals in wireless headphones, using headphone sensors and head-related transformation function databases, the delay and rendering imbalance of wireless headphones in high-quality sound effects are solved, and high-quality binaural sound field effect is achieved.

CN111918176BActive Publication Date: 2025-07-04WAVARTS TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010762073.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-07-31
Publication Date
2025-07-04
Estimated Expiration
2040-07-31

AI Technical Summary

Technical Problem

Due to data transmission delay and unbalanced rendering, existing wireless headphones cannot meet the requirements of high-quality sound effects, especially in high-standard surround sound and all-round immersive three-dimensional panoramic sound effects.

Method used

By setting the first and second wireless headphones in the wireless headphones to receive and render audio signals respectively, the headphone sensor and head-related transformation function database are used to perform audio processing, reducing dependence on the playback device and realizing independent rendering.

Benefits of technology

It greatly reduces delay, improves the sound quality of the headphones, achieves a binaural sound field effect with nearly 0 delays, and improves the sound performance of wireless headphones.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN111918176B_ABST
    Figure CN111918176B_ABST
Patent Text Reader

Abstract

The present invention provides an audio processing method, apparatus, wireless earphone and storage medium. The first wireless earphone receives a first audio signal to be presented sent by a playback device, and the second wireless earphone receives a second audio signal to be presented sent by the playback device. Then, the first wireless earphone performs rendering processing on the first audio signal to be presented to obtain a first playback audio signal, and the second wireless earphone performs rendering processing on the second audio signal to be presented to obtain a second playback audio signal. Finally, the first wireless earphone plays the first playback audio signal, and the second wireless earphone plays the second playback audio signal. Thus, the technical effect that the wireless earphone can render the audio signal independently without relying on the playback device is achieved, thereby greatly reducing the delay and improving the audio quality of the earphone.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of electronic technology, and in particular to an audio processing method, device, wireless headset and storage medium. Background Art

[0002] With the development of smart mobile devices, headphones have become a must-have for people to listen to sounds in their daily lives. Wireless headphones are becoming more and more popular in the market due to their convenience, and have even gradually become mainstream headphone products. As a result, people's requirements for sound quality are getting higher and higher. Not only are they gradually pursuing lossless sound quality, but they are also gradually pursuing a sense of space and immersion in sound, from the initial mono and stereo to now more pursuing 360° surround sound and truly immersive three-dimensional panoramic sound.

[0003] At present, existing wireless headphones, such as traditional wireless Bluetooth headphones and TWS true wireless headphones, transmit head movement information to the playback device for processing. This method has a large data transmission delay compared to the high standard requirements of high-quality surround sound or all-round immersive three-dimensional panoramic sound effects, resulting in unbalanced rendering between the two headphones, or poor real-time rendering effect, causing the rendered sound effect to fail to meet the ideal high-quality requirements.

[0004] Therefore, the existing wireless headsets have a technical problem that the data interaction with the playback terminal cannot meet the requirements of high-quality sound effects. Summary of the invention

[0005] The present invention provides an audio processing method, device, wireless headset and storage medium to solve the technical problem that the data interaction between the existing wireless headset and the playback device cannot meet the requirements of high-quality sound effects.

[0006] In a first aspect, the present invention provides an audio processing method, which is applied to a wireless headset, wherein the wireless headset comprises a first wireless headset and a second wireless headset, wherein the first wireless headset and the second wireless headset are used to establish a wireless connection with a playback device; the method comprises:

[0007] The first wireless headset renders the first audio signal to be presented to obtain a first playback audio signal, and the second wireless headset renders the second audio signal to be presented to obtain a second playback audio signal;

[0008] The first wireless headset plays the first playback audio signal, and the second wireless headset plays the second playback audio signal.

[0009] In a possible design, if the first wireless earphone is a left ear wireless earphone and the second wireless earphone is a right ear wireless earphone, the first playback audio signal is used to present a left ear audio effect, and the second playback audio signal is used to present a right ear audio effect, so as to form a binaural sound field when the first wireless earphone plays the first playback audio signal and the second wireless earphone plays the second playback audio signal.

[0010] In a possible design, before the first wireless earphone performs rendering processing on the first audio signal to be presented, it further includes:

[0011] The first wireless earphone performs decoding processing on the first audio signal to be presented to obtain a first decoded audio signal;

[0012] Correspondingly, the first wireless earphone performing rendering processing on the first audio signal to be presented includes:

[0013] The first wireless earphone performs rendering processing according to the first decoded audio signal and rendering metadata to obtain the first playback audio signal; and

[0014] Before the second wireless earphone performs rendering processing on the second audio signal to be presented, it further includes:

[0015] The second wireless earphone performs decoding processing on the second audio signal to be presented to obtain a second decoded audio signal;

[0016] Correspondingly, the second wireless earphone performing rendering processing on the second audio signal to be presented includes:

[0017] The second wireless earphone performs rendering processing according to the second decoded audio signal and rendering metadata to obtain the second playback audio signal.

[0018] In a possible design, the rendering metadata includes at least one of first wireless earphone metadata, second wireless earphone metadata, and playback device metadata.

[0019] In a possible design, the first wireless earphone metadata includes first earphone sensor metadata and a head-related transfer function (HRTF) database, where the first earphone sensor metadata is used to characterize the motion characteristics of the first wireless earphone;

[0020] The second wireless earphone metadata includes second earphone sensor metadata and a head-related transfer function (HRTF) database, where the second earphone sensor metadata is used to characterize the motion characteristics of the second wireless earphone;

[0021] The playback device metadata includes playback device sensor metadata, where the playback device sensor metadata is used to characterize the motion characteristics of the playback device.

[0022] In a possible design, before performing the rendering process, it further includes:

[0023] The first wireless earphone synchronizes the rendering metadata with the second wireless earphone.

[0024] In a possible design, if the first wireless earphone is provided with an earphone sensor, the second wireless earphone is not provided with an earphone sensor, and the playback device is not provided with a playback device sensor, then the first wireless earphone synchronizes the rendering metadata with the second wireless earphone, including:

[0025] The first wireless earphone sends the first earphone sensor metadata to the second wireless earphone, and the second wireless earphone uses the first earphone sensor metadata as the second earphone sensor metadata.

[0026] In a possible design, if both the first wireless earphone and the second wireless earphone are provided with earphone sensors, and the playback device is not provided with a playback device sensor, then the first wireless earphone synchronizes the rendering metadata with the second wireless earphone, including:

[0027] The first wireless earphone sends the first earphone sensor metadata to the second wireless earphone, and the second wireless earphone sends the second earphone sensor metadata to the first wireless earphone;

[0028] The first wireless earphone and the second wireless earphone respectively determine the rendering metadata according to the first earphone sensor metadata, the second earphone sensor metadata, and a preset numerical algorithm; or,

[0029] The first wireless earphone sends the first earphone sensor metadata to the playback device, and the second wireless earphone sends the second earphone sensor metadata to the playback device, so that the playback device determines the rendering metadata according to the first earphone sensor metadata, the second earphone sensor metadata, and a preset numerical algorithm;

[0030] The first wireless earphone and the second wireless earphone respectively receive the rendering metadata.

[0031] In a possible design, if the first wireless earphone is provided with an earphone sensor, the second wireless earphone is not provided with an earphone sensor, and the playback device is provided with a playback device sensor, then the first wireless earphone synchronizes the rendering metadata with the second wireless earphone, including:

[0032] The first wireless earphone sends the first earphone sensor metadata to the playback device, so that the playback device determines the rendering metadata according to the first earphone sensor metadata, the playback device sensor metadata, and a preset numerical algorithm;

[0033] The first wireless earphone and the second wireless earphone respectively receive the rendering metadata; or,

[0034] The first wireless earphone receives the playback device sensor metadata sent by the playback device;

[0035] The first wireless earphone determines the rendering metadata according to the first earphone sensor metadata, the playback device sensor metadata, and a preset numerical algorithm;

[0036] The first wireless earphone sends the rendering metadata to the second wireless earphone.

[0037] In a possible design, if earphone sensors are provided on both the first wireless earphone and the second wireless earphone, and a playback device sensor is provided on the playback device, the synchronization of the rendering metadata by the first wireless earphone and the second wireless earphone includes:

[0038] The first wireless earphone sends the first earphone sensor metadata to the playback device, and the second wireless earphone sends the second earphone sensor metadata to the playback device, so that the playback device determines the rendering metadata according to the first earphone sensor metadata, the second earphone sensor metadata, the playback device sensor metadata, and a preset numerical algorithm;

[0039] The first wireless earphone and the second wireless earphone respectively receive the rendering metadata; or,

[0040] The first wireless earphone sends the first earphone sensor metadata to the second wireless earphone, and the second wireless earphone sends the second earphone sensor metadata to the first wireless earphone;

[0041] The first wireless earphone and the second wireless earphone respectively receive the playback device sensor metadata;

[0042] The first wireless earphone and the second wireless earphone respectively determine the rendering metadata according to the first earphone sensor metadata, the second earphone sensor metadata, the playback device sensor metadata, and a preset numerical algorithm.

[0043] Optionally, the earphone sensor includes at least one of a gyroscope sensor, a head size sensor, a ranging sensor, a geomagnetic sensor, and an acceleration sensor; and / or,

[0044] The playback device sensor includes at least one of a gyroscope sensor, a head size sensor, a ranging sensor, a geomagnetic sensor, and an acceleration sensor.

[0045] Optionally, the first audio signal to be presented includes at least one of a channel-based audio signal, an object-based audio signal, and a scene-based audio signal; and / or,

[0046] The second audio signal to be presented includes at least one of a channel-based audio signal, an object-based audio signal, and a scene-based audio signal.

[0047] Optionally, the wireless connection includes: Bluetooth connection, infrared connection, WIFI connection, LIFI visible light connection.

[0048] In a second aspect, the present invention provides an audio processing apparatus, including:

[0049] A first audio processing apparatus and a second audio processing apparatus;

[0050] The first audio processing apparatus includes:

[0051] A first receiving module, configured to receive a first audio signal to be presented sent by a playback device;

[0052] A first rendering module, configured to perform rendering processing on the first audio signal to be presented to obtain a first playback audio signal;

[0053] A first playback module, configured to play the first playback audio signal;

[0054] The second audio processing apparatus includes:

[0055] A second receiving module, configured to receive a second audio signal to be presented sent by the playback device;

[0056] A second rendering module, configured to perform rendering processing on the second audio signal to be presented to obtain a second playback audio signal;

[0057] A second playback module, configured to play the second playback audio signal.

[0058] In a possible design, the first audio processing device is a left-ear audio processing device, and the second audio processing device is a right-ear audio processing device. Then, the first playback audio signal is used to present a left-ear audio effect, and the second playback audio signal is used to present a right-ear audio effect, so as to form a binaural sound field when the first audio processing device plays the first playback audio signal and the second audio processing device plays the second playback audio signal.

[0059] In a possible design, the first audio processing device further includes:

[0060] A first decoding module, configured to perform decoding processing on the first audio signal to be presented to obtain a first decoded audio signal;

[0061] The first rendering module is specifically configured to: perform rendering processing according to the first decoded audio signal and rendering metadata to obtain the first playback audio signal;

[0062] The second audio processing device further includes:

[0063] A second decoding module, configured to perform decoding processing on the second audio signal to be presented to obtain a second decoded audio signal;

[0064] The second rendering module is specifically configured to: perform rendering processing according to the second decoded audio signal and rendering metadata to obtain the second playback audio signal.

[0065] In a possible design, the rendering metadata includes at least one of first wireless headphone metadata, second wireless headphone metadata, and playback device metadata.

[0066] In a possible design, the first wireless headphone metadata includes first headphone sensor metadata and a head-related transfer function (HRTF) database, where the first headphone sensor metadata is used to characterize the motion characteristics of the first wireless headphone;

[0067] The second wireless headphone metadata includes second headphone sensor metadata and a head-related transfer function (HRTF) database, where the second headphone sensor metadata is used to characterize the motion characteristics of the second wireless headphone;

[0068] The playback device metadata includes playback device sensor metadata, where the playback device sensor metadata is used to characterize the motion characteristics of the playback device.

[0069] In a possible design, the first audio processing device further includes:

[0070] A first synchronization module for synchronizing the rendering metadata with the second wireless earphone; and / or,

[0071] The second audio processing device further includes:

[0072] A second synchronization module for synchronizing the rendering metadata with the first wireless earphone.

[0073] In a possible design, the first synchronization module is specifically configured to: send the first earphone sensor metadata to the second wireless earphone, so that the second synchronization module uses the first earphone sensor metadata as the second earphone sensor metadata.

[0074] In a possible design, the first synchronization module is specifically configured to:

[0075] Send the first earphone sensor metadata;

[0076] Receive the second earphone sensor metadata;

[0077] Determine the rendering metadata according to the first earphone sensor metadata, the second earphone sensor metadata, and a preset numerical algorithm;

[0078] The second synchronization module is specifically configured to:

[0079] Send the second earphone sensor metadata;

[0080] Receive the first earphone sensor metadata;

[0081] Determine the rendering metadata according to the first earphone sensor metadata, the second earphone sensor metadata, and a preset numerical algorithm; or,

[0082] The first synchronization module is specifically configured to:

[0083] Send the first earphone sensor metadata;

[0084] Receive the rendering metadata;

[0085] The second synchronization module is specifically configured to:

[0086] Send the second earphone sensor metadata;

[0087] Receive the rendering metadata.

[0088] In a possible design, the first synchronization module is specifically configured to:

[0089] Receive playback device sensor metadata;

[0090] Determine the rendering metadata according to the first headphone sensor metadata, the playback device sensor metadata, and a preset numerical algorithm;

[0091] Send the rendering metadata.

[0092] In a possible design, the first synchronization module is specifically configured to:

[0093] Send the first headphone sensor metadata;

[0094] Receive the second headphone sensor metadata;

[0095] Receive the playback device sensor metadata;

[0096] Determine the rendering metadata according to the first headphone sensor metadata, the second headphone sensor metadata, the playback device sensor metadata, and a preset numerical algorithm;

[0097] The second synchronization module is specifically configured to:

[0098] Send the second headphone sensor metadata;

[0099] Receive the first headphone sensor metadata;

[0100] Receive the playback device sensor metadata;

[0101] Determine the rendering metadata according to the first headphone sensor metadata, the second headphone sensor metadata, the playback device sensor metadata, and a preset numerical algorithm.

[0102] Optionally, the first audio signal to be presented includes at least one of a channel-based audio signal, an object-based audio signal, and a scene-based audio signal; and / or,

[0103] The second audio signal to be presented includes at least one of a channel-based audio signal, an object-based audio signal, and a scene-based audio signal.

[0104] In a third aspect, the present invention provides a pair of wireless headphones, including:

[0105] A first wireless headphone and a second wireless headphone;

[0106] The first wireless headphone includes:

[0107] A first processor; and

[0108] A first memory for storing the computer program of the processor;

[0109] Among them, the processor is configured to implement the steps of the first wireless earphone in any possible audio processing method in the first aspect by executing the computer program;

[0110] The second wireless earphone includes:

[0111] A second processor; and

[0112] A second memory for storing the computer program of the processor;

[0113] Among them, the processor is configured to implement the steps of the second wireless earphone in any possible audio processing method in the first aspect by executing the computer program.

[0114] In a fourth aspect, the present invention further provides a storage medium, in which a computer program is stored, and the computer program is used to execute any possible audio processing method provided in the first aspect.

[0115] The present invention provides an audio processing method, device, wireless earphone and storage medium. The first wireless earphone receives a first audio signal to be presented sent by a playback device, and the second wireless earphone receives a second audio signal to be presented sent by the playback device; then the first wireless earphone performs rendering processing on the first audio signal to be presented to obtain a first playback audio signal, and the second wireless earphone performs rendering processing on the second audio signal to be presented to obtain a second playback audio signal; finally, the first wireless earphone plays the first playback audio signal, and the second wireless earphone plays the second playback audio signal. Thus, the technical effect that the wireless earphone can render the audio signal by itself without relying on the playback device is achieved, thereby greatly reducing the delay and improving the audio quality of the earphone. BRIEF DESCRIPTION OF THE DRAWINGS

[0116] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for describing the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0117] Figure 1 It is a schematic structural diagram of a wireless earphone shown according to an exemplary embodiment of the present invention;

[0118] Figure 2 It is a schematic application scenario diagram of an audio processing method shown according to an exemplary embodiment of the present invention;

[0119] Figure 3 It is a schematic flowchart of an audio processing method shown according to an exemplary embodiment of the present invention;

[0120] Figure 4 A schematic diagram of a data link for audio signal processing provided by an embodiment of the present invention;

[0121] Figure 5 A schematic diagram of a HRTF rendering method provided by an embodiment of the present invention;

[0122] Figure 6 A schematic diagram of another HRTF rendering method provided by an embodiment of the present invention;

[0123] Figure 7 A schematic diagram of an application scenario where multiple pairs of wireless earphones are connected to a playback device provided by an embodiment of the present invention;

[0124] Figure 8 A schematic diagram of the structure of an audio processing device provided by an embodiment of the present invention;

[0125] Figure 9 A schematic diagram of the structure of a wireless earphone provided by an embodiment of the present invention. Detailed implementation manners

[0126] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without making creative efforts shall fall within the protection scope of the present invention.

[0127] The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein, for example, can be implemented in an order different from those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0128] The following uses specific embodiments to elaborate in detail on the technical solutions of the present invention and how the technical solutions of this application solve the above technical problems. These several specific embodiments below can be combined with each other, and for the same or similar concepts or processes, they may not be repeated in some embodiments. The following will describe the embodiments of the present invention in conjunction with the accompanying drawings.

[0129] Figure 1 FIG. is a schematic structural diagram of a wireless earphone shown according to an exemplary embodiment of the present invention. Figure 2 FIG. is a schematic diagram of an application scenario of an audio processing method shown according to an exemplary embodiment of the present invention. As Figure 1 - Figure 2 shown, the wireless transceiver device group communication method provided in this embodiment is applied to the wireless earphone 10. Among them, the wireless earphone 10 includes a first wireless earphone 101 and a second wireless earphone 102, and the wireless transceiver devices in the wireless earphone 10 are communicatively connected through a first wireless link 103. It is worth noting that the communication connection between the wireless earphone 101 and the wireless earphone 102 in the wireless earphone 10 can be bidirectional or unidirectional, and no specific limitation is made in this embodiment. In addition, it is understandable that the above wireless earphone 10 and the playback device 20 can be wireless transceiver devices that communicate according to a standard wireless protocol. Among them, the standard wireless protocol can be a Bluetooth protocol, a Wifi protocol, a Lifi protocol, an infrared wireless transmission protocol, etc. In this embodiment, the specific form of its wireless protocol is not limited. In order to specifically illustrate the application scenario of the wireless connection method provided in this embodiment, it can be exemplified that the standard wireless protocol can be the Bluetooth protocol. Here, the wireless earphone 10 can be a TWS (True Wireless Stereo) true wireless earphone or a traditional Bluetooth earphone, etc.

[0130] Figure 3 FIG. is a schematic flowchart of an audio processing method shown according to an exemplary embodiment of the present invention. As Figure 3 shown, the audio processing method provided in this embodiment is applied to a wireless earphone. The wireless earphone includes a first wireless earphone and a second wireless earphone. The method includes:

[0131] S301. The first wireless earphone receives a first audio signal to be presented sent by the playback device, and the second wireless earphone receives a second audio signal to be presented sent by the playback device.

[0132] In this step, the playback device sends the first audio signal to be presented and the second audio signal to be presented to the first wireless earphone and the second wireless earphone respectively.

[0133] It can be understood that in this embodiment, the wireless connection includes: Bluetooth connection, infrared connection, WIFI connection, LIFI visible light connection.

[0134] Optionally, if the first wireless earphone is a left ear wireless earphone and the second wireless earphone is a right ear wireless earphone, the first playback audio signal is used to present a left ear audio effect, and the second playback audio signal is used to present a right ear audio effect, so as to form a binaural sound field when the first wireless earphone plays the first playback audio signal and the second wireless earphone plays the second playback audio signal.

[0135] It should be noted that the first audio signal to be presented and the second audio signal to be presented are two audio signals that can form a complete binaural sound field in terms of audio signal characteristics, or can form a stereo surround sound or a three-dimensional panoramic sound, after the original audio signal is distributed according to a preset distribution model.

[0136] The first audio signal to be presented or the second audio signal to be presented includes scene information such as the number of microphones for collecting HOA / FOA signals, the order of HOA, and the type of HOA virtual sound field. It should be noted that when the first audio signal to be presented or the second audio signal to be presented is an audio signal based on channels or "channels + objects", if the first audio signal to be presented or the second audio signal to be presented contains control signals that do not require subsequent binaural processing, the corresponding channels are directly assigned to the left earphone or the right earphone, that is, the first wireless earphone or the second wireless earphone, according to the instructions. It should also be noted that the first audio signal to be presented and the second audio signal to be presented are both unprocessed signals, while the existing technologies are generally processed signals; in addition, the first audio signal to be presented and the second audio signal to be presented can be the same or different.

[0137] When the first audio signal to be presented or the second audio signal to be presented is an audio signal of other types, such as "stereo + object", the first audio signal to be presented and the second audio signal to be presented need to be sent to the first wireless earphone and the second wireless earphone at the same time. If the above stereo two-channel signal control instruction indicates that the two-channel signal does not require further subsequent binaural processing, the left-channel compressed audio signal, that is, the first audio signal to be presented, is transmitted to the left earphone end, that is, the first wireless earphone, and the right-channel compressed audio signal, that is, the second audio signal to be presented, is transmitted to the right earphone end, that is, the second wireless earphone. The object information still needs to be transmitted to the processing units at both the left and right earphone ends, and finally the playback signals provided to the first wireless earphone and the second wireless earphone are the mixture of the object-rendered signals and the corresponding channel signals.

[0138] It should be noted that in a possible design, the first audio signal to be presented includes at least one of an audio signal based on channels, an audio signal based on objects, and an audio signal based on scenes; and / or,

[0139] The second audio signal to be presented includes at least one of a channel-based audio signal, an object-based audio signal, and a scene-based audio signal.

[0140] It should also be noted that the first audio signal to be presented or the second audio signal to be presented includes metadata information that determines how audio is presented in a specific playback scene, or is related to such metadata information. When the first audio signal to be presented or the second audio signal to be presented is a channel-based audio signal, the first audio signal to be presented or the second audio signal to be presented.

[0141] Further optionally, the playback device can re-encode the rendered audio data and the rendered metadata, and output the encoded audio stream as the audio signal to be presented and transmit it wirelessly to the wireless headset.

[0142] S302. The first wireless headset performs a rendering process on the first audio signal to be presented to obtain a first playback audio signal, and the second wireless headset performs a rendering process on the second audio signal to be presented to obtain a second playback audio signal.

[0143] In this step, the first wireless headset and the second wireless headset respectively perform a rendering process on the first audio signal to be presented and the second audio signal to be presented received by each of them, so as to obtain a first playback audio signal and a second playback audio signal.

[0144] Optionally, before the first wireless headset performs a rendering process on the first audio signal to be presented, it further includes:

[0145] The first wireless headset performs a decoding process on the first audio signal to be presented to obtain a first decoded audio signal;

[0146] Correspondingly, the first wireless headset performing a rendering process on the first audio signal to be presented includes:

[0147] The first wireless headset performs a rendering process according to the first decoded audio signal and the rendering metadata to obtain the first playback audio signal; and

[0148] Before the second wireless headset performs a rendering process on the second audio signal to be presented, it further includes:

[0149] The second wireless headset performs a decoding process on the second audio signal to be presented to obtain a second decoded audio signal;

[0150] Correspondingly, the second wireless headset performing a rendering process on the second audio signal to be presented includes:

[0151] The second wireless earphone performs rendering processing according to the second decoded audio signal and rendering metadata to obtain the second playback audio signal.

[0152] It can be understood that some of the signals to be presented transmitted from the playback device to the wireless earphone can be directly rendered without decoding, while some are compressed bitstreams that need to be decoded before rendering processing.

[0153] To specifically illustrate the above rendering processing, the following will be combined with Figure 4 to specifically illustrate.

[0154] Figure 4 It is a schematic diagram of a data link for audio signal processing provided by an embodiment of the present invention. As Figure 4 shown, the signal S0 to be presented output by the playback device includes two parts, the first signal S01 to be presented and the second signal S02 to be presented, which are respectively received by the first wireless earphone and the second wireless earphone, and then the first wireless earphone and the second wireless earphone decode them respectively to obtain the first decoded audio signal S1 and the second decoded audio signal S2.

[0155] It should be noted that the first signal S01 to be presented and the second signal S02 to be presented can be the same, different, or partially overlapped, but the first signal S01 to be presented and the second signal S02 to be presented can be combined into the signal S0 to be presented.

[0156] Specifically, the first signal to be presented or the second signal to be presented includes audio signals based on channels, such as AAC / AC3 bitstreams, etc., audio signals based on objects, such as ATMOS / MPEG-H bitstreams, etc., audio signals based on scenes, such as MPEG-H HOA bitstreams, or audio signals that are any combination of the above 3 types, such as WANOS bitstreams.

[0157] When the first signal to be presented or the second signal to be presented is an audio signal based on channels, such as AAC / AC3 bitstreams, etc., the audio bitstream is fully decoded to obtain the audio content signals of each channel, and channel characteristic information such as: sound field type, sampling rate, bit rate, etc., and also includes control instructions such as whether binaural processing is required.

[0158] When the first signal to be presented or the second signal to be presented is an audio signal based on objects, such as ATMOS / MPEG-H bitstreams, etc., after decoding the audio signal, the audio content signals of each channel, and channel characteristic information, such as sound field type, sampling rate, bit rate, etc., are obtained, and the audio content signals of the objects, and the metadata of the objects, such as the size of the objects, three-dimensional space information, etc., are obtained.

[0159] When the first audio signal to be presented or the second audio signal to be presented is a scene-based audio signal, such as an MPEG-H HOA bitstream, the audio bitstream is fully decoded to obtain the audio content signals of each channel and channel characteristic information, such as sound field type, sampling rate, bit rate, etc.

[0160] When the first audio signal to be presented or the second audio signal to be presented is a bitstream based on the above three signals, such as a WANOS bitstream, the audio bitstream is decoded according to the decoding description of the bitstreams of the above three signals to obtain the audio content signals of each channel and channel characteristic information, such as sound field type, sampling rate, bit rate, etc., and obtain the audio content signal of the object and the metadata of the object, such as the size of the object, three-dimensional space information, etc.

[0161] Next, as Figure 4 shown, the first wireless earphone uses the first decoded audio signal and the rendering metadata D3 to perform a rendering operation, thereby obtaining the first playback audio signal. Similarly, the second wireless earphone uses the first decoded audio signal and the rendering metadata D5 to perform a rendering operation, thereby obtaining the second playback audio signal. And the first playback audio signal and the second playback audio signal are not separated, but are closely related according to the distribution of the audio signal to be presented and the associated parameters used in the rendering process, such as the HRTF (Head Related Transfer Function) database. It should be noted that those skilled in the art can select the associated parameters according to the actual situation, and the associated parameters can also be an association algorithm, which is not limited in this application.

[0162] The first playback audio signal and the second playback audio signal, which have an inseparable relationship, can form a complete three-dimensional stereo binaural sound field after being played by wireless earphones such as TWS true wireless earphones, so as to achieve a binaural sound field with almost 0 delay without the need for too much participation of the playback device in the rendering, which can greatly improve the sound quality of the earphone playback.

[0163] During the rendering process, during the rendering process of the first playback audio signal, the first decoded audio signal and the rendering metadata D3 play a very important role in the entire rendering process. Similarly, during the rendering process of the second playback audio signal, the second decoded audio signal and the rendering metadata D5 play a very important role in the entire rendering process.

[0164] To facilitate the description that the first wireless earphone and the second wireless earphone are still associated rather than isolated during rendering, the following is combined with Figure 5 and Figure 6Illustrate the implementation methods of synchronous rendering of two first wireless earphones and second wireless earphones by way of examples. The so-called synchronization does not mean simultaneous, but rather coordination with each other to achieve the best rendering effect.

[0165] It should be noted that the first decoded audio signal and the second decoded audio signal may include, but are not limited to, audio content signals of channels, audio content signals of objects, and / or scene content audio signals; the metadata may include, but is not limited to, channel characteristic information, such as sound field type, sampling rate, bit rate, etc., three-dimensional spatial information of objects, and rendering metadata at the earphone end. For example, it may include, but is not limited to, sensor metadata and HRTF databases. Since scene content audio signals such as FOA / HOA can be regarded as special spatially structured channel signals, the following rendering of channel information is equally applicable to scene content audio signals.

[0166] Figure 5 It is a schematic diagram of an HRTF rendering method provided by an embodiment of the present invention. As Figure 5 shown, when the input first decoded audio signal and second decoded audio signal are audio signals regarding channel information, as Figure 5 shown, the specific rendering process is as follows:

[0167] The audio receiving unit 301 receives the incoming left headphone channel information D31 and the content S31(i), i.e., the first decoded audio signal, where 1 ≤ i ≤ N and N is the number of channels received by the left headphone; the audio receiving unit 302 receives the incoming right headphone channel information D32 and the content S32(j), i.e., the second decoded audio signal, where 1 ≤ j ≤ M and M is the number of channels received by the right headphone. The information S31(i) and S32(j) can be completely the same or partially the same. Among them, S31(i) includes the signal S37(i1) to be processed by HRTF filtering, where 1 ≤ i1 ≤ N1 ≤ N and N1 represents the number of channels of the left headphone that need to be processed by HRTF filtering; it can also include S35(i2) that does not require filtering, where 1 ≤ i2 ≤ N2 and N2 represents the number of channels of the left headphone that do not require HRTF filtering, and N2 = N - N1. S32(j) includes the signal S38(j1) to be processed by HRTF filtering, where 1 ≤ j1 ≤ M1 ≤ M and M1 represents the number of channels of the right headphone that need to be processed by HRTF filtering; it can also include S36(j2) that does not require filtering, where 1 ≤ j2 ≤ M2 and M2 represents the number of channels of the right headphone that do not require HRTF filtering, and M2 = M - M1. In theory, N2 can be 0, indicating that there is no channel signal S35 that does not require HRTF filtering for the left ear; similarly, M2 can also be 0, indicating that there is no channel signal S36 that does not require HRTF filtering for the right ear; N2 and M2 can be equal or not equal; and for the channels that need to be processed by HRTF filtering, they must be the same, i.e., N1 = M1, and the corresponding signal contents also need to be the same, i.e., S37 = S38, where S37 is the collection of the signals S37(i1) that need to be filtered for the left ear, and similarly, S38 is the collection of the signals S38(j1) that need to be filtered for the right ear. In addition, the audio receiving units 301 and 302 respectively transmit the channel characteristic information D31 and D32 to the three-dimensional space coordinate construction units 303 and 304.

[0168] After receiving their respective channel information, the space coordinate construction units 303 and 304 construct the three-dimensional space position distributions (X1(i1), Y1(i1), Z1(i1)) and (X2(j1), Y2(j1), Z2(j1)) of each channel, and then transmit the space positions of each channel to the space coordinate conversion units 307 and 308 respectively.

[0169] The metadata unit 305 provides the rendering metadata for the left ear to the entire rendering system, which may include the sensor metadata sensor33 (transmitted to 307) and the HRTF database Data_L for the left ear (transmitted to the filtering processing unit 309); similarly, the metadata unit 306 provides the rendering metadata for the right ear to the entire rendering system, which may include the sensor metadata sensor34 (transmitted to 308) and the HRTF database Data_R for the right ear (transmitted to the filtering processing unit 310). Among them, before transmitting the sensor metadata sensor33 and sensor34 to 307 and 308 respectively, it is necessary to synchronize the sensor metadata.

[0170] In a possible design, before performing the rendering process, it further includes:

[0171] The first wireless earphone synchronizes the rendering metadata with the second wireless earphone.

[0172] Optionally, if the first wireless earphone is provided with an earphone sensor, the second wireless earphone is not provided with an earphone sensor, and the playback device is not provided with a playback device sensor, then the first wireless earphone synchronizing the rendering metadata with the second wireless earphone includes:

[0173] The first wireless earphone sends the first earphone sensor metadata to the second wireless earphone, and the second wireless earphone uses the first earphone sensor metadata as the second earphone sensor metadata.

[0174] In another possible design, if both the first wireless earphone and the second wireless earphone are provided with earphone sensors and the playback device is not provided with a playback device sensor, then the first wireless earphone synchronizing the rendering metadata with the second wireless earphone includes:

[0175] The first wireless earphone sends the first earphone sensor metadata to the second wireless earphone, and the second wireless earphone sends the second earphone sensor metadata to the first wireless earphone;

[0176] The first wireless earphone and the second wireless earphone respectively determine the rendering metadata according to the first earphone sensor metadata, the second earphone sensor metadata, and a preset numerical algorithm; or,

[0177] The first wireless earphone sends the first earphone sensor metadata to the playback device, and the second wireless earphone sends the second earphone sensor metadata to the playback device, so that the playback device determines the rendering metadata according to the first earphone sensor metadata, the second earphone sensor metadata, and a preset numerical algorithm;

[0178] The first wireless earphone and the second wireless earphone respectively receive the rendering metadata.

[0179] Further, if a headphone sensor is provided on the first wireless earphone, no headphone sensor is provided on the second wireless earphone, and a playback device sensor is provided on the playback device, then the first wireless earphone and the second wireless earphone synchronize the rendering metadata, including:

[0180] The first wireless earphone sends the first headphone sensor metadata to the playback device, so that the playback device determines the rendering metadata according to the first headphone sensor metadata, the playback device sensor metadata, and a preset numerical algorithm;

[0181] The first wireless earphone and the second wireless earphone respectively receive the rendering metadata; or,

[0182] The first wireless earphone receives the playback device sensor metadata sent by the playback device;

[0183] The first wireless earphone determines the rendering metadata according to the first headphone sensor metadata, the playback device sensor metadata, and a preset numerical algorithm;

[0184] The first wireless earphone sends the rendering metadata to the second wireless earphone.

[0185] In another possible design, if headphone sensors are provided on both the first wireless earphone and the second wireless earphone, and a playback device sensor is provided on the playback device, then the first wireless earphone and the second wireless earphone synchronize the rendering metadata, including:

[0186] The first wireless earphone sends the first headphone sensor metadata to the playback device, and the second wireless earphone sends the second headphone sensor metadata to the playback device, so that the playback device determines the rendering metadata according to the first headphone sensor metadata, the second headphone sensor metadata, the playback device sensor metadata, and a preset numerical algorithm;

[0187] The first wireless earphone and the second wireless earphone respectively receive the rendering metadata; or,

[0188] The first wireless earphone sends the first headphone sensor metadata to the second wireless earphone, and the second wireless earphone sends the second headphone sensor metadata to the first wireless earphone;

[0189] The first wireless earphone and the second wireless earphone respectively receive the sensor metadata of the playback device;

[0190] The first wireless earphone and the second wireless earphone respectively determine the rendering metadata according to the first earphone sensor metadata, the second earphone sensor metadata, the playback device sensor metadata, and a preset numerical algorithm.

[0191] Optionally, the rendering metadata includes at least one of the first wireless earphone metadata, the second wireless earphone metadata, and the playback device metadata.

[0192] Specifically, the first wireless earphone metadata includes the first earphone sensor metadata and a head-related transfer function (HRTF) database, where the first earphone sensor metadata is used to characterize the movement characteristics of the first wireless earphone;

[0193] The second wireless earphone metadata includes the second earphone sensor metadata and a head-related transfer function (HRTF) database, where the second earphone sensor metadata is used to characterize the movement characteristics of the second wireless earphone;

[0194] The playback device metadata includes the playback device sensor metadata, where the playback device sensor metadata is used to characterize the movement characteristics of the playback device.

[0195] Specifically, as Figure 5 stated, the synchronization implementation schemes include but are not limited to the following:

[0196] (1) When only one of the earphones has a sensor that can provide human head rotation metadata, the synchronization method includes but is not limited to transmitting the metadata in that earphone to the other earphone. For example, when only the left ear has a sensor, at this time, the left ear side generates head rotation metadata sensor33, and at the same time, transmits this metadata wirelessly to the right ear, generating sensor34. At this time, sensor33 = sensor34, and after synchronization, sensor35 = sensor33.

[0197] (2) When both earphones have sensors, sensor data sensor33 and sensor34 are generated on both sides respectively. At this time, the synchronization methods include but are not limited to: a. The metadata on both sides of the earphones are wirelessly transmitted to each other (sensor33 on the left side is transmitted to the right earphone; sensor34 on the right side is transmitted to the left earphone), and then synchronous numerical processing is performed on both sides of the earphone end to generate sensor35; b. Or the sensor metadata on both sides of the earphone are transmitted to the pre-stage device, and after synchronous data processing is performed on the pre-stage device, the processed sensor35 is wirelessly transmitted to both sides of the earphone respectively for 307 and 308 to use.

[0198] (3) When the current - stage device can also provide the corresponding sensor metadata sensor0, when only one earphone has a sensor, for example, only the left ear has a sensor and generates sensor33, the synchronization method includes but is not limited to: a. Transmit sensor33 to the previous - stage device. The previous - stage device performs numerical processing based on sensor0 and sensor33, and then wirelessly returns the processed sensor35 to the left and right earphones for 307 and 308 to use. b. Transmit the sensor metadata sensor0 of the previous - stage device to the earphone side. Combine sensor0 and sensor33 at the left - earphone side for numerical processing to obtain sensor35, and at the same time, wirelessly transmit sensor35 to the right - earphone side; finally, it is used by 307 and 308.

[0199] (4) When the current - stage device can provide the corresponding sensor metadata sensor0, and both bilateral earphones have sensors and generate the corresponding metadata sensor33 and sensor34, the synchronization method includes but is not limited to: a. Transmit the metadata sensor33 and sensor34 on both sides of the earphone to the previous - stage device. In the previous - stage device, combine the three sets of metadata for data integration and calculation to obtain the finally synchronized metadata sensor35, and then send this data to both sides of the earphone for 307 and 308 to use; b. Wirelessly transmit the metadata sensor0 of the previous - stage device to both sides of the earphone, and at the same time, the metadata on the left and right sides of the earphone are transmitted to each other. Then, perform data integration and calculation on the three sets of metadata respectively at both sides of the earphone to obtain sensor35 for 307 and 308 to use.

[0200] In this embodiment, the sensor metadata sensor33 or sensor34 can be provided in a combination manner including but not limited to a gyroscope sensor, a geomagnetic device, and an accelerometer; the HRTF refers to a head - related transfer function; the HRTF database can be based on, but not limited to, other sensor metadata at the earphone side (such as a head - size sensor), or based on a front - end device with a camera or photographing function for intelligent identification of the human head. After considering the physical characteristics of the listener's head, ears, etc., perform personalized selection, processing, and adjustment to achieve a personalized effect; the HRTF database can be pre - stored in the earphone side, or later, a new HRTF database can be imported through wired or wireless means to update the HRTF database to achieve the above - mentioned personalized purpose.

[0201] After receiving the synchronized metadata sensor35, the spatial coordinate conversion units 307 and 308 respectively perform rotation transformations on the spatial positions (X1(i1), Y1(i1), Z1(i1)) and (X2(j1), Y2(j1), Z2(j1)) of each channel of the left and right earphones to obtain the rotated spatial positions (X3(i1), Y3(i1), Z3(i1)) and (X4(j1), Y4(j1), Z4(j1)). The rotation method is based on the general three-dimensional coordinate system rotation method and will not be elaborated here; then they are converted into polar coordinates (ρ1(i1), α1(i1), β1(i1)) and (ρ2(j1), α2(j1), β2(j1)) centered on the human head. The specific conversion method can be calculated according to the conversion method between the general Cartesian coordinate system and the polar coordinate system and will not be elaborated here.

[0202] Based on the angles α1(i1), β1(i1) and α2(j1), β2(j1) in the polar coordinate system, the filtering processing units 309 and 310 respectively select the corresponding HRTF data groups HRTF_L(i1) and HRTF_R(j1) from the left ear HRTF database Data_L input from the metadata unit 305 and the right ear HRTF database Data_R input from 306. Then, HRTF filtering is performed on the channel signals S37(i1) and S38(j1) to be virtually processed input from the audio receiving units 301 and 302, and the virtual signals S33(i1) of each channel at the left earphone end and the virtual signals S34(j1) of each channel at the right earphone end after filtering are obtained.

[0203] The downmixing unit 311 receives the filtered and rendered data S33(i1) from 309 above and the channel signal S35(i2) that does not require HRTF filtering input from 301, and performs downmixing on the N-channel information to obtain the final audio signal S39 that can be used for left ear playback. Similarly, the downmixing unit 312 receives the filtered and rendered data S34(j1) from 310 above and the channel signal S36(j2) that does not require HRTF filtering input from 302, and performs downmixing on the M-channel information to obtain the final audio signal S310 that can be used for right ear playback.

[0204] In this embodiment, since the HRTF database may have limited accuracy, interpolation can be considered to obtain the HRTF data group corresponding to the angle during calculation [2]; in addition, subsequent processing steps can be further added to 311 and 312, including but not limited to equalization (EQ), delay, reverberation, etc.

[0205] Further, optionally, before HRTF virtual rendering (i.e., before 301 and 302), preprocessing can be added, which can include but is not limited to other rendering methods such as channel rendering, object rendering, scene rendering, etc.

[0206] In addition, when the audio signals of the input rendering part, namely the first decoded audio signal and the second decoded audio signal, are objects, the processing method and process are as Figure 6 shown.

[0207] Figure 6 It is a schematic diagram of another HRTF rendering method according to an embodiment of the present invention. As Figure 6 shown, the audio receiving units 401 and 402 both receive the object content S41(k) and the corresponding three-dimensional coordinates (X41(k), Y41(k), Z41(k)), 1 ≤ k ≤ K, where K is the number of objects.

[0208] The metadata unit 403 provides metadata for the left earphone rendering of the entire object, including sensor metadata sensor43 and the left ear HRTF database Data_L; similarly, the metadata unit 404 provides metadata for the right earphone rendering of the entire object, including sensor metadata sensor44 and the right ear HRTF database Data_R. When the sensor metadata is transmitted to the spatial coordinate conversion units 405 or 406, data synchronization processing is required, and the processing methods include but are not limited to the 4 methods described in the metadata units 305 and 306. Finally, the synchronized sensor metadata sensor45 is respectively transmitted into 405 and 406;

[0209] In this embodiment, the sensor metadata sensor43 or sensor44 can be provided by a combination method of a gyroscope sensor, a geomagnetic device, and an accelerometer, etc.; the HRTF database can be based on but is not limited to other sensor metadata at the earphone end (such as a head size sensor), or based on a front-end device with a camera or photographing function for intelligent identification of the human head, and after personalizing and adjusting according to the physical characteristics of the listener's head, ears, etc., to achieve a personalized effect; the HRTF database can be pre-stored at the earphone end, or new HRTF databases can be imported into it subsequently by wired or wireless means to update the HRTF database to achieve the above-mentioned personalized purpose.

[0210] The spatial coordinate conversion units 405 and 406, after receiving the sensor metadata sensor45, respectively perform a rotation transformation on the object spatial coordinates (X41(k), Y41(k), Z41(k)) to obtain the spatial coordinates (X42(k), Y42(k), Z42(k)) in the new coordinate system, and then perform a conversion to the polar coordinate system to obtain the polar coordinates (ρ41(k), α41(k), β41(k)) centered on the human head.

[0211] The filtering processing units 407 and 408, after receiving the polar coordinates (ρ41(k), α41(k), β41(k)) of each object, select the corresponding HRTF data groups HRTF_L(k) and HRTF_R(k) from Data_L transmitted to 407 from 403 and Data_R transmitted to 408 according to their distance and angle information.

[0212] The downmixing unit 409, after receiving the virtual signals S42(k) of each object transmitted from 407, performs downmixing to obtain the audio signal S44 that can finally be used for playback on the left earphone; similarly, the downmixing unit 410, after receiving the virtual signals S43(k) of each object transmitted from 408, performs downmixing to obtain the audio signal S45 that can finally be used for playback on the right earphone. The S44 and S45 played by the left and right earphone ends jointly create the target sound and effect.

[0213] In this embodiment, since the HRTF database may have limited accuracy, interpolation can be considered to obtain the HRTF data group corresponding to the angle during calculation [2]; in addition, subsequent processing steps can be further added to the downmixing units 409 and 410, including but not limited to equalization (EQ), delay, reverberation and other processing.

[0214] Further, optionally, before HRTF virtual rendering (i.e., before 301 and 302), preprocessing can be added, which can include but not limited to other rendering methods such as channel rendering, object rendering, scene rendering, etc.

[0215] This form of binaural separate processing has never been implemented.

[0216] Although it is binaural separate processing, it is not that each does its own thing. The audio after binaural processing can be organically combined into a complete binaural sound field; not only the sensor data needs to be synchronized, but also the audio data needs to be synchronized.

[0217] After binaural separate processing, since each earphone only processes the data of its own channel, the total time consumption is halved, saving computing power; at the same time, the requirements for the memory, speed, etc. of each earphone chip are also halved, meaning that more chips can be competent for the processing work.

[0218] In terms of reliability, in the prior art, if the processing module fails to work, the final output may be silence or noise; when any one of the headphone processing modules in the embodiments of the present invention fails to work, the other headphone can still continue to work, and can communicate with the pre-stage device to obtain, process, and output the audio of both channels simultaneously.

[0219] It should be noted that, optionally, the headphone sensor includes at least one of a gyroscope sensor, a head size sensor, a ranging sensor, a geomagnetic sensor, and an acceleration sensor; and / or,

[0220] The playback device sensor includes at least one of a gyroscope sensor, a head size sensor, a ranging sensor, a geomagnetic sensor, and an acceleration sensor.

[0221] S303. The first wireless headphone plays the first playback audio signal, and the second wireless headphone plays the second playback audio signal.

[0222] In this step, the first playback audio signal and the second playback audio signal jointly construct a complete sound field to form a three-dimensional stereo surround. And since the first wireless headphone and the second wireless headphone are relatively independent with respect to the playback device, that is, there will not be a long delay as in the existing wireless headphone technology between the wireless headphone and the playback device. That is, the technical solution of the present invention transfers the function of audio signal rendering from the playback device side to the wireless headphone side, which can greatly shorten the delay, thereby improving the response speed of the wireless headphone to head movement, and further improving the sound effect of the wireless headphone.

[0223] This embodiment provides an audio processing method. The first wireless headphone receives the first audio signal to be presented sent by the playback device, and the second wireless headphone receives the second audio signal to be presented sent by the playback device; then the first wireless headphone performs rendering processing on the first audio signal to be presented to obtain the first playback audio signal, and the second wireless headphone performs rendering processing on the second audio signal to be presented to obtain the second playback audio signal; finally, the first wireless headphone plays the first playback audio signal, and the second wireless headphone plays the second playback audio signal. Thus, the technical effect that the wireless headphone can render the audio signal independently without relying on the playback device is achieved, thereby greatly reducing the delay and improving the sound quality of the headphone.

[0224] The above content is described for a pair of headphones. When the playback device works together with multiple pairs of wireless headphones such as TWS headphones, the way of rendering the channel information and / or object information in the above-mentioned pair of headphones can be similarly referred to. The differences are as Figure 7 shown.

[0225] Figure 7Schematic diagram of an application scenario for connecting multiple pairs of wireless earphones to a playback device provided by an embodiment of the present invention. As Figure 7 shown, the sensor metadata generated by different pairs of TWS earphones can be different, and the metadata sensor1, sensor2, … sensorN generated after being coupled and synchronized with the sensor metadata of the playback device can be the same, partially the same, or even completely different, where N is the number of pairs of TWS earphones. Therefore, when rendering for channel or object information as described above, with other things remaining unchanged, the only change is the different rendering metadata input at the earphone end. Consequently, the three-dimensional spatial positions of each channel or object presented at different earphone ends will also be different, and ultimately, the sound fields presented at different TWS earphone ends will vary according to the user's location or direction.

[0226] Figure 8 Schematic diagram of the structure of an audio processing device provided by an embodiment of the present invention. As Figure 8 shown, the audio processing device 800 provided in this embodiment includes:

[0227] A first audio processing device and a second audio processing device;

[0228] The first audio processing device includes:

[0229] A first receiving module, configured to receive a first audio signal to be presented sent by a playback device;

[0230] A first rendering module, configured to perform rendering processing on the first audio signal to be presented to obtain a first playback audio signal;

[0231] A first playback module, configured to play the first playback audio signal;

[0232] The second audio processing device includes:

[0233] A second receiving module, configured to receive a second audio signal to be presented sent by the playback device;

[0234] A second rendering module, configured to perform rendering processing on the second audio signal to be presented to obtain a second playback audio signal;

[0235] A second playback module, configured to play the second playback audio signal.

[0236] In a possible design, the first audio processing device is a left ear audio processing device, and the second audio processing device is a right ear audio processing device. Then, the first playback audio signal is used to present a left ear audio effect, and the second playback audio signal is used to present a right ear audio effect, so as to form a binaural sound field when the first audio processing device plays the first playback audio signal and the second audio processing device plays the second playback audio signal.

[0237] In a possible design, the first audio processing device 801 further includes:

[0238] A first decoding module, configured to perform decoding processing on the first audio signal to be presented to obtain a first decoded audio signal;

[0239] The first rendering module is specifically configured to: perform rendering processing according to the first decoded audio signal and rendering metadata to obtain the first playback audio signal;

[0240] The second audio processing device further includes:

[0241] A second decoding module, configured to perform decoding processing on the second audio signal to be presented to obtain a second decoded audio signal;

[0242] The second rendering module is specifically configured to: perform rendering processing according to the second decoded audio signal and rendering metadata to obtain the second playback audio signal.

[0243] In a possible design, the rendering metadata includes at least one of first wireless headphone metadata, second wireless headphone metadata, and playback device metadata.

[0244] In a possible design, the first wireless headphone metadata includes first headphone sensor metadata and a head-related transfer function (HRTF) database, where the first headphone sensor metadata is used to characterize the motion characteristics of the first wireless headphone;

[0245] The second wireless headphone metadata includes second headphone sensor metadata and a head-related transfer function (HRTF) database, where the second headphone sensor metadata is used to characterize the motion characteristics of the second wireless headphone;

[0246] The playback device metadata includes playback device sensor metadata, where the playback device sensor metadata is used to characterize the motion characteristics of the playback device.

[0247] In a possible design, the first audio processing device further includes:

[0248] A first synchronization module, configured to synchronize the rendering metadata with the second wireless headphone; and / or,

[0249] The second audio processing device further includes:

[0250] A second synchronization module, configured to synchronize the rendering metadata with the first wireless headphone.

[0251] In a possible design, the first synchronization module is specifically configured to: send the first headphone sensor metadata to the second wireless headphone, so that the second synchronization module uses the first headphone sensor metadata as the second headphone sensor metadata.

[0252] In a possible design, the first synchronization module is specifically configured to:

[0253] Send the first headphone sensor metadata;

[0254] Receive the second headphone sensor metadata;

[0255] Determine the rendering metadata according to the first headphone sensor metadata, the second headphone sensor metadata, and a preset numerical algorithm;

[0256] The second synchronization module is specifically configured to:

[0257] Send the second headphone sensor metadata;

[0258] Receive the first headphone sensor metadata;

[0259] Determine the rendering metadata according to the first headphone sensor metadata, the second headphone sensor metadata, and a preset numerical algorithm; or,

[0260] The first synchronization module is specifically configured to:

[0261] Send the first headphone sensor metadata;

[0262] Receive the rendering metadata;

[0263] The second synchronization module is specifically configured to:

[0264] Send the second headphone sensor metadata;

[0265] Receive the rendering metadata.

[0266] In a possible design, the first synchronization module is specifically configured to:

[0267] Receive playback device sensor metadata;

[0268] Determine the rendering metadata according to the first headphone sensor metadata, the playback device sensor metadata, and a preset numerical algorithm;

[0269] Send the rendering metadata.

[0270] In a possible design, the first synchronization module is specifically configured to:

[0271] Send the first headphone sensor metadata;

[0272] Receive the second headphone sensor metadata;

[0273] Receive the playback device sensor metadata;

[0274] Determine the rendering metadata according to the first headphone sensor metadata, the second headphone sensor metadata, the playback device sensor metadata, and a preset numerical algorithm;

[0275] The second synchronization module is specifically configured to:

[0276] Send the second headphone sensor metadata;

[0277] Receive the first headphone sensor metadata;

[0278] Receive the playback device sensor metadata;

[0279] Determine the rendering metadata according to the first headphone sensor metadata, the second headphone sensor metadata, the playback device sensor metadata, and a preset numerical algorithm.

[0280] Optionally, the first audio signal to be presented includes at least one of a channel-based audio signal, an object-based audio signal, and a scene-based audio signal; and / or,

[0281] The second audio signal to be presented includes at least one of a channel-based audio signal, an object-based audio signal, and a scene-based audio signal.

[0282] It should be noted that Figure 8 The audio processing device 800 provided in the illustrated embodiment can execute the corresponding method on the playback device side provided in any of the above method embodiments. The specific implementation principles, technical features, professional term explanations, and technical effects are similar and will not be elaborated here.

[0283] Figure 9 This is a schematic structural diagram of a wireless headphone provided by an embodiment of the present invention. As Figure 9 shown, the wireless headphone 900 may include: a first wireless headphone 901 and a second wireless headphone 902.

[0284] The first wireless headphone 901 includes:

[0285] A first processor 9011; and

[0286] A first memory 9012 for storing the computer program of the processor;

[0287] Among them, the processor 9011 is configured to implement the steps of the first wireless earphone in any possible audio processing method in the above method embodiments by executing the computer program;

[0288] A second wireless earphone 902, comprising:

[0289] A second processor 9021; and

[0290] A second memory 9022 for storing the computer program of the processor;

[0291] Among them, the processor is configured to implement the steps of the second wireless earphone in any possible audio processing method in the above method embodiments by executing the computer program.

[0292] Both the first processor 901 and the second processor 902 include at least one processor and memory. Figure 9 The illustrated electronic device takes one processor as an example.

[0293] The first memory 9012 and the second memory 9022 are used to store programs. Specifically, the program may include program codes, and the program codes include computer operation instructions.

[0294] The first memory 9012 and the second memory 9022 may include high-speed RAM memories, and may also include non-volatile memories, such as at least one disk memory.

[0295] The first processor 9011 is used to execute the computer execution instructions stored in the first memory 9012 to implement the steps of the first wireless earphone in the audio processing method described in the above method embodiments.

[0296] The first processor 9011 and the second processor 9021 are respectively used to execute the computer execution instructions stored in the first memory 9012 and the second memory 9022 to implement the steps of the second wireless earphone in the audio processing method described in the above method embodiments.

[0297] Among them, the first processor 9011 or the second processor 9021 may be a central processing unit (CPU for short), or an application specific integrated circuit (ASIC for short), or one or more integrated circuits configured to implement the embodiments of the present application.

[0298] Optionally, the first memory 9012 can be either independent or integrated with the first processor 9011. When the first memory 9012 is a device independent of the first processor 9011, the first wireless earphone 901 may further include:

[0299] A first bus 9013 for connecting the first processor 9011 and the first memory 9012. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc., but it does not mean that there is only one bus or one type of bus.

[0300] Optionally, the second memory 9022 can be either independent or integrated with the second processor 9021. When the second memory 9022 is a device independent of the second processor 9021, the second wireless earphone 902 may further include:

[0301] A second bus 9023 for connecting the second processor 9021 and the second memory 9022. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc., but it does not mean that there is only one bus or one type of bus.

[0302] Optionally, in specific implementation, if the first memory 9012 and the first processor 9011 are integrated on a single chip, the first memory 9012 and the first processor 9011 can communicate through an internal interface.

[0303] Optionally, in specific implementation, if the second memory 9022 and the second processor 9021 are integrated on a single chip, the second memory 9022 and the second processor 9021 can communicate through an internal interface.

[0304] The present application also provides a computer-readable storage medium, which may include: various media capable of storing program codes such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs. Specifically, program instructions are stored in the computer-readable storage medium, and the program instructions are used for the methods in the above embodiments.

[0305] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.

Claims

1. An audio processing method, characterized in that, Applied to wireless earphones, the wireless earphones include a first wireless earphone and a second wireless earphone, wherein the first wireless earphone and the second wireless earphone are used to establish a wireless connection with a playback device; The method includes: The first wireless earphone receives a first audio signal to be presented sent by the playback device, and the second wireless earphone receives a second audio signal to be presented sent by the playback device; the first audio signal to be presented and the second audio signal to be presented are two audio signals obtained after the playback device distributes the original audio signal according to a preset distribution model, and the two audio signals can form a complete binaural sound field in terms of audio signal characteristics; The first wireless earphone and the second wireless earphone synchronize rendering metadata; The rendering metadata includes first wireless earphone metadata, second wireless earphone metadata, and playback device metadata; The first wireless earphone metadata includes first earphone sensor metadata and a head-related transfer function (HRTF) database, wherein the first earphone sensor metadata is used to characterize the movement characteristics of the first wireless earphone; The second wireless earphone metadata includes second earphone sensor metadata and a head-related transfer function (HRTF) database, wherein the second earphone sensor metadata is used to characterize the movement characteristics of the second wireless earphone; The playback device metadata includes playback device sensor metadata, wherein the playback device sensor metadata is used to characterize the movement characteristics of the playback device; The first wireless earphone performs rendering processing on the first audio signal to be presented based on the rendering metadata to obtain a first playback audio signal, and the second wireless earphone performs rendering processing on the second audio signal to be presented based on the rendering metadata to obtain a second playback audio signal; The first wireless earphone plays the first playback audio signal, and the second wireless earphone plays the second playback audio signal; Headphone sensors are provided on both the first wireless earphone and the second wireless earphone, and a playback device sensor is provided on the playback device; the first wireless earphone and the second wireless earphone synchronize rendering metadata, including: The first wireless earphone sends the first earphone sensor metadata to the playback device, and the second wireless earphone sends the second earphone sensor metadata to the playback device, so that the playback device determines the rendering metadata according to the first earphone sensor metadata, the second earphone sensor metadata, the playback device sensor metadata, and a preset numerical algorithm; the first wireless earphone and the second wireless earphone respectively receive the rendering metadata; or, The first wireless earphone sends the first earphone sensor metadata to the second wireless earphone, and the second wireless earphone sends the second earphone sensor metadata to the first wireless earphone; the first wireless earphone and the second wireless earphone respectively receive the playback device sensor metadata; both the first wireless earphone and the second wireless earphone determine the rendering metadata according to the first earphone sensor metadata, the second earphone sensor metadata, the playback device sensor metadata, and a preset numerical algorithm.

2. The audio processing method according to claim 1, wherein The earphone sensor includes at least one of a gyroscope sensor, a head size sensor, a distance measurement sensor, a geomagnetic sensor, and an acceleration sensor; and / or, The playback device sensor includes at least one of a gyroscope sensor, a head size sensor, a distance measurement sensor, a geomagnetic sensor, and an acceleration sensor.

3. An audio processing method, characterized in that, Applied to wireless earphones, the wireless earphones include a first wireless earphone and a second wireless earphone, wherein the first wireless earphone and the second wireless earphone are used to establish a wireless connection with a playback device; The method includes: The first wireless earphone receives a first audio signal to be presented sent by the playback device, and the second wireless earphone receives a second audio signal to be presented sent by the playback device; the first audio signal to be presented and the second audio signal to be presented are two audio signals obtained after the playback device distributes an original audio signal according to a preset distribution model, and the two audio signals can form a complete binaural sound field in terms of audio signal characteristics. The first wireless earphone and the second wireless earphone synchronize the rendering metadata. The rendering metadata includes first wireless earphone metadata and second wireless earphone metadata. The first wireless earphone metadata includes first earphone sensor metadata and a head-related transfer function (HRTF) database, wherein the first earphone sensor metadata is used to characterize the motion characteristics of the first wireless earphone. The second wireless earphone metadata includes second earphone sensor metadata and a head-related transfer function (HRTF) database, wherein the second earphone sensor metadata is used to characterize the motion characteristics of the second wireless earphone. The first wireless earphone performs a rendering process on the first audio signal to be presented based on the rendering metadata to obtain a first playback audio signal, and the second wireless earphone performs a rendering process on the second audio signal to be presented based on the rendering metadata to obtain a second playback audio signal. The first wireless earphone plays the first playback audio signal, and the second wireless earphone plays the second playback audio signal. When earphone sensors are provided on both the first wireless earphone and the second wireless earphone, and no playback device sensor is provided on the playback device, the synchronization of the rendering metadata by the first wireless earphone and the second wireless earphone includes: The first wireless earphone sends the first earphone sensor metadata to the second wireless earphone, and the second wireless earphone sends the second earphone sensor metadata to the first wireless earphone. Both the first wireless earphone and the second wireless earphone determine the rendering metadata according to the first earphone sensor metadata, the second earphone sensor metadata, and a preset numerical algorithm; or, The first wireless earphone sends the first earphone sensor metadata to the playback device, and the second wireless earphone sends the second earphone sensor metadata to the playback device, so that the playback device determines the rendering metadata according to the first earphone sensor metadata, the second earphone sensor metadata, and a preset numerical algorithm; The first wireless earphone and the second wireless earphone respectively receive the rendering metadata.

4. The audio processing method according to claim 3, wherein The earphone sensor includes at least one of a gyroscope sensor, a head size sensor, a ranging sensor, a geomagnetic sensor, and an acceleration sensor.

5. An audio processing method, characterized in that, Applied to wireless earphones, the wireless earphones include a first wireless earphone and a second wireless earphone, wherein the first wireless earphone and the second wireless earphone are used to establish a wireless connection with a playback device; The method includes: The first wireless earphone receives a first audio signal to be presented sent by the playback device, and the second wireless earphone receives a second audio signal to be presented sent by the playback device; the first audio signal to be presented and the second audio signal to be presented are two audio signals obtained after the playback device distributes the original audio signal according to a preset distribution model, and the two audio signals can form a complete binaural sound field in terms of audio signal characteristics; The first wireless earphone and the second wireless earphone synchronize the rendering metadata; The rendering metadata includes first wireless earphone metadata and playback device metadata; The first wireless earphone metadata includes first earphone sensor metadata and a head-related transfer function (HRTF) database, wherein the first earphone sensor metadata is used to characterize the motion characteristics of the first wireless earphone; The playback device metadata includes playback device sensor metadata, wherein the playback device sensor metadata is used to characterize the motion characteristics of the playback device; The first wireless earphone performs rendering processing on the first audio signal to be presented based on the rendering metadata to obtain a first playback audio signal, and the second wireless earphone performs rendering processing on the second audio signal to be presented based on the rendering metadata to obtain a second playback audio signal; The first wireless earphone plays the first playback audio signal, and the second wireless earphone plays the second playback audio signal; If a headphone sensor is provided on the first wireless earphone, no headphone sensor is provided on the second wireless earphone, and a playback device sensor is provided on the playback device, then the synchronization of the rendering metadata by the first wireless earphone and the second wireless earphone includes: The first wireless earphone sends the first earphone sensor metadata to the playback device, so that the playback device determines the rendering metadata according to the first earphone sensor metadata, the playback device sensor metadata, and a preset numerical algorithm; The first wireless earphone and the second wireless earphone respectively receive the rendering metadata; or, The first wireless earphone receives the playback device sensor metadata sent by the playback device; The first wireless earphone determines the rendering metadata according to the first earphone sensor metadata, the playback device sensor metadata, and a preset numerical algorithm; The first wireless earphone sends the rendering metadata to the second wireless earphone.

6. The audio processing method according to claim 5, wherein The earphone sensor includes at least one of a gyroscope sensor, a head size sensor, a ranging sensor, a geomagnetic sensor, and an acceleration sensor; and / or, The playback device sensor includes at least one of a gyroscope sensor, a head size sensor, a ranging sensor, a geomagnetic sensor, and an acceleration sensor.

7. An audio processing method, characterized in that, Applied to wireless earphones, the wireless earphones include a first wireless earphone and a second wireless earphone, wherein the first wireless earphone and the second wireless earphone are used to establish a wireless connection with a playback device; The method includes: The first wireless earphone receives a first audio signal to be presented sent by the playback device, and the second wireless earphone receives a second audio signal to be presented sent by the playback device; the first audio signal to be presented and the second audio signal to be presented are two audio signals obtained after the playback device distributes the original audio signal according to a preset distribution model, and the two audio signals can form a complete binaural sound field in terms of audio signal characteristics; The first wireless earphone and the second wireless earphone synchronize the rendering metadata; The rendering metadata is first wireless earphone metadata; the first wireless earphone metadata includes first earphone sensor metadata and a head-related transfer function (HRTF) database, wherein the first earphone sensor metadata is used to characterize the motion characteristics of the first wireless earphone; The first wireless earphone performs rendering processing on the first audio signal to be presented based on the rendering metadata to obtain a first playback audio signal, and the second wireless earphone performs rendering processing on the second audio signal to be presented based on the rendering metadata to obtain a second playback audio signal; The first wireless earphone plays the first playback audio signal, and the second wireless earphone plays the second playback audio signal; If a headset sensor is provided on the first wireless earphone, no headset sensor is provided on the second wireless earphone, and no playback device sensor is provided on the playback device, then the first wireless earphone and the second wireless earphone synchronize the rendering metadata, including: The first wireless earphone sends the first earphone sensor metadata to the second wireless earphone, and the second wireless earphone uses the first earphone sensor metadata as the second earphone sensor metadata.

8. The audio processing method according to claim 7, wherein The earphone sensor includes at least one of a gyroscope sensor, a head size sensor, a ranging sensor, a geomagnetic sensor, and an acceleration sensor; and / or, The playback device sensor includes at least one of a gyroscope sensor, a head size sensor, a ranging sensor, a geomagnetic sensor, and an acceleration sensor.

9. The audio processing method according to any one of claims 1-8, characterized in that, If the first wireless earphone is a left ear wireless earphone and the second wireless earphone is a right ear wireless earphone, the first playback audio signal is used to present a left ear audio effect, and the second playback audio signal is used to present a right ear audio effect, so as to form a binaural sound field when the first wireless earphone plays the first playback audio signal and the second wireless earphone plays the second playback audio signal.

10. The audio processing method according to any one of claims 1-8, characterized in that, Before the first wireless earphone performs rendering processing on the first audio signal to be presented based on the rendering metadata, it further includes: The first wireless earphone performs decoding processing on the first audio signal to be presented to obtain a first decoded audio signal; Correspondingly, the first wireless earphone performing rendering processing on the first audio signal to be presented based on the rendering metadata includes: The first wireless earphone performs rendering processing according to the first decoded audio signal and the rendering metadata to obtain the first playback audio signal; and Before the second wireless earphone performs rendering processing on the second audio signal to be presented based on the rendering metadata, it further includes: The second wireless earphone performs decoding processing on the second audio signal to be presented to obtain a second decoded audio signal; Correspondingly, the second wireless earphone performing rendering processing on the second audio signal to be presented based on the rendering metadata includes: The second wireless earphone performs rendering processing according to the second decoded audio signal and the rendering metadata to obtain the second playback audio signal.

11. The audio processing method according to any one of claims 1-8, characterized in that The first audio signal to be presented includes at least one of a channel-based audio signal, an object-based audio signal, and a scene-based audio signal; and / or, The second audio signal to be presented includes at least one of a channel-based audio signal, an object-based audio signal, and a scene-based audio signal.

12. The audio processing method according to any one of claims 1-8, characterized in that The wireless connection includes: Bluetooth connection, infrared connection, WIFI connection, LIFI visible light connection.

13. An audio processing device, characterized in that, It includes: A first audio processing device and a second audio processing device; The first audio processing device includes: A first receiving module for receiving the first audio signal to be presented sent by a playback device; A first synchronization module for synchronizing the rendering metadata with the second wireless earphone; A first rendering module for performing rendering processing on the first audio signal to be presented based on the rendering metadata to obtain a first playback audio signal; A first playback module for playing the first playback audio signal; The second audio processing device includes: A second receiving module for receiving the second audio signal to be presented sent by the playback device; A second synchronization module for synchronizing the rendering metadata with the first wireless earphone; A second rendering module for performing rendering processing on the second audio signal to be presented based on the rendering metadata to obtain a second playback audio signal; A second playback module for playing the second playback audio signal; Wherein, the first audio signal to be presented and the second audio signal to be presented are two audio signals obtained after the playback device distributes the original audio signal according to a preset distribution model, and the two audio signals can form a complete binaural sound field in terms of audio signal characteristics; the rendering metadata includes first wireless headphone metadata, second wireless headphone metadata, and playback device metadata; the first wireless headphone metadata includes first headphone sensor metadata and a head-related transfer function (HRTF) database, wherein the first headphone sensor metadata is used to characterize the motion characteristics of the first wireless headphone; the second wireless headphone metadata includes second headphone sensor metadata and a head-related transfer function (HRTF) database, wherein the second headphone sensor metadata is used to characterize the motion characteristics of the second wireless headphone; the playback device metadata includes playback device sensor metadata, wherein the playback device sensor metadata is used to characterize the motion characteristics of the playback device; The first synchronization module is specifically configured to: Send the first headphone sensor metadata; Receive the second headphone sensor metadata; Receive the playback device sensor metadata; Determine the rendering metadata according to the first headphone sensor metadata, the second headphone sensor metadata, the playback device sensor metadata, and a preset numerical algorithm; The second synchronization module is specifically configured to: Send the second headphone sensor metadata; Receive the first headphone sensor metadata; Receive the playback device sensor metadata; Determine the rendering metadata according to the first headphone sensor metadata, the second headphone sensor metadata, the playback device sensor metadata, and a preset numerical algorithm; Or, The first synchronization module is specifically configured to: Send the first headphone sensor metadata; Receive the rendering metadata; The second synchronization module is specifically configured to: Send the second headphone sensor metadata; Receive the rendering metadata.

14. An audio processing device, characterized in that, It includes: A first audio processing device and a second audio processing device; The first audio processing device includes: A first receiving module, configured to receive the first audio signal to be presented sent by the playback device; A first synchronization module, configured to synchronize the rendering metadata with the second wireless headphone; A first rendering module, configured to perform rendering processing on the first audio signal to be presented based on the rendering metadata to obtain a first playback audio signal; A first playback module, configured to play the first playback audio signal; The second audio processing device includes: A second receiving module, configured to receive the second audio signal to be presented sent by the playback device; A second synchronization module, configured to synchronize the rendering metadata with the first wireless headphone; A second rendering module, configured to perform rendering processing on the second audio signal to be presented based on the rendering metadata to obtain a second playback audio signal; A second playback module, configured to play the second playback audio signal; Wherein, the first audio signal to be presented and the second audio signal to be presented are two audio signals obtained after the playback device distributes the original audio signal according to a preset distribution model, and the two audio signals can form a complete binaural sound field in terms of audio signal characteristics; the rendering metadata includes first wireless headphone metadata and second wireless headphone metadata; the first wireless headphone metadata includes first headphone sensor metadata and a head-related transfer function (HRTF) database, wherein the first headphone sensor metadata is used to characterize the motion characteristics of the first wireless headphone; the second wireless headphone metadata includes second headphone sensor metadata and a head-related transfer function (HRTF) database, wherein the second headphone sensor metadata is used to characterize the motion characteristics of the second wireless headphone. The first synchronization module is specifically configured to: Send the first headphone sensor metadata to the second audio processing device; Receive the second headphone sensor metadata; Determine the rendering metadata according to the first headphone sensor metadata, the second headphone sensor metadata, and a preset numerical algorithm; The second synchronization module is specifically configured to: Send the second headphone sensor metadata to the first audio processing device; Receive the first headphone sensor metadata; Determine the rendering metadata according to the first headphone sensor metadata, the second headphone sensor metadata, and a preset numerical algorithm; Or, The first synchronization module is specifically configured to: Send the first headphone sensor metadata to the playback device; Receive the rendering metadata; The second synchronization module is specifically configured to: Send the second headphone sensor metadata to the playback device; Receive the rendering metadata.

15. An audio processing device, characterized in that, It includes: A first audio processing device and a second audio processing device; The first audio processing device includes: A first receiving module, configured to receive the first audio signal to be presented sent by the playback device; A first synchronization module, configured to synchronize the rendering metadata with the second wireless headphone; A first rendering module, configured to perform rendering processing on the first audio signal to be presented based on the rendering metadata to obtain a first playback audio signal; A first playback module, configured to play the first playback audio signal; The second audio processing device includes: A second receiving module, configured to receive the second audio signal to be presented sent by the playback device; A second synchronization module, configured to synchronize the rendering metadata with the first wireless headphone; A second rendering module, configured to perform rendering processing on the second audio signal to be presented based on the rendering metadata to obtain a second playback audio signal; A second playback module, configured to play the second playback audio signal; Wherein, the first audio signal to be presented and the second audio signal to be presented are two audio signals obtained after the playback device distributes the original audio signal according to a preset distribution model, and the two audio signals can form a complete binaural sound field in terms of audio signal characteristics; the rendering metadata includes first wireless headphone metadata and playback device metadata; the first wireless headphone metadata includes first headphone sensor metadata and a head-related transfer function (HRTF) database, wherein the first headphone sensor metadata is used to characterize the movement characteristics of the first wireless headphone; the playback device metadata includes playback device sensor metadata, wherein the playback device sensor metadata is used to characterize the movement characteristics of the playback device. The first synchronization module is specifically configured to: Receive the playback device sensor metadata; Determine the rendering metadata according to the first headphone sensor metadata, the playback device sensor metadata, and a preset numerical algorithm; Send the rendering metadata; The second synchronization module is specifically configured to receive the rendering metadata; Or; The first synchronization module is specifically configured to: Send the first headphone sensor metadata to the playback device; Receive the rendering metadata; The second synchronization module is specifically configured to receive the rendering metadata.

16. An audio processing device, characterized in that, It includes: A first audio processing device and a second audio processing device; The first audio processing device includes: A first receiving module, configured to receive the first audio signal to be presented sent by the playback device; A first synchronization module, configured to synchronize the rendering metadata with the second wireless headphone; A first rendering module, configured to perform rendering processing on the first audio signal to be presented based on the rendering metadata to obtain a first playback audio signal; A first playback module, configured to play the first playback audio signal; The second audio processing device includes: A second receiving module, configured to receive the second audio signal to be presented sent by the playback device; A second synchronization module, configured to synchronize the rendering metadata with the first wireless headphone; A second rendering module, configured to perform rendering processing on the second audio signal to be presented based on the rendering metadata to obtain a second playback audio signal; A second playback module, configured to play the second playback audio signal; Wherein, the first audio signal to be presented and the second audio signal to be presented are two audio signals obtained after the playback device distributes the original audio signal according to a preset distribution model, and the two audio signals can form a complete binaural sound field in terms of audio signal characteristics; the rendering metadata includes first wireless headphone metadata; the first wireless headphone metadata includes first headphone sensor metadata and a head-related transfer function (HRTF) database, wherein the first headphone sensor metadata is used to characterize the movement characteristics of the first wireless headphone; the first synchronization module is specifically configured to send the first headphone sensor metadata to the second wireless headphone; The second synchronization module is specifically configured to use the first headphone sensor metadata as the second headphone sensor metadata.

17. The audio processing device according to any one of claims 13-16, characterized in that, If the first audio processing device is a left-ear audio processing device and the second audio processing device is a right-ear audio processing device, then the first playback audio signal is used to present a left-ear audio effect, and the second playback audio signal is used to present a right-ear audio effect, so as to form a binaural sound field when the first audio processing device plays the first playback audio signal and the second audio processing device plays the second playback audio signal.

18. The audio processing device according to any one of claims 13-16, characterized in that, The first audio processing device further includes: A first decoding module, configured to perform decoding processing on the first audio signal to be presented to obtain a first decoded audio signal; The first rendering module is specifically configured to: perform rendering processing according to the first decoded audio signal and rendering metadata to obtain the first playback audio signal; The second audio processing device further includes: A second decoding module, configured to perform decoding processing on the second audio signal to be presented to obtain a second decoded audio signal; The second rendering module is specifically configured to: perform rendering processing according to the second decoded audio signal and rendering metadata to obtain the second playback audio signal.

19. The audio processing device according to any one of claims 13-16, characterized in that, The first audio signal to be presented includes at least one of a channel-based audio signal, an object-based audio signal, and a scene-based audio signal; and / or, The second audio signal to be presented includes at least one of a channel-based audio signal, an object-based audio signal, and a scene-based audio signal.

20. A wireless earphone, characterized in that, It includes: A first wireless earphone and a second wireless earphone; The first wireless earphone includes: A first processor; and A first memory for storing a computer program of the processor; Wherein, the processor is configured to implement the steps of the first wireless earphone in the audio processing method according to any one of claims 1-12 by executing the computer program; The second wireless earphone includes: A second processor; and A second memory for storing a computer program of the processor; Wherein, the processor is configured to implement the steps of the second wireless earphone in the audio processing method according to any one of claims 1-12 by executing the computer program.

21. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the audio processing method according to any one of claims 1-12.

Citation Information

Patent Citations

  • Information processing device and information processing method

    EP3687193A1

  • Binaural Audio Navigation Using Short Range Wireless Transmission from Bilateral Earpieces to Receptor Device System and Method

    US20180073886A1