Audio processing method and device and storage medium
By rendering and modulating audio data based on motion information, the method enhances spatial audio accuracy and quality in wearable devices, addressing delays and inaccuracies in existing technologies.
Patent Information
- Application Number
- CN202410050953.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-12
- Publication Date
- 2025-07-15
AI Technical Summary
In the prior art, due to chip size and power consumption limitations, wearable devices have low audio rendering accuracy and insufficient sense of space, resulting in poor sound quality and strong user experience delay.
The rotation information is transmitted between the electronic device and the wearable device, and the initial audio data is rendered and modulated through the electronic device, modulated audio data carrying the rotation information is generated, and secondary rendering is performed in the wearable device to ensure that the audio data matches the rotation information of the device.
Improves the accuracy and spatial sense of audio rendering, reduces user delay, and improves the audio texture, while maintaining practicality without modifying existing Bluetooth encoding.
Smart Images

Figure CN120321578A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of information technology, and in particular, to an audio processing method, apparatus, and storage medium. Background Art
[0002] With the progress of technology and the development of diverse electronic devices, wearable devices have emerged. To achieve diverse functions of wearable devices, an audio playback module can be configured for the wearable device. To improve the user's auditory experience and enable the user to perceive the source and location of sound, the sound can be processed and adjusted so that the audio played by the wearable device presents a specific spatial effect, that is, spatial audio is generated.
[0003] In related technologies, audio can be rendered on an electronic device (e.g., a mobile phone), and the rendered spatial audio is sent to a wearable device (e.g., headphones) for playback, which may cause the spatial sense of sound to lag significantly behind the actual movement of the wearable device, creating an obvious sense of delay for the user. Additionally, audio can also be rendered on the wearable device. However, due to the limitations of the chip size and power consumption of the wearable device, the computing power is relatively small, resulting in insufficient accuracy of audio rendering, insufficient spatial sense, and difficulty in performing complex tone adjustment processing, and the sound quality is also affected. Summary of the Invention
[0004] To overcome the problems of the sense of delay caused to users in related technologies, as well as the low accuracy, insufficient spatial sense, and poor sound quality of the rendered audio, the present disclosure provides an audio processing method, apparatus, and storage medium to improve the accuracy and spatial sense of audio rendering, and thereby enhance the texture of the generated target spatial audio data.
[0005] According to a first aspect of an embodiment of the present disclosure, there is provided an audio processing method, including:
[0006] Receiving first rotation information sent by a wearable device;
[0007] Rendering initial audio data based on the first rotation information to obtain rendered spatial audio data;
[0008] Performing modulation processing on the rendered spatial audio data based on the first rotation information to obtain modulated audio data; wherein the modulated audio data carries the first rotation information;
[0009] Sending the modulated audio data to the wearable device, so that the wearable device performs rendering processing on the rendered spatial audio data based on the first rotation information and the obtained second rotation information to obtain target spatial audio data.
[0010] In some embodiments, modulating the rendered spatial audio data based on the first rotation information to obtain modulated audio data includes:
[0011] Preprocessing the rendered spatial audio data to obtain preprocessed spatial audio data; wherein, the preprocessed spatial audio data does not include audio data having a target frequency band;
[0012] Mapping the first rotation information to the target frequency band in the preprocessed spatial audio data to obtain the modulated audio data.
[0013] In some embodiments, preprocessing the rendered spatial audio data to obtain preprocessed spatial audio data includes:
[0014] Filtering the audio data in the rendered spatial audio data that is in the target frequency band to obtain the preprocessed spatial audio data.
[0015] In some embodiments, mapping the first rotation information to the target frequency band in the preprocessed spatial audio data to obtain the modulated audio data includes:
[0016] Dividing the target frequency band into at least one sub - frequency band based on the frame length of the first rotation information;
[0017] Converting the format of the first rotation information based on the carrying parameters of the target frequency band to obtain encoded information corresponding to the first rotation information;
[0018] Splitting the encoded information based on the number of sub - frequency bands to obtain at least one sub - encoded information;
[0019] Mapping each sub - encoded information to the corresponding sub - frequency band based on the arrangement of the encoded information to obtain the modulated audio data.
[0020] In some embodiments, the method further includes:
[0021] Determining the audio data in the rendered spatial audio data that is within the highest frequency range as the audio data in the target frequency band.
[0022] In some embodiments, preprocessing the rendered spatial audio data to obtain preprocessed spatial audio data includes:
[0023] Performing frequency shift processing on the audio data in the rendered spatial audio data that is in the target frequency band to obtain the preprocessed spatial audio data.
[0024] According to a second aspect of the embodiments of the present disclosure, there is provided an audio processing method, including:
[0025] Send the first rotation information of the wearable device to the electronic device;
[0026] Receive the modulated audio data sent by the electronic device; wherein, the modulated audio data is obtained by the electronic device modulating the rendered spatial audio data based on the first rotation information, and the rendered spatial audio data is obtained by the electronic device rendering the initial audio data based on the first rotation information;
[0027] Demodulate the modulated audio data to obtain the rendered spatial audio data and the first rotation information;
[0028] Obtain the second rotation information of the wearable device;
[0029] Based on the first rotation information and the second rotation information, perform rendering processing on the rendered spatial audio data to obtain target spatial audio data.
[0030] In some embodiments, the demodulating the modulated audio data to obtain the first rotation information includes:
[0031] Detect the sub-encoding information in each sub-band located in the target frequency band of the modulated audio data;
[0032] Based on the arrangement manner of each sub-band, combine each sub-encoding information to obtain the encoding information;
[0033] Perform format conversion on the encoding information to obtain the first rotation information.
[0034] In some embodiments, the obtaining the rendered spatial audio data includes:
[0035] Perform filtering processing on each sub-encoding information located in the target frequency band to obtain the preprocessed spatial audio data;
[0036] Based on the preprocessed spatial audio data, determine the rendered spatial audio data.
[0037] In some embodiments, the based on the preprocessed spatial audio data, determining the rendered spatial audio data includes:
[0038] Determine the preprocessed spatial audio data as the rendered spatial audio data; or
[0039] Perform frequency shift processing on the preprocessed spatial audio data to obtain the rendered spatial audio data.
[0040] According to a third aspect of embodiments of the present disclosure, there is provided an audio processing device, including:
[0041] A first receiving module, configured to receive first rotation information sent by a wearable device;
[0042] A first rendering module, configured to render initial audio data based on the first rotation information to obtain rendered spatial audio data;
[0043] A modulation module, configured to perform modulation processing on the rendered spatial audio data based on the first rotation information to obtain modulated audio data; wherein, the modulated audio data carries the first rotation information;
[0044] A first sending module, configured to send the modulated audio data to the wearable device, so that the wearable device renders the rendered spatial audio data based on the first rotation information and acquired second rotation information to obtain target spatial audio data.
[0045] In some embodiments, the modulation module includes:
[0046] A preprocessing module, configured to preprocess the rendered spatial audio data to obtain preprocessed spatial audio data; wherein, the preprocessed spatial audio data does not include audio data having a target frequency band;
[0047] A mapping module, configured to map the first rotation information to the target frequency band in the preprocessed spatial audio data to obtain the modulated audio data.
[0048] In some embodiments, the preprocessing module is specifically configured to:
[0049] Perform filtering processing on the audio data in the target frequency band in the rendered spatial audio data to obtain the preprocessed spatial audio data.
[0050] In some embodiments, the mapping module is specifically configured to:
[0051] Divide the target frequency band into at least one sub-frequency band based on the frame length of the first rotation information;
[0052] Perform format conversion on the first rotation information based on the carrying parameter of the target frequency band to obtain encoded information corresponding to the first rotation information;
[0053] Split the encoded information based on the number of the sub-frequency bands to obtain at least one sub-encoded information;
[0054] Based on the arrangement of the encoded information, map each of the sub-encoded information to the corresponding sub-band to obtain the modulated audio data.
[0055] In some embodiments, the modulation module further includes:
[0056] A first determination module configured to determine the audio data within the highest frequency range in the rendered spatial audio data as the audio data within the target band.
[0057] In some embodiments, the preprocessing module is specifically configured to:
[0058] Perform a frequency shift process on the audio data within the target band in the rendered spatial audio data to obtain the preprocessed spatial audio data.
[0059] According to a fourth aspect of the embodiments of the present disclosure, there is provided an audio processing device, including:
[0060] A second sending module configured to send the first rotation information of the wearable device to the electronic device;
[0061] A second receiving module configured to receive the modulated audio data sent by the electronic device; wherein the modulated audio data is obtained by the electronic device performing modulation processing on the rendered spatial audio data based on the first rotation information, and the rendered spatial audio data is obtained by the electronic device performing rendering on the initial audio data based on the first rotation information;
[0062] A demodulation module configured to perform demodulation processing on the modulated audio data to obtain the rendered spatial audio data and the first rotation information;
[0063] An acquisition module configured to acquire the second rotation information of the wearable device;
[0064] A second rendering module configured to perform rendering processing on the rendered spatial audio data based on the first rotation information and the second rotation information to obtain the target spatial audio data.
[0065] In some embodiments, the demodulation module is specifically configured to:
[0066] Detect the sub-encoded information in each sub-band within the target band in the modulated audio data;
[0067] Based on the arrangement of each sub-band, combine each sub-encoded information to obtain the encoded information;
[0068] Perform format conversion on the encoded information to obtain the first rotation information.
[0069] In some embodiments, the demodulation module includes:
[0070] A filtering module configured to filter each of the sub-encoded information located in the target frequency band to obtain the preprocessed spatial audio data;
[0071] A second determination module configured to determine the rendered spatial audio data based on the preprocessed spatial audio data.
[0072] In some embodiments, the second determination module is specifically configured to:
[0073] Determine the preprocessed spatial audio data as the rendered spatial audio data; or
[0074] Perform a frequency shift process on the preprocessed spatial audio data to obtain the rendered spatial audio data.
[0075] According to a fifth aspect of the embodiments of the present disclosure, there is provided an electronic device, including:
[0076] A processor;
[0077] A memory configured to store instructions executable by the processor;
[0078] Wherein, the processor is configured to: when executed, implement the steps in any one of the audio processing methods in the first aspect and the second aspect above.
[0079] According to a sixth aspect of the embodiments of the present disclosure, there is provided a non-transitory computer-readable storage medium, when the instructions in the storage medium are executed by a processor of an audio processing device, enabling the device to execute any one of the audio processing methods in the first aspect and the second aspect above.
[0080] The technical solutions provided by the embodiments of the present disclosure may include the following beneficial effects:
[0081] In the embodiments of the present disclosure, the first rotation information sent by the wearable device is received; the initial audio data is rendered based on the first rotation information to obtain the rendered spatial audio data; the rendered spatial audio data is modulated based on the first rotation information to obtain the modulated audio data; wherein, the modulated audio data carries the first rotation information; finally, the modulated audio data is sent to the wearable device, so that the wearable device renders the rendered spatial audio data based on the first rotation information and the obtained second rotation information to obtain the target spatial audio data.
[0082] On the one hand, the electronic device can directly modulate the rendered spatial audio based on the first rotation information to obtain modulated audio data carrying the first rotation information. In this way, without modifying the existing Bluetooth encoding, the modulated audio data can be sent to the wearable device, thereby reducing the latency brought to the user by audio rendering while improving the practicability. On the other hand, audio rendering can be performed at both ends of the electronic device and the wearable device simultaneously, which is beneficial to improving the accuracy and sense of space of audio rendering, and further improving the texture of the generated target spatial audio data.
[0083] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0084] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure.
[0085] Figure 1 is a flowchart of an audio processing method shown according to an exemplary embodiment Figure 1 .
[0086] Figure 2A is Scenario 1 of listening to audio while wearing headphones shown according to an exemplary embodiment.
[0087] Figure 2B is Scenario 2 of listening to audio while wearing headphones shown according to an exemplary embodiment.
[0088] Figure 3A is a schematic diagram of the frequency spectrum of audio shown according to an exemplary embodiment Figure 1 .
[0089] Figure 3B is Schematic Diagram 2 of the frequency spectrum of audio shown according to an exemplary embodiment.
[0090] Figure 4 is Flowchart 2 of the audio processing method shown according to an exemplary embodiment.
[0091] Figure 5 is Flowchart 3 of the audio processing method shown according to an exemplary embodiment.
[0092] Figure 6 is a block diagram of an audio processing device shown according to an exemplary embodiment Figure 1 .
[0093] Figure 7 is Block Diagram 2 of an audio processing device shown according to an exemplary embodiment.
[0094] Figure 8It is a block diagram of a device 800 shown according to an exemplary embodiment.
[0095] Figure 9 It is a block diagram of a device 1900 for audio processing shown according to an exemplary embodiment. Detailed implementation manners
[0096] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0097] Figure 1 It is a flow of an audio processing method shown according to an exemplary embodiment Figure 1 , as Figure 1 shown, the method mainly includes the following steps:
[0098] In step 101, receive the first rotation information sent by the wearable device;
[0099] In step 102, render the initial audio data based on the first rotation information to obtain rendered spatial audio data;
[0100] In step 103, perform modulation processing on the rendered spatial audio data based on the first rotation information to obtain modulated audio data; wherein, the modulated audio data carries the first rotation information;
[0101] In step 104, send the modulated audio data to the wearable device, so that the wearable device performs rendering processing on the rendered spatial audio data based on the first rotation information and the obtained second rotation information to obtain target spatial audio data.
[0102] Here, the method can be applied to an electronic device. The electronic device may include a terminal device. Among them, the terminal device may include a mobile terminal and a fixed terminal, for example, a mobile phone, a tablet computer, a handheld computer, a laptop computer, a desktop computer, a wearable device, a smart speaker, a television, and a vehicle-mounted terminal, etc.
[0103] Among them, the wearable device may include: a portable device directly worn on the user's body or integrated into the user's accessories. For example, it may include: a wearable device supported by the wrist (including watches, wristbands, etc.); a wearable device supported by the foot (including shoes, socks, etc.), a wearable device supported by the head (including headphones, glasses, helmets, headbands, etc.). It may also include: smart clothing, schoolbags, crutches, accessories, etc.
[0104] In some embodiments, the wearable device may have an audio playback function. Since the wearable device can be directly worn on the user's body, when the user's body rotates, the wearable device will also rotate along with the user's body.
[0105] Taking the wearable device as an earphone as an example, when the user's head rotates, the earphone will also rotate along with the user's head. Figure 2A FIG. 1 is a diagram showing a first scenario of wearing an earphone to listen to audio according to an exemplary embodiment, as Figure 2A shown, the audio played by the earphone will also rotate along with the rotation of the head; Figure 2B FIG. 2 is a diagram showing a second scenario of wearing an earphone to listen to audio according to an exemplary embodiment, as Figure 2B shown, the audio played by the earphone will maintain the original playback direction.
[0106] In the embodiments of the present disclosure, when the wearable device rotates, the first rotation information of the wearable device can be obtained and the first rotation information can be sent to the electronic device. Here, the first rotation information may include: information generated when the wearable device rotates. For example, it may include the rotation angle, rotation orientation, etc. of the wearable device. In some embodiments, a detection module (such as an angle detection module, an orientation detection module, etc.) may be set in the wearable device. When it is detected that the wearable device rotates, the detection module can sense the first rotation information of the wearable device.
[0107] Taking the first rotation information as the first rotation angle as an example, an angular motion detection module may be set in the wearable device, and the first rotation angle of the wearable device can be determined through the angular motion detection module. Among them, the angular motion detection module may include: an angular velocity sensor (gyroscope). Taking the first rotation information as the first rotation orientation as an example, an orientation detection module may be set in the wearable device, and the first rotation orientation of the wearable device can be determined through the orientation detection module. Among them, the orientation detection module may include: an orientation sensor.
[0108] In the embodiments of the present disclosure, after obtaining the first rotation information of the wearable device, the first rotation information may be sent to the electronic device. In some embodiments, a communication connection may be established between the wearable device and the electronic device. For example, a wired connection may be established between the wearable device and the electronic device. The wired connection may include: a data line connection; or a wireless connection may be established between the wearable device and the electronic device. The wireless connection may include: a short-range wireless connection (e.g., a Bluetooth connection) and a long-range wireless connection (e.g., a Wi-Fi connection). For example, a Bluetooth communication may be established between the electronic device and the wearable device.
[0109] In some embodiments, the electronic device is a device that outputs initial audio data, and the wearable device is a device that outputs target spatial audio data, where the target spatial audio data is spatial audio obtained by rendering the initial audio data.
[0110] After receiving the first rotation information, the electronic device may render the initial audio data based on the first rotation information to obtain rendered spatial audio data. In some embodiments, the initial audio data may be rendered using a spatial transfer function. For example, the initial audio data may be rendered using a head-related transfer function (HRTF) to obtain rendered spatial audio data. When constructing the head-related transfer function, the relevant impulse responses corresponding to each angle in space may be collected in advance, and a mapping relationship between each angle and the relevant impulse response corresponding to each angle may be established, and then a data set may be constructed.
[0111] In the embodiments of the present disclosure, after obtaining the first rotation information, the electronic device may obtain the relevant impulse response corresponding to the first rotation information from the data set corresponding to the spatial transfer function based on the first rotation information, and then perform convolution processing on the initial audio data and the relevant impulse response to obtain rendered spatial audio data.
[0112] Taking the first rotation information as the first rotation angle as an example, the relevant impulse response corresponding to the first rotation angle may be obtained based on the first rotation angle, and convolution processing may be performed on the initial audio data and the relevant impulse response to obtain rendered spatial audio data. Here, the rendered spatial audio data is spatial audio obtained by the electronic device rendering the initial audio data.
[0113] It can be understood that after obtaining the rendered spatial audio data, the electronic device modulates the rendered spatial audio data based on the first rotation information to obtain modulated audio data carrying the first rotation information.
[0114] In some embodiments, modulation is a process of processing and loading the rendered spatial audio data onto a carrier wave to transform it into a form suitable for channel transmission.
[0115] Here, the modulation includes one of frequency modulation, amplitude modulation, or phase modulation, and the embodiments of the present disclosure do not limit this.
[0116] Taking amplitude modulation as an example, it is to modulate the rendered spatial audio data into modulated audio data carrying the first rotation information based on the first rotation information. For example, using the presence of amplitude to represent carrying the first rotation information and the absence of amplitude to represent not carrying the first rotation information, then the first rotation information can be output as an analog signal with or without amplitude to obtain the modulated audio data, which is beneficial to improving the efficiency of obtaining the modulated audio data.
[0117] Taking frequency modulation as an example, it is to modulate the rendered spatial audio data into modulated audio data carrying the first rotation information based on the first rotation information. For example, using a preset first frequency change per unit time to represent carrying the first rotation information and a preset second frequency change per unit time to represent not carrying the first rotation information, where the first frequency and the second frequency are different, then the first rotation information can be output as an analog signal with different frequency combinations to obtain the modulated audio data, which is beneficial to improving the anti-interference ability of obtaining the modulated audio data.
[0118] Correspondingly, after receiving the modulated audio data, the wearable device can perform demodulation processing on the modulated audio data to obtain the rendered spatial audio data and the first rotation information.
[0119] In the embodiments of the present disclosure, after the wearable device obtains the rendered spatial audio data and the first rotation information, it can obtain the second rotation information. It can be understood that the first rotation information is the rotation information collected when the wearable device first rotates, and the second rotation information is the rotation information collected when the wearable device currently rotates.
[0120] In the embodiments of the present disclosure, when it is detected that the wearable device rotates, the first rotation information can be collected and the audio rendering mechanism can be triggered, so that the electronic device can perform rendering processing on the initial audio data. After the wearable device obtains the rendered spatial audio data, the second rotation information can be collected, and then the current rotation information of the wearable device can be determined more accurately. Furthermore, based on the first rotation information and the second rotation information, the rendered spatial audio data can be rendered again to obtain the target spatial audio data, which can ensure that the obtained target spatial audio data is closer to the current rotation information of the wearable device, and further ensure the accuracy and real-time performance of the obtained target spatial audio data.
[0121] In an embodiment of the present disclosure, first rotation information sent by a wearable device is received; initial audio data is rendered based on the first rotation information to obtain rendered spatial audio data; the rendered spatial audio data is then modulated based on the first rotation information to obtain modulated audio data; wherein the modulated audio data carries the first rotation information; and finally, the modulated audio data is sent to the wearable device, so that the wearable device renders the rendered spatial audio data based on the first rotation information and acquired second rotation information to obtain target spatial audio data.
[0122] On the one hand, the electronic device can directly modulate the rendered spatial audio based on the first rotation information to obtain modulated audio data carrying the first rotation information. In this way, without modifying the existing Bluetooth encoding, the modulated audio data can be sent to the wearable device, thereby reducing the latency brought to the user by audio rendering while improving practicability. On the other hand, audio rendering can be performed at both ends of the electronic device and the wearable device, which is conducive to improving the accuracy and sense of space of audio rendering, and further improving the texture of the generated target spatial audio data.
[0123] In some embodiments, the modulating the rendered spatial audio data based on the first rotation information to obtain modulated audio data includes:
[0124] Preprocessing the rendered spatial audio data to obtain preprocessed spatial audio data; wherein the preprocessed spatial audio data does not include audio data having a target frequency band;
[0125] Mapping the first rotation information to the target frequency band in the preprocessed spatial audio data to obtain the modulated audio data.
[0126] It should be noted that, in order to reduce the amount of computation in the process of obtaining the modulated audio data, the rendered spatial audio data can be preprocessed first to obtain preprocessed spatial audio data that does not include audio data having a target frequency band, and then based on the first rotation information and the preprocessed spatial audio data, the modulated audio data is further obtained.
[0127] Here, the preprocessing may include frequency shift processing or filtering processing, etc. The embodiments of the present disclosure do not limit this, as long as the preprocessed spatial audio data does not have audio data of the target frequency band.
[0128] In some embodiments, the rendered spatial audio data can be first subjected to Fourier transform to obtain the spectrum corresponding to the rendered spatial audio data, and normalization processing is performed to facilitate preprocessing of the rendered spatial audio data to obtain preprocessed spatial audio data that does not include audio data having a target frequency band.
[0129] It can be understood that after obtaining the preprocessed spatial audio data, the first rotation information can be mapped to the target frequency band in the preprocessed spatial audio data, so as to obtain accurate modulated audio data.
[0130] Taking the first rotation information as the first rotation angle as an example, after obtaining the preprocessed spatial audio data, the first rotation angle can be loaded into the target frequency band in the preprocessed spatial audio data to obtain the modulated audio data; or the first rotation angle can be split to obtain multiple first angle information, and then according to the number of the first angle information, the target frequency band in the preprocessed spatial audio data is split to obtain multiple first frequency bands, and then one first angle information is loaded into each first frequency band to obtain the modulated audio data.
[0131] In the embodiments of the present disclosure, the rendered spatial audio data can be preprocessed first to obtain the preprocessed spatial audio data; then the first rotation information is mapped to the target frequency band in the preprocessed spatial audio data to obtain the modulated audio data, so that the electronic device can directly generate the modulated audio data carrying the first rotation information, and can transmit the modulated audio data by means of wireless communication, thereby improving the practicability and being beneficial to ensuring the texture of the target spatial audio data generated by the wearable device.
[0132] In some embodiments, the preprocessing of the rendered spatial audio data to obtain the preprocessed spatial audio data includes:
[0133] Filtering the audio data in the target frequency band of the rendered spatial audio data to obtain the preprocessed spatial audio data.
[0134] It can be understood that in order to facilitate mapping the first rotation information to the target frequency band in the preprocessed spatial audio data, the audio data in the target frequency band of the rendered spatial audio data can be filtered to obtain the preprocessed spatial audio data.
[0135] In some embodiments, a filter model can be established according to the signal characteristics of the rendered spatial audio data and the components to be filtered (i.e., the audio data in the target frequency band of the rendered spatial audio data). Then, according to the signal characteristics and the filter model, corresponding filtering parameters are selected, where the filtering parameters can affect the performance of the filter, such as the cut-off frequency, the filter order, etc. Then the filter model is applied to the audio data in the target frequency band, and time-domain filtering or frequency-domain filtering is realized through convolution operation, frequency-domain transformation, etc. After filtering the audio data in the target frequency band, the filtered audio data can be evaluated based on evaluation metrics, such as signal-to-noise ratio, spectral analysis, or time-domain waveform, etc.
[0136] In some other embodiments, if the filtered audio data does not meet the preset conditions, the filtering parameters of the filter can also be adjusted according to the evaluation results, and the audio data in the target frequency band in the rendered spatial audio data can be filtered again, so as to form the preprocessed spatial audio data that meets the preset conditions.
[0137] Here, the filter can be linear or non-linear; the filter model can be a low-pass filter, a high-pass filter, a band-pass filter, etc., and the embodiments of the present disclosure do not limit this.
[0138] In the embodiments of the present disclosure, the preprocessed spatial audio data is obtained by filtering the audio data in the target frequency band in the rendered spatial audio data. In this way, by filtering the audio data in the target frequency band, it is convenient to map the first rotation information to the target frequency band, which is beneficial to reducing the amount of calculation while obtaining the modulated audio data carrying the first rotation information, and can improve the practicability.
[0139] In some embodiments, the mapping of the first rotation information to the target frequency band in the preprocessed spatial audio data to obtain the modulated audio data includes:
[0140] Based on the frame length of the first rotation information, the target frequency band is divided into at least one sub-frequency band;
[0141] Based on the carrying parameter of the target frequency band, the format of the first rotation information is converted to obtain the encoded information corresponding to the first rotation information;
[0142] Based on the number of the sub-frequency bands, the encoded information is split to obtain at least one sub-encoded information;
[0143] Based on the arrangement mode of the encoded information, each sub-encoded information is mapped to the corresponding sub-frequency band to obtain the modulated audio data.
[0144] It should be noted that, in order to improve the accuracy of mapping the first rotation information to the target frequency band, the target frequency band can be divided into at least one sub-frequency band based on the frame length of the first rotation information, and then the first rotation information is mapped to each sub-frequency band.
[0145] Exemplarily, based on the frame length of the first rotation information, the frequency interval for dividing the target frequency band can be obtained; based on the frequency interval, the target frequency band can be divided into multiple sub-frequency bands of equal length. Therefore, the calculation formula of the frequency interval can be as follows:
[0146] f r =f s / N (1);
[0147] In formula (1), f s is the sampling rate, N is the frame length, and f r is the frequency interval.
[0148] In some embodiments, when the sampling rate is 48 kilohertz (kHz) and the frame length of the first rotation information is 512 bytes, based on the above formula (1), the frequency interval can be obtained as 93.75 hertz (Hz). Therefore, the interval between two adjacent sub-bands is 93.75 Hz.
[0149] Here, the longer the frame length of the first rotation information, the more first rotation information carried by the modulated audio data; the shorter the frame length of the first rotation information, the less first rotation information carried by the modulated audio data.
[0150] Exemplarily, taking the target frequency band of 18 kHz - 22 kHz as an example, if the calculated frequency interval is 375 Hz, the target frequency band can be divided into 11 sub-bands [18000, 18375, 18750, 19125, 19500, 19875, 20250, 20625, 21000, 21375, 21750], and the interval between two adjacent sub-bands is 375 Hz.
[0151] In the embodiments of the present disclosure, the format of the first rotation information can be converted based on the carrying parameter of the target frequency band to obtain the coding information corresponding to the first rotation information, so as to map the coding information to each sub-band.
[0152] Here, the carrying parameter represents the length that can load the coding information. With different carrying parameters of the target frequency band, the format of the first rotation information is converted, and the coding information corresponding to the first rotation information is also different. The embodiments of the present disclosure do not limit this. For example, if the carrying parameter of the target frequency band is set to load 11-bit binary coding, the first rotation information can be converted into 11-bit binary data.
[0153] Taking the first rotation information as the first rotation angle as an example, assuming the first rotation angle is 50°, and the carrying parameter of the target frequency band is set to load 11-bit binary coding, the format of the first rotation angle can be converted, that is, from decimal data to binary data, to obtain the 11-bit binary coding information 00000110010.
[0154] It can be understood that the coding information corresponding to the first rotation information can be split based on the number of sub-bands to obtain at least one sub-coding information, so as to map each sub-coding information to the corresponding sub-band.
[0155] Taking the number of sub - bands as 11 and the coding information corresponding to the first rotation information as 00000110010 as an example, the coding information can be split into 11 sub - information, each sub - information being 0 or 1, such that each sub - band loads one binary number, that is, loads 0 or 1.
[0156] After obtaining each sub - coding information, based on the arrangement of the coding information, each of the sub - coding information can be mapped to the corresponding sub - band to obtain the modulated audio data.
[0157] Exemplarily, if the coding information is 00000110010 and the sub - bands are [18000, 18375, 18750, 19125, 19500, 19875, 20250, 20625, 21000, 21375, 21750], based on the arrangement of the coding information, the first 0 can be mapped to the first sub - band 18000, the second 0 can be mapped to the second sub - band 18375... the tenth 1 can be mapped to the tenth sub - band 21375, etc.
[0158] Another exemplarily, Figure 3A is a spectrum schematic diagram of audio shown according to an exemplary embodiment Figure 1 , such as Figure 3A shown. In the figure, the solid line 303 is the pre - processed spatial audio data, and the dashed line 304 is the modulated audio data. It can be seen that there is no difference in the low - frequency signal part (i.e., the rendered spatial audio data), but in the high - frequency signal part of 18KHZ - 22KHZ (i.e., the target band), the coding information of the first rotation information is added to the modulated audio data. For example, it is 1 around - 55 decibels (dB) and 0 around - 75 dB, and the coding information of the first rotation information can be obtained as 10111010111.
[0159] In the embodiments of the present disclosure, based on the frame length of the first rotation information, the target band can be divided into at least one sub - band; then, based on the carrying parameter of the target band, the first rotation information is format - converted to obtain the coding information corresponding to the first rotation information; and based on the number of the sub - bands, the coding information is split to obtain at least one sub - coding information; finally, based on the arrangement of the coding information, each of the sub - coding information is mapped to the corresponding sub - band to obtain the modulated audio data. In this way, the coding information corresponding to the first rotation information can be mapped to the target band to generate the modulated audio data carrying the first rotation information, so that without modifying the existing coding scheme of the Bluetooth encoder, the modulated audio data can be transmitted to the wearable device, improving the practicability.
[0160] In some embodiments, the method further includes:
[0161] Determine the audio data within the highest frequency range in the rendered spatial audio data as the audio data within the target frequency band.
[0162] It should be noted that, in order to reduce the impact on the rendered spatial audio data and ensure the texture of the target spatial audio data generated by the wearable device, the audio data within the highest frequency range in the rendered spatial audio data can be determined as the audio data within the target frequency band, and then the audio data within the target frequency band is filtered to obtain preprocessed spatial audio data; finally, the first rotation information is mapped to the target frequency band in the preprocessed spatial audio data to obtain modulated audio data.
[0163] Here, the size of the highest frequency range can be set arbitrarily, and the embodiments of the present disclosure do not limit this.
[0164] In some embodiments, the highest frequency range is a set of high-frequency frequencies exceeding the user's perception. For example, since the hearing limit of the human ear is the frequency band of 20KHZ, but in fact most people can only hear the frequency band of 16KHZ, therefore, the frequency band above 17KHZ can be determined as the highest frequency range, the audio data within the frequency band above 17KHZ is determined as the audio data within the target frequency band, and the audio data within the target frequency band is filtered to facilitate storing the encoded information corresponding to the first rotation information.
[0165] In other embodiments, the highest frequency range is a set of audio signals exceeding those generated by conventional audio devices. For example, it is difficult for conventional musical instruments to emit audio signals above 18KHZ, so the frequency band above 18KHZ can be determined as the highest frequency range, the audio data within the frequency band above 18KHZ is determined as the audio data within the target frequency band, and the audio data within the target frequency band is filtered to facilitate storing the encoded information corresponding to the first rotation information.
[0166] Exemplarily, the audio data corresponding to the frequency band of 18KHZ - 22KHZ in the rendered spatial audio data is determined as the audio data within the target frequency band, and the audio data within the target frequency band is filtered using a low-pass filter of 18KHZ to clear the audio data in the frequency band of 18KHZ - 22KHZ, thereby reducing the interference caused by the audio data within the target frequency band to the modulation of the first rotation information.
[0167] Figure 3B Figure 2 shows a schematic diagram of the audio spectrum according to an exemplary embodiment, as Figure 3B shown, the solid line 301 is the audio signal corresponding to the rendered spatial audio data, and the dashed line 302 is the audio signal corresponding to the preprocessed spatial audio data obtained through filtering. By filtering the audio data within the highest frequency range in the solid line 301, the dashed line 301 can be obtained.
[0168] In the embodiments of the present disclosure, the audio data within the highest frequency range in the rendered spatial audio data is determined as the audio data within the target frequency band. In this way, even if the first rotation information is mapped to the target frequency band, it is not easy to affect the rendered spatial audio data, which is beneficial to ensuring the texture of the target spatial audio data generated by the wearable device.
[0169] In some embodiments, the preprocessing the rendered spatial audio data to obtain preprocessed spatial audio data includes:
[0170] Performing a frequency shift process on the audio data within the target frequency band in the rendered spatial audio data to obtain the preprocessed spatial audio data.
[0171] It can be understood that in addition to filtering the audio data within the target frequency band to obtain the preprocessed spatial audio data, a frequency shift process can also be performed on the audio data within the target frequency band such that there is no audio data within the target frequency band in the rendered spatial audio data, and the preprocessed spatial audio data can also be obtained.
[0172] Exemplarily, a time-domain frequency doubling method and / or a spectral shift method can be used to perform a frequency shift process on the audio data within the target frequency band in the rendered spatial audio data.
[0173] In some embodiments, a full-wave repetition time-domain frequency doubling method or a half-wave repetition (such as half-wave reverse repetition or half-wave forward repetition) time-domain frequency doubling method can be used to perform a frequency shift process on the audio data within the target frequency band.
[0174] In other embodiments, the spectrum of the audio data within the target frequency band can be highly faithfully shifted to any frequency band without audio data such that the first rotation information can be mapped to the target frequency band after the spectral shift process.
[0175] In the embodiments of the present disclosure, performing a frequency shift process on the audio data within the target frequency band in the rendered spatial audio data can also obtain the preprocessed spatial audio data. Thus, by performing a frequency shift process on the audio data within the target frequency band, it is convenient to map the first rotation information to the target frequency band, which is beneficial to obtaining modulated audio data carrying the first rotation information.
[0176] Figure 4 is Flowchart 2 of an audio processing method shown according to an exemplary embodiment. As Figure 4 shown, this method is applied to a wearable device, and this method mainly includes the following steps:
[0177] In step 401, send the first rotation information of the wearable device to the electronic device;
[0178] In step 402, receive the modulated audio data sent by the electronic device; wherein, the modulated audio data is obtained by the electronic device modulating the rendered spatial audio data based on the first rotation information, and the rendered spatial audio data is obtained by the electronic device rendering the initial audio data based on the first rotation information;
[0179] In step 403, demodulate the modulated audio data to obtain the rendered spatial audio data and the first rotation information;
[0180] In step 404, obtain the second rotation information of the wearable device;
[0181] In step 405, based on the first rotation information and the second rotation information, render the rendered spatial audio data to obtain the target spatial audio data.
[0182] Here, this method can be applied to an electronic device. The electronic device may include a terminal device. Among them, the terminal device may include a mobile terminal and a fixed terminal. For example, mobile phones, tablet computers, personal digital assistants, laptop computers, desktop computers, wearable devices, smart speakers, televisions, and in-vehicle terminals, etc.
[0183] Among them, the wearable device may include: a portable device directly worn on the user's body or integrated into the user's accessories. For example, it may include: wearable devices supported by the wrist (including watches, wristbands, etc.); wearable devices supported by the feet (including shoes, socks, etc.), wearable devices supported by the head (including headphones, glasses, helmets, headbands, etc.). It may also include: smart clothing, schoolbags, crutches, accessories, etc.
[0184] In some embodiments, the wearable device may have an audio playback function. Since the wearable device can be directly worn on the user's body, when the user's body rotates, the wearable device will also rotate with the user's body.
[0185] In the embodiments of the present disclosure, when the wearable device rotates, the first rotation information of the wearable device can be obtained and sent to the electronic device. Here, the first rotation information may include: information generated when the wearable device rotates. For example, it may include the rotation angle, rotation orientation, etc. of the wearable device. In some embodiments, a detection module (such as an angle detection module, an orientation detection module, etc.) may be set in the wearable device. When it is detected that the wearable device rotates, the detection module can sense the first rotation information of the wearable device.
[0186] Taking the first rotation information as the first rotation angle as an example, an angular motion detection module can be set in the wearable device, and the first rotation angle of the wearable device can be determined through the angular motion detection module. Among them, the angular motion detection module can include: an angular velocity sensor (gyroscope). Taking the first rotation information as the first rotation orientation as an example, an orientation detection module can be set in the wearable device, and the first rotation orientation of the wearable device can be determined through the orientation detection module. Among them, the orientation detection module can include: an orientation sensor.
[0187] In the embodiments of the present disclosure, after obtaining the first rotation information of the wearable device, the first rotation information can be sent to the electronic device. In some embodiments, a communication connection can be established between the wearable device and the electronic device. For example, a wired connection can be established between the wearable device and the electronic device. Among them, the wired connection can include: a data line connection; a wireless connection can also be established between the wearable device and the electronic device. Among them, the wireless connection can include: a short-range wireless connection (for example, a Bluetooth connection) and a long-range wireless connection (for example, a Wi-Fi connection). For example, a Bluetooth communication can be established between the electronic device and the wearable device.
[0188] In some embodiments, the electronic device is a device that outputs initial audio data, and the wearable device is a device that outputs target spatial audio data. Among them, the target spatial audio data is spatial audio obtained by rendering the initial audio data.
[0189] It can be understood that after receiving the modulated audio data, the wearable device can demodulate the modulated audio data to obtain the rendered spatial audio data and the first rotation information.
[0190] In some embodiments, after obtaining the first rotation information, the wearable device can obtain the second rotation information in real time, and determine the perception deviation parameter based on the first rotation information and the second rotation information; then render the rendered spatial audio data based on the perception deviation parameter to obtain the target spatial audio data.
[0191] Among them, the perception deviation parameter can be a deviation parameter used to characterize the deviation of the audio when the wearable device rotates.
[0192] In some embodiments, the perception deviation parameter can include: a delay deviation parameter and / or a sound pressure deviation parameter.
[0193] Therefore, based on the first rotation information and the second rotation information, determining the perception deviation parameter includes: determining the compensation rotation information based on the first rotation information and the second rotation information, and determining the delay deviation parameter based on the compensation rotation information and a preset correlation coefficient; and / or, obtaining the first sound pressure parameter corresponding to the first rotation information and the second sound pressure parameter corresponding to the second rotation information, and determining the sound pressure deviation parameter based on the first sound pressure parameter and the second sound pressure parameter.
[0194] Here, the delay deviation parameter may include a time parameter related to the rotation information. For example, it may include: Interaural Time Difference (ITD). The sound pressure deviation parameter may include a sound pressure parameter related to the rotation information. For example, it may include: Interaural Level Difference (ILD).
[0195] Exemplarily, the compensation rotation information may be determined based on the difference between the first rotation information and the second rotation information. For example, if the first rotation information is the first rotation angle, the second rotation information is the second rotation angle, and the compensation rotation information is the compensation rotation angle, then the difference between the first rotation angle and the second rotation angle may be determined as the compensation rotation angle.
[0196] In another example, the first rotation information may be weighted based on the first weight coefficient to obtain the first weighted rotation information; then the second rotation information may be weighted based on the second weight coefficient to obtain the second weighted rotation information; finally, the compensation rotation information may be determined based on the difference between the first weighted rotation information and the second rotation information.
[0197] In some embodiments, the first weight coefficient and the second weight coefficient may be set as needed. For example, they may be set based on experience, or based on historical data, etc., and no specific limitation is made here.
[0198] In some embodiments, the first rotation information may be the first rotation angle, the second rotation information may be the second rotation angle, and the compensation rotation information may be the compensation rotation angle. In this way, the first rotation angle may be weighted based on the first weight coefficient to obtain the first weighted rotation angle, and the second rotation angle may be weighted based on the second weight coefficient to obtain the second weighted rotation angle. Then, the compensation rotation angle may be determined based on the difference between the first weighted rotation angle and the second rotation angle. For example, the difference between the first weighted rotation angle and the second rotation angle may be directly determined as the compensation rotation angle. Or, for another example, the difference between the first weighted rotation angle and the second rotation angle may be weighted, and the result of the weighting may be determined as the compensation rotation angle, etc.
[0199] Since the first rotation information can be weighted based on the first weight coefficient, the second rotation information can be weighted based on the second weight coefficient, and the compensation rotation information can be obtained based on the difference between the weighted first weighted rotation information and the weighted second weighted rotation information, it is possible to adjust the proportions of the first rotation information and the second rotation information in the compensation rotation information by adjusting the first weight coefficient and the second weight coefficient. That is, the user can adjust the proportions of the first rotation information and the second rotation information during the audio rendering process as needed, thereby improving the flexibility of audio rendering.
[0200] After obtaining the compensation rotation information, the delay deviation parameter can be determined based on the compensation rotation information and a preset correlation coefficient. Among them, the preset correlation coefficient can be set as needed or obtained through experiments. For example, it can be determined based on the sound propagation speed and the head radius during the testing process. In some embodiments, the calculation formula of the delay deviation parameter is as follows:
[0201]
[0202]
[0203] In Formulas (2) and (3), ITD represents the delay deviation parameter, λ represents the preset correlation coefficient, θ represents the compensation rotation information, c represents the sound speed, and a represents the head radius.
[0204] In some embodiments, the mapping relationship between the rotation information and the sound pressure parameter can be preset. In this way, when the rotation information is determined, the corresponding sound pressure parameter can be determined according to the rotation information and this mapping relationship. Taking the rotation information as the rotation angle as an example, during the testing process, the impulse responses corresponding to all rotation angles of the HRTF can be convolved with the white noise signal w to obtain the rendered white noise signal c corresponding to the rotation angle, and then the sound pressure parameters corresponding to each rotation angle can be determined based on the rendered white noise signal c, and the mapping relationship between each rotation angle and the sound pressure parameter can be established.
[0205] In some embodiments, the calculation formula of the sound pressure deviation parameter is as follows:
[0206]
[0207] In Formula (4), ILD represents the sound pressure deviation parameter, represents the first sound pressure parameter, represents the second sound pressure parameter.
[0208] Among them, the sound pressure parameter can be determined according to the rotation information. In some embodiments, the calculation formula of the sound pressure parameter is as follows:
[0209]
[0210] In formula (5), represents the sound pressure parameter, represents the rendered white noise, and m is the audio data length of c.
[0211] In some embodiments, the impulse responses for all rotation angles corresponding to the HRTF can be convolved with the white noise signal w to obtain the rendered white noise signal c corresponding to the rotation angle. The calculation formula for the rendered white noise signal is as follows:
[0212]
[0213] In formula (6), represents the rendered white noise signal, w represents the initial white noise signal, and φ n represents the rotation angle corresponding to the HRTF.
[0214] In the embodiments of the present disclosure, the delay deviation parameter can be determined based on the compensation rotation information determined based on the first rotation information and the second rotation information, and the sound pressure deviation parameter can be determined based on the first sound pressure parameter corresponding to the first rotation information and the second sound pressure parameter corresponding to the second rotation information. The first rotation information and the second rotation information can be fully considered in the audio rendering process. In this way, even if the wearable device rotates multiple times during the audio rendering process, it does not affect the spatial sense of the finally output target spatial audio data.
[0215] In the embodiments of the present disclosure, the rendered spatial audio data can be delayed based on the delay deviation parameter, and / or the sound pressure of the rendered spatial audio data can be adjusted based on the sound pressure deviation parameter to obtain the target spatial audio data.
[0216] Here, taking the perception deviation parameter as the delay deviation parameter as an example, the rendered spatial audio data can be delayed based on the delay deviation parameter. In some embodiments, delaying the rendered spatial audio data may include: delaying the rendered spatial audio data corresponding to the channels with a longer distance. For example, if the left ear is farther from the sound source, the rendered spatial audio data corresponding to the left ear channel can be delayed.
[0217] Taking the perception deviation parameter as the sound pressure deviation parameter as an example, the sound pressure of the rendered spatial audio data can be adjusted based on the sound pressure deviation parameter to obtain the target spatial audio data. Taking the perception deviation parameter including the delay deviation parameter and the sound pressure deviation parameter as an example, the rendered spatial audio data can be delayed based on the delay deviation parameter, and the sound pressure of the rendered spatial audio data can be adjusted based on the sound pressure deviation parameter to obtain the target spatial audio data.
[0218] In an embodiment of the present disclosure, first, the first rotation information of the wearable device is sent to the electronic device; then, the modulated audio data sent by the electronic device is received; and the modulated audio data is demodulated to obtain the rendered spatial audio data and the first rotation information; then, the second rotation information of the wearable device is obtained; and finally, based on the first rotation information and the second rotation information, the rendered spatial audio data is rendered to obtain the target spatial audio data.
[0219] In the technical solution of the present disclosure, on the one hand, the modulated audio data sent by the electronic device is received by encoding, and then the wearable device can demodulate the modulated audio data to obtain the rendered spatial audio data and the first rotation information, thereby avoiding the secondary delay caused by the return of the first rotation information in other solutions and without modifying the original encoding scheme, which is beneficial to improving the practicability; on the other hand, audio rendering can be performed at both ends of the electronic device and the wearable device, which can improve the accuracy and sense of space of audio rendering, and further improve the texture of the generated target spatial audio data.
[0220] In some embodiments, the demodulating the modulated audio data to obtain the first rotation information includes:
[0221] Detecting sub-encoding information in each sub-band located in the target band of the modulated audio data;
[0222] Combining the sub-encoding information based on the arrangement of each sub-band to obtain the encoding information;
[0223] Converting the format of the encoding information to obtain the first rotation information.
[0224] It should be noted that since the electronic device sends the modulated audio data carrying the first rotation information to the wearable device, in order for the wearable device to obtain the first rotation information, it can first detect the sub-encoding information in each sub-band located in the target band of the modulated audio data, and then obtain the first rotation information based on each sub-encoding information.
[0225] In some embodiments, it is possible to detect whether there is a signal in each sub-band of the target band to determine the sub-encoding information in each sub-band. For example, when it is detected that there is a signal in a sub-band of the target band, it can be determined that the sub-encoding information of this sub-band is 1; when it is detected that there is no signal in a sub-band of the target band, it can be determined that the sub-encoding information of this sub-band is 0.
[0226] It can be understood that, for the convenience of calculating the first rotation information, after obtaining the sub-encoding information in each sub-band of the target frequency band, the sub-encoding information can be combined according to the arrangement of each sub-band to obtain the encoding information corresponding to the first rotation information.
[0227] Exemplarily, if the target frequency band includes 11 sub-bands, i.e., [18000, 18375, 18750, 19125, 19500, 19875, 20250, 20625, 21000, 21375, 21750], and it is detected that there is no signal in the first sub-band 18000, then the first sub-encoding information is 0; if there is no signal in the second sub-band 18375, then the second sub-encoding information is also 0... if there is a signal in the eleventh sub-band 21750, then the eleventh sub-encoding information is 1; then the obtained binary data is combined according to the arrangement of each sub-band to obtain the encoding information.
[0228] Here, after obtaining the encoding information corresponding to the first rotation information, the first encoding information is format-converted to obtain the first rotation information.
[0229] Exemplarily, taking the first rotation information as the first rotation angle, by detecting each sub-band in the target frequency band, the encoding information is obtained, such as 00000101001, and the encoding information is format-converted to obtain the first rotation angle of 41°.
[0230] In the embodiments of the present disclosure, by detecting the sub-encoding information in each sub-band of the modulation audio data located in the target frequency band; and then combining the detected sub-encoding information based on the arrangement of each sub-band to obtain the encoding information. In this way, without modifying the Bluetooth decoding scheme, the wearable device can demodulate the modulation audio data to obtain the first rotation information, which is beneficial to improving the practicability.
[0231] In some embodiments, the obtaining of the rendered spatial audio data includes:
[0232] Performing filtering processing on each of the sub-encoding information located in the target frequency band to obtain the preprocessed spatial audio data;
[0233] Determining the rendered spatial audio data based on the preprocessed spatial audio data.
[0234] It should be noted that since the modulated audio data carries the first rotation information, after determining the first rotation information from the modulated audio data, the first rotation information in the modulated audio data can be filtered to facilitate obtaining the rendered spatial audio data. Therefore, in the embodiments of the present disclosure, each sub-encoded information located in the target frequency band is filtered to obtain preprocessed spatial audio data, and then the rendered spatial audio data is further obtained.
[0235] In some embodiments, a filter model can be established according to the signal characteristics of the rendered spatial audio data and the components to be filtered (i.e., the audio data located in the target frequency band in the rendered spatial audio data). Then, according to the signal characteristics and the filter model, corresponding filtering parameters are selected. Among them, the filtering parameters can affect the performance of the filter, such as the cut-off frequency, the filter order, etc. Then, the filter model is applied to the audio data in the target frequency band, and time-domain filtering or frequency-domain filtering is implemented through methods such as convolution operation and frequency-domain transformation. After filtering the audio data in the target frequency band, the filtered audio data can be evaluated based on evaluation metrics, such as signal-to-noise ratio, spectral analysis, or time-domain waveform.
[0236] In some other embodiments, if the filtered audio data does not meet the preset conditions, the filtering parameters of the filter can also be adjusted according to the evaluation results, and the audio data located in the target frequency band in the rendered spatial audio data is filtered again, so as to form preprocessed spatial audio data that meets the preset conditions.
[0237] Here, the filter can be linear or non-linear; the filter model can be a low-pass filter, a high-pass filter, or a band-pass filter, etc., and the embodiments of the present disclosure do not limit this.
[0238] In the embodiments of the present disclosure, by filtering each sub-encoded information located in the target frequency band, preprocessed audio data can be obtained. Then, based on the preprocessed audio data, the rendered spatial audio data can be determined. In this way, by filtering each sub-encoded information in the target frequency band, it is beneficial to obtain accurate preprocessed audio data, and further obtain accurate rendered spatial audio data.
[0239] In some embodiments, the determining the rendered spatial audio data based on the preprocessed spatial audio data includes:
[0240] Determining the preprocessed spatial audio data as the rendered spatial audio data; or
[0241] Performing a frequency shift process on the preprocessed spatial audio data to obtain the rendered spatial audio data.
[0242] It can be understood that since the electronic device can directly filter the audio data in the target frequency band in the rendered spatial audio data to obtain preprocessed spatial audio data; or it can perform frequency shift processing on the audio data in the target frequency band in the rendered spatial audio data to obtain preprocessed spatial audio data. Therefore, after the wearable device filters each sub-encoding information in the target frequency band to obtain preprocessed spatial audio data, it can directly determine the preprocessed spatial audio data as the rendered spatial audio data; or it can perform frequency shift processing on the preprocessed spatial audio data to obtain the rendered spatial audio data.
[0243] In the embodiments of the present disclosure, the preprocessed spatial audio data can be directly determined as the rendered spatial audio data; or frequency shift processing can be performed on the preprocessed spatial audio data to obtain the rendered spatial audio data, so as to be able to obtain accurate rendered spatial audio data and accurate target spatial audio data.
[0244] Figure 5 It is the flowchart three of the audio processing method shown according to an exemplary embodiment, as Figure 5 shown, when the wearable device 501 rotates, the wearable device 501 can use the motion detection module to obtain the first rotation information and send the first rotation information to the electronic device 502 through wireless communication. After receiving the first rotation information, the electronic device 502 can use the spatial transfer function to render the initial audio data based on the first rotation information to obtain the rendered spatial audio data, and perform modulation processing on the rendered spatial audio data based on the first rotation information to obtain the modulated audio data, and send the encoded data to the wearable device 501 through wireless communication. After receiving the encoded data, the wearable device 501 can demodulate the modulated audio data to obtain the rendered spatial audio data and the first rotation information, and determine the perception deviation parameter based on the first rotation information and the obtained second rotation information, and then perform audio rendering on the rendered spatial audio data based on the perception deviation parameter to obtain the target spatial audio data, and output the target spatial audio data to the user.
[0245] On the one hand, the technical solution of the present disclosure enables the electronic device to directly modulate the rendered spatial audio based on the first rotation information to obtain the modulated audio data carrying the first rotation information. In this way, without modifying the existing Bluetooth encoding, the modulated audio data can be sent to the wearable device, thereby reducing the latency brought to the user by audio rendering while improving the practicality. On the other hand, audio rendering can be performed at both ends of the electronic device and the wearable device simultaneously, which is beneficial to improving the accuracy and spatial sense of audio rendering, and further improving the texture of the generated target spatial audio data.
[0246] In the embodiments of the present disclosure, since the HRTF rendering of the first rotation information (e.g., the first rotation angle) is performed in an electronic device (e.g., a mobile phone), the externalization effect and the positioning of the first rotation information are both guaranteed. At the same time, after the time (e.g., ITD) and sound pressure (e.g., ILD) are rendered in a wearable device (e.g., headphones), the azimuth can be closer to the second rotation information, avoiding obvious delays. At the same time, the rendered spatial audio data and the first rotation information are directly transmitted in an encoded manner, and there will be no delay between the two, avoiding the secondary delay caused by other solutions returning the first rotation information, thereby reducing the sense of delay brought to the user by audio rendering. The technical solution of the present disclosure can ensure that the delay of the head tracking function is small enough, the tracking accuracy is higher, the sound quality is better, and the sense of space is better without increasing the power consumption computing power of the wearable device and modifying the system delay.
[0247] Figure 6 is a block diagram of an audio processing device shown according to an exemplary embodiment. As Figure 6 shown, the audio processing device 600 mainly includes:
[0248] A first receiving module 601, configured to receive the first rotation information sent by the wearable device;
[0249] A first rendering module 602, configured to render the initial audio data based on the first rotation information to obtain rendered spatial audio data;
[0250] A modulation module 603, configured to perform modulation processing on the rendered spatial audio data based on the first rotation information to obtain modulated audio data; wherein, the modulated audio data carries the first rotation information;
[0251] A first sending module 604, configured to send the modulated audio data to the wearable device, so that the wearable device performs rendering processing on the rendered spatial audio data based on the first rotation information and the obtained second rotation information to obtain target spatial audio data.
[0252] In some embodiments, the modulation module 603 includes:
[0253] A preprocessing module, configured to preprocess the rendered spatial audio data to obtain preprocessed spatial audio data; wherein, the preprocessed spatial audio data does not include audio data with a target frequency band;
[0254] A mapping module, configured to map the first rotation information to the target frequency band in the preprocessed spatial audio data to obtain the modulated audio data.
[0255] In some embodiments, the preprocessing module is specifically configured to:
[0256] Perform filtering on the audio data in the target frequency band in the rendered spatial audio data to obtain the preprocessed spatial audio data.
[0257] In some embodiments, the mapping module is specifically configured to:
[0258] Divide the target frequency band into at least one sub - frequency band based on the frame length of the first rotation information;
[0259] Perform format conversion on the first rotation information based on the carrying parameters of the target frequency band to obtain the encoded information corresponding to the first rotation information;
[0260] Split the encoded information based on the number of sub - frequency bands to obtain at least one sub - encoded information;
[0261] Map each of the sub - encoded information to the corresponding sub - frequency band based on the arrangement mode of the encoded information to obtain the modulated audio data.
[0262] In some embodiments, the modulation module 603 further includes:
[0263] The first determination module is configured to determine the audio data in the highest frequency range in the rendered spatial audio data as the audio data in the target frequency band.
[0264] In some embodiments, the preprocessing module is specifically configured to:
[0265] Perform frequency shift on the audio data in the target frequency band in the rendered spatial audio data to obtain the preprocessed spatial audio data.
[0266] Figure 7 It is a block diagram of an audio processing device shown according to an exemplary embodiment. As Figure 7 shown, the audio processing device 700 mainly includes:
[0267] The second sending module 701 is configured to send the first rotation information of the wearable device to the electronic device;
[0268] The second receiving module 702 is configured to receive the modulated audio data sent by the electronic device; wherein, the modulated audio data is obtained by the electronic device performing modulation processing on the rendered spatial audio data based on the first rotation information, and the rendered spatial audio data is obtained by the electronic device performing rendering on the initial audio data based on the first rotation information;
[0269] A demodulation module 703, configured to demodulate the modulated audio data to obtain the rendered spatial audio data and the first rotation information;
[0270] An acquisition module 704, configured to acquire second rotation information of the wearable device;
[0271] A second rendering module 705, configured to perform rendering processing on the rendered spatial audio data based on the first rotation information and the second rotation information to obtain target spatial audio data.
[0272] In some embodiments, the demodulation module 703 is specifically configured to:
[0273] Detect sub-encoding information in each sub-band located in a target frequency band of the modulated audio data;
[0274] Based on the arrangement mode of each sub-band, combine each sub-encoding information to obtain the encoding information;
[0275] Perform format conversion on the encoding information to obtain the first rotation information.
[0276] In some embodiments, the demodulation module 703 includes:
[0277] A filtering module, configured to perform filtering processing on each sub-encoding information located in the target frequency band to obtain the preprocessed spatial audio data;
[0278] A second determination module, configured to determine the rendered spatial audio data based on the preprocessed spatial audio data.
[0279] In some embodiments, the second determination module is specifically configured to:
[0280] Determine the preprocessed spatial audio data as the rendered spatial audio data; or
[0281] Perform frequency shift processing on the preprocessed spatial audio data to obtain the rendered spatial audio data.
[0282] Regarding the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated here.
[0283] Figure 8 It is a structural block diagram of a device 800 shown according to an exemplary embodiment. For example, the device 800 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.
[0284] With reference to Figure 8 , device 800 may include one or more of the following components: processing component 802, memory 804, power component 806, multimedia component 808, audio component 810, input / output (I / O) interface 812, sensor component 814, and communication component 816.
[0285] Processing component 802 generally controls the overall operation of device 800, such as operations associated with at least one of display, telephone call, data communication, camera operation, and recording operation. Processing component 802 may include one or more processors 820 to execute instructions to complete all or part of the steps of the above-described methods. In addition, processing component 802 may include one or more modules to facilitate interaction between processing component 802 and other components. For example, processing component 802 may include a multimedia module to facilitate interaction between multimedia component 808 and processing component 802.
[0286] Memory 804 is configured to store various types of data to support the operation of device 800. Examples of such data include at least one of the following: instructions for any application or method operating on device 800, contact data, phone book data, messages, pictures, and videos. Memory 804 may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read only memory (EEPROM), erasable programmable read only memory (EPROM), programmable read only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disk.
[0287] Power component 806 provides power to the various components of device 800. Power component 806 may include at least one of the following: a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for device 800.
[0288] The multimedia component 808 includes a screen that provides an output interface between the device 800 and the user. In some embodiments, the screen may include a Liquid Crystal Display (LCD) and a Touch Panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can sense not only the boundaries of the touch or swipe actions but also detect the duration and pressure associated with the touch or swipe operations. In some embodiments, the multimedia component 808 includes a front camera and / or a rear camera. When the device 800 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each of the front camera and the rear camera can be a fixed optical lens system or have focal length and optical zoom capabilities.
[0289] The audio component 810 is configured to output and / or input audio signals. For example, the audio component 810 includes a Microphone (MIC) that is configured to receive external audio signals when the device 800 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory 804 or transmitted via the communication component 816. In some embodiments, the audio component 810 further includes a speaker for outputting audio signals.
[0290] The I / O interface 812 provides an interface between the processing component 802 and the peripheral interface module, and the peripheral interface module can be a keyboard, a click wheel, buttons, etc. These buttons can include, but are not limited to: a home button, a volume button, a power button, and a lock button.
[0291] The sensor assembly 814 includes one or more sensors for providing a status assessment of various aspects of the device 800. For example, the sensor assembly 814 can detect the on / off state of the device 800, the relative positioning of components, such as the display and keypad of the device 800. The sensor assembly 814 can also detect a change in the position of the device 800 or a component of the device 800, the presence or absence of user contact with the device 800, the orientation or acceleration / deceleration of the device 800, and a change in the temperature of the device 800. The sensor assembly 814 can include proximity sensors configured to detect the presence of nearby objects without any physical contact. The sensor assembly 814 can also include light sensors, such as complementary metal oxide semiconductor (CMOS) or charge coupled device (CCD) image sensors, for use in imaging applications. In some embodiments, the sensor assembly 814 can also include, but is not limited to, at least one of the following: an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, and a temperature sensor.
[0292] The communication component 816 is configured to facilitate communication between the device 800 and other devices in a wired or wireless manner. The device 800 can access a wireless network based on communication standards, such as Wi-Fi, 4G, 5G, or a combination thereof. In an exemplary embodiment, the communication component 816 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 816 also includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra wide band (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0293] In an exemplary embodiment, the apparatus 800 may be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components.
[0294] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 804 including executable instructions or a computer program, and the above instructions or computer program can be executed by a processor 820 of the apparatus 800 to complete the above method. For example, the non-transitory computer-readable storage medium may be a ROM, a random access memory (RAM), a compact disc read-only memory (CD-ROM), a magnetic tape, a floppy disk, and an optical data storage device, etc.
[0295] A non-transitory computer-readable storage medium, when the instructions in the storage medium are executed by a processor of a mobile terminal, enables the mobile terminal to execute an audio processing method, and the method includes:
[0296] Receiving first rotation information sent by a wearable device;
[0297] Rendering initial audio data based on the first rotation information to obtain rendered spatial audio data;
[0298] Performing modulation processing on the rendered spatial audio data based on the first rotation information to obtain modulated audio data; wherein, the modulated audio data carries the first rotation information;
[0299] Sending the modulated audio data to the wearable device, so that the wearable device performs rendering processing on the rendered spatial audio data based on the first rotation information and the obtained second rotation information to obtain target spatial audio data.
[0300] Figure 9 It is a block diagram of an apparatus 1900 for audio processing shown according to an exemplary embodiment. For example, the apparatus 1900 may be provided as a server. Refer to Figure 9, the apparatus 1900 includes a processing component 1922, which further includes one or more processors, and memory resources represented by a memory 1932 for storing instructions executable by the processing component 1922, such as application programs. The application programs stored in the memory 1932 may include one or more modules each corresponding to a set of instructions. In addition, the processing component 1922 is configured to execute instructions to perform the above audio processing method, the method including:
[0301] Receiving first rotation information sent by the wearable device;
[0302] Rendering initial audio data based on the first rotation information to obtain rendered spatial audio data;
[0303] Performing modulation processing on the rendered spatial audio data based on the first rotation information to obtain modulated audio data; wherein, the modulated audio data carries the first rotation information;
[0304] Sending the modulated audio data to the wearable device, so that the wearable device performs rendering processing on the rendered spatial audio data based on the first rotation information and the obtained second rotation information to obtain target spatial audio data.
[0305] The apparatus 1900 may further include a power component 1926 configured to perform power management of the apparatus 1900, a wired or wireless network interface 1950 configured to connect the apparatus 1900 to a network, and an input / output (I / O) interface 1958. The apparatus 1900 may operate based on an operating system stored in the memory 1932, such as Windows ServerTM, MacOS XTM, UnixTM, LinuxTM, FreeBSDTM or the like.
[0306] Those skilled in the art will readily conceive of other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include known common knowledge or conventional technical means in the technical field not disclosed by the present disclosure. The specification and examples are only illustrative, and the true scope and spirit of the present disclosure are pointed out by the following claims.
[0307] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.
Claims
1. An audio processing method, characterized in that, Including: Receiving first rotation information sent by a wearable device; Rendering initial audio data based on the first rotation information to obtain rendered spatial audio data; Performing modulation processing on the rendered spatial audio data based on the first rotation information to obtain modulated audio data; wherein, the modulated audio data carries the first rotation information; Sending the modulated audio data to the wearable device, so that the wearable device performs rendering processing on the rendered spatial audio data based on the first rotation information and the obtained second rotation information to obtain target spatial audio data.
2. The method according to claim 1, characterized in that, The performing modulation processing on the rendered spatial audio data based on the first rotation information to obtain modulated audio data includes: Performing preprocessing on the rendered spatial audio data to obtain preprocessed spatial audio data; wherein, the preprocessed spatial audio data does not include audio data having a target frequency band; Mapping the first rotation information to the target frequency band in the preprocessed spatial audio data to obtain the modulated audio data.
3. The method according to claim 2, wherein The performing preprocessing on the rendered spatial audio data to obtain preprocessed spatial audio data includes: Performing filtering processing on the audio data in the rendered spatial audio data that is located in the target frequency band to obtain the preprocessed spatial audio data.
4. The method according to claim 2, characterized in that, The mapping the first rotation information to the target frequency band in the preprocessed spatial audio data to obtain the modulated audio data includes: Dividing the target frequency band into at least one sub-frequency band based on the frame length of the first rotation information; Performing format conversion on the first rotation information based on the carrying parameter of the target frequency band to obtain encoded information corresponding to the first rotation information; Splitting the encoded information based on the number of the sub-frequency bands to obtain at least one sub-encoded information; Mapping each of the sub-encoded information to the corresponding sub-frequency band based on the arrangement manner of the encoded information to obtain the modulated audio data.
5. The method according to claim 2, wherein The method further includes: Determining the audio data in the rendered spatial audio data that is located in the highest frequency range as the audio data located in the target frequency band.
6. The method according to claim 2, characterized in that, The performing preprocessing on the rendered spatial audio data to obtain preprocessed spatial audio data includes: Performing frequency shift processing on the audio data in the rendered spatial audio data that is located in the target frequency band to obtain the preprocessed spatial audio data.
7. An audio processing method, characterized in that, Including: Sending the first rotation information of the wearable device to an electronic device; Receiving the modulated audio data sent by the electronic device; wherein, the modulated audio data is obtained by the electronic device performing modulation processing on the rendered spatial audio data based on the first rotation information, and the rendered spatial audio data is obtained by the electronic device rendering the initial audio data based on the first rotation information; Performing demodulation processing on the modulated audio data to obtain the rendered spatial audio data and the first rotation information; Obtaining the second rotation information of the wearable device; Performing rendering processing on the rendered spatial audio data based on the first rotation information and the second rotation information to obtain target spatial audio data.
8. The method according to claim 7, wherein Demodulating the modulated audio data to obtain the first rotation information includes: Detecting sub-encoding information in each sub-band within a target frequency band of the modulated audio data; Combining the sub-encoding information based on the arrangement of the sub-bands to obtain the encoding information; Converting the format of the encoding information to obtain the first rotation information.
9. The method according to claim 8, characterized in that, Obtaining the rendered spatial audio data includes: Filtering each sub-encoding information within the target frequency band to obtain the preprocessed spatial audio data; Determining the rendered spatial audio data based on the preprocessed spatial audio data.
10. The method according to claim 9, characterized in that, Determining the rendered spatial audio data based on the preprocessed spatial audio data includes: Determining the preprocessed spatial audio data as the rendered spatial audio data; or Performing a frequency shift on the preprocessed spatial audio data to obtain the rendered spatial audio data.
11. An audio processing device, characterized in that, Includes: A first receiving module configured to receive the first rotation information sent by a wearable device; A first rendering module configured to render initial audio data based on the first rotation information to obtain rendered spatial audio data; A modulation module configured to perform modulation processing on the rendered spatial audio data based on the first rotation information to obtain modulated audio data; wherein the modulated audio data carries the first rotation information; A first sending module configured to send the modulated audio data to the wearable device, so that the wearable device renders the rendered spatial audio data based on the first rotation information and the acquired second rotation information to obtain target spatial audio data.
12. An audio processing device, characterized in that, Includes: A second sending module configured to send the first rotation information of the wearable device to an electronic device; A second receiving module configured to receive the modulated audio data sent by the electronic device; wherein the modulated audio data is obtained by the electronic device performing modulation processing on the rendered spatial audio data based on the first rotation information, and the rendered spatial audio data is obtained by the electronic device rendering initial audio data based on the first rotation information; A demodulation module configured to perform demodulation processing on the modulated audio data to obtain the rendered spatial audio data and the first rotation information; An acquisition module configured to acquire the second rotation information of the wearable device; A second rendering module configured to perform rendering processing on the rendered spatial audio data based on the first rotation information and the second rotation information to obtain target spatial audio data.
13. An electronic device, characterized in that, Includes: A processor; A memory configured to store instructions executable by the processor; Wherein the processor is configured to: when executed, implement the steps in any one of the audio processing methods in claims 1 to 10 above.
14. A non-transitory computer-readable storage medium, when the instructions in the storage medium are executed by the processor of the audio processing device, enabling the device to execute any one of the audio processing methods in claims 1 to 10 above.