A method and apparatus for determining a sound effect of a spatial audio

By extracting the amplitude and time delay characteristics of the left and right channel audio data and calculating the changing trends, the problem of the inability to quantify and compare sound effects caused by subjective evaluation in the existing technology is solved, and the quantitative evaluation and algorithm optimization of spatial audio effects are realized.

CN116528142BActive Publication Date: 2026-08-25HISENSE VISUAL TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210082265.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-24
Publication Date
2026-08-25
Estimated Expiration
2042-01-24

AI Technical Summary

Technical Problem

Existing spatial audio sound effect evaluation methods rely on subjective auditory perception and cannot be quantitatively compared, resulting in large differences in perception among different people. They also lack objective quantitative analysis and graphical display, making it difficult to optimize spatial audio algorithms.

Method used

By acquiring audio data from the left and right channels, audio features such as amplitude and time delay are extracted, the changing trends are calculated, and the sound effects are judged to be normal. Quantitative data analysis is provided to optimize spatial audio algorithms.

Benefits of technology

It enables quantitative comparison and graphical display of spatial audio effects, improving the objectivity and accuracy of sound effect evaluation and guiding algorithm optimization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116528142B_ABST
    Figure CN116528142B_ABST
Patent Text Reader

Abstract

The application discloses a method and device for determining the sound effect of spatial audio, to solve the problem that the sound effect cannot be quantitatively compared due to the subjective judgment of the sound effect of spatial audio. The method provided by the application comprises: acquiring left channel audio data and right channel audio data output by a first electronic device; extracting audio features from the left channel audio data, right channel audio data and original audio data respectively; determining a first change trend of the audio features of the left channel audio data relative to the audio features of the original audio data within a first pose change range, and determining a second change trend of the audio features of the right channel audio data relative to the audio features of the original audio data within the first pose change range; and determining whether the sound effect of spatial audio is normal according to whether the first change trend and the second change trend satisfy a change condition corresponding to the first pose change range.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to audio processing technology, and more particularly to a method and apparatus for determining the sound effects of spatial audio. Background Technology

[0002] Spatial audio technology can be applied to Virtual Reality (VR) or Augmented Reality (AR) devices. It achieves spatial audio effects through head tracking using sensors such as laser positioning and gyroscopes. With the increasing use of VR devices, 360° videos are gradually dominating traditional media distribution channels, and the demand for realistic audio is stronger than ever before. Current methods for evaluating the sound effects of spatial audio are based on subjective auditory perception. This involves a person wearing the device moving up, down, left, and right to assess the effect and analyze the quality of the spatial audio algorithm. This method optimizes the spatial audio algorithm through personal subjective judgment. Different people have vastly different perceptions, making quantitative comparison impossible. At a microscopic level, it can only provide a rough judgment on the trend of location and distance, lacking objective quantitative comparison of trends and intuitive graphical representation. Summary of the Invention

[0003] This application provides a method and apparatus for determining the sound effects of spatial audio, in order to solve the problem that the sound effects cannot be quantitatively compared due to relying on subjective judgment to determine the sound effects of spatial audio.

[0004] In a first aspect, embodiments of this application provide a method for determining the sound effects of spatial audio, including:

[0005] The left channel audio data and right channel audio data output by the first electronic device are acquired; the first electronic device carries the spatial audio algorithm to be tested; the left channel audio data is obtained by the first electronic device using the spatial audio algorithm to be tested to simulate the original audio data played by the audio playback device heard by the left ear under multiple pose change ranges; the right channel audio data is obtained by the first electronic device using the spatial audio algorithm to be tested to simulate the original audio data played by the audio playback device heard by the right ear under the multiple pose change ranges.

[0006] Audio features are extracted from the left channel audio data, the right channel audio data, and the original audio data respectively; a first trend of change of the audio features of the left channel audio data relative to the audio features of the original audio data is determined under the first pose change range; and a second trend of change of the audio features of the right channel audio data relative to the audio features of the original audio data is determined under the first pose change range; the first pose change range is one of the plurality of pose change ranges.

[0007] Based on whether the first and second trends of change meet the change conditions corresponding to the first pose change range, determine whether the sound effect of the spatial audio is normal.

[0008] Based on the above scheme, audio features can be extracted from the original audio data, left channel audio data, and right channel audio data, and the sound effects of spatial audio can be evaluated through these features. This scheme can calculate quantified data and present comparative analysis results in a digital and graphical format.

[0009] In one possible implementation, the audio features include amplitude and / or time delay.

[0010] Based on the above scheme, the sound effects of the left and right channel audio data can be judged from two aspects: amplitude and time delay.

[0011] In one possible implementation, the audio feature includes amplitude, the first trend of change is the first difference between the amplitude of the left channel audio data under the first pose change range and the amplitude of the original audio data under the first pose change range, and the second trend of change is the second difference between the amplitude of the right channel audio data under the first pose change range and the amplitude of the original audio data under the first pose change range.

[0012] The step of determining whether the change conditions corresponding to the first pose change range are met based on the first change trend and the second change trend includes: when it is determined that the spatial audio sound effect is normal under the first pose change range, the first difference is located in the first amplitude range corresponding to the first pose change range and the second difference is located in the second amplitude range corresponding to the first pose change range.

[0013] Based on the above scheme, when the audio features include amplitude, the amplitudes of the left channel audio data and the right channel audio data can be extracted, and data analysis can be performed through amplitude to determine the sound effects of spatial audio.

[0014] In another possible implementation, the audio feature includes time delay, the first trend being a third difference between the time delay of the left channel audio data under the first pose change range and the time delay of the original audio data under the first pose change range, and the second trend being a fourth difference between the time delay of the right channel audio data under the first pose change range and the time delay of the original audio data under the first pose change range.

[0015] The step of determining whether the change conditions corresponding to the first pose change range are met based on the first change trend and the second change trend includes: when it is determined that the third difference is located within the first time delay range corresponding to the first pose change range and the fourth difference is located within the second time delay range corresponding to the first pose change range, it is determined that the sound effect of the spatial audio is normal under the first pose change range.

[0016] Based on the above scheme, when the audio features include time delay, the time delay between the left channel audio data and the right channel audio data can be extracted, and data analysis can be performed through time delay to determine the sound effects of spatial audio.

[0017] In another possible implementation, the audio features include amplitude and delay, wherein the first trend is a fifth difference between the amplitude of the left channel audio data under the first pose change range and the amplitude of the original audio data under the first pose change range, and a sixth difference between the delay of the left channel audio data under the first pose change range and the delay of the original audio data under the first pose change range; the second trend is a seventh difference between the amplitude of the right channel audio data under the first pose change range and the amplitude of the original audio data under the first pose change range, and an eighth difference between the delay of the right channel audio data under the first pose change range and the delay of the original audio data under the first pose change range.

[0018] The step of determining whether the change conditions corresponding to the first pose change range are met based on the first change trend and the second change trend includes: when it is determined that the fifth difference of the amplitude is located in the first amplitude range corresponding to the first pose change range and the seventh difference is located in the second amplitude range corresponding to the first pose change range, and the sixth difference of the delay is located in the first delay range corresponding to the first pose change range and the eighth difference is located in the second delay range corresponding to the first pose change range, the spatial audio sound effect is determined to be normal under the first pose change range.

[0019] Based on the above scheme, when the audio features include amplitude and time delay, the amplitude and time delay of the left channel audio data and the amplitude and time delay of the right channel audio data can be extracted respectively, thereby determining the sound effects of spatial audio and improving accuracy.

[0020] In one possible implementation, the method further includes: when it is determined that the first change trend and the second change trend do not satisfy the change conditions corresponding to the first pose change range, recording the first pose change range and outputting the first pose change range.

[0021] Based on the above scheme, the spatial audio algorithm under the obtained first pose change range can be optimized, thereby optimizing the sound effect of the spatial audio.

[0022] Secondly, embodiments of this application provide an apparatus for determining the sound effects of spatial audio, comprising:

[0023] The acquisition unit is used to acquire left channel audio data and right channel audio data output by a first electronic device; the first electronic device carries a spatial audio algorithm to be tested; the left channel audio data is obtained by the first electronic device using the spatial audio algorithm to be tested to simulate the original audio data played by the audio playback device heard by the left ear under multiple pose change ranges; the right channel audio data is obtained by the first electronic device using the spatial audio algorithm to be tested to simulate the original audio data played by the audio playback device heard by the right ear under the multiple pose change ranges.

[0024] The processing unit is configured to extract audio features from the left channel audio data, the right channel audio data, and the original audio data respectively; determine a first trend of change of the audio features of the left channel audio data relative to the audio features of the original audio data under a first pose change range; and determine a second trend of change of the audio features of the right channel audio data relative to the audio features of the original audio data under the first pose change range; wherein the first pose change range is one of the plurality of pose change ranges;

[0025] The processing unit is further configured to determine whether the spatial audio sound effect is normal based on whether the first change trend and the second change trend meet the change conditions corresponding to the first pose change range.

[0026] In one possible implementation, the audio features include amplitude and / or time delay.

[0027] In one possible implementation, the audio feature includes amplitude, the first trend of change being a first difference between the amplitude of the left channel audio data under the first pose change range and the amplitude of the original audio data under the first pose change range, and the second trend of change being a second difference between the amplitude of the right channel audio data under the first pose change range and the amplitude of the original audio data under the first pose change range.

[0028] The processing unit, when determining whether the change conditions corresponding to the first pose change range are met based on the first change trend and the second change trend, is specifically used to: determine that the spatial audio sound effect is normal under the first pose change range when it is determined that the first difference is located in the first amplitude range corresponding to the first pose change range and the second difference is located in the second amplitude range corresponding to the first pose change range.

[0029] In another possible implementation, the audio feature includes time delay, the first trend being a third difference between the time delay of the left channel audio data under the first pose change range and the time delay of the original audio data under the first pose change range, and the second trend being a fourth difference between the time delay of the right channel audio data under the first pose change range and the time delay of the original audio data under the first pose change range.

[0030] The processing unit, when determining whether the change conditions corresponding to the first pose change range are met based on the first change trend and the second change trend, is specifically used to: determine that the spatial audio sound effect is normal under the first pose change range when it is determined that the third difference is located in the first time delay range corresponding to the first pose change range and the fourth difference is located in the second time delay range corresponding to the first pose change range.

[0031] In another possible implementation, the audio features include amplitude and delay, the first trend being a fifth difference between the amplitude of the left channel audio data under the first pose change range and the amplitude of the original audio data under the first pose change range, and a sixth difference between the delay of the left channel audio data under the first pose change range and the delay of the original audio data under the first pose change range; the second trend being a seventh difference between the amplitude of the right channel audio data under the first pose change range and the amplitude of the original audio data under the first pose change range, and an eighth difference between the delay of the right channel audio data under the first pose change range and the delay of the original audio data under the first pose change range.

[0032] The processing unit, when determining whether the change conditions corresponding to the first pose change range are met based on the first change trend and the second change trend, is specifically configured to: determine that the spatial audio sound effect is normal under the first pose change range when it is determined that the fifth difference of the amplitude is located in the first amplitude range corresponding to the first pose change range and the seventh difference is located in the second amplitude range corresponding to the first pose change range, and the sixth difference of the delay is located in the first delay range corresponding to the first pose change range and the eighth difference is located in the second delay range corresponding to the first pose change range.

[0033] In one possible implementation, the processing unit is further configured to: when it is determined that the first change trend and the second change trend do not satisfy the change conditions corresponding to the first pose change range, record the first pose change range and output the first pose change range.

[0034] Thirdly, embodiments of this application provide a sound effect device for determining spatial audio, including a memory and a processor;

[0035] The memory is used for storing program instructions;

[0036] A processor is configured to invoke program instructions stored in the memory and execute the method of the first aspect and any one of the first aspects according to the obtained program.

[0037] Fourthly, embodiments of this application provide a computer-readable storage medium storing computer instructions that, when executed on a computer, cause the computer to perform the methods described in the first aspect and any one of the first aspects.

[0038] Furthermore, the technical effects of any of the implementation methods in the second to fourth aspects can be found in the first aspect and the technical effects of different implementation methods of the first aspect, which will not be repeated here. Attached Figure Description

[0039] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0040] Figure 1A This is a schematic diagram of the time difference between the two ears and the sound level difference between the two ears;

[0041] Figure 1BThis is a schematic diagram illustrating the filtering effect of sound on the human body.

[0042] Figure 1C A diagram illustrating how head movements can be used to sense the location of a sound.

[0043] Figure 2A A service system architecture diagram provided for an embodiment of this application;

[0044] Figure 2B A server structure diagram provided for an embodiment of this application;

[0045] Figure 3 A schematic flowchart illustrating a method for determining spatial audio sound effects, provided in an embodiment of this application;

[0046] Figure 4 A schematic diagram illustrating a process for determining spatial audio provided in an embodiment of this application;

[0047] Figure 5 A flowchart of spatial audio analysis provided in this application embodiment;

[0048] Figure 6A A schematic diagram illustrating spatial audio data comparison provided in an embodiment of this application;

[0049] Figure 6B A schematic diagram illustrating another spatial audio data comparison provided in an embodiment of this application;

[0050] Figure 7 A schematic diagram illustrating another spatial audio data comparison provided in an embodiment of this application;

[0051] Figure 8 A schematic diagram illustrating another spatial audio data comparison provided in an embodiment of this application;

[0052] Figure 9 A schematic diagram of a device for determining the sound effects of spatial audio provided in an embodiment of this application;

[0053] Figure 10 This is a schematic diagram of another device for determining the sound effects of spatial audio provided in an embodiment of this application. Detailed Implementation

[0054] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. The components of the embodiments of this application described and shown in the accompanying drawings can be arranged and designed in various different configurations.

[0055] Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of the application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.

[0056] The terms "first" and "second" in the specification, claims, and drawings of this application are used to distinguish different objects, not to describe a specific order. Furthermore, the term "comprising" and any variations thereof are intended to cover non-exclusive protection. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or devices. The term "multiple" in this application can mean at least two, for example, two, three, or more, and is not limited by the embodiments of this application.

[0057] Furthermore, the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this article, unless otherwise specified, generally indicates that the preceding and following related objects have an "or" relationship.

[0058] Spatial audio technology can be applied to VR / AR devices, using the device's sensors to track the head and achieve spatial audio effects. As VR device usage continues to grow, 360° videos are gradually dominating traditional media distribution channels, and the demand for realistic audio is stronger than ever. XR services include Augmented Reality (AR), Virtual Reality (VR), and Mixed Reality (MR). Spatial audio technology fixes a sound source, attached to a virtual object, at a specific location in XR space. When the user turns their head, changes position, or their head position changes, the relative position of their ears and the sound source changes. Algorithms simulate this change in sound intensity. People's perception of sound location is mainly based on four factors: binaural time difference, binaural sound level difference, human body filtering effect, and head movement.

[0059] (1) Interaural Time Difference (ITD): Due to the different distances between the sound source and the left and right ears, there will be a difference in the time it takes for the sound to reach the left and right ears. This difference is called the time difference. Figure 1A As shown.

[0060] (2) Interaural Level Difference (ILD): Due to head obstruction, the sound pressure levels reaching the left and right ears are different, creating a sound level difference. Figure 1A As shown.

[0061] (3) Human body filtering effect: The outer ear, head, shoulders, neck, and torso of a person will have different effects on sounds coming from different directions, forming reflections, blocking, or diffractions. For example, the left and right ears can filter sounds from the front and back, and can also filter sounds at different heights, such as... Figure 1B As shown.

[0062] (4) Head shaking: When the location of a sound source is difficult to determine, people often unconsciously shake their heads slightly, causing changes in time difference, sound level difference, or human body filtering effect, and quickly relocating the source based on these changes, such as... Figure 1C As shown.

[0063] Currently, the evaluation of spatial audio effects is based on subjective auditory perception. Algorithms are optimized through individual judgment, meaning a person wearing an XR device moves and rotates to assess the effect and analyze the algorithm's strengths and weaknesses. However, individual perceptions vary greatly, making quantitative comparison impossible. At a microscopic level, only general judgments about location and distance trends can be made. There is a lack of objective quantitative comparisons of these trends, and no intuitive graphical representation, which hinders algorithm feedback and optimization.

[0064] To address the aforementioned issues, this application provides a method and apparatus for determining the sound effects of spatial audio. It employs spatial audio algorithms to obtain test audio data for various scenarios, including left and right channel audio data. By analyzing the amplitude and delay of the obtained test audio data on the output audio data, it objectively compares the sound effects, provides effect analysis conclusions, offers data analysis, evaluates the spatial audio algorithm, guides algorithm optimization, and ultimately optimizes the sound effects.

[0065] Figure 2A An exemplary service system architecture used in an embodiment of this application is illustrated. This system architecture may include one or more servers 100, which may be hosts or various electronic devices. See, for example, [link to example]. Figure 1B As shown, server 100 may include processor 110, communication interface 120, and memory 130. Of course, server 100 may also include other components. Figure 2B Not shown in the image.

[0066] The communication interface 120 is used to communicate with multiple electronic devices, receive raw audio data, left channel audio data and right channel audio data, and realize communication.

[0067] Processor 110 is the control center of server 100. It connects various parts of server 100 through various interfaces and routes. By running or executing software programs and / or modules stored in memory 130, and by calling data stored in memory 130, it performs various functions of server 100 and processes data. Processor 110 may be a processor, microprocessor, controller, or other control component. For example, it may be a general-purpose central processing unit (CPU), a general-purpose processor, a digital signal processing unit (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof.

[0068] The memory 130 can be used to store software programs and modules. The processor 110 executes various functional applications and data processing by running the software programs and modules stored in the memory 130. The memory 130 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function, etc.; the data storage area may store data created according to business processing, etc. In addition, the memory 130 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device.

[0069] It should be noted that the above Figure 1A and Figure 1B The structure shown is merely an example, and the embodiments of this application are not limited thereto.

[0070] The solutions provided in the embodiments of this application are described below with reference to the accompanying drawings.

[0071] This application provides a method for determining the sound effects of spatial audio. Figure 3 The flowchart of a method for determining spatial sound effects is illustrated exemplarily. This method can be executed by server 100 or by processor 110 within the server. The following description will use server 100 as an example; for ease of description, server 100 will not be identified thereafter.

[0072] 301, Obtain the left channel audio data and right channel audio data output by the first electronic device.

[0073] The first electronic device carries the spatial audio algorithm to be tested. The left channel audio data is obtained by the first electronic device using the spatial audio algorithm to be tested to simulate the original audio data played by the audio playback device as heard by the left ear under multiple pose changes. The right channel audio data is obtained by the first electronic device using the spatial audio algorithm to be tested to simulate the original audio data played by the audio playback device as heard by the right ear under multiple pose changes.

[0074] In some embodiments, the raw audio data must first be acquired. This raw audio data is either source audio data or source audio data recorded using audio editing software. In some scenarios, the raw audio data includes raw left channel audio data and raw right channel audio data, with the raw left channel audio data being the same as the raw right channel audio data. After obtaining the raw audio data, it can be decoded, for example, by decoding the raw audio data format into PCM format, and then saving the PCM format raw audio data.

[0075] In some embodiments, after acquiring the raw audio data, the pose data for testing needs to be determined. In some scenarios, this pose data can be understood as the pose changes of the first electronic device and the device playing the raw audio data. This pose data can be understood as the pose changes of the first electronic device relative to the device playing the raw audio data during rotation around the first electronic device as the origin. In some scenarios, the first electronic device can remain stationary as the origin, and the pose data required by the spatial audio algorithm can be constructed based on the various directional axes of the first electronic device. For example, the directional axes can be divided into the x-axis, y-axis, and z-axis. The device playing the raw audio data plays PCM format audio source data, and the device playing the raw audio data rotates around the various coordinate axes of the first electronic device to generate pose data. As an example, rotating the raw audio data 360 degrees around the x-axis is used as the pose data for forward and backward rotation; rotating the raw audio data 360 degrees around the y-axis is used as the pose data for horizontal rotation; and rotating the raw audio data 360 degrees around the z-axis is used as the pose data for vertical rotation. The first electronic device generates spatial audio data based on the original audio data and pose data. This spatial audio data can simulate changes in sound position. The spatial audio data includes left channel audio data and right channel audio data, such as… Figure 4 As shown.

[0076] 302. Extract audio features from the left channel audio data, right channel audio data, and original audio data respectively.

[0077] In some embodiments, audio features include amplitude and / or time delay. In some scenarios, raw audio data, left channel audio data, and right channel audio data can be input into audio analysis software to obtain waveforms of the raw audio data, left channel audio data, and right channel audio data. Audio features can be extracted from the audio data. For example, the amplitude change over a certain time period can be obtained from the audio data, where the amplitude represents the volume of the audio data within that time period and the corresponding pose change range. Furthermore, the time delay at a certain timestamp can be obtained from the audio data, representing the time difference between the reception time of the left channel audio data at that timestamp and the transmission time of the raw audio data at that timestamp.

[0078] 303, determine the first trend of change of the audio features of the left channel audio data relative to the audio features of the original audio data under the first pose change range, and determine the second trend of change of the audio features of the right channel audio data relative to the audio features of the original audio data under the first pose change range.

[0079] The first pose variation range is one of multiple pose variation ranges. The first variation trend is the difference between the audio features of the original audio data and the audio features of the left channel audio data within the first pose variation range. For ease of distinction, this difference is called the first difference. The first difference represents the first variation trend. The second variation trend is the second difference between the audio features of the original audio data and the audio features of the right channel audio data within the first pose variation range. The second difference represents the second variation range.

[0080] In some embodiments, when the audio features include amplitude, the first trend of change can be represented by the difference between the amplitude of the left channel audio data under the first pose change and the amplitude of the original audio data within the first pose change range, and the second trend of change can be represented by the difference between the amplitude of the right channel audio data under the first pose change and the amplitude of the original audio data within the first pose change range. As an example, when the angle of pose change between the device playing the original audio data and the first electronic device changes from 30 degrees to 60 degrees, the first pose change range is from 30 degrees to 60 degrees. Assuming that the pose data corresponding to every 10 degrees from 30 degrees to 60 degrees corresponds to times T1, T2, T3, and T4, the amplitudes of the original audio data and the left channel audio data corresponding to times T1, T2, T3, and T4 are extracted respectively. At each time point, the amplitude of the left channel audio data is subtracted from the amplitude of the original audio data to obtain four differences. For example, the amplitude differences at times T1, T2, T3, and T4 are 0.1, 0.2, 0.3, and 0.4, respectively. The increasing difference between the original audio data and the left channel audio data indicates that the amplitude of the left channel audio data is decreasing. Therefore, the first range of change is from large to small (or from near to far). Similarly, extracting the amplitudes of the original audio data and the right channel audio data at times T1, T2, T3, and T4, and subtracting the amplitude of the right channel audio data from the amplitude of the original audio data at each time point, yields four differences: 0.4, 0.3, 0.2, and 0.1. The decreasing difference between the original audio data and the right channel audio data indicates that the amplitude of the right channel audio data is increasing. Therefore, the second range of change is from small to large (or from far to near).

[0081] In other embodiments, when the audio features include time delay, the first trend can be represented by the difference between the time delay of the original audio data within the first pose change range and the time delay of the left channel audio data within the first pose change range. The second trend can be represented by the difference between the time delay of the original audio data within the first pose change range and the time delay of the right channel audio data within the first pose change range. For example, when the angle of pose change between the device playing the original audio data and the first electronic device changes from 30 degrees to 60 degrees, the first pose change range is from 30 degrees to 60 degrees. Assuming that the time corresponding to each 10-degree pose data from 30 degrees to 60 degrees is T1, T2, T3, and T4, the time delays of the original audio data and the left channel audio data corresponding to time T1, T2, T3, and T4 are extracted respectively. At each time point, the time delay of the original audio data is subtracted from the time delay of the left channel audio data to obtain four differences. For example, the time delay differences corresponding to T1, T2, T3, and T4 are -0.4, -0.3, -0.2, and -0.1, respectively. The difference between the original audio data delay and the left channel audio data delay under the first pose change shows an increasing trend, indicating that the delay of the left channel audio data is decreasing. Therefore, the first range of change is from far to near. Similarly, extracting the time delays of the original audio data and the right channel audio data corresponding to T1, T2, T3, and T4, and subtracting the delay of the right channel audio data from the original audio data delay at each time point, yields four differences: -0.1, -0.2, -0.3, and -0.4. The difference between the original audio data delay and the right channel audio data delay under the first pose change range shows a decreasing trend, indicating that the delay of the right channel audio data is increasing. Therefore, the second range of change is from near to far.

[0082] In some embodiments, when the audio features include amplitude and delay, the first trend can be represented by the difference between the amplitude of the left channel audio data under the first pose change range and the amplitude of the original audio data under the first pose change range, and by the difference between the delay of the left channel audio data under the first pose change range and the delay of the original audio data under the first pose change range. The second trend can be represented by the difference between the amplitude of the original audio data under the first pose change range and the amplitude of the right channel audio data, and by the difference between the delay of the original audio data under the first pose change range and the delay of the right channel audio data under the first pose change range. Therefore, the first trend includes both the amplitude and delay trends of the left channel audio data; similarly, the second trend includes both the amplitude and delay trends of the right channel audio data. Specific methods can be found in the above embodiments and will not be repeated here.

[0083] 304. Determine whether the change conditions corresponding to the range of the first pose change are met based on the first and second change trends.

[0084] The changing conditions can be determined based on the range of the first posture change and the original audio data. For example, based on the judgment of sound location perception, the changing conditions of the left and right channel audio data of the first electronic device within the range of the first posture change can be determined according to the range of the first posture change and the original audio data. The changing conditions of the left and right channel audio data within the range of the first posture change can be determined based on the distance between the left and right channel outputs of the first electronic device, the range of the first posture change, and the original audio data.

[0085] In some embodiments, when the audio data includes amplitude, a first difference between the amplitude of the original audio data and the amplitude of the left channel audio data within the first pose change range is obtained according to step 303, and a second difference between the amplitude of the original audio data and the amplitude of the right channel audio data within the first pose change range. Furthermore, a first amplitude range corresponding to the first pose change range and a second amplitude range corresponding to the first pose change range are determined. When the first difference is within the first amplitude range and the second difference is within the second amplitude range, it is determined that the spatial audio effect within the first pose change range is normal. For example, within the time period T1-T4, differences 1, 2, 3, and 4 between the amplitude of the original audio data and the amplitude of the left channel audio data corresponding to times T1 to T4 are obtained, and the first trend of change in the amplitude of the left channel audio data within the time period T1-T4 is determined by the changing trend of differences 1 to 4. Similarly, the differences 5, 6, 7, and 8 between the amplitude of the original audio data and the amplitude of the right channel audio data at times T1 to T4 are obtained respectively. The second trend of the amplitude change of the right channel audio data in the T1-T4 time period is determined by the changing trend of the difference 5 to the difference 8. When the amplitude difference 1 to the difference 4 is within the first amplitude range corresponding to the pose change range in the T1-T4 time period, and the first trend determined by the difference 1 to the difference 4 meets the amplitude change condition corresponding to the pose change range in the T1-T4 time period, and when the amplitude difference 5 to the difference 8 is within the second amplitude range corresponding to the pose change range in the T1-T4 time period, and the second trend determined by the difference 5 to the difference 8 meets the amplitude change condition corresponding to the pose change range in the T1-T4 time period, then the sound effect of the spatial audio in that pose change range is normal.

[0086] In other embodiments, when the audio data includes a time delay, a third difference between the time delay of the original audio data and the time delay of the left channel audio data within the first pose change range is obtained according to step 303, and a fourth difference between the time delay of the original audio data and the time delay of the right channel audio data within the first pose change range. Then, a first time delay range corresponding to the first pose change range and a second time delay range corresponding to the first pose change range are determined. When the third difference is within the first time delay range and the fourth difference is within the second time delay range, it is determined that the spatial audio sound effect within the first pose change range is normal. For example, within the time period T1-T4, differences 1, 2, 3, and 4 between the time delay of the original audio data and the time delay of the left channel audio data corresponding to times T1 to T4 are obtained, and the first trend of the time delay of the left channel audio data within the time period T1-T4 is determined by the changing trend of differences 1 to 4. Similarly, the differences 5, 6, 7, and 8 between the delay of the original audio data and the delay of the right channel audio data are obtained from times T1 to T4 respectively. The second trend of the delay of the right channel audio data in the T1-T4 time period is determined by the changing trend of differences 5 to 8. When differences 1 to 4 are within the first delay range corresponding to the pose change range in the T1-T4 time period, and the first trend determined by differences 1 to 4 meets the change condition of the delay corresponding to the pose change range in the T1-T4 time period, and when differences 5 to 8 are within the second delay range corresponding to the pose change range in the T1-T4 time period, and the second trend determined by differences 5 to 8 meets the change condition of the delay corresponding to the pose change range in the T1-T4 time period, then the sound effect of the spatial audio in that pose change range is normal.

[0087] In some embodiments, when the audio features include amplitude and time delay, a fifth difference between the amplitude of the original audio data and the amplitude of the left channel audio data within the first pose change range, and a seventh difference between the amplitude of the original audio data and the amplitude of the right channel audio data within the first pose change range are obtained according to step 303. A sixth difference between the time delay of the original audio data and the time delay of the left channel audio data within the first pose change range, and an eighth difference between the time delay of the original audio data and the time delay of the channel audio data within the first pose change range are obtained through step 303. Then, a first amplitude range and a second amplitude range, as well as a first time delay range and a second time delay range, corresponding to the first pose change range are determined. When the fifth difference is within the first amplitude range and the seventh difference is within the second amplitude range, and the sixth difference is within the first time delay range and the eighth difference is within the second time delay range, the spatial audio sound effect within the first pose change range is determined to be normal. Based on this, when the audio features include amplitude and time delay, the difference and trend of amplitude and the difference and trend of time delay within a certain pose change range are obtained. When the difference in amplitude is within the amplitude range corresponding to the pose change range, the amplitude change trend satisfies the amplitude change condition corresponding to the pose change range, and the difference in time delay is within the time delay change range corresponding to the pose change range, and the time delay change trend satisfies the time delay change condition corresponding to the pose change range, then the spatial audio effect of the pose change range is determined to be normal. Specific methods can be found in the above embodiments and will not be repeated here.

[0088] In some embodiments, the original audio data, left channel audio data, and right channel audio data can be opened using audio analysis software. The waveforms of the left and right channel audio data are compared with the original audio data, and the waveforms of the left and right channel audio data are also compared. Specifically, the waveforms of the original audio data, left channel audio data, and right channel audio data can be analyzed and compared based on timestamps. Analysis is performed at the corresponding time and pose to determine whether the left and right channel audio data meet the variation range. When the amplitude and delay of the audio data at a certain timestamp meet the variation range, it is determined whether that timestamp is the last timestamp. If it is not the last timestamp, the audio data at the next timestamp is analyzed; if it is the last timestamp, the process ends. When the amplitude and delay of the audio data at a certain timestamp do not meet the variation range, the pose data corresponding to that timestamp is recorded and output. The deviation from the variation range is analyzed and compared to optimize the spatial audio algorithm's processing of the pose, and the audio data at the next timestamp is analyzed. Figure 5 As shown.

[0089] As an example, taking the upward direction of the first electronic device as the y-axis, assume the device playing the original audio data rotates 360 degrees around the y-axis of the first electronic device to obtain the corresponding pose data. Specifically, for the first 10 seconds, the device playing the original audio data remains stationary, and the first electronic device determines the left and right channel audio data using the original audio data and pose data. Then, the original audio data, left channel audio data, and right channel audio data are input into audio analysis software to obtain waveforms of the original audio data, left channel audio data, and right channel audio data, as shown below. Figure 6A As shown. The first and second columns are waveforms of the original left and right channel audio data presented by audio analysis software. The third and fourth columns are waveforms of the left and right channel audio data obtained through a spatial audio algorithm, presented by the same software. As an example, it is assumed that the time interval between two adjacent waveform nodes in each column of the waveforms presented by the audio analysis software is 1 second. From... Figure 6A It can be seen that the amplitude and delay of the left and right channel audio data in the first 10 seconds are basically the same, and the amplitude and delay of the original left and right channel audio data in the original audio data are also the same, thus satisfying the change condition. Then, starting from the 11th second, assume that the audio data playback device rotates 360 degrees counterclockwise around the y-axis of the first electronic device until it returns to its original position at the 20th second. The pose changes of the original audio data playback device and the first electronic device are used as pose data. The left and right channel audio data from the 11th to the 20th second are obtained through the original audio data and pose data. The original audio data, left channel audio data, and right channel audio data from the 11th to the 20th second are input into audio analysis software to obtain waveforms of the original audio data, left channel audio data, and right channel audio data, as shown below. Figure 6B As shown. From the 11th to the 14th second, the pose change of the left channel of the device playing the original audio data and the first electronic device is first closer and then farther away, while the pose change of the right channel of the device playing the original audio data and the first electronic device is first farther away and then closer. Therefore, the trend of the first difference between the amplitude and delay of the left channel audio data is first decreasing and then increasing, while the trend of the first difference between the amplitude and delay of the right channel audio data is first increasing and then decreasing. In addition, the amplitude of the left channel audio data should be greater than that of the right channel audio data, and the delay of the left channel audio data relative to the original left channel audio data should be less than the delay of the right channel audio data relative to the original right channel audio data. From Figure 6BAs can be seen, the amplitude and delay of the left channel audio data, as well as the amplitude and delay of the right channel, satisfy the change conditions. Therefore, the spatial audio sound effect from the 11th to the 14th second is normal. At the 15th second, the audio data playback device rotates 180 degrees counterclockwise around the first electronic device. The distance between the left and right channels and the original audio data playback device is the same. Therefore, the amplitude and delay of the left and right channel audio data at the 15th second are basically the same. Figure 6B The left and right channel audio data at 15 seconds meet the change conditions. From 16 to 19 seconds, the position change of the right channel of the device playing the original audio data and the first electronic device is first closer and then farther away, while the position change of the left channel of the device playing the original audio data and the first electronic device is first farther away and then closer. Therefore, the trend of the first difference in amplitude and delay of the right channel audio data is first decreasing and then increasing, while the trend of the first difference in amplitude and delay of the left channel audio data is first increasing and then decreasing. In addition, the amplitude of the right channel audio data should be greater than that of the left channel audio data, and the delay of the right channel audio data relative to the original right channel audio data should be less than the delay of the left channel audio data relative to the original left channel audio data. Figure 6B As can be seen, the amplitude and delay of the left channel audio data, as well as the amplitude and delay of the right channel, satisfy the change conditions. Therefore, the spatial audio sound effect from the 16th to the 19th second is normal. At the 20th second, the audio data playback device rotates 360 degrees counterclockwise around the first electronic device. The distance between the left and right channels and the original audio data playback device is the same. Therefore, the amplitude and delay of the left and right channel audio data at the 20th second are basically the same. Figure 6B The left and right channel audio data at the 20th second meet the change condition. In summary, the left and right channel audio data from the 1st to the 20th second meet the change condition, therefore it can be determined that the spatial audio sound effect is normal during this time period.

[0090] In some scenarios, the left-to-right direction of the first electronic device can be used as the x-axis. Assuming the device playing the original audio data rotates 360 degrees clockwise around the x-axis of the first electronic device, the corresponding pose data is obtained. Specifically, the first electronic device determines the left and right channel audio data using the original audio data and the pose data. As an example, assuming the time taken for the device playing the original audio data to rotate 360 ​​degrees around the x-axis of the first electronic device is 12 seconds, the original audio data, left channel audio data, and right channel audio data are input into audio analysis software. The waveforms of the original audio data, left channel audio data, and right channel audio data presented by the audio analysis software are obtained, such as... Figure 7As shown. When the audio data playback device is located directly in front of the first electronic device and rotates around the x-axis of the first electronic device, the volume of the left channel audio data and the right channel audio data should be the same. Therefore, the amplitude of the left channel audio data and the right channel audio data should be the same. Figure 7 In the data at the 7th second, the amplitude of the left channel audio data is significantly greater than that of the right channel audio data. The audio data at the 7th second does not meet the change condition, therefore, the spatial audio algorithm for this pose needs to be optimized. The optimization direction is to balance the data so that the volume of the left channel audio data and the right channel audio data are the same in this pose.

[0091] In some scenarios, the back-to-forehead direction of the first electronic device can be used as the z-axis. Assume the device playing the original audio data rotates 360 degrees clockwise around the z-axis of the first electronic device, starting at a predetermined distance to the left front of the device carrying the spatial audio algorithm, and obtains the corresponding pose data. Specifically, the first electronic device determines the left and right channel audio data using the original audio data and the pose data. As an example, assuming the time taken for the device playing the original audio data to rotate 360 ​​degrees around the z-axis of the first electronic device is 12 seconds, the original audio data, left channel audio data, and right channel audio data are input into audio analysis software. The resulting waveforms of the original audio data, left channel audio data, and right channel audio data are then obtained through the audio analysis software, such as... Figure 8 As shown. When the audio data playback device rotates around the z-axis of the first electronic device, at the 7th second, the audio data playback device rotates 180 degrees around the z-axis of the first electronic device. At this time, the amplitude change of the left channel audio data relative to the right channel audio data at the 7th second should be exactly the opposite of the amplitude change of the left channel audio data relative to the right channel audio data at the 1st second. In the 1st second, the amplitude of the left channel audio data is slightly larger than the amplitude of the right channel audio data. Therefore, at the 7th second, the amplitude of the left channel audio data should be smaller than the amplitude of the right channel audio data. Figure 8 As shown, the amplitude of the left channel audio data in the 7th second is significantly greater than that of the right channel audio data. Therefore, the audio data in the 7th second does not meet the change condition, indicating that the spatial audio algorithm has a defect in this pose. Therefore, the spatial audio algorithm needs to be optimized for irregular movements.

[0092] Based on the same technical concept, embodiments of this application provide a device 900 for determining the sound effects of spatial audio, see [link to related document]. Figure 9 As shown. The device 900 can perform each step of the above-described method for determining spatial audio sound effects; to avoid repetition, it will not be described in detail here. The device 900 includes an acquisition unit 901 and a processing unit 902.

[0093] The acquisition unit 901 is used to acquire left channel audio data and right channel audio data output by the first electronic device; the first electronic device carries a spatial audio algorithm to be tested; the left channel audio data is obtained by the first electronic device using the spatial audio algorithm to be tested to simulate the original audio data played by the audio playback device heard by the left ear under multiple pose change ranges; the right channel audio data is obtained by the first electronic device using the spatial audio algorithm to be tested to simulate the original audio data played by the audio playback device heard by the right ear under the multiple pose change ranges.

[0094] Processing unit 902 is configured to extract audio features from the left channel audio data, the right channel audio data, and the original audio data respectively; determine a first trend of change of the audio features of the left channel audio data relative to the audio features of the original audio data under a first pose change range; and determine a second trend of change of the audio features of the right channel audio data relative to the audio features of the original audio data under the first pose change range; wherein the first pose change range is one of the plurality of pose change ranges;

[0095] The processing unit 902 is further configured to determine whether the sound effect of the spatial audio is normal based on whether the first change trend and the second change trend meet the change conditions corresponding to the first pose change range.

[0096] In some embodiments, the audio features include amplitude and / or time delay.

[0097] In some embodiments, the audio feature includes amplitude, the first trend of change being a first difference between the amplitude of the left channel audio data under the first pose change range and the amplitude of the original audio data under the first pose change range, and the second trend of change being a second difference between the amplitude of the right channel audio data under the first pose change range and the amplitude of the original audio data under the first pose change range.

[0098] The processing unit 902, when determining whether the change conditions corresponding to the first pose change range are met based on the first change trend and the second change trend, is specifically used to: determine that the sound effect of the spatial audio under the first pose change range is normal when it is determined that the first difference is located in the first amplitude range corresponding to the first pose change range and the second difference is located in the second amplitude range corresponding to the first pose change range.

[0099] In other embodiments, the audio feature includes time delay, the first trend being a third difference between the time delay of the left channel audio data under the first pose change range and the time delay of the original audio data under the first pose change range, and the second trend being a fourth difference between the time delay of the right channel audio data under the first pose change range and the time delay of the original audio data under the first pose change range.

[0100] The processing unit 902, when determining whether the change conditions corresponding to the first pose change range are met based on the first change trend and the second change trend, is specifically used to: determine that the spatial audio sound effect is normal under the first pose change range when it is determined that the third difference is located in the first time delay range corresponding to the first pose change range and the fourth difference is located in the second time delay range corresponding to the first pose change range.

[0101] In some other embodiments, the audio features include amplitude and delay, the first trend being a fifth difference between the amplitude of the left channel audio data under the first pose change range and the amplitude of the original audio data under the first pose change range, and a sixth difference between the delay of the left channel audio data under the first pose change range and the delay of the original audio data under the first pose change range; the second trend being a seventh difference between the amplitude of the right channel audio data under the first pose change range and the amplitude of the original audio data, and an eighth difference between the delay of the right channel audio data under the first pose change range and the delay of the original audio data under the first pose change range.

[0102] The processing unit 902, when determining whether the change conditions corresponding to the first pose change range are met based on the first change trend and the second change trend, is specifically configured to: determine that the spatial audio sound effect is normal under the first pose change range when it is determined that the fifth difference of the amplitude is located in the first amplitude range corresponding to the first pose change range and the seventh difference is located in the second amplitude range corresponding to the first pose change range, and the sixth difference of the delay is located in the first delay range corresponding to the first pose change range and the eighth difference is located in the second delay range corresponding to the first pose change range.

[0103] In some embodiments, the processing unit 902 is further configured to: when it is determined that the first change trend and the second change trend do not satisfy the change conditions corresponding to the first pose change range, record the first pose change range and output the first pose change range.

[0104] Based on the same technical concept, embodiments of this application provide a device 1000 for determining the sound effects of spatial audio, see [link to previous document]. Figure 10 As shown. The device 1000 can perform the various steps in the method for determining the sound effects of spatial audio described above. The device 1000 includes a memory 1001 and a processor 1002.

[0105] The memory 1001 is used to store program instructions;

[0106] The processor 1002 is used to call the program instructions stored in the memory and execute the above-mentioned method for determining the sound effects of spatial audio according to the obtained program.

[0107] In the embodiments of this application, the processor may be a general-purpose processor, a digital signal processor, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components, capable of implementing or executing the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly manifested as being executed by a hardware processor, or executed by a combination of hardware and software modules within the processor.

[0108] Memory, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules. Memory can include at least one type of storage medium, such as flash memory, hard disk, multimedia card, card-type memory, random access memory (RAM), static random access memory (SRAM), programmable read-only memory (PROM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), magnetic memory, magnetic disk, optical disk, etc. Memory is any other medium capable of carrying or storing desired program code in the form of instructions or data structures that can be accessed by a computer, but is not limited to this. The memory in the embodiments of this application can also be a circuit or any other device capable of implementing storage functions for storing program instructions and / or data.

[0109] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0110] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions specified in one or more blocks of the flowchart illustrations and / or one or more blocks of the block diagrams.

[0111] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means that implement the functions specified in one or more flowcharts and / or one or more block diagrams.

[0112] These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, such that the instructions, which execute on the computer or other programmable apparatus, provide steps for implementing the functions specified in one or more flowcharts and / or one or more block diagrams.

[0113] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.

Claims

1. A method for determining the sound effects of spatial audio, characterized in that, include: Acquire the left channel audio data and right channel audio data output by the first electronic device; The first electronic device carries the spatial audio algorithm to be tested; The left channel audio data is obtained by the first electronic device using the spatial audio algorithm to be tested to simulate the original audio data played by the audio playback device heard by the left ear under multiple pose change ranges. The right channel audio data is obtained by the first electronic device using the spatial audio algorithm to be tested to simulate the original audio data played by the audio playback device heard by the right ear under multiple pose change ranges. The left channel audio data and the right channel audio data support the simulation of sound position changes. Audio features are extracted from the left channel audio data, the right channel audio data, and the original audio data, respectively. Determine a first trend of change in the audio features of the left channel audio data relative to the audio features of the original audio data within the first pose change range, and determine a second trend of change in the audio features of the right channel audio data relative to the audio features of the original audio data within the first pose change range; the first pose change range is one of the plurality of pose change ranges; Based on whether the first and second trends of change meet the change conditions corresponding to the first pose change range, determine whether the sound effect of the spatial audio is normal.

2. The method according to claim 1, characterized in that, The audio features include amplitude and / or time delay.

3. The method according to claim 1 or 2, characterized in that, The audio features include amplitude, the first trend of change is the first difference between the amplitude of the left channel audio data under the first pose change range and the amplitude of the original audio data under the first pose change range, and the second trend of change is the second difference between the amplitude of the right channel audio data under the first pose change range and the amplitude of the original audio data under the first pose change range. The step of determining whether the change conditions corresponding to the first pose change range are met based on the first change trend and the second change trend includes: When it is determined that the first difference is located within the first amplitude range corresponding to the first pose change range, and the second difference is located within the second amplitude range corresponding to the first pose change range, it is determined that the spatial audio sound effect is normal within the first pose change range.

4. The method as described in claim 1 or 2, characterized in that, The audio features include time delay, the first trend of change is the third difference between the time delay of the left channel audio data under the first pose change range and the time delay of the original audio data under the first pose change range, and the second trend of change is the fourth difference between the time delay of the right channel audio data under the first pose change range and the time delay of the original audio data under the first pose change range. The step of determining whether the change conditions corresponding to the first pose change range are met based on the first change trend and the second change trend includes: When it is determined that the third difference is located within the first time delay range corresponding to the first pose change range, and the fourth difference is located within the second time delay range corresponding to the first pose change range, it is determined that the spatial audio sound effect is normal within the first pose change range.

5. The method as described in claim 1 or 2, characterized in that, The audio features include amplitude and time delay. The first change trend is the fifth difference between the amplitude of the left channel audio data under the first pose change range and the amplitude of the original audio data under the first pose change range, and the sixth difference between the time delay of the left channel audio data under the first pose change range and the time delay of the original audio data under the first pose change range. The second trend of change is the seventh difference between the amplitude of the right channel audio data under the first pose change range and the amplitude of the original audio data under the first pose change range, and the eighth difference between the delay of the right channel audio data under the first pose change range and the delay of the original audio data under the first pose change range. The step of determining whether the change conditions corresponding to the first pose change range are met based on the first change trend and the second change trend includes: When it is determined that the fifth difference of the amplitude is located in the first amplitude range corresponding to the first pose change range and the seventh difference is located in the second amplitude range corresponding to the first pose change range, and the sixth difference of the delay is located in the first delay range corresponding to the first pose change range and the eighth difference is located in the second delay range corresponding to the first pose change range, it is determined that the sound effect of the spatial audio under the first pose change range is normal.

6. The method as described in claim 1 or 2, characterized in that, The method further includes: When it is determined that the first change trend and the second change trend do not meet the change conditions corresponding to the first pose change range, the first pose change range is recorded and output.

7. A device for determining the sound effects of spatial audio, characterized in that, include: The acquisition unit is used to acquire the left channel audio data and the right channel audio data output by the first electronic device; The first electronic device carries the spatial audio algorithm to be tested; The left channel audio data is obtained by the first electronic device using the spatial audio algorithm to be tested to simulate the original audio data played by the audio playback device heard by the left ear under multiple pose change ranges. The right channel audio data is obtained by the first electronic device using the spatial audio algorithm to be tested to simulate the original audio data played by the audio playback device heard by the right ear under multiple pose change ranges. The left channel audio data and the right channel audio data support the simulation of sound position changes. The processing unit is used to extract audio features from the left channel audio data, the right channel audio data, and the original audio data, respectively; Determine a first trend of change in the audio features of the left channel audio data relative to the audio features of the original audio data within the first pose change range, and determine a second trend of change in the audio features of the right channel audio data relative to the audio features of the original audio data within the first pose change range; the first pose change range is one of the plurality of pose change ranges; The processing unit is further configured to determine whether the spatial audio sound effect is normal based on whether the first change trend and the second change trend meet the change conditions corresponding to the first pose change range.

8. The apparatus as claimed in claim 7, characterized in that, The processing unit is further configured to: When it is determined that the first change trend and the second change trend do not meet the change conditions corresponding to the first pose change range, the first pose change range is recorded and output.

9. A sound effect device for determining spatial audio, characterized in that, Including memory and processor; The memory is used for storing program instructions; The processor is configured to call program instructions stored in the memory and execute the method of any one of claims 1 to 6 according to the obtained program.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that, when executed on a computer, cause the computer to perform the method of any one of claims 1 to 6.

Citation Information

Patent Citations

  • Audio playing method and device, storage medium and electronic equipment

    CN111654806A

  • Audio chip test method, storage device and computer equipment

    CN112291696A

  • Multichannel sound reproduction method and device

    US20130010970A1