Method and device for evaluating sound space impression, and electronic device

By acquiring the spatial audio signals of the vehicle audio system and calculating the angle matrix and objective spatial perception indicators, the problem of evaluating in-vehicle spatial audio in the prior art has been solved, achieving efficient and convenient objective evaluation and assisting in the optimization of the audio system.

CN119729296BActive Publication Date: 2026-03-27NIO TECH ANHUI CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-13
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing audio evaluation standards are insufficient to effectively assess the spatial perception characteristics of in-vehicle spatial audio systems. Subjective evaluation methods are time-consuming and labor-intensive, while objective evaluation methods are not applicable, resulting in a lack of an efficient in-vehicle spatial audio evaluation system.

Method used

By acquiring the spatial audio signal output from the speaker, determining the angle matrix of the spatial audio signal, and calculating objective indicators of spatial perception such as positioning and immersion, including static positioning, dynamic positioning, and envelopment, a reasonable objective evaluation index is constructed.

Benefits of technology

It provides a non-invasive and easy-to-use objective evaluation method that can quickly reflect the subjective listening experience of spatial audio signals, assist in the optimization of audio systems, save time and effort, and improve efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119729296B_ABST
    Figure CN119729296B_ABST
Patent Text Reader

Abstract

The application is suitable for the field of audio evaluation, and provides an evaluation method and device for sound space feeling and electronic equipment. The evaluation method for sound space feeling provided by the application comprises: acquiring a spatial audio signal output by sound; determining an angle matrix of the spatial audio signal according to the spatial audio signal; and determining a spatial feeling objective index according to the angle matrix of the spatial audio signal, wherein the spatial feeling objective index comprises at least any one of a positioning feeling objective index and an immersion feeling objective index. The application constructs a set of reasonable objective evaluation indexes, and can be close to the subjective feeling of users as much as possible, combines the advantages of the objective evaluation method, such as convenient operation and strong repeatability, and the subjective evaluation method, such as high correlation with hearing feeling, and can bring greater benefits to the evaluation of sound (for example, a vehicle-mounted immersive sound system).
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the technical field of audio evaluation, and particularly relates to a method and device for evaluating a sound space sense and an electronic device. BACKGROUND

[0002] With the continuous development of electric vehicles, its endurance, reliability and other aspects are constantly improving. The potential of electric vehicles as long-distance travel vehicles is emerging. The car cabin, as a window for the interaction between the car and the user, its comfort directly affects the user's experience. A comfortable car cabin can greatly improve the experience of long-distance travel, which promotes the progress of in-car multimedia systems, and the car audio system has become the focus of consumers. The car cabin has the disadvantages of small volume, many obstacles and asymmetric seats compared with the common listening space. Compared with the traditional multi-channel sound system, the in-car speaker system has fixed position, unsatisfactory cabinet condition and limited power. Therefore, the design of the car audio system is a challenging technical problem. Evaluating the audio output by the car audio system is also another challenging technical problem.

[0003] At present, the relatively mature audio evaluation standards in the world mainly focus on speech quality evaluation, especially the audio quality evaluation of coding and transmission damage, such as subjective MOS (Mean Opinion Score) score, objective PESQ (Perceptual Evaluation of Speech Quality), POLQA (Perceptual Objective Listening Quality Assessment), and the like. The typical application scenario is the evaluation of smart phone calls. The audio quality evaluation is still mainly subjective, such as ITU-R BS.1116 (a standard standard number, ITU-R refers to the radio communication department of International Telecommunication Union), subjective MUSHRA (Multi-Stimulus Test with Hidden Reference and Anchor) score, and the like. The typical application scenario is the evaluation of indoor multi-channel surround sound or headphone audio quality, and the only objective audio evaluation standard PEAQ (Perceptual Evaluation of Audio Quality) in the world is only applicable to monaural music signals, and the effect is poor at present. The above standard methods all need to compare the original audio with the reference and the audio to be tested, but for immersive audio systems, it is difficult to find the original audio with reference. Although ITU-R recently proposed BS.2132 (another standard standard number) for subjective evaluation standard for comparison of multiple renderers, but there are still many constraints in actual implementation, such as the need to repeatedly listen to the comparison of multiple different comparison items, which is not suitable for single test object.

[0004] The current subjective evaluation method is time-consuming and laborious. The subjective evaluation must be performed manually in the vehicle, and only relying on the listening score will cause the work efficiency of technical development or product comparison to be reduced. The current objective evaluation method cannot reflect the perception characteristics of the vehicle space audio system. In summary, the existing evaluation standards or methods are not suitable for efficient measurement of the vehicle space audio playback effect, and it is urgent to establish an evaluation system for the vehicle space audio. SUMMARY

[0005] The embodiments of the present application provide an evaluation method, device and electronic equipment for sound space feeling, which can solve the technical problem of lack of effective evaluation method for space audio effect in the prior art.

[0006] In a first aspect, the embodiments of the present application provide an evaluation method for sound space feeling, comprising:

[0007] obtain a spatial audio signal of an acoustic output;

[0008] determine an angle matrix of the spatial audio signal according to the spatial audio signal;

[0009] determine a spatial sense objective index according to the angle matrix of the spatial audio signal, the spatial sense objective index comprising at least any one of a localization sense objective index and an immersion sense objective index.

[0010] In a possible implementation manner of the first aspect, the spatial audio signal comprises a first-order Ambisonics A-format spatial audio signal, and the determining the angle matrix of the spatial audio signal according to the spatial audio signal comprises:

[0011] performing format conversion on the first-order Ambisonics A-format spatial audio signal to obtain a FOA-B-format spatial audio signal;

[0012] determining the angle matrix of the spatial audio signal according to the FOA-B-format spatial audio signal.

[0013] In a possible implementation manner of the first aspect, the spatial audio signal comprises a single-source stationary spatial audio signal, and the localization sense objective index comprises a static localization sense objective index, and the determining the spatial sense objective index according to the angle matrix of the spatial audio signal comprises:

[0014] determining a sound source direction of the single-source stationary spatial audio signal according to the angle matrix of the single-source stationary spatial audio signal;

[0015] calculating a variance of the angle matrix as the static localization sense objective index according to the angle matrix of the single-source stationary spatial audio signal and the sound source direction.

[0016] In a possible implementation manner of the first aspect, the spatial audio signal comprises a single-source moving spatial audio signal, and the localization sense objective index comprises a dynamic localization sense objective index, and the determining the spatial sense objective index according to the angle matrix of the spatial audio signal comprises:

[0017] determining a position coordinate of the single-source moving spatial audio signal in each time period according to the angle matrix of the single-source moving spatial audio signal;

[0018] Based on the azimuth coordinates of the single-source motion spatial audio signal in various time periods, a dynamic positioning objective index is determined. The dynamic positioning objective index includes at least one of the following: a first dynamic positioning objective index and a second dynamic positioning objective index. The first dynamic positioning objective index is the error of the single-source motion spatial audio signal, and the second dynamic positioning objective index is the first-order difference of the single-source motion spatial audio signal.

[0019] In one possible implementation of the first aspect, the dynamic positioning objective index includes a first dynamic positioning objective index, and determining the dynamic positioning objective index based on the azimuth coordinates of the single-source motion spatial audio signal in various time periods includes:

[0020] Based on the temporal distribution of the azimuth coordinates and reference trajectory of the single-source motion spatial audio signal in various time periods, the error of the single-source motion spatial audio signal is determined as the first objective index of dynamic positioning sense.

[0021] In one possible implementation of the first aspect, the dynamic positioning objective index includes a second dynamic positioning objective index, and determining the dynamic positioning objective index based on the azimuth coordinates of the single-source motion spatial audio signal in various time periods includes:

[0022] Based on the azimuth coordinates of the single-source motion spatial audio signal at various time periods, the first-order difference of the single-source motion spatial audio signal is determined as the second objective index of dynamic positioning.

[0023] In one possible implementation of the first aspect, the spatial audio signal includes a multi-channel orthogonal spatial audio signal, the immersion objective index includes an envelopment objective index, and determining the spatial objective index based on the angle matrix of the spatial audio signal includes:

[0024] Based on the angle matrix of the multi-channel orthogonal spatial audio signal, at least two of the following objective energy distribution indicators are determined: front energy distribution objective indicator, rear energy distribution objective indicator, left energy distribution objective indicator, right energy distribution objective indicator, upper energy distribution objective indicator, and lower energy distribution objective indicator.

[0025] Based on the at least two objective indicators of energy distribution, an objective indicator of sense of enclosure is determined, wherein the objective indicator of sense of enclosure includes at least one of the following: objective indicator of front-back sense of enclosure, objective indicator of horizontal sense of enclosure, and objective indicator of vertical sense of enclosure.

[0026] In one possible implementation of the first aspect, the method further includes:

[0027] Based on the first-order high-fidelity stereo sound, the multi-channel orthogonal spatial audio signal in format A is reproduced, and the forward component and backward component of the spatial audio signal are determined.

[0028] Based on the forward component and the backward component of the spatial audio signal, correlation information is determined as an objective indicator of information independence;

[0029] The objective index of immersion is determined based on the objective index of sense of immersion and the objective index of information independence.

[0030] Secondly, embodiments of this application provide a device for evaluating the spatial sense of sound, comprising:

[0031] The signal acquisition module is used to acquire the spatial audio signal output by the speaker.

[0032] The matrix determination module is used to determine the angle matrix of the spatial audio signal based on the spatial audio signal.

[0033] The index determination module is used to determine the objective index of spatial sense based on the angle matrix of the spatial audio signal. The objective index of spatial sense includes at least one of the following: objective index of positioning sense and objective index of immersion sense.

[0034] Thirdly, embodiments of this application provide an electronic device, including a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the electronic device enables the evaluation method for acoustic spatiality as described in the first aspect above.

[0035] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the sound spatial evaluation method as described in the first aspect above.

[0036] Fifthly, embodiments of this application provide a computer program product that, when run on an electronic device, causes the electronic device to execute the sound spatial perception evaluation method described in the first aspect.

[0037] The beneficial effects of the embodiments in this application compared with the prior art are:

[0038] This application acquires the spatial audio signal output by an audio system, determines the angle matrix of the spatial audio signal based on the spatial audio signal, and determines the objective spatial perception index based on the angle matrix of the spatial audio signal. The objective spatial perception index includes at least one of the following: objective positioning index and objective immersion index. This application constructs a set of reasonable objective evaluation indicators that closely approximate the user's subjective experience. It combines the advantages of objective evaluation methods—ease of operation and high repeatability—with the advantages of subjective evaluation methods—high correlation with listening experience—and can bring greater benefits to the evaluation of audio systems (such as in-vehicle immersive sound systems). This application can reflect the subjective listening experience of spatial audio signals through a process-oriented, objectively measured spatial perception index. This application more clearly reveals the relationship between subjective and objective perception evaluations of spatial audio signals, assisting in the optimization and iteration of audio systems. Furthermore, the audio spatial perception evaluation method proposed in this application is non-invasive, requiring no modification to the hardware environment of the audio system (e.g., no modification to the vehicle environment where the audio system is located), avoiding significant modification costs. It is convenient and process-oriented, enabling rapid signal analysis and providing an important tool for the design, tuning, acceptance, and comparison of related products of audio systems. It can be easily used for comparison of related products. It can significantly save engineers' time and energy and improve efficiency during the design and tuning stages.

[0039] It is understood that the beneficial effects of the second to fifth aspects mentioned above can be found in the relevant descriptions in the first aspect mentioned above, and will not be repeated here. Attached Figure Description

[0040] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0041] Figure 1 This is a schematic flowchart of a method for evaluating the spatial sense of sound provided in an embodiment of this application;

[0042] Figure 2 This is a flowchart illustrating the method for evaluating the spatial sense of sound provided in an application embodiment of this application;

[0043] Figure 3 This is a schematic diagram of the structure of an acoustic spatial perception evaluation device provided in an embodiment of this application;

[0044] Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0045] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.

[0046] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.

[0047] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.

[0048] As used in this application specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if detected [the described condition or event]" may be interpreted, depending on the context, as meaning "once determined," "in response to determination," "once detected [the described condition or event]," or "in response to detection [the described condition or event]."

[0049] Furthermore, in the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0050] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.

[0051] The sound spatial perception evaluation method provided in this application embodiment can be applied to electronic devices such as mobile phones, tablets, wearable devices, in-vehicle devices, augmented reality (AR) / virtual reality (VR) devices, laptops, ultra-mobile personal computers (UMPCs), netbooks, and personal digital assistants (PDAs). This application embodiment does not impose any restrictions on the specific type of electronic device.

[0052] For example, the electronic device may be a station (STAION, ST) in a WLAN, a cellular phone, a cordless phone, a Session Initiation Protocol (SIP) phone, a Wireless Local Loop (WLL) station, a Personal Digital Assistant (PDA) device, a handheld device with wireless communication capabilities, a computing device or other processing device connected to a wireless modem, an in-vehicle device, a vehicle networking terminal, a computer, a laptop computer, a handheld communication device, a handheld computing device, a satellite wireless device, a wireless modem card, a set-top box (STB), customer premises equipment (CPE), and / or other devices for communication over a wireless system, as well as next-generation communication systems, such as mobile terminals in 5G networks or mobile terminals in future evolved Public Land Mobile Network (PLMN) networks.

[0053] Figure 1 This is a schematic flowchart of an embodiment of the method for evaluating the spatial sense of sound provided in this application.

[0054] S11, acquire the spatial audio signal output by the speaker.

[0055] Here, spatial audio signals from the speaker output can be recorded using devices such as microphones. Alternatively, the recorded spatial audio signals from the speaker output can be acquired via file transfer. Microphones include, but are not limited to, FOA (First-Order Ambisonics) microphones. FOA microphones have mature technology and corresponding commercial plugins, enabling the acquisition of sound field information for the entire space.

[0056] Spatial audio signals are synonymous with three-dimensional sound specifications. Compared to traditional multi-channel audio, spatial audio signals not only have a greater number of channels than stereo, but also may allow object-based sound information. These sounds are not simply panned into a channel, but are precisely recorded in cylindrical coordinates relative to the listener's location (static) or even displacement (dynamic). On the user's terminal, in conjunction with local or cloud-based decoders and renderers, the amplitude-frequency response (timbre) and phase-frequency response (spatial sense) of these sound objects in the channels (sound bed) can be accurately calculated and reproduced.

[0057] Spatial audio signals can provide more precise sound source location information and more immersive environmental information; all spatial audio signal content can be considered as a combination of sound source and environment. Based on this, the spatial audio signals acquired for audio output include, but are not limited to, mono-source spatial audio signals and surround spatial audio signals. Mono-source spatial audio signals include, but are not limited to, mono-source stationary spatial audio signals and mono-source moving spatial audio signals. Surround spatial audio signals include, but are not limited to, multi-channel orthogonal spatial audio signals and other non-perfectly orthogonal spatial audio signals.

[0058] Since spatial perception is primarily determined by the spatial distribution of spatial audio signals, it is not significantly related to sound loudness. However, sufficient loudness is a necessary condition for improving the signal-to-noise ratio and fully revealing audio details. Therefore, the loudness of all spatial audio signals should be increased as much as possible without exceeding the microphone's range.

[0059] The definitions of single-source static spatial audio signals, single-source moving spatial audio signals, and multi-channel orthogonal spatial audio signals are shown in Table 1 below:

[0060]

[0061] Table 1

[0062] Those skilled in the art will understand that the above-described methods of using a vehicle audio system for playback are merely examples. While this application primarily uses in-vehicle audio environments as examples, it can also be applied to other environments where audio systems play audio, such as home environments, cinema environments, and various other indoor and outdoor scenarios.

[0063] S12, Determine the angle matrix of the spatial audio signal based on the spatial audio signal.

[0064] The angular matrix of the spatial audio signal includes, but is not limited to, the planar azimuth matrix and the pitch matrix.

[0065] Wherein, the planar azimuth matrix Θ is derived from θ k,nComposition, θ represents the plane azimuth angle, θ∈[-π, π), k represents the frequency point, n represents the time frame sequence (or simply time sequence), θ k,n This represents the planar azimuth angle value at frequency k and time frame n. The planar azimuth angle values ​​θ at different frequencies k and different time frame sequences n are also represented. k,n Form the planar azimuth matrix Θ.

[0066] The pitch angle matrix Φ is composed of φ k,n Composition, φ represents the pitch angle, k represents the frequency point, n represents the time frame sequence (or simply time sequence), and φ k,n This represents the elevation angle value at frequency k and time frame n. The elevation angle values ​​φ at different frequency k and different time frame sequences n are also represented. k,n Form the pitch angle matrix Φ.

[0067] In one embodiment, the spatial audio signal includes a first-order high-fidelity stereo reproduction A-format spatial audio signal, and determining the angle matrix of the spatial audio signal based on the spatial audio signal includes:

[0068] The spatial audio signal in format A of the first-order high-fidelity stereo reproduction is converted to obtain a spatial audio signal in format B of the first-order high-fidelity stereo reproduction; the angle matrix of the spatial audio signal is determined based on the spatial audio signal in format B of the first-order high-fidelity stereo reproduction.

[0069] The spatial audio signals for first-order high-fidelity stereo reproduction in A-format (FOA-A format) include, but are not limited to, at least one of the following: front left signal (LF), rear left signal (LB), front right signal (RF), and rear right signal (RB). The spatial audio signals for first-order high-fidelity stereo reproduction in B-format (FOA-B format) include, but are not limited to, at least one of the following: omnidirectional signal (W), front-to-back depth signal (X), left-to-right width signal (Y), and vertical height signal (Z).

[0070] The following two methods are included but are not limited to converting FOA-A format spatial audio signals to obtain FOA-B format spatial audio signals.

[0071] (1) Substitute the FOA-A format spatial audio signal as the input into the following calculation formula to obtain the FOA-B format spatial audio signal.

[0072]

[0073] Among them, LF, LB, RF, and RB are spatial audio signals in FOA-A format, where LF represents the left front signal, LB represents the left rear signal, RF represents the right front signal, and RB represents the right rear signal. W, X, Y, and Z are spatial audio signals in FOA-B format, where W represents the omnidirectional signal, X represents the front-to-back depth signal, Y represents the left-to-right width signal, and Z represents the vertical height signal.

[0074] (2) Convert the spatial audio signal in FOA-A format to FOA-B format by using the Audition plugin of the FOA microphone.

[0075] After acquiring the spatial audio signal in FOA-B format, the angle matrix of the spatial audio signal can be determined. The angle matrix of the spatial audio signal includes, but is not limited to, the plane azimuth matrix and the elevation matrix.

[0076] Planar azimuth angle values ​​θ at different frequency points k and different time frame sequences n k,n This forms the planar azimuth matrix Θ. Where θ k,n It can be calculated based on the following formula:

[0077]

[0078] in, x represents the conjugate result of the omnidirectional signal values ​​at frequency k and time frame n. y (k,n) represents the left and right width signal values ​​at frequency point k and time frame number n, x x (k,n) represents the depth signal values ​​before and after frequency point k and time frame number n. Indicates to The complex result (which includes both real and imaginary parts) takes the result calculated from the real part. Indicates to The complex result (which includes both real and imaginary parts) is calculated using the real part.

[0079] Pitch angle values ​​φ at different frequency points k and different time frame sequences n k,n This forms the pitch angle matrix Φ. Where φ k,n It can be calculated based on the following formula:

[0080]

[0081] in, x represents the conjugate result of the omnidirectional signal values ​​at frequency k and time frame n. z (k,n) represents the vertical height signal values ​​at frequency point k and time frame number n, x x (k,n) represents the depth signal values ​​before and after frequency point k and time frame number n, xy (k,n) represents the left and right width signal values ​​at frequency k and time frame n. Indicates to The complex result (which includes both real and imaginary parts) takes the result calculated from the real part. Indicates to The complex result (which includes both real and imaginary parts) takes the result calculated from the real part. Indicates to The complex result (which includes both real and imaginary parts) takes the result calculated from the real part. Indicates to The complex result (which includes both real and imaginary parts) is calculated using the real part.

[0082] S13. Determine the objective indicators of spatial perception based on the angle matrix of the spatial audio signal. The objective indicators of spatial perception include at least one of the following: objective indicators of positioning and objective indicators of immersion.

[0083] Through research on spatial hearing, the information that spatial hearing can provide has been obtained. That is, when the human auditory system hears a sound, in addition to the sound content itself, how much information about the listening space can it obtain? Based on this information, spatial hearing is decoupled and decomposed into two main categories: localization and immersion. Therefore, the objective spatial perception indicators determined in this application include at least one of the following: objective localization indicators and objective immersion indicators. Among them, the objective localization indicators are used to describe the objective representation capability of the position of each sound element when a loudspeaker system plays music. Objective localization indicators include static and dynamic localization indicators. Objective localization indicators can be expressed as a comprehensive manifestation of the above-mentioned static and dynamic objective localization indicators.

[0084] In one embodiment, the spatial audio signal includes a single-source stationary spatial audio signal, the objective index of spatial positioning includes a static objective index of spatial positioning, and the step of determining the objective index of spatial positioning based on the angle matrix of the spatial audio signal includes:

[0085] Based on the angle matrix of the single-source static spatial audio signal, the sound source direction of the single-source static spatial audio signal is determined; based on the angle matrix of the single-source static spatial audio signal and the sound source direction, the variance of the angle matrix is ​​calculated as an objective index of static positioning.

[0086] Static positioning accuracy is an objective metric used to measure directional precision. By defining the difference between the channel direction perceived by a listener inside the vehicle and a standard (reference) direction, it measures the difference in sound location perception between the in-vehicle spatial audio system and a standard listening environment. This metric can be used to assess the accuracy of the location of sound elements during multi-channel in-vehicle audio playback. Because the speaker layout in a car differs from that of a standard multi-channel home theater or listening room, it's necessary to create a location that conforms to the standard layout using synthesized virtual sources. This means that directly measuring speaker location does not represent the perceived location of the sound source. Accurate channel location is fundamental to ensuring accurate overall spatial sound location, as multi-channel audio content is typically produced in a standard recording studio. The more standardized the location of each channel in the car, the better the spatial audio signal content can be presented. A synthesized virtual source refers to a virtual sound source synthesized from two or more real sources; the sound is imaged at that location, but the actual sound source does not exist.

[0087] During measurement, a monophonic static spatial audio signal can be recorded using a FOA microphone, based on the number of audio channels supported by the in-vehicle audio system. This results in a FOA-A format monophonic static spatial audio signal output from the in-vehicle spatial audio system. The FOA-A format monophonic static spatial audio signal is then converted to FOA-B format. Based on the FOA-B format monophonic static spatial audio signal, the angle matrix of the monophonic static spatial audio signal is determined. Test points can be selected, for example, the center of the car or the driver's seat, and the height of the measurement points can be aligned with the user's head position.

[0088] The planar azimuth matrix Θ of a single-source static spatial audio signal is composed of the planar azimuth angle values ​​θ at different frequency points k and different time frame sequences n. k,n Composition. Because angle data exhibits cyclical symmetry, the variance cannot be directly calculated; therefore, the angle values ​​θ are... k,n Transform to the complex field, i.e., θ k,n →(cosθ k,n sinθ k,n ).

[0089] Based on the angle matrix of the single-source stationary spatial audio signal, the probability of the single-source stationary spatial audio signal coming from each direction can be obtained by statistically analyzing the angle distribution. The source direction of the single-source stationary spatial audio signal can then be obtained by calculating the mean.

[0090]

[0091] Where mean(Θ) represents the direction of the sound source of a single-source static spatial audio signal.

[0092] The variance of the angle matrix can be calculated using the following formula based on the angle matrix of a single-source static spatial audio signal and the direction of the sound source, serving as an objective indicator of static positioning.

[0093]

[0094] Where θ represents the plane azimuth angle, θ∈[-π, π), k represents the frequency point, and n represents the time frame sequence (or simply time sequence). k,n This represents the planar azimuth angle value at frequency k and time frame n. The planar azimuth angle values ​​θ at different frequencies k and different time frame sequences n are also represented. k,n This forms the planar azimuth matrix Θ. K is the total number of frequency bands, and N is the total number of time frames. var(Θ) represents the variance of the angle matrix.

[0095] Using the variance of the angle matrix as an objective index of static positioning sense, this index reflects the concentration of sound direction in a car audio system and characterizes the quality of static sound source location sense. A smaller static positioning sense index value indicates a more concentrated sound image and better static positioning sense.

[0096] In one embodiment, the spatial audio signal includes a single-source motion spatial audio signal, the objective index of spatial positioning includes a dynamic objective index of spatial positioning, and determining the objective index of spatial positioning based on the angle matrix of the spatial audio signal includes:

[0097] Based on the angle matrix of the single-source motion spatial audio signal, the azimuth coordinates of the single-source motion spatial audio signal in each time period are determined; based on the azimuth coordinates of the single-source motion spatial audio signal in each time period, a dynamic positioning objective index is determined, wherein the dynamic positioning objective index includes at least one of the following: a first dynamic positioning objective index and a second dynamic positioning objective index, wherein the first dynamic positioning objective index is the error of the single-source motion spatial audio signal, and the second dynamic positioning objective index is the first-order difference of the single-source motion spatial audio signal.

[0098] Dynamic positioning sense defines whether the in-vehicle spatial audio system can ensure the accuracy and clarity of the trajectory when outputting moving sound sources.

[0099] During measurement, a monophonic motion spatial audio signal can be used, based on the number of audio channels supported by the in-vehicle audio system. This signal is measured using a FOA microphone to obtain a monophonic motion spatial audio signal in FOA-A format output by the in-vehicle spatial audio system. The FOA-A format monophonic motion spatial audio signal is then converted to FOA-B format. Based on the FOA-B format monophonic motion spatial audio signal, the angle matrix of the monophonic motion spatial audio signal is determined. Test points can be selected, for example, the center of the car or the driver's seat, and the height of the measurement points can be aligned with the user's head level.

[0100] Because single-source motion spatial audio signals require the sound source to move slowly, the sound source can be considered stationary within a short time frame. Therefore, the single-source motion spatial audio signal can be segmented to determine its azimuth coordinates for each time period. The planar azimuth matrix Θ of the single-source motion spatial audio signal is composed of the planar azimuth angle values ​​θ at different frequency points k and different time frame sequences n. k,n Composition. Because angle data exhibits cyclical symmetry, the variance cannot be directly calculated; therefore, the angle values ​​θ are... k,n Transform to the complex field, i.e., θ k,n →(cosθ k,n sinθ k,n Based on the angle matrix of the single-source moving spatial audio signal, the azimuth coordinates of the single-source moving spatial audio signal in each time period can be determined using the following formula.

[0101]

[0102] Where, Θ t Let N' represent the angle matrix for each time period t, and let N' represent the number of frames in each segment. t The mean(Θ) represents the azimuth coordinates of a single-source moving spatial audio signal at various time intervals t. t This can be written as trace(X(t)) = mean(Θ). t ).

[0103] The angular distribution variance of a single-source moving spatial audio signal at various time periods can be determined based on the angular matrix of the single-source moving spatial audio signal and the following formula.

[0104]

[0105] Where, Θ t This represents the angle matrix for each time period t, and N′ represents the number of frames in each segment. var(Θ) t) represents the variance of the angular distribution of a single-source moving spatial audio signal at various time intervals t.

[0106] In one embodiment, determining the objective index of dynamic positioning sense based on the azimuth coordinates of the single-source motion spatial audio signal at various time periods includes:

[0107] Based on the temporal distribution of the azimuth coordinates and reference trajectory of the single-source motion spatial audio signal in various time periods, the error of the single-source motion spatial audio signal is determined as the first objective index of dynamic positioning sense.

[0108] Here, based on the azimuth coordinates and temporal distribution of the single-source motion spatial audio signal in various time periods and the reference trajectory, the error of the single-source motion spatial audio signal can be determined as the first objective indicator of dynamic positioning sense using the following formula.

[0109]

[0110] in, The temporal distribution of the reference trajectory is represented by Err, which represents the error of the single-source motion spatial audio signal. The error characterizes the accuracy of the dynamic sound source trajectory. The error of the single-source motion spatial audio signal is used as the first objective index of dynamic positioning. The smaller the value of the first objective index of dynamic positioning, the better the continuity of dynamic changes and the better the dynamic positioning.

[0111] In one embodiment, determining the objective index of dynamic positioning sense based on the azimuth coordinates of the single-source motion spatial audio signal at various time periods includes:

[0112] Based on the azimuth coordinates of the single-source motion spatial audio signal at various time periods, the first-order difference of the single-source motion spatial audio signal is determined as the second objective index of dynamic positioning.

[0113] Here, the first-order difference of the single-source motion spatial audio signal can be determined as the second objective index of dynamic positioning sense based on the azimuth coordinates of the single-source motion spatial audio signal at various time periods, according to the following formula.

[0114]

[0115] Here, Diff represents the first-order difference of the single-source motion spatial audio signal, which can reflect the continuity of change. The first-order difference of the single-source motion spatial audio signal is used as the second objective index of dynamic localization, and the smaller the value of the second objective index of dynamic localization, the better the continuity.

[0116] In one embodiment, the spatial audio signal includes a multi-channel orthogonal spatial audio signal, the immersion objective index includes an envelopment objective index, and determining the spatial objective index based on the angle matrix of the spatial audio signal includes:

[0117] Based on the angle matrix of the multi-channel orthogonal spatial audio signal, at least two of the following objective energy distribution indicators are determined: front energy distribution objective indicator, rear energy distribution objective indicator, left energy distribution objective indicator, right energy distribution objective indicator, upper energy distribution objective indicator, and lower energy distribution objective indicator.

[0118] Based on the at least two objective indicators of energy distribution, an objective indicator of sense of enclosure is determined, wherein the objective indicator of sense of enclosure includes at least one of the following: objective indicator of front-back sense of enclosure, objective indicator of horizontal sense of enclosure, and objective indicator of vertical sense of enclosure.

[0119] A crucial experience in spatial audio signals is the feeling of being "immersed in the sound field." This feeling stems from: 1) the sensation of being surrounded by ambient sound energy, also known as immersion; and 2) the independence of sound information from different directions. Among these, the immersion effect of sound energy distribution is relatively important, while information independence, after ensuring a basic level of uniformity in the distribution, is used to assess the potential for in-head effects. Immersion can be expressed as a combined manifestation of these two characteristics: immersion and information independence.

[0120] Among them, the sense of immersion defines the uniformity of sound energy distribution in all directions of the car cabin when the car audio system plays surround sound sources.

[0121] During measurement, a multi-channel orthogonal spatial audio signal can be used, based on the number of audio channels supported by the in-vehicle audio system. This signal is measured through a FOA microphone to obtain the FOA-A format multi-channel orthogonal spatial audio signal output by the in-vehicle spatial audio system. The FOA-A format multi-channel orthogonal spatial audio signal is then converted to FOA-B format. Based on the FOA-B format multi-channel orthogonal spatial audio signal, the angle matrix of the multi-channel orthogonal spatial audio signal is determined. Test points can be selected, for example, the center of the car or the driver's seat, and the height of the measurement points can be aligned with the user's head level.

[0122] Among them, the angle matrix of the multi-channel orthogonal spatial audio signal includes, but is not limited to, the planar azimuth matrix and the pitch matrix.

[0123] Wherein, the planar azimuth matrix Θ is derived from θ k,n Composition, θ represents the plane azimuth angle, θ∈[-π, π), k represents the frequency point, n represents the time frame sequence (or simply time sequence), θ k,nThis represents the planar azimuth angle value at frequency k and time frame n. The planar azimuth angle values ​​θ at different frequencies k and different time frame sequences n are also represented. k,n Form the planar azimuth matrix Θ.

[0124] The pitch angle matrix Φ is composed of φ k,n Composition, φ represents the pitch angle, k represents the frequency point, n represents the time frame sequence (or simply time sequence), and φ k,n This represents the elevation angle value at frequency k and time frame n. The elevation angle values ​​φ at different frequency k and different time frame sequences n are also represented. k,n Form the pitch angle matrix Φ.

[0125] θ is calculated using P(·). k,n The statistical probability is used to obtain the angular distribution probability P(θ) of the plane azimuth angle. φ is then calculated using P(·). k,n The statistical probability is used to obtain the angular distribution probability P(φ) of the plane azimuth angle.

[0126] Based on the angle matrix of the multi-channel orthogonal spatial audio signal, at least two of the following objective energy distribution indicators are determined using the following formulas: front energy distribution objective indicator, rear energy distribution objective indicator, left energy distribution objective indicator, right energy distribution objective indicator, upper energy distribution objective indicator, and lower energy distribution objective indicator.

[0127]

[0128] Among them, P Front P represents an objective indicator of the energy distribution ahead. Back P represents an objective indicator of energy distribution at the rear. Left P represents an objective indicator of energy distribution on the left. Right P represents an objective indicator of energy distribution on the right. Up P represents an objective indicator of the energy distribution above. Down This represents an objective indicator of the energy distribution below.

[0129] Based on at least two objective energy distribution indicators, the objective indicators of sense of enclosure are determined according to the following formula, wherein the objective indicators of sense of enclosure include at least one of the following: front-back sense of enclosure objective indicator, horizontal sense of enclosure objective indicator, and vertical sense of enclosure objective indicator.

[0130]

[0131]

[0132] Among them, Ev F-B Ev represents an objective indicator of front and rear surround feel. A-R Ev represents an objective indicator of horizontal enclosure.U-D This refers to an objective indicator of vertical enclosure. The objective indicator of front-to-back enclosure is Ev. F-B Horizontal sense of enclosure objective index Ev A-R Vertical sense of enclosure objective index Ev U-D The closer the value is to 1, the stronger the sense of immersion.

[0133] In one embodiment, the method for evaluating acoustic spatiality further includes:

[0134] Based on the first-order high-fidelity stereo sound, the multi-channel orthogonal spatial audio signal in format A is reproduced, and the forward component and backward component of the spatial audio signal are determined.

[0135] Based on the forward component and the backward component of the spatial audio signal, correlation information is determined as an objective indicator of information independence;

[0136] The objective index of immersion is determined based on the objective index of sense of immersion and the objective index of information independence.

[0137] Because the objective index of immersion uses a weighted summation of sound source angle distributions (with weights calculated using trigonometric functions), it can also reflect energy proportions. Under certain energy uniformity conditions, sounds from all directions can be heard by the listener. At this point, the in-head effect needs to be considered. Since coherent sound sources are rare in natural conditions, when sound content from opposite directions is very similar, the listener will feel a sense of strangeness, resulting in the in-head effect. If the objective index of immersion is close to 1, the existence of the in-head effect can be determined based on the objective index of information independence, thus more accurately determining the objective index of immersion and better assessing immersion.

[0138] During measurement, a multi-channel orthogonal spatial audio signal can be used, based on the number of audio channels supported by the in-vehicle audio system. This signal is measured through a FOA microphone to obtain the FOA-A format multi-channel orthogonal spatial audio signal output by the in-vehicle spatial audio system. Test points can be selected, for example, the center of the car or the driver's seat, and the height of the measurement point can be level with the user's head.

[0139] The FOA-A format (First-Order Ambisonics A-format, first-order high-fidelity stereo reproduction A-format) multichannel orthogonal spatial audio signal includes, but is not limited to, at least one of the following: left front signal LF, left rear signal LB, right front signal RF, and right rear signal RB. Since the FOA microphone uses a cardioid microphone, and the content measured by the cardioid microphone can be approximated as a sound vector pointing in the same direction, the forward and backward components of the spatial audio signal can be determined using vector summation based on the following formula:

[0140]

[0141] Among them, f Front f represents the forward component of the spatial audio signal. Back This represents the backward component of the spatial audio signal.

[0142] The IACC (Interaural Cross-Correlation) index, adjusted by logarithmic weighting, characterizes the correlation of information. Based on the forward and backward components of the spatial audio signal, the correlation information is determined as an objective indicator of information independence using the following formula:

[0143]

[0144] FIACC = max|FIACF|

[0145] Where k1 represents the lower limit of the frequency band and k2 represents the upper limit of the frequency band, determined by the upper and lower limits of the time difference between the two-channel signals, and is generally taken as ±1ms. Front f represents the forward component of the spatial audio signal. Back Represents the backward component of the spatial audio signal. This represents the result of the conjugate operation on the backward component of the spatial audio signal. Indicates to The complex result (which includes both real and imaginary parts) is calculated using the real part.

[0146] If the immersion objective index is close to 1 and the information independence objective index FIACC is large, it indicates that the immersion objective index decreases, and the immersion weakens. If the immersion objective index is large (e.g., greater than 3), the head-in-the-head effect represented by the information independence objective index FIACC has a small impact and can be considered to have no effect on the immersion objective index.

[0147] This application acquires the spatial audio signal output by an audio system, determines the angle matrix of the spatial audio signal based on the spatial audio signal, and determines the objective spatial perception index based on the angle matrix of the spatial audio signal. The objective spatial perception index includes at least one of the following: objective positioning index and objective immersion index. This application constructs a set of reasonable objective evaluation indicators that closely approximate the user's subjective experience. It combines the advantages of objective evaluation methods—ease of operation and high repeatability—with the advantages of subjective evaluation methods—high correlation with listening experience—and can bring greater benefits to the evaluation of audio systems (such as in-vehicle immersive sound systems). This application can reflect the subjective listening experience of spatial audio signals through a process-oriented, objectively measured spatial perception index. This application more clearly reveals the relationship between subjective and objective perception evaluations of spatial audio signals, assisting in the optimization and iteration of audio systems. Furthermore, the audio spatial perception evaluation method proposed in this application is non-invasive, requiring no modification to the hardware environment of the audio system (e.g., no modification to the vehicle environment where the audio system is located), avoiding significant modification costs. It is convenient and process-oriented, enabling rapid signal analysis and providing an important tool for the design, tuning, acceptance, and comparison of related products of audio systems. It can be easily used for comparison of related products. It can significantly save engineers' time and energy and improve efficiency during the design and tuning stages.

[0148] The applicant, through research on the perceptual mechanism of the auditory system, found a mapping from physical acoustics to perceptual indicators, deriving the final perceptual indicators. These indicators include: static directional accuracy, dynamic directional accuracy, static sound source intelligibility, dynamic sound source intelligibility, energy distribution uniformity, and information independence. The proposed indicators address the problem of overly general evaluation methods for spatial audio signals. The objective indicators of spatial positioning and immersion, decoupled from spatial perception, can accurately and comprehensively describe the quality of spatial audio signals.

[0149] The applicant, after comprehensively considering the needs, costs, and computational requirements of the perceived metrics, determined the necessary test signals and sound acquisition equipment. The test signals include single-source static spatial audio signals, single-source moving spatial audio signals, and multi-channel orthogonal spatial audio signals. Playback only requires an in-vehicle audio playback system; no additional speakers or vehicle modifications are needed. The sound acquisition environment setup is simple and inexpensive.

[0150] The applicant derived specific calculation methods for the aforementioned indicators through research on cognitive and psychoacoustic models, and calculated perceptual indicators using measured data. This evaluation method retains the advantages of traditional objective evaluation methods—repeatability and low cost—while also incorporating the advantages of subjective evaluation (i.e., high correlation with auditory perception), thus accurately reflecting the user experience of the product.

[0151] Figure 2This is a flowchart illustrating the method for evaluating the spatial sense of sound provided in an application embodiment of this application.

[0152] like Figure 2 As shown, the FOA (First-Order Ambisonics) microphone device can acquire the spatial audio signal output by the audio system through signal acquisition. The spatial audio signal includes the orthogonal pure test tones played by the in-vehicle speakers (i.e., multi-channel orthogonal spatial audio signals), the single-source stationary test tones played by the in-vehicle speakers (i.e., single-source stationary spatial audio signals), and the single-source motion test tones played by the in-vehicle speakers (i.e., single-source motion spatial audio signals). Next, the angle matrix of the spatial audio signal is determined based on the spatial audio signal. Subsequently, the objective index of information independence is determined through auditory cross-correlation coefficient calculation. The objective indices of surround feeling, static positioning, and dynamic positioning are determined through sound energy-direction distribution calculation. Then, the objective index of positioning is obtained based on the objective indices of static and dynamic positioning. The objective index of immersion is obtained based on the objective indices of surround feeling and information independence. The objective index of spatial perception is determined based on the objective indices of positioning and immersion for objective evaluation of spatial perception.

[0153] Figure 3 This is a schematic diagram of the structure of an acoustic spatial perception evaluation device provided in an embodiment of this application.

[0154] like Figure 3 As shown, the evaluation device 3 includes:

[0155] Signal acquisition module 31 is used to acquire the spatial audio signal output by the speaker;

[0156] The matrix determination module 32 is used to determine the angle matrix of the spatial audio signal based on the spatial audio signal.

[0157] The indicator determination module 33 is used to determine the objective indicators of spatial perception based on the angle matrix of the spatial audio signal. The objective indicators of spatial perception include at least one of the following: objective indicators of positioning and objective indicators of immersion.

[0158] Another embodiment of the present invention discloses an evaluation device 3. This embodiment is based on the above... Figure 3 Based on the corresponding embodiment, the spatial audio signal includes a first-order high-fidelity stereo reproduction A-format spatial audio signal, and the matrix determination module 32 is used for:

[0159] The spatial audio signal in the first-order high-fidelity stereo copy A format is converted to obtain the spatial audio signal in the first-order high-fidelity stereo copy B format.

[0160] Based on the first-order high-fidelity stereo reproduction of the B-format spatial audio signal, the angle matrix of the spatial audio signal is determined.

[0161] Another embodiment of the present invention discloses an evaluation device 3. This embodiment is based on the above... Figure 3 Based on the corresponding embodiment, the spatial audio signal includes a single-source static spatial audio signal, the objective positioning index includes a static positioning index, and the index determination module 33 is used for:

[0162] Based on the angle matrix of the single-source static spatial audio signal, determine the source direction of the single-source static spatial audio signal;

[0163] Based on the angle matrix of the single-source static spatial audio signal and the direction of the sound source, the variance of the angle matrix is ​​calculated as an objective indicator of static positioning.

[0164] Another embodiment of the present invention discloses an evaluation device 3. This embodiment is based on the above... Figure 3 Based on the corresponding embodiment, the spatial audio signal includes a single-source motion spatial audio signal, the objective positioning index includes a dynamic positioning index, and the index determination module 33 is used for:

[0165] Based on the angle matrix of the single-source motion spatial audio signal, determine the azimuth coordinates of the single-source motion spatial audio signal in each time period;

[0166] Based on the azimuth coordinates of the single-source motion spatial audio signal in various time periods, a dynamic positioning objective index is determined. The dynamic positioning objective index includes at least one of the following: a first dynamic positioning objective index and a second dynamic positioning objective index. The first dynamic positioning objective index is the error of the single-source motion spatial audio signal, and the second dynamic positioning objective index is the first-order difference of the single-source motion spatial audio signal.

[0167] Another embodiment of the present invention discloses an evaluation device 3. This embodiment is based on the above... Figure 3 Based on the corresponding embodiment, the objective index of dynamic positioning sense includes a first objective index of dynamic positioning sense, and the index determination module 33 is used for:

[0168] Based on the temporal distribution of the azimuth coordinates and reference trajectory of the single-source motion spatial audio signal in various time periods, the error of the single-source motion spatial audio signal is determined as the first objective indicator of dynamic positioning.

[0169] Another embodiment of the present invention discloses an evaluation device 3. This embodiment is based on the above... Figure 3Based on the corresponding embodiment, the dynamic positioning sense objective index includes a second dynamic positioning sense objective index, and the index determination module 33 is used for:

[0170] Based on the azimuth coordinates of the single-source motion spatial audio signal at various time periods, the first-order difference of the single-source motion spatial audio signal is determined as the second objective index of dynamic positioning.

[0171] Another embodiment of the present invention discloses an evaluation device 3. This embodiment is based on the above... Figure 3 Based on the corresponding embodiment, the spatial audio signal includes a multi-channel orthogonal spatial audio signal, the immersion objective index includes an envelopment objective index, and the index determination module 33 is used for:

[0172] Based on the angle matrix of the multi-channel orthogonal spatial audio signal, at least two of the following objective energy distribution indicators are determined: front energy distribution objective indicator, rear energy distribution objective indicator, left energy distribution objective indicator, right energy distribution objective indicator, upper energy distribution objective indicator, and lower energy distribution objective indicator.

[0173] Based on the at least two objective indicators of energy distribution, an objective indicator of sense of enclosure is determined, wherein the objective indicator of sense of enclosure includes at least one of the following: objective indicator of front-back sense of enclosure, objective indicator of horizontal sense of enclosure, and objective indicator of vertical sense of enclosure.

[0174] Another embodiment of the present invention discloses an evaluation device 3. This embodiment is based on the above... Figure 3 Based on the corresponding embodiment, the evaluation device 3 further includes:

[0175] The front and rear component determination module is used to determine the forward component and the backward component of the spatial audio signal based on the multi-channel orthogonal spatial audio signal in A format copied from the first-order high-fidelity stereo sound system.

[0176] The independent index determination module is used to determine correlation information as an objective index of information independence based on the forward component and the backward component of the spatial audio signal.

[0177] The immersion index determination module is used to determine the immersion objective index based on the envelopment objective index and the information independence objective index.

[0178] It should be noted that the information interaction and execution process between the above-mentioned devices / units are based on the same concept as the method embodiments of this application. For details on their specific functions and technical effects, please refer to the method embodiments section, and they will not be repeated here.

[0179] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0180] This application also provides an electronic device, such as... Figure 4 As shown, the electronic device 4 includes: at least one processor 40, a memory 41, and a computer program 42 stored in the memory 41 and executable on the at least one processor 40, wherein the processor 40 executes the computer program 42 to implement the steps in any of the above-described method embodiments.

[0181] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps described in the various method embodiments above.

[0182] This application provides a computer program product that, when run on a mobile terminal, enables the mobile terminal to implement the steps described in the above-described method embodiments.

[0183] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of this application can be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include at least: any entity or device capable of carrying computer program code to a photographic device / electronic device, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium. Examples include USB flash drives, portable hard drives, magnetic disks, or optical disks. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electrical carrier signals or telecommunication signals.

[0184] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0185] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0186] In the embodiments provided in this application, it should be understood that the disclosed devices / electronic devices and methods can be implemented in other ways. For example, the device / electronic device embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual couplings or direct couplings or communication connections may be through some interfaces; indirect couplings or communication connections between devices or units may be electrical, mechanical, or other forms.

[0187] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0188] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.

Claims

1. A method for evaluating the spatial sense of sound, characterized in that, include: Acquire the spatial audio signal output from the speaker; Based on the spatial audio signal, determine the angle matrix of the spatial audio signal; Based on the angle matrix of the spatial audio signal, determine the objective indicators of spatial perception, which include at least one of the following: objective indicators of positioning and objective indicators of immersion. The spatial audio signal includes at least one of a single-source static spatial audio signal and a multi-channel orthogonal spatial audio signal; the objective index of positioning includes an objective index of static positioning; the objective index of immersion includes an objective index of surround feeling; and the determination of the objective index of spatial feeling based on the angle matrix of the spatial audio signal includes at least one of the following: Based on the angle matrix of the single-source static spatial audio signal, the sound source direction of the single-source static spatial audio signal is determined; based on the angle matrix of the single-source static spatial audio signal and the sound source direction, the variance of the angle matrix is ​​calculated as an objective index of static positioning; and / or, Based on the angle matrix of the multi-channel orthogonal spatial audio signal, at least two of the following objective energy distribution indicators are determined: front energy distribution objective indicator, rear energy distribution objective indicator, left energy distribution objective indicator, right energy distribution objective indicator, upper energy distribution objective indicator, and lower energy distribution objective indicator; based on the at least two objective energy distribution indicators, an objective envelopment indicator is determined, which includes at least one of the following: front-to-back envelopment objective indicator, horizontal envelopment objective indicator, and vertical envelopment objective indicator.

2. The method for evaluating acoustic spatiality as described in claim 1, characterized in that, The spatial audio signal includes a first-order high-fidelity stereo reproduction A-format spatial audio signal. Determining the angle matrix of the spatial audio signal based on the spatial audio signal includes: The spatial audio signal in the first-order high-fidelity stereo copy A format is converted to obtain the spatial audio signal in the first-order high-fidelity stereo copy B format. Based on the first-order high-fidelity stereo reproduction of the B-format spatial audio signal, the angle matrix of the spatial audio signal is determined.

3. The method for evaluating acoustic spatiality as described in claim 1 or 2, characterized in that, The spatial audio signal includes a single-source motion spatial audio signal, the objective index of positioning includes a dynamic positioning objective index, and the step of determining the objective index of spatial positioning based on the angle matrix of the spatial audio signal includes: Based on the angle matrix of the single-source motion spatial audio signal, determine the azimuth coordinates of the single-source motion spatial audio signal in each time period; Based on the azimuth coordinates of the single-source motion spatial audio signal in various time periods, a dynamic positioning objective index is determined. The dynamic positioning objective index includes at least one of the following: a first dynamic positioning objective index and a second dynamic positioning objective index. The first dynamic positioning objective index is the error of the single-source motion spatial audio signal, and the second dynamic positioning objective index is the first-order difference of the single-source motion spatial audio signal.

4. The method for evaluating acoustic spatiality as described in claim 3, characterized in that, The objective indicators of dynamic positioning include a first objective indicator of dynamic positioning. Determining the objective indicators of dynamic positioning based on the azimuth coordinates of the single-source motion spatial audio signal at various time periods includes: Based on the temporal distribution of the azimuth coordinates and reference trajectory of the single-source motion spatial audio signal in various time periods, the error of the single-source motion spatial audio signal is determined as the first objective index of dynamic positioning sense.

5. The method for evaluating acoustic spatiality as described in claim 3, characterized in that, The dynamic positioning objective index includes a second dynamic positioning objective index. Determining the dynamic positioning objective index based on the azimuth coordinates of the single-source motion spatial audio signal at various time periods includes: Based on the azimuth coordinates of the single-source motion spatial audio signal at various time periods, the first-order difference of the single-source motion spatial audio signal is determined as the second objective index of dynamic positioning.

6. The method for evaluating acoustic spatiality as described in claim 2, characterized in that, The method also includes: Based on the first-order high-fidelity stereo sound, the multi-channel orthogonal spatial audio signal in format A is reproduced, and the forward component and backward component of the spatial audio signal are determined. Based on the forward component and the backward component of the spatial audio signal, correlation information is determined as an objective indicator of information independence; The objective index of immersion is determined based on the objective index of sense of immersion and the objective index of information independence.

7. A device for evaluating the spatial sense of sound, characterized in that, include: The signal acquisition module is used to acquire the spatial audio signal output by the speaker. The matrix determination module is used to determine the angle matrix of the spatial audio signal based on the spatial audio signal. The index determination module is used to determine the objective index of spatial sense based on the angle matrix of the spatial audio signal. The objective index of spatial sense includes at least one of the following: objective index of positioning sense and objective index of immersion sense. The spatial audio signal includes at least one of a single-source static spatial audio signal and a multi-channel orthogonal spatial audio signal; the objective index of positioning includes an objective index of static positioning; the objective index of immersion includes an objective index of surround sound; and the index determination module is used for at least one of the following: Based on the angle matrix of the single-source static spatial audio signal, the sound source direction of the single-source static spatial audio signal is determined; based on the angle matrix of the single-source static spatial audio signal and the sound source direction, the variance of the angle matrix is ​​calculated as an objective index of static positioning; and / or, Based on the angle matrix of the multi-channel orthogonal spatial audio signal, at least two of the following objective energy distribution indicators are determined: front energy distribution objective indicator, rear energy distribution objective indicator, left energy distribution objective indicator, right energy distribution objective indicator, upper energy distribution objective indicator, and lower energy distribution objective indicator; based on the at least two objective energy distribution indicators, an objective envelopment indicator is determined, which includes at least one of the following: front-to-back envelopment objective indicator, horizontal envelopment objective indicator, and vertical envelopment objective indicator.

8. An electronic device comprising a processor, a memory, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it causes the electronic device to implement the method as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Method and device for processing audio data of sound field

    CN106993249A

  • Method and system for objectively evaluating sense of space of sound equipment in vehicle, equipment and storage medium

    CN111935624A