A binaural rendering spatial perceptual quality test method, device, equipment and storage medium

By defining the test paradigm and sequence, calibrating the signal, and receiving listener scores, a valid score dataset is generated, solving the systematic evaluation problem of binaural rendering spatial perception quality testing and improving testing efficiency and user experience.

CN122435952APending Publication Date: 2026-07-21MALANSHAN AUDIO & VIDEO LABORATORY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
MALANSHAN AUDIO & VIDEO LABORATORY
Filing Date
2026-05-13
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Existing technologies lack a systematic evaluation of the spatial perception quality of binaural rendering, cannot fully reflect the real user experience, and lack standardized and automated testing systems.

Method used

A method for testing spatial perception quality using binaural rendering is provided. By determining the test paradigm and sequence, calibrating the signal, and receiving multi-dimensional scoring data from listeners, an effective scoring dataset is generated, including the quality test results of sound quality score, spatial perception score, and defect distribution.

Benefits of technology

It improves the efficiency of spatial perception quality testing for binaural rendering audio materials and enhances user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122435952A_ABST
    Figure CN122435952A_ABST
Patent Text Reader

Abstract

The application discloses a binaural rendering spatial perception quality test method and device, equipment and a storage medium, and relates to the technical field of audio processing. The method comprises the following steps: determining a corresponding test paradigm based on a binaural rendering audio material to be tested and a test type, determining a test sequence based on the test paradigm, calibrating the test sequence to obtain a calibrated test signal, receiving original score data input by a target listener based on the calibrated test signal, an audio quality basic attribute dimension and a spatial perception attribute dimension, the target listener being a listener who has passed a spatial perception special training test, verifying the original score data to obtain an effective score data set, and generating a quality test result comprising an audio quality score, a spatial perception score and a defect distribution based on the effective score data set, so that the efficiency of spatial perception quality test on the binaural rendering audio material is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of audio processing technology, and in particular to a method, apparatus, device, and storage medium for testing spatial perception quality using binaural rendering. Background Technology

[0002] Currently, AI-enhanced binaural rendering technology based on HRTF / BRIR is widely used in consumer terminals such as mobile phones, headphones, and VR / AR devices. However, the industry generally faces the following key technical challenges: Existing subjective evaluation systems only focus on basic sound quality attributes such as low frequency, clarity, dynamic range, timbre, noise, and vocal quality, completely lacking a systematic evaluation of spatial perception dimensions such as spatial orientation, sound source movement, and spatial immersion, and thus cannot fully reflect the real user experience of binaural rendering.

[0003] There is currently no dedicated testing system for spatial perception in binaural rendering. For spatial perception characteristics such as sound source localization accuracy, sound source motion continuity, motion uniformity, and spatial immersion, there is no dedicated system that can support standardized and automated testing.

[0004] As can be seen from the above, improving the efficiency of spatial perception quality testing for binaural rendered audio materials is an urgent problem to be solved. Summary of the Invention

[0005] In view of this, the purpose of this invention is to provide a method, apparatus, device, and storage medium for testing the spatial perception quality of binaural rendering, which can improve the efficiency of testing the spatial perception quality of binaural rendered audio materials during the spatial perception quality testing process. The specific solution is as follows: Firstly, this application provides a method for testing the spatial perception quality of binaural rendering, including: The corresponding test paradigm is determined based on the binaural rendered audio material to be tested and the test type, and the test sequence is determined based on the test paradigm. The test sequence is calibrated to obtain a calibrated test signal, and the original scoring data input by the target listener based on the calibrated test signal, the basic sound quality attribute dimension, and the spatial perception attribute dimension is received; the target listener is a listener who has passed the spatial perception specialized training test; The original scoring data is examined to obtain a valid scoring dataset, and quality test results including sound quality score, spatial perception score and defect distribution are generated based on the valid scoring dataset.

[0006] Optionally, determining the corresponding test paradigm based on the binaural rendered audio material to be tested and the test type includes: Determine the test type corresponding to the binaural rendered audio material to be tested; the test type includes positioning test, motion test, and immersion and sound quality test. Obtain reference material; the reference material is the original audio with the same content as the binaural rendering audio material to be tested; the reference material is used to measure the reliability of the target listener's score; Acquire spatial anchor point materials; the spatial perception quality level corresponding to the spatial anchor point materials meets preset stability conditions, and the spatial anchor point materials cover each preset segment region in a preset grading scale; the spatial anchor point materials are used to distinguish the unacceptable and acceptable quality boundaries of the binaural rendering audio materials to be tested; the preset segment regions include preset low-score segment regions that meet preset low-score conditions, preset medium-score segment regions that meet preset medium-score conditions, and preset high-score segment regions that meet preset high-score conditions; The corresponding test paradigm is determined using the material management unit based on the binaural rendered audio material to be tested, the reference material, the spatial anchor point material, and the test type; the test paradigm includes a first test paradigm, a second test paradigm, and a third test paradigm; the first test paradigm is a three-excitation paradigm; the second test paradigm is a multi-excitation hidden reference and anchor point paradigm; and the third test paradigm is a multiple comparison method paradigm.

[0007] Optionally, obtaining reference materials includes: Obtain reference materials; wherein, the reference materials include any one of the first reference materials, the second reference materials, and the third reference materials; The first reference material is an ideal spatial rendering reference material obtained by binaural recording in an anechoic chamber using HATS or by a real person wearing a miniature microphone; the second reference material is material generated using a preset high-quality binaural rendering system; the rendering quality of the preset high-quality binaural rendering system meets the preset high-quality conditions; the third reference material is material generated by directly feeding the original signal to the headphones without binaural rendering.

[0008] Optionally, determining the test sequence based on the test paradigm includes: Based on the test paradigm, the corresponding signal combination method is determined, and the binaural rendering audio material to be tested is randomly sorted to obtain the sorted material. Based on the signal combination method, the reference material, and the spatial anchor point material, the sorted material is subjected to the insertion operation of the reference material and the spatial anchor point and sorting combination to obtain the test sequence.

[0009] Optionally, the positioning test corresponds to the three-excitation paradigm; the motion test corresponds to the multi-excitation hidden reference and anchor point paradigm; and the immersion and sound quality test corresponds to the multiple comparison method paradigm.

[0010] Optionally, the reference material is inserted into the binaural rendering audio material to be tested based on the reference material, including: Determine the test requirements corresponding to the test paradigm, and insert the reference material into the first preset position in the sorted material based on the test requirements to obtain the inserted material.

[0011] Optionally, spatial anchoring is performed on the binaural rendering audio material to be tested based on the spatial anchoring material, including: Based on the test requirements, insert the spatial anchor point material into the second preset position in the inserted material; The spatial anchor point material is set for the sound source localization accuracy dimension; the spatial anchor point material is a rendering output result with localization deviation and before-after confusion, or the spatial anchor point material is a material obtained by applying quantized spatial perturbation to the reference material; Set the spatial anchor point material for the continuity dimension of the sound source motion; the spatial anchor point material is the rendering output result of inserting periodic jumps or breaks in the continuous motion trajectory; The spatial anchor point material is set for the uniformity dimension of the sound source motion; the spatial anchor point material is the rendering output result of applying random speed fluctuations in uniform motion, or the spatial anchor point material is the output result of interpolating the azimuth angle to meet the preset low frame rate condition. The spatial anchor point material is set for the spatial immersion dimension; the spatial anchor point material is the lowest score anchor point of spatial immersion loss determined by the mono submix version and the externalized loss anchor point material determined by the stereo signal without binaural rendering; wherein, the number of anchor point signals corresponding to each evaluation dimension is a preset number.

[0012] Optionally, calibrating the test sequence to obtain the calibrated test signal includes: Time delay calibration is performed on each of the binaural rendered audio materials to be tested, the reference material, and the spatial anchor point material in the test sequence to obtain the time delay calibrated test signal; Phase calibration is performed on each of the binaural rendered audio materials to be tested, the reference material, and the spatial anchor point material in the test sequence to obtain phase-calibrated test signals; the phase response of each phase-calibrated test signal at the binaural ears of the target listener is consistent; The sound level of each of the binaural rendered audio materials to be tested, the reference material, and the spatial anchor point material in the test sequence is calibrated to obtain a sound level calibrated test signal; the perceived loudness of the sound level calibrated test signal at the binaural ears of the target listener is not less than a preset standard sound pressure level; The calibrated test signal is determined based on the time delay calibrated test signal, the phase calibrated test signal, and the sound level calibrated test signal.

[0013] Optionally, before receiving the raw scoring data input by the target listener based on the calibrated test signal, the basic sound quality attribute dimension, and the spatial perception attribute dimension, the method further includes: The test method corresponding to the quality test is determined. If the test method is the ITU-R BS.1116 method, a first preset number of first expert listeners with experience in audio impairment judgment and spatial perception identification are determined. The first preset number is greater than a first preset threshold. If the testing method is the MUSHRA method, then a second preset number of second expert listeners with spatial perception discrimination experience are determined; the second preset number is greater than a second preset threshold; the first preset threshold is greater than the second preset threshold; the spatial perception training background of the first expert listener and the second expert listener is consistent; Based on the first expert listeners and the second expert listeners, the listeners to be screened are determined, and the listeners to be screened are screened for hearing thresholds based on a preset frequency range. The frequency hearing thresholds corresponding to the listeners to be screened and each test frequency point are obtained, and it is determined whether the hearing thresholds of each frequency point are greater than a preset decibel value. If the hearing threshold at each frequency point is not greater than the preset decibel value, the candidate to be screened is determined to have passed the hearing screening. If the hearing threshold at each frequency point is less than the preset decibel value, the candidate to be screened is determined to have failed the hearing screening, and the testing process for the candidate to be screened is terminated.

[0014] Optionally, the process of training listeners may include: The listener is trained on sound source localization accuracy using training materials to obtain the first trained listener; the sound source localization accuracy training includes identifying azimuth deviation, elevation deviation, distance perception error and front-back confusion, and determining the perceptual characteristics of the head-in-the-head effect; The first trained listener is trained in the continuity dimension of sound source motion to obtain the second trained listener; the continuity dimension training of sound source motion includes distinguishing the continuity and discontinuity of motion trajectory and identifying perceptual defects including trajectory jumps and jitters. The second trained listener is trained in the uniformity dimension of sound source motion to obtain the third trained listener; the uniformity dimension training of sound source motion includes perceiving speed fluctuations in uniform motion and recognizing speed non-uniformity phenomena. The third trained listener is then trained in the spatial immersion dimension to obtain the target listener; the spatial immersion dimension training includes experiencing the differences in the degree of externalization of sound and spatial envelopment, as well as understanding the perceptual characteristics of each spatial type; wherein, the training time for training the listener is not less than a preset training threshold. Record the cumulative training time of the target listener in the spatial perception-specific training, and compare the cumulative training time with a preset training time threshold; If the comparison result indicates that the cumulative training time is not less than the preset training time threshold, then a verification pass instruction is generated.

[0015] Optionally, before receiving the raw scoring data input by the target listener based on the calibrated test signal, the basic sound quality attribute dimension, and the spatial perception attribute dimension, the method further includes: Based on the preset requirements for sound source localization accuracy, single-point sound source fixed-point emission materials are determined; the sound source azimuth distribution corresponding to the single-point sound source fixed-point emission materials covers the front, side, rear and elevation directions. Based on the preset requirements for the continuity of sound source motion, audio material with a continuous motion trajectory is determined, including a continuous motion trajectory; the continuous motion trajectory includes sweeping motion from left to right and surround motion; the motion speed corresponding to the continuous motion trajectory satisfies the preset motion speed condition; The uniform motion scene material is determined based on the preset sound source motion uniformity requirements; the uniform motion scene material does not include natural acceleration and deceleration content. Based on the preset requirements for spatial immersion, multi-channel materials including spatial information are determined; the multi-channel materials include 3D sound movie audio and symphonic music. Training materials are constructed based on the single-point sound source fixed-point sound material, the continuous motion trajectory audio material, the uniform motion scene material, and the multi-channel material; the training materials include examples of binaural rendering spatial perception defects; the examples of binaural rendering spatial perception defects include positioning deviations of various degrees, motion discontinuity, speed unevenness, and loss of immersion; the number and duration of the training materials meet the material conditions corresponding to the subjective evaluation method for minor defects in the audio system.

[0016] Optionally, the raw scoring data input by the target listener based on the calibrated test signal, the basic sound quality attribute dimension, and the spatial perception attribute dimension includes: After obtaining the verification pass instruction, the calibrated test signals are played sequentially to the target listener, and a two-dimensional scoring interface is generated based on the basic sound quality attribute dimension and the spatial perception attribute dimension. A two-dimensional scoring interface is provided to the target listener so that the target listener can generate raw scoring data based on the calibrated test signal in the two-dimensional scoring interface.

[0017] Optionally, the basic sound quality attributes include low-frequency performance, clarity, dynamic range, timbre naturalness, noise, background noise, and vocal quality; the spatial perception attributes include sound source localization accuracy, motion continuity, motion uniformity, and spatial immersion; the sound source localization accuracy includes horizontal positioning accuracy, elevation angle positioning accuracy, distance perception accuracy, and front-back discrimination accuracy; the spatial immersion includes surround feeling, spatial size perception, degree of sound externalization, and degree of head-in-the-head effect.

[0018] Optionally, the process of generating raw rating data includes: Determine the preset grading scale corresponding to each sub-item in the dual-dimensional scoring interface, so that the target listener can generate a scoring value corresponding to each sub-item based on the value corresponding to the preset grading scale. Based on the aforementioned score values, generate original score data corresponding to the binaural rendered audio material to be tested.

[0019] Optionally, generating the original score data corresponding to the binaural rendered audio material to be tested based on each of the score values ​​includes: If any of the sub-items has a defect, a defect marker is generated based on the defect, and a score value is generated based on the defect marker. Then, based on each score value, unprocessed score data corresponding to the binaural rendered audio material to be tested is generated. The defect marker includes defect type information. A number of unprocessed score data generated by the same target listener in repeated tests of binaural rendered audio material to be tested are determined, and a score variance is generated based on each of the unprocessed score data. The rating variance is compared with a preset variance threshold, and the rating data to be processed corresponding to the rating variances that are greater than the preset variance threshold in the comparison results are removed to obtain an effective rating dataset.

[0020] Optionally, verifying the original scoring data to obtain a valid scoring dataset includes: Extract the reference score data of the target listener for each reference material, and determine the deviation between each reference score data and the preset hidden reference standard score. Then compare the deviation with the preset reference deviation threshold to remove the original score data of the target listener corresponding to the deviation greater than the preset reference deviation threshold, so as to obtain the score dataset to be processed. The scoring data of each target listener are subjected to group statistical analysis, and outliers that deviate from the group statistical distribution in the group statistical analysis results are identified. The outliers are then removed from the scoring dataset to be processed to obtain a valid scoring dataset.

[0021] Optionally, before generating the quality test results including sound quality score, spatial awareness score, and defect distribution based on the effective scoring dataset, the method further includes: The scores of each sub-item corresponding to the basic audio quality attribute dimension in the effective scoring dataset are weighted to obtain the audio quality score; The scores of each sub-item corresponding to the spatial awareness attribute dimension in the effective scoring dataset are weighted to obtain the spatial awareness score.

[0022] Optionally, generating quality test results based on the effective scoring dataset, including sound quality score, spatial awareness score, and defect distribution, includes: A radar chart is generated based on the sound quality score and the spatial perception score; the radar chart is used to display the performance profile of each attribute dimension. The defect labels marked by each target listener in the effective scoring dataset are statistically analyzed, and defect distribution information is generated according to the defect type and frequency of occurrence in the statistical results. Quality test results are generated based on the sound quality score, spatial perception score, radar chart, defect distribution information, and preset comprehensive rating classification standards.

[0023] Optionally, the process of receiving the raw scoring data generated by the listener includes: Control the switching of test dimensions. When switching from the basic sound quality attribute dimension to the spatial perception attribute dimension or vice versa, a preset rest interval is forcibly inserted, and the total test time for a single day is accumulated. When the total daily duration exceeds the preset maximum duration, the test reception for that day will be automatically terminated.

[0024] Secondly, this application provides a spatial perception quality testing device for binaural rendering, comprising: The test sequence determination module is used to determine the corresponding test paradigm based on the binaural rendered audio material to be tested and the test type, and to determine the test sequence based on the test paradigm. The scoring data generation module is used to calibrate the test sequence, obtain the calibrated test signal, and receive the original scoring data input by the target listener based on the calibrated test signal, the basic sound quality attribute dimension, and the spatial perception attribute dimension; the target listener is a listener who has passed the spatial perception specialized training test; The quality test result generation module is used to verify the original scoring data to obtain a valid scoring dataset, and to generate quality test results including sound quality score, spatial perception score and defect distribution based on the valid scoring dataset.

[0025] Thirdly, this application provides an electronic device, comprising: Memory, used to store computer programs; A processor is used to execute the computer program to implement the aforementioned binaural rendering spatial perception quality testing method.

[0026] Fourthly, this application provides a computer-readable storage medium for storing a computer program, wherein the computer program, when executed by a processor, implements the aforementioned binaural rendering spatial perception quality testing method.

[0027] As can be seen from the above, before conducting the spatial perception quality test of binaural rendering, this application needs to determine the corresponding test paradigm based on the binaural rendering audio material to be tested and the test type, and determine the test sequence based on the test paradigm; calibrate the test sequence to obtain the calibrated test signal, and receive the original scoring data input by the target listener based on the calibrated test signal, the basic sound quality attribute dimension and the spatial perception attribute dimension; verify the original scoring data to obtain a valid scoring dataset, and generate a quality test result including sound quality score, spatial perception score and defect distribution based on the valid scoring dataset.

[0028] Therefore, this application first needs to determine the corresponding test paradigm based on the binaural rendered audio material to be tested and the test type, and then determine the test sequence based on the test paradigm; secondly, the test sequence is calibrated to obtain a calibrated test signal, and the original scoring data input by the target listener based on the calibrated test signal, the basic sound quality attribute dimension, and the spatial perception attribute dimension is received; finally, the original scoring data is verified to obtain a valid scoring dataset, and a quality test result including sound quality score, spatial perception score, and defect distribution is generated based on the valid scoring dataset. In this way, the efficiency of spatial perception quality testing of binaural rendered audio materials is improved during the process of binaural rendered spatial perception quality testing, thereby enhancing the user experience. Attached Figure Description

[0029] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0030] Figure 1 This application discloses a flowchart of a spatial perception quality testing method for binaural rendering. Figure 2 This is a schematic diagram of a spatial perception quality testing device for binaural rendering disclosed in this application. Figure 3 This is a structural diagram of an electronic device disclosed in this application. Detailed Implementation

[0031] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0032] The current binaural rendering technology industry generally suffers from the following problems: existing subjective evaluation systems only focus on basic sound quality attributes such as low frequency, clarity, dynamic range, timbre, noise, and vocal quality; furthermore, there is no dedicated testing system for spatial perception in binaural rendering. Therefore, this application provides a spatial perception quality testing method for binaural rendering, which improves the efficiency of spatial perception quality testing for binaural rendered audio materials.

[0033] See Figure 1 As shown, this embodiment of the invention discloses a spatial perception quality testing method for binaural rendering, including: Step S11: Determine the corresponding test paradigm based on the binaural rendering audio material to be tested and the test type, and determine the test sequence based on the test paradigm.

[0034] In this embodiment, the application selects a binaural rendering spatial perception subjective quality testing system to divide the evaluation dimensions into two categories: basic sound quality attributes and spatial perception attributes. Subsequently, the system works collaboratively through modules such as acoustic playback calibration, dual-dimensional scoring interaction, listener management, automated process control, test signal scheduling, data reliability verification, and comprehensive evaluation report: it executes a five-stage process of pre-screening → training → pre-testing → formal testing → post-screening according to standard requirements; then, it automatically matches the ITU R BS.1116 / BS.1534 / BS.1284 test methods according to the dimensions; and it automatically inserts hidden references and spatial anchor points to stabilize the scoring scale. Then, it simultaneously completes the scoring of six sound quality items and four spatial perception dimensions; finally, it automatically removes abnormal data and outputs the comprehensive evaluation results.

[0035] It is worth mentioning that the system modules include: a binaural rendering acoustic playback and calibration module; a dual-dimensional scoring and interaction module for sound quality and spatial perception; a module for selecting and training listeners; a module for automated control of the testing process; a module for scheduling test materials and managing reference / anchor points; a module for verifying data reliability and post-screening; a module for comprehensive quality assessment and report generation; and a binaural rendering acoustic playback and calibration module.

[0036] That is, the embodiments of this application support the playback of HRTF / BRIR binaural rendering signals, reference signals, and spatial anchor signals; subsequently, time delay, phase, and sound level calibration are performed to ensure the true reproduction of spatial perception cues. At the same time, the embodiments of this application are compatible with ITU standard test paradigms such as three-excitation, MUSHRA, and multiple comparison methods.

[0037] Specifically, determining the corresponding test paradigm based on the binaural rendering audio material to be tested and the test type may include: determining the test type corresponding to the binaural rendering audio material to be tested; the test type includes positioning test, motion test, and immersion and sound quality test; obtaining reference material; the reference material is the original audio with the same content as the binaural rendering audio material to be tested; the reference material is used to measure the reliability of the target listener's scoring; obtaining spatial anchor point material; the spatial perception quality level corresponding to the spatial anchor point material meets the preset stability conditions, and the spatial anchor point material covers each preset segment region in the preset grading scale; the preset segment regions include It includes a preset low-score segment region that meets preset low-score conditions, a preset middle-score segment region that meets preset middle-score conditions, and a preset high-score segment region that meets preset high-score conditions; spatial anchor point materials are used to distinguish the unacceptable and acceptable quality boundaries of the binaural rendered audio material to be tested; the corresponding test paradigm is determined using the material management unit based on the binaural rendered audio material to be tested, reference materials, spatial anchor point materials, and test type; the test paradigm includes a first test paradigm, a second test paradigm, and a third test paradigm; the first test paradigm is a three-excitation paradigm; the second test paradigm is a multi-excitation hidden reference and anchor point paradigm; the third test paradigm is a multiple comparison method paradigm.

[0038] It is worth mentioning that the material selection principle in this application embodiment is as follows: in order to accurately evaluate each spatial perception dimension, the materials selected in this application embodiment should be able to fully stimulate the perceptual cues of the corresponding dimension, so that the listener can stably judge the differences between the tested systems. The quantity and duration of the materials should comply with the requirements of ITU-R BS.1116.

[0039] The key points for selecting materials in each dimension are as follows: Accuracy of sound source localization: Select materials with single-point sound sources and fixed-point sound emission. The distribution of sound sources should cover the front, side, rear, and elevation directions; Continuity of sound source movement: Select audio with continuous movement trajectories, such as sweeping from left to right, surround motion, etc., and the movement speed should be moderate; Uniformity of sound source movement: Select materials with preset uniform motion scenes and avoid using content with natural acceleration and deceleration; Spatial immersion: Select multi-channel content with rich spatial information, such as 7.1.4 channel 3D sound movie audio, symphony, etc., to stimulate differences in spatial immersion.

[0040] For binaural rendering spatial awareness testing, the methods for acquiring the reference signal should include: Method A (Ideal Reference): Binaural recordings are made in an anechoic chamber using HATS or by a person wearing a miniature microphone, serving as a reference for rendering the ideal space. HATS recordings accurately capture the HRTF characteristics of real space and are an ideal benchmark for evaluating the performance of a rendering system.

[0041] Method B (Rendering Reference): Uses a signal generated by a known high-quality binaural rendering system as a comparison benchmark. This method is suitable for situations where HATS recordings are unavailable, but it must be ensured that the rendering quality of the selected reference system truly meets the "high-quality" standard.

[0042] Method C (Direct Feed): The raw signal is fed directly to the headphones without binaural rendering (used to assess spatial impairments introduced by rendering, such as the in-head effect). This method helps listeners identify specific perceptual defects introduced by the rendering system.

[0043] It is worth mentioning that when assessing the spatial perception dimension, it is recommended to use method A as a reference first, so as to ensure that the reference signal has real and natural spatial perception characteristics and provides a reliable scoring benchmark for listeners.

[0044] Specifically, obtaining reference materials may include: obtaining reference materials; wherein, the reference materials include any one of the first reference materials, the second reference materials, and the third reference materials; wherein, the first reference material is the reference material of ideal spatial rendering obtained by binaural recording in an anechoic chamber using HATS or by a real person wearing a miniature microphone; the second reference material is the material generated using a preset high-quality binaural rendering system; the rendering quality of the preset high-quality binaural rendering system meets the preset high-quality conditions; the third reference material is the material generated by directly feeding the original signal to the headphones without binaural rendering.

[0045] Furthermore, the embodiments of this application require determining the test sequence based on the test paradigm: determining the corresponding signal combination method based on the test paradigm, randomly sorting the audio materials to be tested for binaural rendering to obtain sorted materials; and performing insertion operations on reference materials and spatial anchor points and sorting combinations on the sorted materials based on the signal combination method, reference materials, and spatial anchor point materials to obtain the test sequence. Specifically, the positioning test corresponds to the three-excitation paradigm; the motion test corresponds to the multi-excitation hidden reference and anchor point paradigm; and the immersion and sound quality test corresponds to the multiple comparison method paradigm.

[0046] In this embodiment, the material scheduling and reference / anchor point management module is used to locate: BS.1116 corresponds to the three excitation paradigms, motion corresponds to MUSHRA; immersion / sound quality corresponds to the multiple comparison method, and performs automatic random sorting, hiding reference and spatial anchor point material insertion.

[0047] Specifically, inserting reference material into the binaural rendering audio material to be tested based on reference material can include: determining the test requirements corresponding to the test paradigm, inserting reference material into the first preset position of the sorted material based on the test requirements, and obtaining the inserted material.

[0048] The anchor point selection principle is as follows: to stabilize the scoring scale and ensure the comparability of scores across listeners and laboratories, appropriate anchor point conditions should be set in the test of each evaluation dimension. In one specific implementation, the anchor point may have a known and stable spatial perception quality level and be able to cover each segment area of ​​the scoring scale based on needs (usually used to distinguish the quality boundary between "unacceptable" and "acceptable"), without being specifically limited here.

[0049] Further recommendations for anchor point selection across various dimensions: Sound source localization accuracy: Use low-quality rendering output with significant localization bias or confusion, or apply quantized spatial perturbations to the reference signal. Sound source motion continuity: Use low-quality rendering output with periodic jumps or breaks inserted into continuous motion trajectories. Sound source motion uniformity: Use rendering output with random speed fluctuations applied to uniform motion, or azimuth interpolation rendering with a low frame rate. Spatial immersion: Use a mono submixed version as the lowest anchor point for loss of spatial immersion. Unrendered stereo signals can be used as anchor points for externalized loss.

[0050] All anchor signals should have their attributes clearly stated in the test instructions and explanatory materials, but they should not be labeled as anchors in the test interface. It is recommended to set 1-2 anchor signals for each evaluation dimension.

[0051] Specifically, spatial anchoring based on spatial anchor materials for the binaural rendered audio material to be tested can include: inserting spatial anchor materials at a second preset position in the inserted material based on test requirements; setting spatial anchor materials for the sound source localization accuracy dimension; spatial anchor materials are rendering output results with localization deviation and confusion, or spatial anchor materials are materials obtained after applying quantized spatial perturbation to reference materials; setting spatial anchor materials for the sound source motion continuity dimension; spatial anchor materials are rendering output results with periodic jumps or breaks inserted in a continuous motion trajectory; setting spatial anchor materials for the sound source motion uniformity dimension; spatial anchor materials are rendering output results with random speed fluctuations applied in uniform motion, or spatial anchor materials are output results with interpolation rendering of azimuth angles that meet preset low frame rate conditions; setting spatial anchor materials for the spatial immersion dimension; spatial anchor materials are the lowest score anchor points for spatial immersion loss using the mono downmixed version and the externalized loss anchor points determined by the stereo signal without binaural rendering; wherein, the number of anchor signals corresponding to each evaluation dimension is a preset number.

[0052] Step S12: Calibrate the test sequence to obtain a calibrated test signal, and receive the original scoring data input by the target listener based on the calibrated test signal, the basic sound quality attribute dimension, and the spatial perception attribute dimension; the target listener is a listener who has passed the spatial perception specialized training test.

[0053] In this embodiment, the test sequence needs to be calibrated to obtain calibrated test signals: Time delay calibration is performed on each of the binaural rendered audio materials, reference materials, and spatial anchor point materials in the test sequence to obtain time delay calibrated test signals; Phase calibration is performed on each of the binaural rendered audio materials, reference materials, and spatial anchor point materials in the test sequence to obtain phase calibrated test signals; The phase response of each phase calibrated test signal is consistent at the target listener's ears; Sound level calibration is performed on each of the binaural rendered audio materials, reference materials, and spatial anchor point materials in the test sequence to obtain sound level calibrated test signals; The perceived loudness of the sound level calibrated test signals at the target listener's ears is not less than a preset standard sound pressure level; The calibrated test signal is determined based on the time delay calibrated test signal, the phase calibrated test signal, and the sound level calibrated test signal.

[0054] Furthermore, this application embodiment requires auditor screening: auditors should pass a standardized hearing screening, and their hearing threshold should not exceed 20 dBHL in the frequency range of 125 Hz to 8 kHz. Specifically, tests using the ITU-R BS.1116 method should be conducted using expert auditors with experience in diagnosing audio impairments and identifying spatial perception. The MUSHRA test can use trained auditors, but it should be ensured that they are familiar with the perceptual characteristics of the spatial perception dimension.

[0055] The required number of listeners for this embodiment is as follows: ITU-R BS.1116 Method: No fewer than 20 expert listeners; MUSHRA method: No fewer than 15 trained listeners; Furthermore, the listeners in the same batch of tests all had the same spatial perception training background.

[0056] Specifically, before receiving the raw scoring data input by the target listener based on the calibrated test signal, the basic sound quality attribute dimension, and the spatial perception attribute dimension, it may also include: determining the test method corresponding to the quality test, if the test method is ITU-R. The BS.1116 method determines a first preset number of first expert listeners with experience in audio impairment judgment and spatial perception discrimination; the first preset number is greater than a first preset threshold. If the testing method is the MUSHRA method, a second preset number of second expert listeners with experience in spatial perception discrimination is determined; the second preset number is greater than a second preset threshold; the first preset threshold is greater than the second preset threshold; the spatial perception training backgrounds of the first and second expert listeners are consistent; based on each first and second expert listener, listeners to be screened are determined, and hearing threshold screening is performed on the listeners to be screened based on a preset frequency range to obtain the frequency hearing thresholds corresponding to each test frequency point for each listener to be screened, and it is determined whether the hearing thresholds at each frequency point are all greater than a preset decibel value; if the hearing thresholds at each frequency point are not greater than the preset decibel value, the listener to be screened is determined to have passed the hearing screening; if the hearing thresholds at each frequency point are all less than the preset decibel value, the listener to be screened is determined to have failed the hearing screening, and the testing process for the listener to be screened is terminated.

[0057] Furthermore, this application embodiment requires spatial perception training for listeners, that is, using the listener screening and specialized training module to screen hearing thresholds of 125Hz–8kHz (≤20 dBHL); spatial perception specialized training: positioning deviation, motion defects, head-in-the-head effect, and front-back confusion; training duration ≥2 hours with automatic verification.

[0058] That is, the listeners in this application embodiment should receive specialized training in the spatial perception dimensions of this standard before formal testing, including: sound source localization accuracy dimension training: identifying azimuth deviation, elevation deviation, distance perception error and confusion between front and back, and understanding the perceptual characteristics of the in-head effect; sound source motion continuity dimension training: used to distinguish the continuity and discontinuity of motion trajectory, and identify perceptual defects such as trajectory jumps and jitters; sound source motion uniformity dimension training: 3. perceiving speed fluctuations in uniform motion, and identifying speed unevenness phenomena such as sudden speed changes; spatial immersion dimension training: experiencing different degrees of sound externalization and spatial immersion, and understanding the perceptual characteristics of different space types (such as living rooms, bathrooms, and concert halls).

[0059] It is worth noting that the training materials should cover typical examples of binaural rendering spatial perception defects, including varying degrees of localization deviation, discontinuous motion, uneven speed, and loss of immersion, to help listeners establish stable scoring benchmarks. The training time should be no less than 2 hours and should be completed within one week before the formal test.

[0060] Specifically, the training process for listeners includes: training listeners on sound source localization accuracy using training materials to obtain a first-stage listener; sound source localization accuracy training includes identifying azimuth deviation, elevation deviation, distance perception error, and front-back confusion, and determining the perceptual characteristics of the mid-head effect; training the first-stage listener on sound source motion continuity to obtain a second-stage listener; sound source motion continuity training includes distinguishing the continuity and discontinuity of motion trajectories, and identifying perceptual defects including trajectory jumps and jitters; and training the second-stage listener on sound source motion uniformity to obtain a third-stage listener. The training includes: 1) Training on the uniformity of sound source motion, which involves perceiving speed fluctuations in uniform motion and recognizing speed non-uniformity; 2) Training on spatial immersion, which involves the third training session, to obtain the target listener; 3) Training on spatial immersion, which includes experiencing the differences in the degree of sound externalization and spatial envelopment, and understanding the perceptual characteristics of various spatial types; 4) The training time for the listener is not less than a preset training threshold; 5) Recording the cumulative training time of the target listener in spatial perception training and comparing the cumulative training time with the preset training time threshold; 6) If the comparison result indicates that the cumulative training time is not less than the preset training time threshold, a verification pass instruction is generated.

[0061] Furthermore, before receiving the raw scoring data input by the target listener based on the calibrated test signal, the basic sound quality attribute dimension, and the spatial perception attribute dimension, the process may further include: determining single-point sound source fixed-point emission material based on preset sound source localization accuracy requirements; the sound source azimuth distribution corresponding to the single-point sound source fixed-point emission material covering the front, side, rear, and elevation directions; determining continuous motion trajectory audio material including continuous motion trajectories based on preset sound source motion continuity requirements; the continuous motion trajectory includes sweeping and circling motion from left to right; the motion speed corresponding to the continuous motion trajectory meets preset motion speed conditions; and determining uniform speed based on preset sound source motion uniformity requirements. Motion scene materials; uniform motion scene materials excluding natural acceleration and deceleration content; multi-channel materials including spatial information are determined based on preset spatial immersion requirements; multi-channel materials include 3D sound movie audio and symphony; training materials are constructed based on single-point sound source fixed-point sound materials, continuous motion trajectory audio materials, uniform motion scene materials and multi-channel materials; training materials include examples of spatial perception defects in binaural rendering; examples of spatial perception defects in binaural rendering include various degrees of positioning deviation, motion discontinuity, speed unevenness and loss of immersion; the number and duration of training materials meet the material conditions corresponding to the subjective evaluation method for minor damage in the audio system.

[0062] In this embodiment, the present application requires the use of a two-dimensional scoring interaction module. The first dimension is the basic sound quality attribute dimension: low-frequency performance, clarity, dynamic range, timbre naturalness, noise and background noise, and vocal quality. The second dimension is the spatial perception attribute dimension: sound source localization accuracy (horizontal / elevation angle / distance / foreground / background discrimination); sound source motion continuity (smooth trajectory, no jumps or breaks); sound source motion uniformity (uniform and stable speed, no sudden changes in speed); spatial immersion (sense of envelopment, spatial size, externalization, head-in-the-head effect). The two-dimensional scoring interaction module supports synchronous scoring, defect marking, and scale-based grading display.

[0063] Specifically, receiving raw scoring data from the target listener based on calibrated test signals, fundamental sound quality attributes, and spatial perception attributes can include: after obtaining a verification pass instruction, sequentially playing each calibrated test signal to the target listener and generating a two-dimensional scoring interface based on the fundamental sound quality attributes and spatial perception attributes; providing the target listener with the two-dimensional scoring interface so that they can generate raw scoring data based on the calibrated test signals within the two-dimensional scoring interface. The fundamental sound quality attributes include low-frequency performance, clarity, dynamic range, timbre naturalness, noise, background noise, and vocal quality; the spatial perception attributes include sound source localization accuracy, motion continuity, motion uniformity, and spatial immersion; sound source localization accuracy includes horizontal localization accuracy, elevation localization accuracy, distance perception accuracy, and front-back discrimination accuracy; spatial immersion includes surround feeling, spatial size perception, degree of sound externalization, and degree of in-head effect.

[0064] The process of generating raw scoring data includes: determining the preset grading scales corresponding to each sub-item in the dual-dimensional scoring interface, so that the target listener can generate scoring values ​​corresponding to each sub-item based on the values ​​corresponding to the preset grading scales; and generating raw scoring data corresponding to the binaural rendered audio material to be tested based on each scoring value.

[0065] Furthermore, in this embodiment of the application, a data filtering module is required to remove data with excessive variance, large deviation from the hidden reference, or abnormal population deviation.

[0066] Specifically, generating raw score data corresponding to the binaural rendered audio material to be tested based on each score value can include: if a sub-item has a defect, generating a defect marker based on the defect, generating a score value based on the defect marker, and then generating unprocessed score data corresponding to the binaural rendered audio material to be tested based on each score value; the defect marker includes defect type information; identifying several unprocessed score data generated by the same target listener in repeated tests of the binaural rendered audio material to be tested, and generating a score variance based on each unprocessed score data; comparing the score variance with a preset variance threshold, and removing the unprocessed score data corresponding to the score variance greater than the preset variance threshold in the comparison results to obtain a valid score dataset.

[0067] Step S13: Verify the original scoring data to obtain a valid scoring dataset, and generate quality test results including sound quality score, spatial perception score and defect distribution based on the valid scoring dataset.

[0068] In this embodiment, the requirement to verify the original scoring data to obtain a valid scoring dataset may include: extracting reference scoring data of the target listeners for each reference material, determining the deviation between each reference scoring data and a preset hidden reference standard score, then comparing the deviation with a preset reference deviation threshold to remove the original scoring data of the target listeners corresponding to deviations greater than the preset reference deviation threshold, thus obtaining a scoring dataset to be processed; performing group statistical analysis on the scoring data of each target listener, identifying outliers in the group statistical analysis results that deviate from the group statistical distribution, and removing outliers from the scoring dataset to be processed, thus obtaining a valid scoring dataset.

[0069] Before generating quality test results including sound quality score, spatial perception score and defect distribution based on the effective scoring dataset, the process may further include: weighting the scores of each sub-item corresponding to the basic sound quality attribute dimension in the effective scoring dataset to obtain the sound quality score; and weighting the scores of each sub-item corresponding to the spatial perception attribute dimension in the effective scoring dataset to obtain the spatial perception score.

[0070] In this embodiment, the present application requires the use of a report generation module to output sound quality score, spatial perception score, radar chart, defect distribution, and comprehensive rating.

[0071] Specifically, generating quality test results based on the effective scoring dataset, including sound quality score, spatial perception score, and defect distribution, may include: drawing a radar chart based on the sound quality score and spatial perception score; the radar chart is used to display the performance profile of each attribute dimension; statistically analyzing the defect markers marked by each target listener in the effective scoring dataset, and generating defect distribution information according to the defect type and frequency of occurrence in the statistical results; and generating quality test results based on the sound quality score, spatial perception score, radar chart, defect distribution information, and preset comprehensive rating classification standards.

[0072] It is worth mentioning that the embodiments of this application utilize the test process automation control module to fully automatically execute the five-stage standard process; wherein, in one specific implementation, the rest between dimensions is 5–10 minutes, and the rest time per day is ≤45 minutes.

[0073] Specifically, the process of receiving raw scoring data generated by listeners includes: controlling the switching of test dimensions; when switching from the basic sound quality attribute dimension to the spatial perception attribute dimension or vice versa, a preset rest interval is forcibly inserted, and the total test duration for the day is accumulated; when the total test duration for the day exceeds the preset duration limit, the test reception for that day is automatically terminated.

[0074] It is worth mentioning that the highlights of the technical solution in this application embodiment are: a pioneering dual-dimensional fusion testing system integrating basic sound quality attributes and spatial perception attributes, enabling comprehensive subjective quality assessment within the same system, which is more in line with actual binaural rendering usage scenarios. The system incorporates a quantitative scoring module for four dimensions of spatial perception, transforming sound source localization, motion continuity, motion uniformity, and spatial immersion into scoreable and statistically verifiable system functions. Furthermore, the system possesses automatic matching and scheduling capabilities for ITU standard methods, automatically executing corresponding test logic according to different dimensions without manual configuration or switching. Further, the system achieves five-stage fully automated closed-loop control, from listener selection, training, and testing to data quality control and report generation, all running automatically. The system is equipped with a hidden reference and spatial anchor intelligent scheduling module, which can automatically and randomly insert and complete scoring calibration, improving the consistency and reliability of test results.

[0075] As can be seen from the above, the embodiments of this application first need to determine the corresponding test paradigm based on the binaural rendered audio material to be tested and the test type, and then determine the test sequence based on the test paradigm; secondly, the test sequence is calibrated to obtain the calibrated test signal, and the original scoring data input by the target listener based on the calibrated test signal, the basic sound quality attribute dimension, and the spatial perception attribute dimension is received; finally, the original scoring data is verified to obtain a valid scoring dataset, and a quality test result including sound quality score, spatial perception score, and defect distribution is generated based on the valid scoring dataset. In this way, the efficiency of spatial perception quality testing of binaural rendered audio material is improved during the process of binaural rendered spatial perception quality testing, thereby enhancing the user experience.

[0076] Accordingly, see Figure 2 As shown, this application also provides a spatial perception quality testing device for binaural rendering, comprising: The test sequence determination module 11 is used to determine the corresponding test paradigm based on the binaural rendered audio material to be tested and the test type, and to determine the test sequence based on the test paradigm. The scoring data generation module 12 is used to calibrate the test sequence, obtain the calibrated test signal, and receive the original scoring data input by the target listener based on the calibrated test signal, the basic sound quality attribute dimension, and the spatial perception attribute dimension; the target listener is a listener who has passed the spatial perception specialized training test; The quality test result generation module 13 is used to verify the original scoring data to obtain a valid scoring dataset, and to generate quality test results including sound quality score, spatial perception score and defect distribution based on the valid scoring dataset.

[0077] In some specific embodiments, the test sequence determination module 11 may specifically include: The test type determination unit is used to determine the test type corresponding to the binaural rendered audio material to be tested; the test types include positioning test, motion test, and immersion and sound quality test. The first reference material acquisition unit is used to acquire reference materials; the reference materials are original audio with the same content as the binaural rendering audio material to be tested; the reference materials are used to measure the reliability of the target listener's score; A spatial anchor point material acquisition unit is used to acquire spatial anchor point materials; the spatial perception quality level corresponding to the spatial anchor point materials meets preset stability conditions, and the spatial anchor point materials cover each preset segment region in a preset grading scale; the spatial anchor point materials are used to distinguish the unacceptable and acceptable quality boundaries of the binaural rendering audio material to be tested; the preset segment regions include a preset low-score segment region that meets preset low-score conditions, a preset medium-score segment region that meets preset medium-score conditions, and a preset high-score segment region that meets preset high-score conditions; The test paradigm determination unit is used to determine the corresponding test paradigm based on the material management unit and the binaural rendered audio material to be tested, the reference material, the spatial anchor point material, and the test type; the test paradigm includes a first test paradigm, a second test paradigm, and a third test paradigm; the first test paradigm is a three-excitation paradigm; the second test paradigm is a multi-excitation hidden reference and anchor point paradigm; and the third test paradigm is a multiple comparison method paradigm.

[0078] In some specific embodiments, the test sequence determination module 11 may specifically include: The second reference material acquisition unit is used to acquire reference materials; wherein, the reference materials include any one of the first reference materials, the second reference materials, and the third reference materials; wherein, the first reference material is reference material for ideal spatial rendering obtained by binaural recording in an anechoic chamber using HATS or by a real person wearing a miniature microphone; the second reference material is material generated using a preset high-quality binaural rendering system; the rendering quality of the preset high-quality binaural rendering system meets preset high-quality conditions; the third reference material is material generated by directly feeding the original signal to the headphones without binaural rendering.

[0079] In some specific embodiments, the test sequence determination module 11 may specifically include: The signal combination method determination unit is used to determine the corresponding signal combination method based on the test paradigm, and to randomly sort the binaural rendered audio material to be tested to obtain the sorted material. The test sequence determination subunit is used to perform insertion operations and sorting combinations of reference materials and spatial anchors on the sorted materials based on the signal combination method, the reference materials and the spatial anchor materials, to obtain the test sequence.

[0080] In some specific embodiments, the test sequence determination module 11 may specifically include: The test requirement determination unit is used to determine the test requirements corresponding to the test paradigm, and to insert the reference material into the first preset position of the sorted material based on the test requirements to obtain the inserted material.

[0081] In some specific embodiments, the test sequence determination module 11 may specifically include: A spatial anchor point material insertion unit is used to insert the spatial anchor point material into the second preset position in the inserted material based on the test requirements; The first spatial anchor point material setting unit is used to set the spatial anchor point material for the sound source localization accuracy dimension; the spatial anchor point material is a rendering output result with localization deviation and front-back confusion, or the spatial anchor point material is a material obtained by applying quantized spatial perturbation to the reference material; The second spatial anchor point material setting unit is used to set the spatial anchor point material for the continuity dimension of the sound source motion; the spatial anchor point material is the rendering output result of inserting periodic jumps or breaks in the continuous motion trajectory; The third spatial anchor point material setting unit is used to set the spatial anchor point material for the uniformity dimension of sound source motion; the spatial anchor point material is the rendering output result of applying random speed fluctuations in uniform motion, or the spatial anchor point material is the output result of interpolating the azimuth angle that meets the preset low frame rate condition. The fourth spatial anchor point material setting unit is used to set the spatial anchor point material for the spatial immersion dimension; the spatial anchor point material is the lowest score anchor point of spatial immersion loss determined by the mono downmix version and the externalized loss anchor point material determined by the stereo signal without binaural rendering; wherein, the number of anchor point signals corresponding to each evaluation dimension is a preset number.

[0082] In some specific embodiments, the scoring data generation module 12 may specifically include: The material delay calibration unit is used to perform delay calibration on each of the binaural rendered audio materials to be tested, the reference material and the spatial anchor point material in the test sequence to obtain the delay-calibrated test signal; The material phase calibration unit is used to perform phase calibration on each of the test-to-be-tested binaural rendered audio materials, the reference material, and the spatial anchor point material in the test sequence to obtain a phase-calibrated test signal; the phase response of each of the phase-calibrated test signals is consistent at the binaural ears of the target listener; The material sound level calibration unit is used to perform sound level calibration on each of the binaural rendered audio materials to be tested, the reference materials, and the spatial anchor point materials in the test sequence to obtain a sound level calibrated test signal; the perceived loudness of the sound level calibrated test signal at the binaural ears of the target listener is not less than a preset standard sound pressure level; The calibration post-test signal determination unit is used to determine the calibration post-test signal based on the time delay calibration post-test signal, the phase calibration post-test signal, and the sound level calibration post-test signal.

[0083] In some specific embodiments, the binaural rendering spatial perception quality testing device may further include: The test method determination unit is used to determine the test method corresponding to the quality test. If the test method is the ITU-RBS.1116 method, then a first preset number of first expert listeners with experience in audio damage judgment and spatial perception identification are determined; the first preset number is greater than a first preset threshold. An expert listener determination unit is used to determine a second preset number of second expert listeners with spatial perception discrimination experience if the testing method is the MUSHRA method; the second preset number is greater than a second preset threshold; the first preset threshold is greater than the second preset threshold; and the spatial perception training backgrounds of the first expert listeners and the second expert listeners are consistent. The listener to be screened unit is used to determine the listeners to be screened based on each of the first expert listeners and each of the second expert listeners, and to screen the listeners to be screened for hearing thresholds based on a preset frequency range, to obtain the frequency hearing thresholds corresponding to each test frequency point of the listeners to be screened, and to determine whether each frequency hearing threshold is greater than a preset decibel value. The listener selection unit is used to determine if the listener passes the hearing screening if the hearing threshold at each frequency point is not greater than the preset decibel value, and to determine if the listener fails the hearing screening if the hearing threshold at each frequency point is less than the preset decibel value, and to terminate the testing process of the listener.

[0084] In some specific embodiments, the scoring data generation module 12 may specifically include: The first listener training unit is used to train the listener on the accuracy dimension of sound source localization using training materials to obtain the first trained listener; the sound source localization accuracy dimension training includes identifying azimuth deviation, elevation deviation, distance perception error and front-back confusion phenomenon, and determining the perception characteristics of the head-in-the-head effect. The second listener training unit is used to train the first trained listener in the continuity dimension of sound source motion to obtain the second trained listener; the continuity dimension training of sound source motion includes distinguishing the continuity and discontinuity of motion trajectory and identifying perceptual defects including trajectory jumps and jitters. The third listener training unit is used to train the second trained listener on the uniformity of sound source motion to obtain the third trained listener; the uniformity of sound source motion training includes perceiving speed fluctuations in uniform motion and recognizing speed non-uniformity phenomena. The fourth listener training unit is used to train the listeners after the third training in the spatial immersion dimension to obtain the target listener; the spatial immersion dimension training includes experiencing the differences in the degree of externalization of sound and spatial envelopment, as well as understanding the perceptual characteristics of each spatial type; wherein, the training time for training the listeners is not less than a preset training threshold. The cumulative training time determination unit is used to record the cumulative training time of the target listener in the spatial perception-specific training, and compare the cumulative training time with a preset training time threshold. The verification pass instruction generation unit is used to generate a verification pass instruction if the comparison result indicates that the cumulative training time is not less than the preset training time threshold.

[0085] In some specific embodiments, the binaural rendering spatial perception quality testing device may further include: The single-point sound source fixed-point sound material determination unit is used to determine the single-point sound source fixed-point sound material based on the preset sound source positioning accuracy requirements; the sound source azimuth distribution corresponding to the single-point sound source fixed-point sound material covers the front, side, rear and elevation directions. The continuous motion trajectory audio material determination unit is used to determine continuous motion trajectory audio material, including continuous motion trajectory, based on preset sound source motion continuity requirements; the continuous motion trajectory includes sweeping motion from left to right and surround motion; the motion speed corresponding to the continuous motion trajectory satisfies preset motion speed conditions; The unit for uniform motion scene material is used to determine uniform motion scene material based on the preset sound source motion uniformity requirements; the uniform motion scene material does not include natural acceleration and deceleration content; A multi-channel material determination unit is used to determine multi-channel materials including spatial information based on preset spatial immersion requirements; the multi-channel materials include 3D sound movie audio and symphonic music. The training material construction unit is used to construct training materials based on the single-point sound source fixed-point sound material, the continuous motion trajectory audio material, the uniform motion scene material, and the multi-channel material; the training materials include examples of binaural rendering spatial perception defects; the examples of binaural rendering spatial perception defects include positioning deviations of various degrees, motion discontinuity, speed unevenness, and loss of immersion; the number and duration of the training materials meet the material conditions corresponding to the subjective evaluation method for minor defects in the audio system.

[0086] In some specific embodiments, the scoring data generation module 12 may specifically include: The dual-dimensional scoring interface generation unit is used to sequentially play each of the calibrated test signals to the target listener after obtaining the verification pass instruction, and generate a dual-dimensional scoring interface based on the sound quality basic attribute dimension and the spatial perception attribute dimension. The raw score data generation unit is used to provide the target listener with a two-dimensional score interface, so that the target listener can generate raw score data based on the calibrated test signal in the two-dimensional score interface.

[0087] In some specific embodiments, the scoring data generation module 12 may specifically include: A preset grading scale determination unit is used to determine the preset grading scale corresponding to each sub-item in the dual-dimensional scoring interface, so that the target listener can generate a scoring value corresponding to each sub-item based on the value corresponding to the preset grading scale. The scoring data generation subunit is used to generate original scoring data corresponding to the binaural rendered audio material to be tested based on each of the scoring values.

[0088] In some specific embodiments, the scoring data generation module 12 may specifically include: The scoring value generation unit is used to generate a defect marker based on the defect if the sub-item has a defect, generate a scoring value based on the defect marker, and then generate unprocessed scoring data corresponding to the binaural rendered audio material to be tested based on each scoring value; the defect marker includes defect type information; The scoring variance generation unit is used to determine several score data to be processed generated by the same target listener in repeated tests of the binaural rendered audio material to be tested, and to generate a scoring variance based on each score data to be processed. The scoring data elimination unit is used to compare the scoring variance with a preset variance threshold, and eliminate the scoring data to be processed corresponding to the scoring variance that is greater than the preset variance threshold in the comparison result, so as to obtain an effective scoring dataset.

[0089] In some specific embodiments, the quality test result generation module 13 may specifically include: The reference scoring data extraction unit is used to extract the reference scoring data of the target listener for each of the reference materials, and determine the deviation between each of the reference scoring data and the preset hidden reference standard score. Then, the deviation is compared with the preset reference deviation threshold to remove the original scoring data of the target listener corresponding to the deviation greater than the preset reference deviation threshold, so as to obtain the scoring dataset to be processed. The group statistical analysis unit is used to perform group statistical analysis on the scoring data of each target listener, identify outlier data that deviates from the group statistical distribution in the group statistical analysis results, and remove the outlier data from the scoring dataset to be processed to obtain a valid scoring dataset.

[0090] In some specific embodiments, the binaural rendering spatial perception quality testing device may further include: The sound quality score generation unit is used to weight the scores of each sub-item corresponding to the basic sound quality attribute dimension in the effective score dataset to obtain the sound quality score. The spatial perception score generation unit is used to weight the scores of each sub-item corresponding to the spatial perception attribute dimension in the effective score dataset to obtain the spatial perception score.

[0091] In some specific embodiments, the quality test result generation module 13 may specifically include: A radar chart drawing unit is used to draw a radar chart based on the sound quality score and the spatial perception score; the radar chart is used to display the performance profile of each attribute dimension. The defect labeling statistics unit is used to count the defect labels marked by each target listener in the effective scoring dataset, and generate defect distribution information according to the defect type and frequency of occurrence in the statistical results. The quality test result generation subunit is used to generate quality test results based on the sound quality score, the spatial perception score, the radar chart, the defect distribution information, and the preset comprehensive rating classification standard.

[0092] In some specific embodiments, the quality test result generation module 13 may specifically include: The total test duration accumulation unit is used to control the switching of test dimensions. When switching from the sound quality basic attribute dimension to the spatial perception attribute dimension or from the spatial perception attribute dimension to the sound quality basic attribute dimension, a preset rest interval is forcibly inserted, and the total test duration for a single day is accumulated. The test reception termination unit is used to automatically terminate the test reception for the day when the total duration of the day exceeds the preset duration limit.

[0093] Furthermore, embodiments of this application also disclose an electronic device, Figure 3 This is a structural diagram of an electronic device 20 according to an exemplary embodiment. The content of the diagram should not be construed as limiting the scope of this application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. The memory 22 stores a computer program, which is loaded and executed by the processor 21 to implement the relevant steps in the binaural rendering spatial perception quality testing method disclosed in any of the foregoing embodiments. Furthermore, the electronic device 20 in this embodiment may specifically be an electronic computer.

[0094] In this embodiment, the power supply 23 is used to provide operating voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows can be any communication protocol applicable to the technical solution of this application, and is not specifically limited here; the input / output interface 25 is used to acquire external input data or output data to the outside world, and its specific interface type can be selected according to specific application needs, and is not specifically limited here.

[0095] In addition, the memory 22, as a carrier for resource storage, can be a read-only memory, random access memory, disk or optical disk, etc. The resources stored thereon can include operating system 221, computer program 222, etc., and the storage method can be temporary storage or permanent storage.

[0096] The operating system 221 is used to manage and control the various hardware devices on the electronic device 20 and the computer program 222, which may be Windows Server, Netware, Unix, Linux, etc. In addition to including a computer program capable of performing the binaural rendering spatial perception quality testing method executed by the electronic device 20 as disclosed in any of the foregoing embodiments, the computer program 222 may further include computer programs capable of performing other specific tasks.

[0097] Furthermore, this application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the aforementioned binaural rendering spatial perception quality testing method. Specific steps of this method can be found in the corresponding content disclosed in the foregoing embodiments, and will not be repeated here.

[0098] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to in the method section.

[0099] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0100] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly by hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.

[0101] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0102] The technical solutions provided in this application have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A method for testing spatial perception quality using binaural rendering, characterized in that, include: The corresponding test paradigm is determined based on the binaural rendered audio material to be tested and the test type, and the test sequence is determined based on the test paradigm. The test sequence is calibrated to obtain a calibrated test signal, and the original scoring data input by the target listener based on the calibrated test signal, the basic sound quality attribute dimension, and the spatial perception attribute dimension is received. The target listener is a listener who has passed the spatial perception-specific training test; The original scoring data is examined to obtain a valid scoring dataset, and quality test results including sound quality score, spatial perception score and defect distribution are generated based on the valid scoring dataset.

2. The spatial perception quality testing method for binaural rendering according to claim 1, characterized in that, The method of determining the corresponding test paradigm based on the binaural rendered audio material to be tested and the test type includes: Determine the test type corresponding to the binaural rendered audio material to be tested; the test type includes positioning test, motion test, and immersion and sound quality test. Obtain reference material; the reference material is the original audio with the same content as the binaural rendering audio material to be tested; the reference material is used to measure the reliability of the target listener's score; Acquire spatial anchor point materials; the spatial perception quality level corresponding to the spatial anchor point materials meets preset stability conditions, and the spatial anchor point materials cover each preset segment region in a preset grading scale; the spatial anchor point materials are used to distinguish the unacceptable and acceptable quality boundaries of the binaural rendering audio materials to be tested; the preset segment regions include preset low-score segment regions that meet preset low-score conditions, preset medium-score segment regions that meet preset medium-score conditions, and preset high-score segment regions that meet preset high-score conditions; The corresponding test paradigm is determined using the material management unit based on the binaural rendered audio material to be tested, the reference material, the spatial anchor point material, and the test type; the test paradigm includes a first test paradigm, a second test paradigm, and a third test paradigm; the first test paradigm is a three-excitation paradigm; the second test paradigm is a multi-excitation hidden reference and anchor point paradigm; and the third test paradigm is a multiple comparison method paradigm.

3. The spatial perception quality testing method for binaural rendering according to claim 2, characterized in that, The acquisition of reference materials includes: Obtain reference materials; wherein, the reference materials include any one of the first reference materials, the second reference materials, and the third reference materials; The first reference material is an ideal spatial rendering reference material obtained by binaural recording in an anechoic chamber using HATS or by a real person wearing a miniature microphone; the second reference material is material generated using a preset high-quality binaural rendering system; the rendering quality of the preset high-quality binaural rendering system meets the preset high-quality conditions; the third reference material is material generated by directly feeding the original signal to the headphones without binaural rendering.

4. The spatial perception quality testing method for binaural rendering according to claim 3, characterized in that, Determining the test sequence based on the test paradigm includes: Based on the test paradigm, the corresponding signal combination method is determined, and the binaural rendering audio material to be tested is randomly sorted to obtain the sorted material. Based on the signal combination method, the reference material, and the spatial anchor point material, the sorted material is subjected to the insertion operation of the reference material and the spatial anchor point and sorting combination to obtain the test sequence.

5. The spatial perception quality testing method for binaural rendering according to claim 2, characterized in that, The positioning test corresponds to the three-excitation paradigm; the motion test corresponds to the multi-excitation hidden reference and anchor point paradigm; and the immersion and sound quality test corresponds to the multiple comparison method paradigm.

6. The spatial perception quality testing method for binaural rendering according to claim 4, characterized in that, Based on the reference material, the reference material is inserted into the binaural rendering audio material to be tested, including: Determine the test requirements corresponding to the test paradigm, and insert the reference material into the first preset position in the sorted material based on the test requirements to obtain the inserted material.

7. The spatial perception quality testing method for binaural rendering according to claim 6, characterized in that, Spatial anchor points are established for the binaural rendering audio material to be tested based on the spatial anchor point material, including: Based on the test requirements, insert the spatial anchor point material into the second preset position in the inserted material; The spatial anchor point material is set for the sound source localization accuracy dimension; the spatial anchor point material is a rendering output result with localization deviation and before-after confusion, or the spatial anchor point material is a material obtained by applying quantized spatial perturbation to the reference material; Set the spatial anchor point material for the continuity dimension of the sound source motion; the spatial anchor point material is the rendering output result of inserting periodic jumps or breaks in the continuous motion trajectory; The spatial anchor point material is set for the uniformity dimension of the sound source motion; the spatial anchor point material is the rendering output result of applying random speed fluctuations in uniform motion, or the spatial anchor point material is the output result of interpolating the azimuth angle to meet the preset low frame rate condition. The spatial anchor point material is set for the spatial immersion dimension; the spatial anchor point material is the lowest score anchor point of spatial immersion loss determined by the mono submix version and the externalized loss anchor point material determined by the stereo signal without binaural rendering; wherein, the number of anchor point signals corresponding to each evaluation dimension is a preset number.

8. The spatial perception quality testing method for binaural rendering according to claim 7, characterized in that, The calibration of the test sequence to obtain the calibrated test signal includes: Time delay calibration is performed on each of the binaural rendered audio materials to be tested, the reference material, and the spatial anchor point material in the test sequence to obtain the time delay calibrated test signal; Phase calibration is performed on each of the binaural rendered audio materials to be tested, the reference material, and the spatial anchor point material in the test sequence to obtain phase-calibrated test signals; the phase response of each phase-calibrated test signal at the binaural ears of the target listener is consistent; The sound level of each of the binaural rendered audio materials to be tested, the reference material, and the spatial anchor point material in the test sequence is calibrated to obtain a sound level calibrated test signal; the perceived loudness of the sound level calibrated test signal at the binaural ears of the target listener is not less than a preset standard sound pressure level; The calibrated test signal is determined based on the time delay calibrated test signal, the phase calibrated test signal, and the sound level calibrated test signal.

9. The spatial perception quality testing method for binaural rendering according to claim 1, characterized in that, Before receiving the raw scoring data input by the target listener based on the calibrated test signal, the basic sound quality attribute dimension, and the spatial perception attribute dimension, the following steps are also included: The test method corresponding to the quality test is determined. If the test method is the ITU-R BS.1116 method, a first preset number of first expert listeners with experience in audio impairment judgment and spatial perception identification are determined. The first preset number is greater than a first preset threshold. If the testing method is the MUSHRA method, then a second preset number of second expert listeners with spatial perception discrimination experience are determined; the second preset number is greater than a second preset threshold; the first preset threshold is greater than the second preset threshold; the spatial perception training background of the first expert listener and the second expert listener is consistent; Based on the first expert listeners and the second expert listeners, the listeners to be screened are determined, and the listeners to be screened are screened for hearing thresholds based on a preset frequency range. The frequency hearing thresholds corresponding to the listeners to be screened and each test frequency point are obtained, and it is determined whether the hearing thresholds of each frequency point are greater than a preset decibel value. If the hearing threshold at each frequency point is not greater than the preset decibel value, the candidate to be screened is determined to have passed the hearing screening. If the hearing threshold at each frequency point is less than the preset decibel value, the candidate to be screened is determined to have failed the hearing screening, and the testing process for the candidate to be screened is terminated.

10. The spatial perception quality testing method for binaural rendering according to claim 1, characterized in that, The process of training listeners includes: The listener is trained on sound source localization accuracy using training materials to obtain the first trained listener; the sound source localization accuracy training includes identifying azimuth deviation, elevation deviation, distance perception error and front-back confusion, and determining the perceptual characteristics of the head-in-the-head effect; The first trained listener is trained in the continuity dimension of sound source motion to obtain the second trained listener; the continuity dimension training of sound source motion includes distinguishing the continuity and discontinuity of motion trajectory and identifying perceptual defects including trajectory jumps and jitters. The second trained listener is trained in the uniformity dimension of sound source motion to obtain the third trained listener; the uniformity dimension training of sound source motion includes perceiving speed fluctuations in uniform motion and recognizing speed non-uniformity phenomena. The third trained listener is then trained in the spatial immersion dimension to obtain the target listener; the spatial immersion dimension training includes experiencing the differences in the degree of externalization of sound and spatial envelopment, as well as understanding the perceptual characteristics of each spatial type; wherein, the training time for training the listener is not less than a preset training threshold. Record the cumulative training time of the target listener in the spatial perception-specific training, and compare the cumulative training time with a preset training time threshold; If the comparison result indicates that the cumulative training time is not less than the preset training time threshold, then a verification pass instruction is generated.

11. The spatial perception quality testing method for binaural rendering according to claim 10, characterized in that, Before receiving the raw scoring data input by the target listener based on the calibrated test signal, the basic sound quality attribute dimension, and the spatial perception attribute dimension, the following steps are also included: Based on the preset requirements for sound source localization accuracy, single-point sound source fixed-point emission materials are determined; the sound source azimuth distribution corresponding to the single-point sound source fixed-point emission materials covers the front, side, rear and elevation directions. Based on the preset requirements for the continuity of sound source motion, audio material with a continuous motion trajectory is determined, including a continuous motion trajectory; the continuous motion trajectory includes sweeping motion from left to right and surround motion; the motion speed corresponding to the continuous motion trajectory satisfies the preset motion speed condition; The uniform motion scene material is determined based on the preset sound source motion uniformity requirements; the uniform motion scene material does not include natural acceleration and deceleration content. Based on the preset requirements for spatial immersion, multi-channel materials including spatial information are determined; the multi-channel materials include 3D sound movie audio and symphonic music. Training materials are constructed based on the single-point sound source fixed-point sound material, the continuous motion trajectory audio material, the uniform motion scene material, and the multi-channel material; the training materials include examples of binaural rendering spatial perception defects; the examples of binaural rendering spatial perception defects include positioning deviations of various degrees, motion discontinuity, speed unevenness, and loss of immersion; the number and duration of the training materials meet the material conditions corresponding to the subjective evaluation method for minor defects in the audio system.

12. The spatial perception quality testing method for binaural rendering according to claim 11, characterized in that, The raw scoring data received by the target listener based on the calibrated test signal, the basic sound quality attribute dimension, and the spatial perception attribute dimension includes: After obtaining the verification pass instruction, the calibrated test signals are played sequentially to the target listener, and a two-dimensional scoring interface is generated based on the basic sound quality attribute dimension and the spatial perception attribute dimension. A two-dimensional scoring interface is provided to the target listener so that the target listener can generate raw scoring data based on the calibrated test signal in the two-dimensional scoring interface.

13. The spatial perception quality testing method for binaural rendering according to claim 12, characterized in that, The basic sound quality attributes include low-frequency performance, clarity, dynamic range, timbre naturalness, noise, background noise, and vocal quality; the spatial perception attributes include sound source localization accuracy, motion continuity, motion uniformity, and spatial immersion; the sound source localization accuracy includes horizontal positioning accuracy, elevation angle positioning accuracy, distance perception accuracy, and front-back discrimination accuracy; the spatial immersion includes surround feeling, spatial size perception, degree of sound externalization, and degree of head-in-the-head effect.

14. The spatial perception quality testing method for binaural rendering according to claim 13, characterized in that, The process of generating raw rating data includes: Determine the preset grading scale corresponding to each sub-item in the dual-dimensional scoring interface, so that the target listener can generate a scoring value corresponding to each sub-item based on the value corresponding to the preset grading scale. Based on the aforementioned score values, generate original score data corresponding to the binaural rendered audio material to be tested.

15. The spatial perception quality testing method for binaural rendering according to claim 14, characterized in that, The process of generating original score data corresponding to the binaural rendered audio material to be tested based on each of the score values ​​includes: If any of the sub-items has a defect, a defect marker is generated based on the defect, and a score value is generated based on the defect marker. Then, based on each score value, unprocessed score data corresponding to the binaural rendered audio material to be tested is generated. The defect marker includes defect type information. A number of unprocessed score data generated by the same target listener in repeated tests of binaural rendered audio material to be tested are determined, and a score variance is generated based on each of the unprocessed score data. The rating variance is compared with a preset variance threshold, and the rating data to be processed corresponding to the rating variances that are greater than the preset variance threshold in the comparison results are removed to obtain an effective rating dataset.

16. The spatial perception quality testing method for binaural rendering according to claim 2, characterized in that, The verification of the original scoring data to obtain a valid scoring dataset includes: Extract the reference score data of the target listener for each reference material, and determine the deviation between each reference score data and the preset hidden reference standard score. Then compare the deviation with the preset reference deviation threshold to remove the original score data of the target listener corresponding to the deviation greater than the preset reference deviation threshold, so as to obtain the score dataset to be processed. The scoring data of each target listener are subjected to group statistical analysis, and outliers that deviate from the group statistical distribution in the group statistical analysis results are identified. The outliers are then removed from the scoring dataset to be processed to obtain a valid scoring dataset.

17. The spatial perception quality testing method for binaural rendering according to claim 1, characterized in that, Before generating the quality test results, including sound quality score, spatial awareness score, and defect distribution, based on the effective scoring dataset, the process also includes: The scores of each sub-item corresponding to the basic audio quality attribute dimension in the effective scoring dataset are weighted to obtain the audio quality score; The scores of each sub-item corresponding to the spatial awareness attribute dimension in the effective scoring dataset are weighted to obtain the spatial awareness score.

18. The spatial perception quality testing method for binaural rendering according to claim 17, characterized in that, The generation of quality test results based on the effective scoring dataset, including sound quality score, spatial awareness score, and defect distribution, includes: A radar chart is generated based on the sound quality score and the spatial perception score; the radar chart is used to display the performance profile of each attribute dimension. The defect labels marked by each target listener in the effective scoring dataset are statistically analyzed, and defect distribution information is generated according to the defect type and frequency of occurrence in the statistical results. Quality test results are generated based on the sound quality score, spatial perception score, radar chart, defect distribution information, and preset comprehensive rating classification standards.

19. The spatial perception quality testing method for binaural rendering according to claim 1, characterized in that, The process of receiving the raw scoring data generated by the listener includes: Control the switching of test dimensions. When switching from the basic sound quality attribute dimension to the spatial perception attribute dimension or vice versa, a preset rest interval is forcibly inserted, and the total test time for a single day is accumulated. When the total daily duration exceeds the preset maximum duration, the test reception for that day will be automatically terminated.

20. A spatial perception quality testing device using binaural rendering, characterized in that, include: The test sequence determination module is used to determine the corresponding test paradigm based on the binaural rendered audio material to be tested and the test type, and to determine the test sequence based on the test paradigm. The scoring data generation module is used to calibrate the test sequence, obtain the calibrated test signal, and receive the original scoring data input by the target listener based on the calibrated test signal, the basic sound quality attribute dimension, and the spatial perception attribute dimension; the target listener is a listener who has passed the spatial perception specialized training test; The quality test result generation module is used to verify the original scoring data to obtain a valid scoring dataset, and to generate quality test results including sound quality score, spatial perception score and defect distribution based on the valid scoring dataset.

21. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor for executing the computer program to implement the spatial perception quality testing method for binaural rendering as described in any one of claims 1 to 19.

22. A computer-readable storage medium, characterized in that, Used to store a computer program, wherein the computer program, when executed by a processor, implements the spatial perception quality testing method of binaural rendering as described in any one of claims 1 to 19.