Audio signal pickup regulation method based on relative distance discrimination threshold
By synthesizing near-field HRTF and virtual spatial audio, combined with air conduction and bone conduction devices, the auditory system's ability to distinguish sound sources at close range was measured and enhanced. This solved the problems of high computational resources and unstable sensitivity of traditional methods at close range, and improved the sound source localization accuracy and experience of the auditory system.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- GUANGZHOU UNIVERSITY
- Filing Date
- 2024-12-17
- Publication Date
- 2026-05-29
AI Technical Summary
In close-range situations, traditional plane wave-based sound source localization methods are no longer effective, and high-precision HRTF requires high computational resources. Furthermore, the auditory system's sensitivity to distance changes is unstable at close range, affecting the auditory system's accurate perception of the sound source's distance.
Virtual spatial audio is constructed by synthesizing near-field HRTF, and equal loudness matching is performed using air conduction and bone conduction devices. The relative distance discrimination threshold is measured by combining the binomial forced selection method and the 2-up-1-down adaptive method. The spherical harmonic domain beamforming algorithm is used to enhance the distance perception capability for minute object movements.
It improves the auditory system's ability to distinguish the distance of sound sources under different directions and distances, reduces the demand for computing resources, and enhances the auditory experience and quality of life.
Smart Images

Figure CN119767228B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of sound source localization technology, and more specifically, to an audio signal pickup and control method based on a relative distance discrimination threshold. Background Technology
[0002] Hearing plays a crucial role in our lives, especially in determining the location of objects. Our ability to locate sound sources encompasses two main aspects: determining the direction of the sound source and estimating its distance. In natural environments, this ability is essential for our spatial orientation and navigation, especially in low light conditions or when vision is limited, where hearing becomes a key tool for judging the distance of sound sources.
[0003] Research has found that sound source distance localization is influenced by a variety of factors, including loudness, the ratio of direct to reverberation sound, frequency distribution, differences in sound received by both ears, and changes caused by head movement. Among these factors, loudness and the spectral characteristics of the sound primarily provide information about the relative distance of the sound source, while in indoor environments, the ratio of direct to reverberation provides clues to the absolute distance of the sound source. Of all these factors, loudness is a primary factor in determining the distance of a sound source.
[0004] The auditory system's ability to perceive changes in distance, known as the Just Noticeable Difference (JND), is related to the reference distance. At greater distances, the JND remains relatively stable, meaning that our auditory system's sensitivity to distance changes is constant within a certain proportion. The pressure discrimination hypothesis posits that the threshold for perceiving a distance change can reach a 5% difference in sound pressure level.
[0005] However, JND changes more significantly at close range. This is mainly because nearby sound sources generate spherical waves, unlike the plane waves produced at greater distances. Therefore, traditional plane wave-based localization methods are no longer effective in the near field. Furthermore, traditional high-precision HRTF (Heat-Resolved Radio Frequency) methods have very high computational resource requirements for sound effect rendering.
[0006] Therefore, an audio signal pickup and control method based on a relative distance discrimination threshold is provided. Summary of the Invention
[0007] To address the aforementioned technical problems, this application is proposed. This application provides an audio signal pickup and control method based on a relative distance discrimination threshold, comprising:
[0008] S1. Synthesize near-field HRTF using sound source information and head parameters;
[0009] S2. Construct virtual spatial audio based on the near-field HRTF;
[0010] S3. Use the equal loudness matching method to ensure that the stimuli presented by the air conduction device and the bone conduction device have the same loudness level within the measurement frequency range;
[0011] S4. Using air conduction and bone conduction devices to play virtual spatial audio, the relative distance discrimination threshold of the subjects in different directions and at different distances is measured by the binomial forced selection method and the 2-up-1-down adaptive method.
[0012] S5. Using the relative distance discrimination threshold measurement results of the subjects, calculate the relative distance discrimination threshold of the audience to obtain the relative distance discrimination threshold of the audience at different directions and different distances;
[0013] S6. Set relative distance discrimination threshold evaluation criteria, and determine whether it is necessary to add a spherical harmonic beamforming algorithm for adjustment based on the audience's ability to perceive the distance of small movements of objects.
[0014] S7. For listeners who have poor distance perception of minute object movements, a beamforming algorithm based on the spherical harmonic domain is used to decouple the angle and frequency in the sound field to enhance the ability to perceive the distance of minute object movements.
[0015] Preferably, in S1, the distance variation function is as follows:
[0016]
[0017]
[0018] in, It is a spherical Hankel function of the first kind, with order . , Is it in radius The first derivative at point, It is the wave number. It is the speed of sound in the air. yes Legendre polynomial, From the center of the ball to a point on the surface of the ball The angle between the vector and the vector from the center of the sphere to the sound source. It is near field distance The pressure of the sound source at the ear. It is the far-field distance The pressure of the sound source at the ear.
[0019] Preferably, step S2 includes: convolving the synthesized near-field HRTF and Gaussian white noise to obtain spatial audio corresponding to different locations. The Gaussian white noise used is generated by Audition software and ranges from 100Hz to 8000Hz.
[0020] Preferably, S3 includes: S31, adjusting the air conduction device to ensure that the spatial audio sound pressure level is 60 dBSPL; S32, the subject wears Sennheiser IE800 air conduction headphones and Radioear B81 bone conduction transducer, and removes the air conduction headphones while controlling the bone conduction transducer to play spatial audio signals; S33, the subject adjusts the gain applied to the bone conduction transducer to match the perceived loudness of the stimulation from the air conduction headphones.
[0021] Preferably, S33 includes: keeping the gain of the air conduction headphone stimulation constant, comparing the perceived loudness of the bone conduction oscillator stimulation with the perceived loudness of the air conduction headphone stimulation, and if the perceived loudness of the bone conduction oscillator is greater than the perceived loudness of the air conduction headphone stimulation, then decreasing the perceived loudness of the bone conduction oscillator; if the perceived loudness of the bone conduction oscillator is less than the perceived loudness of the air conduction headphone stimulation, then increasing the perceived loudness of the bone conduction oscillator.
[0022] Preferably, step S4 includes: S41, conducting a preliminary experiment for each subject; S42, setting a reference distance for the sound source in the formal experiment and configuring initial adaptive virtual distances for sound sources in different directions; S43, playing sound sources in different stimulus sequences, and using a binary forced choice method, having the subject judge two consecutively played sound stimuli; S44, employing a 2-up-1-down adaptive method, where when the subject makes two consecutive correct judgments, the sound source moves away from the reference distance by 1%, and when the subject makes an incorrect judgment, the sound source moves closer to the reference distance by 1%, until the adaptive virtual distance of the sound source matches the reference distance; wherein, before the first incorrect judgment, the size of each step is set to 3%.
[0023] Preferably, S41 includes: having the subject listen to surround sound on a horizontal plane through an air conduction device or a bone conduction device to familiarize them with the virtual sound field; having the subject listen to stimulus signals from near to far at different azimuth angles in sequence so that the subject can become familiar with virtual sound sources at different distances; and giving each subject a short training session consisting of several feedback trials so that they can become familiar with the task and the stimulus.
[0024] Preferably, after eliminating loudness factors, a compensation coefficient K related to the sound source distance r is used to normalize stimulus signals at different distances, so that spatial audio at different distances in the same direction exhibits the same loudness; wherein the calculation formula for the compensation coefficient is as follows:
[0025]
[0026] in It is the distance from the sound source to the right ear. It is the distance from the sound source to the left ear;
[0027]
[0028] in It is a normalized audio signal. It is the original audio signal.
[0029] Preferably, in step S5, the relative distance discrimination threshold is calculated as follows:
[0030]
[0031] in It is the relative distance discrimination threshold at the reference distance r. This is a reference distance. It is the average distance corresponding to the last 5 flip positions.
[0032] Preferably, step S6 includes: setting a relative distance discrimination threshold j to evaluate an individual's spatial perception ability. If the listener's relative distance discrimination threshold is greater than j, it indicates that their ability to perceive the distance of moving objects is poor; even if the object moves a large distance, the listener cannot perceive it. In this case, a beamforming algorithm in the spherical harmonic domain needs to be added for enhancement. Conversely, if the listener's relative distance discrimination threshold is less than j, then a beamforming algorithm in the spherical harmonic domain is not needed for enhancement.
[0033] Preferably, in step S7, the formula for calculating the radial filter coefficient d is:
[0034]
[0035] in These correspond to different orders of response patterns. It is the predetermined radial coefficient matrix;
[0036] Perform filtering:
[0037]
[0038] in It is the wave number. It is the distance to the target sound source. satisfy Where a is the radius, Let be the boundary radius of the near field, and its value is... .
[0039] Compared with existing technologies, this application provides an audio signal pickup and control method based on a relative distance discrimination threshold. This method measures the relative distance discrimination threshold under different directions, reference distances, stimulus sequences, and stimulus types. It analyzes the factors affecting the relative distance discrimination threshold when using bone conduction devices, and predicts the listener's ability to perceive the distance of moving objects under different conditions. This provides a scientific basis for the improvement and optimization of hearing aids and consumer-grade bone conduction headphones, thereby helping hearing-impaired individuals better perceive sound and improve their auditory experience and quality of life. Simultaneously, by combining the relative distance discrimination threshold with a spherical harmonic beamforming algorithm, the algorithm's complexity is minimized while enhancing the listener's ability to perceive sound movement. Attached Figure Description
[0040] The above and other objects, features, and advantages of this application will become more apparent from the more detailed description of the embodiments of this application in conjunction with the accompanying drawings. The drawings are provided to further illustrate the embodiments of this application and form part of the specification. They are used together with the embodiments of this application to explain this application and do not constitute a limitation thereof. In the drawings, the same reference numerals generally represent the same components or steps.
[0041] Figure 1 This is a flowchart illustrating a method of an example of the present invention.
[0042] Figure 2 This is a schematic diagram of the sound source location and distance in an example of the present invention.
[0043] Figure 3 This is a flowchart of step S3 in an example of the present invention.
[0044] Figure 4 This is a flowchart of step S4 in an example of the present invention.
[0045] Figure 5 This represents the average relative distance discrimination threshold results for 10 subjects under fixed stimulus conditions in this invention example.
[0046] Figure 6 This represents the average relative distance discrimination threshold results for nine subjects under normalized stimulus conditions in this invention example. Detailed Implementation
[0047] The technical solution of this application will be further described below with reference to the accompanying drawings and embodiments. Obviously, the described embodiments are merely some embodiments of this application, and not all embodiments of this application. It should be understood that this application is not limited to the exemplary embodiments described herein.
[0048] Example:
[0049] like Figure 1As shown, the audio signal pickup and control method based on relative distance discrimination threshold according to an embodiment of this application includes: S1, synthesizing a near-field HRTF using sound source information and head parameters; S2, constructing virtual spatial audio based on the near-field HRTF; S3, using an equal loudness matching method to ensure that stimuli presented by air conduction and bone conduction devices have the same loudness level within the measurement frequency range; S4, playing the virtual spatial audio using air conduction and bone conduction devices, and measuring the relative distance discrimination threshold of the subject in different directions and at different distances using a binomial forced selection method and a 2-up-1-down adaptive method. S5. Using the relative distance discrimination threshold measurement results of the subjects, calculate the relative distance discrimination threshold of the listeners, and obtain the relative distance discrimination threshold of the listeners at different directions and distances; S6. Set the relative distance discrimination threshold evaluation criteria, and determine whether it is necessary to add a spherical harmonic domain beamforming algorithm for adjustment based on the listeners' ability to perceive the distance of small objects; S7. For listeners with poor ability to perceive the distance of small objects, use a spherical harmonic domain-based beamforming algorithm to decouple the angle and frequency in the sound field to enhance the ability to perceive the distance of small objects.
[0050] In this embodiment, step S1 involves synthesizing a near-field HRTF using sound source information and head parameters. It should be understood that traditional high-precision HRTF (Head-Related Transfer Function) sound effect rendering requires simulating the transmission process of sound waves from the sound source to both ears. This includes air filtering, reverberation from the surrounding environment, and scattering and reflection from the human body (torso, head, auricles, etc.). This process involves complex calculations of physical and environmental characteristics, requiring a large amount of mathematical computation, thus necessitating significant computational resources. Therefore, to reduce this computational burden, a pre-defined method is used to convert the traditional far-field HRTF into an HRTF suitable for near-field sound source environments using sound source information and head parameters. This method simulates sound source localization by pre-measuring and establishing a detailed HRTF database, thereby reducing the need for real-time computation.
[0051] Specifically, in the embodiments of this application, a method is set up. Figure 2 The diagram showing the sound source orientation and distance illustrates two sound source directions (90° and 180°) and two reference distances (50cm and 100cm), from which near-field HRTF data were synthesized. Specifically, within the distance range of 100cm to 50cm, 50 HRTFs were synthesized at 1cm intervals; within the range of 50cm to 25cm, 50 HRTFs were synthesized at 0.5cm intervals; and within the range of 25cm to 12.5cm, 50 HRTFs were synthesized at 0.25cm intervals. A total of 300 HRTF data points were synthesized, covering both directions (90° and 180°).
[0052] Specifically, the far-field HRTF database used is the SCUT HRTF database measured by South China University of Technology. This database provides HRTF data in the range of 20cm to 100cm, specifically including 20cm, 25cm, 30cm, 40cm, 50cm, 60cm, 70cm, 80cm, 90cm, and 100cm, that is, HRTF data at 10 distances in one direction.
[0053] HRTF data at 100cm was selected and combined with a distance variation function (DVF) to synthesize the required near-field HRTF. The distance variation function is as follows:
[0054]
[0055]
[0056] in, It is a spherical Hankel function of the first kind, with order . , Is it in radius The first derivative at point, It is the wave number. It is the speed of sound in the air. yes Legendre polynomials From the center of the ball to a point on the surface of the ball The angle between the vector of the sphere and the vector from the center of the sphere to the sound source. By considering all frequencies of interest, the location of a specific sound source at a given point can be obtained. The transfer function at that point. It is near field distance The pressure of the sound source at the ear. It is the far-field distance The pressure of the sound source at the ear.
[0057] Specifically, the distance variation function is a mathematical model designed based on the distance variation between the sound source and the receiver, which can obtain the required near-field HRTF as needed. Using this method, we can simulate the pressure of the sound source at the ear at near-field distances, which matches the pressure of the sound source at the ear at far-field distances. This method, based on a pre-set database and mathematical model, effectively reduces the computational load of real-time processing, enabling relatively accurate sound source localization even on devices with limited computing resources.
[0058] In this embodiment of the application, step S2 involves constructing virtual spatial audio based on the near-field HRTF. It should be understood that, in order to reproduce the spatial distribution of sound in a virtual environment, including the direction and distance of sound, thereby providing a more realistic auditory experience, a synthesized near-field HRTF is further used to construct a virtual spatial audio environment.
[0059] Specifically, step S2 includes: convolving the synthesized near-field HRTF with Gaussian white noise to obtain spatial audio corresponding to different locations. The Gaussian white noise used is generated by Audition software and ranges from 100Hz to 8000Hz.
[0060] In this embodiment, step S3 uses equal loudness matching to ensure that the stimuli presented by the air conduction device and the bone conduction device have the same loudness level within the measurement frequency range. It should be understood that the human ear has different sensitivities to different frequencies, particularly in the 2000Hz to 5000Hz frequency range, where the human ear is most sensitive to mid-frequency frequencies, while its sensitivity decreases at the low and high frequencies. Therefore, to maintain consistency in the auditory experience of sounds at different frequencies, the loudness of sounds at different frequencies is adjusted using equal loudness matching, so that the loudness perceived by the human ear remains consistent at different volume levels. In other words, equal loudness matching ensures that the stimuli presented by the air conduction device and the bone conduction device have the same loudness level within the measurement frequency range.
[0061] Specifically, such as Figure 3 As shown, step S3 includes: S31, adjusting the air conduction device to ensure that the sound pressure level of the played spatial audio is 60 dB SPL; S32, the subject wears Sennheiser IE800 air conduction headphones and Radioear B81 bone conduction transducer, and removes the air conduction headphones while controlling the bone conduction transducer to play spatial audio signals (i.e., alternately using the air conduction device and the bone conduction device to play spatial audio, and test the loudness perception of the air conduction device and the bone conduction device at the same sound source location respectively); S33, the subject adjusts the gain applied to the bone conduction transducer to match the perceived loudness of the stimulation from the air conduction headphones.
[0062] The specific implementation process of step S31 is as follows: The Sennheiser IE800 air-conduction headphones are worn on the Head and Torso Simulator (HATS, B&K 4128 model); spatial audio at a distance of 50cm in a 180° direction is played through the air-conduction headphones, and the sound pressure level (SPL) is measured using the HATS. The device configuration at a SPL of 60dB SPL is recorded. This step ensures that the audio played by the air-conduction device has a known reference SPL during subsequent loudness matching.
[0063] The specific implementation process of step S33 is as follows: keep the gain of the air conduction headphone stimulation constant, compare the perceived loudness of the bone conduction oscillator stimulation with the perceived loudness of the air conduction headphone stimulation, if the perceived loudness of the bone conduction oscillator is greater than the perceived loudness of the air conduction headphone stimulation, then decrease the perceived loudness of the bone conduction oscillator, if the perceived loudness of the bone conduction oscillator is less than the perceived loudness of the air conduction headphone stimulation, then increase the perceived loudness of the bone conduction oscillator.
[0064] In this embodiment, step S4 involves playing virtual spatial audio using both air-conduction and bone-conduction devices, and measuring the relative distance discrimination threshold of the subject at different directions and distances using a binomial forced selection method and a 2-up-1-down adaptive method. It should be understood that in virtual reality and augmented reality applications, understanding and simulating human perception of sound source distance is crucial for creating a more realistic auditory experience. Therefore, to better predict the performance issues of bone-conduction devices in perceiving minute object movements and to provide a basis and direction for performance improvement, the relative distance discrimination threshold of the subject is further measured using the binomial forced selection method and the 2-up-1-down adaptive method. The differences in the ability of air-conduction and bone-conduction devices to detect subtle sound movements are analyzed, and the results are used to predict the performance issues of bone-conduction devices in perceiving minute object movements, providing a basis and direction for performance improvement.
[0065] Specifically, such as Figure 4 As shown, step S4 includes: S41, conducting a preliminary experiment for each subject; S42, setting the reference distance for the sound source in the formal experiment and configuring initial adaptive virtual distances for sound sources in different directions; S43, playing sound sources with different stimulus sequences, and using a binomial forced choice method, the subject judges two consecutively played sound stimuli; S44, using a 2-up-1-down adaptive method, when the subject makes two consecutive correct judgments, the sound source moves away from the reference distance by 1%, and when the subject makes an incorrect judgment, the sound source moves closer to the reference distance by 1%, until the adaptive virtual distance of the sound source matches the reference distance; wherein, before the first incorrect judgment, the size of each step is set to 3%.
[0066] The specific implementation process of step S41 is as follows:
[0067] 1. Have the subjects listen to the surround sound of the horizontal plane through air conduction or bone conduction devices to familiarize them with the virtual sound field.
[0068] Control the air-conducting headphones or bone-conducting vibrators to play spatial audio at the same distance. The spatial audio consists of surround sound at 7 azimuth angles (0°, 30°, 60°, 90°, 120°, 150°, 180°) at 25cm, followed by surround sound at 50cm and 100cm.
[0069] 2. Allow the subjects to listen to stimulation signals from different azimuth angles, from near to far, so that the subjects can become familiar with virtual sound sources at different distances.
[0070] Control the air-conducting headphones or bone-conducting vibrators to play spatial audio at the same azimuth angle, where the azimuth angle of the spatial audio includes 0°, 90° and 180°, the distance is from 20cm to 100cm, and the interval is 10cm (i.e. 20cm, 30cm, 40cm, 50cm, 60cm, 70cm, 80cm, 90cm, 100cm).
[0071] 3. Each participant will undergo a brief training session consisting of several feedback-based trials to familiarize them with the task and stimuli. No feedback will be provided to participants during the formal test.
[0072] Specifically, in the embodiments of this application, the formal experiment described above covers four main variables: two directions (e.g., 180° and 90°), two reference distances (e.g., 50cm and 100cm), two stimulus sequences (proximity stimulus sequence and distance stimulus sequence), and two conduction methods (air conduction device and bone conduction device). A total of 24 experimental conditions are formed by combining these variables. Among them, playing the spatial audio at the reference distance first and then the spatial audio at the adaptive position is defined as the proximity stimulus sequence, and conversely, playing the spatial audio at the adaptive distance first and then the spatial audio at the reference position is defined as the distance stimulus sequence.
[0073] For sound sources in the 90° direction, the initial adaptive virtual distance was set to 70% of the reference distance (e.g., 35cm if the reference distance is 50cm). A larger initial adaptive virtual distance was set for sound sources in the 180° direction. Simultaneously, two spatial audio samples were combined into one experimental test. One spatial audio sample corresponded to the spatial location at the reference distance, and the other spatial audio sample corresponded to the spatial location at the adaptive distance. The Gaussian white noise used in the construction of the spatial audio samples was 0.5 seconds long, including a 30ms rise time and a 30ms fall time. The rise and fall times were set to ensure a smooth transition of the sound signal and avoid sharp sound wave interference caused by sudden start or end of the sound. This not only helps improve the accuracy and reliability of the measurement but also enhances the auditory experience of the subjects and reduces the influence of external noise on the experimental results.
[0074] Specifically, during the experiment, participants sat in comfortable chairs and held a transponder. The transponder had two buttons: a red button indicated that the second sound was closer than the first, and a blue button indicated that the second sound was farther away. Both the transponder and the sound card were connected to a computer running the experimental program. To avoid visual interference, the computer screen was turned off during the experiment, ensuring participants could focus on the auditory task. During the experiment, participants pressed the corresponding red or blue button to provide feedback based on the differences in the sounds they heard. During breaks, participants could choose to turn on the computer screen, which displayed a countdown to the next test block, helping them understand the progress of the experiment and prepare. The entire experimental environment was kept quiet to ensure data accuracy. Before the experiment began, the transponder and sound card were checked to ensure all equipment was functioning properly and the buttons were working correctly.
[0075] It is worth noting that the aforementioned formal experiment (hereinafter referred to as "Experiment 1") was conducted under conditions of loudness (i.e., the synthesized near-field HRTF was synthesized using a distance variation function and then convolved with Gaussian white noise to obtain spatial audio). Therefore, to investigate the effects of direction, reference distance, stimulus order, and conduction mode on the relative distance discrimination threshold while eliminating the loudness factor, another experimental method is provided (hereinafter referred to as "Experiment 2"). In this experiment, to remove the loudness factor, a compensation coefficient K related to the sound source distance r was used to normalize the stimulus signals at different distances, thereby ensuring that spatial audio at different distances in the same direction exhibits the same loudness. Thus, the original audio signal is compensated by the compensation coefficient to obtain the spatial audio required for Experiment 2. The calculation formula for the compensation coefficient is as follows:
[0076]
[0077] in It is the distance from the sound source to the right ear. It is the distance from the sound source to the left ear. The specific usage method is as follows:
[0078]
[0079] in It is the normalized audio signal. It is the original audio signal.
[0080] Experiment 2 is similar to Experiment 1. Before conducting Experiment 2, the preparatory experiment in step S41 above must also be performed. Experiment 2 also covers four main variables: two directions (such as 180° and 90°), two reference distances (such as 50cm and 100cm), two stimulus sequences (close stimulus sequence and far stimulus sequence), and two conduction methods (air conduction device and bone conduction device), which can yield a total of 24 experimental conditions.
[0081] Optionally, after compensating for the audio signal and obtaining the normalized audio signal, the amplitude of the audio signal is randomized by 15 dB SPL. By randomly adjusting the sound pressure level, a fixed sound pressure level may cause participants to adapt to the stimulus signal, reducing their sensitivity or leading to fixed expectations. Randomizing the sound pressure level can break this adaptation, maintaining participants' high sensitivity and attention to each stimulus, thus improving the effectiveness of the experiment. Simultaneously, sound pressure level randomization increases the diversity of experimental conditions, making the experimental results less limited by specific sound pressure level settings. This helps improve the robustness of the experimental results and their generalization ability under different environments or conditions.
[0082] Through the above experiments one and two, we can obtain Figure 5 The average relative distance discrimination threshold of the 10 subjects under fixed stimulus conditions is shown as follows: Figure 6 The figure shows the average relative distance discrimination thresholds of the nine subjects under normalized stimulus conditions. After obtaining the relative distance discrimination thresholds calculated using air-conduction headphones and bone-conduction transducers, analysis of variance (ANOVA) can be performed on the data. Multivariate repeated measures ANOVA can be conducted to analyze the effects of factors such as direction, reference distance, and stimulus order on the relative distance discrimination thresholds. Further, repeated measures ANOVA can be used to analyze the relative distance discrimination thresholds of bone-conduction transducers and air-conduction headphones to check for significant differences, thus providing a basis and direction for improving and enhancing the near-field spatial sound reproduction performance of bone-conduction devices.
[0083] In this embodiment, step S5 uses the relative distance discrimination threshold measurement results of the subject to calculate the relative distance discrimination threshold of the listener, obtaining the relative distance discrimination threshold of the listener at different locations and distances. It should be understood that by measuring the relative distance discrimination threshold under different sound source directions, different sound source reference distances, different stimulus sequences, different conduction devices, and different stimulus types, the listener's ability to perceive the distance of object movement under different conditions can be estimated, obtaining the minimum distance required for the listener to perceive object movement. This allows for the evaluation of the distance perception capabilities of hearing aids and consumer-grade bone conduction headphones, and consequently, the design of hearing devices that enable listeners to better judge object movement. Therefore, further using the relative distance discrimination threshold measurement results of the subject, the relative distance discrimination threshold of the listener is calculated, obtaining the relative distance discrimination threshold of the listener at different locations and distances.
[0084] Specifically, in step S5, the relative distance discrimination threshold is calculated as follows:
[0085]
[0086] in It is the relative distance discrimination threshold at the reference distance r. This is a reference distance. It is the average distance corresponding to the last 5 flip positions.
[0087] In this embodiment, step S6 involves setting a relative distance discrimination threshold evaluation standard and determining whether a spherical harmonic domain beamforming algorithm is needed to enhance the listener's spatial perception ability based on the listener's ability to perceive the distance of minute object movements. Specifically, in step S6, a relative distance discrimination threshold j is set to evaluate an individual's spatial perception ability. If the listener's relative distance discrimination threshold is greater than j, it indicates that their ability to perceive the distance of object movement is poor; even if the object moves a large distance, the listener cannot perceive it. In this case, a spherical harmonic domain beamforming algorithm is needed for enhancement. Conversely, if the listener's relative distance discrimination threshold is less than j, then a spherical harmonic domain beamforming algorithm is not needed for enhancement.
[0088] In this embodiment, step S7, for listeners with poor distance perception of minute object movements, uses a beamforming algorithm based on the spherical harmonic domain to decouple the angle and frequency in the sound field, thereby enhancing the ability to perceive the distance of minute object movements. Specifically, step S7 employs spherical harmonic domain (SH-Domain) processing to decompose the pressure distribution in the sound field into angularly correlated components and radially correlated components, and completes filtering through a series of weighted calculations of spherical harmonic functions and mode intensity factors.
[0089] Specifically, the spherical harmonic domain-based beamforming algorithm is a radial filter design method based on orthogonal polynomials for separating near-field sound sources. That is, the algorithm constructs the radial filter using a family of orthogonal polynomials, which outperforms previous methods in terms of white noise gain (WNG) and directivity index. Furthermore, some orthogonal polynomials are more effective at separating sound sources located near the microphone array surface. The formula for calculating the radial filter coefficients d is:
[0090]
[0091] in These correspond to different orders of response patterns. It is the predetermined radial coefficient matrix.
[0092] Further filtering can be performed:
[0093]
[0094] in It is the distance to the target sound source. It is the wave number. For an array with radius a, Need to meet Where N is the order of the array. Furthermore, for near-field sound sources, the distance to the sound source... Need to meet ,in Let be the boundary radius of the near field, and its value is... .
[0095] In particular, besides using the aforementioned spherical harmonic beamforming algorithm, other algorithms suitable for near-field positioning can also be used for control, such as near-field... Near field .
[0096] In summary, the audio signal pickup and control method based on relative distance discrimination threshold described in the embodiments of this application is clarified. This method uses a synthesized near-field HRTF to construct spatial audio with different directions, distances, and stimulus sequences to measure and analyze the relative distance discrimination threshold of the subject. The analysis results are used to predict the differences in relative distance discrimination thresholds between air conduction and bone conduction devices in terms of sound source direction, sound source distance, and sound source stimulus sequence within the near-field range. The relative distance discrimination threshold of the listener is then calculated. For listeners with a large relative distance discrimination threshold, a spherical harmonic beamforming method is used to pick up the audio at the corresponding location, increasing their distance perception at that location. This method can be used to evaluate the distance perception capability of hearing aids and consumer-grade bone conduction headphones, and subsequently design hearing devices that enable listeners to better judge the movement of objects.
[0097] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.
[0098] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit of the technical solutions of the present invention.
Claims
1. An audio signal pickup and control method based on a relative distance discrimination threshold, characterized in that, include: S1. Synthesize near-field HRTF using sound source information and head parameters; S2. Construct virtual spatial audio based on the near-field HRTF; S3. Use the equal loudness matching method to ensure that the stimuli presented by the air conduction device and the bone conduction device have the same loudness level within the measurement frequency range; S4. Using air conduction and bone conduction devices to play virtual spatial audio, the relative distance discrimination threshold of the subjects in different directions and at different distances is measured by the binomial forced selection method and the 2-up-1-down adaptive method. S5. Using the relative distance discrimination threshold measurement results of the subjects, calculate the relative distance discrimination threshold of the audience to obtain the relative distance discrimination threshold of the audience at different directions and different distances; S6. Set relative distance discrimination threshold evaluation criteria, and determine whether it is necessary to add a spherical harmonic beamforming algorithm for adjustment based on the audience's ability to perceive the distance of small movements of objects. S7. For listeners who have poor distance perception of minute object movements, a beamforming algorithm based on the spherical harmonic domain is used to decouple the angle and frequency in the sound field to enhance the distance perception of minute object movements. The S4 includes: S41, conduct a preliminary experiment on each subject; S42, set the reference distance of the sound source in the formal experiment, and configure the initial adaptive virtual distance for the sound source in different directions; S43, different sequences of stimuli were played from the sound source, and the subjects judged the two consecutively played sound stimuli using a binary forced choice method; S44 employs a 2-up-1-down adaptive method. When the subject makes two consecutive correct judgments, the sound source moves away from the reference distance by 1% of its original size. When the subject makes an incorrect judgment, the sound source moves closer to the reference distance by 1% of its original size. This process continues until the adaptive virtual distance of the sound source matches the reference distance. Before the first incorrect judgment, the size of each step is set to 3%.
2. The audio signal pickup and control method based on a relative distance discrimination threshold according to claim 1, characterized in that, The required near-field HRTF is synthesized by selecting HRTF data and combining it with a distance variation function; The distance variation function is as follows: in, It is a spherical Hankel function of the first kind, with order . , Is it in radius The first derivative at point, It is the wave number. It is the speed of sound in the air. yes Legendre polynomials From the center of the ball to a point on the surface of the ball The angle between the vector and the vector from the center of the sphere to the sound source. It is near field distance The pressure of the sound source at the ear. It is the far-field distance The pressure of the sound source at the ear.
3. The audio signal pickup and control method based on a relative distance discrimination threshold according to claim 2, characterized in that, S2 includes: convolving the synthesized near-field HRTF and Gaussian white noise to obtain spatial audio corresponding to different positions; wherein the Gaussian white noise used is generated by Audition software and is 100Hz-8000Hz.
4. The audio signal pickup and control method based on a relative distance discrimination threshold according to claim 3, characterized in that, The S3 includes: S31, Adjust the air conduction device to ensure that the sound pressure level of the spatial audio playback is 60dB SPL; S32, the subject wore Sennheiser IE800 air conduction headphones and Radioear B81 bone conduction vibrator, and removed the air conduction headphones while controlling the bone conduction vibrator to play spatial audio signals; S33, the subject adjusts the gain applied to the bone conduction oscillator to match the perceived loudness of the air conduction headphone stimulation.
5. The audio signal pickup and control method based on a relative distance discrimination threshold according to claim 4, characterized in that, S33 includes: keeping the gain of the air conduction headphone stimulation constant, comparing the perceived loudness of the bone conduction oscillator stimulation with the perceived loudness of the air conduction headphone stimulation, and if the perceived loudness of the bone conduction oscillator is greater than the perceived loudness of the air conduction headphone stimulation, then decreasing the perceived loudness of the bone conduction oscillator; if the perceived loudness of the bone conduction oscillator is less than the perceived loudness of the air conduction headphone stimulation, then increasing the perceived loudness of the bone conduction oscillator.
6. The audio signal pickup and control method based on a relative distance discrimination threshold according to claim 5, characterized in that, S41 includes: Subjects were allowed to listen to surround sound on the horizontal plane through air conduction or bone conduction devices to familiarize themselves with the virtual sound field; Subjects were given stimulation signals from different azimuth angles, from near to far, in order to familiarize them with virtual sound sources at different distances. Each participant was given a short training session consisting of several feedback-based trials to familiarize them with the task and stimuli.
7. The audio signal pickup and control method based on a relative distance discrimination threshold according to claim 6, characterized in that, After eliminating loudness factors, a compensation coefficient K, which is related to the sound source distance, is used to normalize stimulus signals at different distances, so that spatial audio at different distances in the same direction exhibits the same loudness. The formula for calculating the compensation coefficient is as follows: in It is the distance from the sound source to the right ear. It is the distance from the sound source to the left ear; in It is a normalized audio signal. It is the original audio signal.
8. The audio signal pickup and control method based on a relative distance discrimination threshold according to claim 7, characterized in that, In step S5, the relative distance discrimination threshold is calculated as follows: in The relative distance discrimination threshold at the sound source distance It is the average distance corresponding to the last 5 flip positions.
9. The audio signal pickup and control method based on a relative distance discrimination threshold according to claim 8, characterized in that, S6 includes: setting a relative distance discrimination threshold j to evaluate an individual's spatial perception ability. If the listener's relative distance discrimination threshold is greater than j, it indicates that their ability to perceive the distance of moving objects is poor. Even if the object moves a large distance, the listener cannot perceive it. In this case, a beamforming algorithm in the spherical harmonic domain needs to be added for enhancement. Conversely, if the listener's relative distance discrimination threshold is less than j, then a beamforming algorithm in the spherical harmonic domain does not need to be added for enhancement.
10. The audio signal pickup and control method based on a relative distance discrimination threshold according to claim 9, characterized in that, In S7, the formula for calculating the radial filter coefficient d is: in These correspond to different orders of response patterns. It is a preset radial coefficient matrix; Perform filtering: in It is the wave number. It is the distance to the target sound source. satisfy Where a is the radius, Let be the boundary radius of the near field, and its value is... .