Audio processing method and device, electronic equipment and storage medium
By acquiring information about the speaker and the listener, and combining virtual sound sources and head-related transfer functions for sound field reconstruction and correction, the problem of ambiguous lateral sound image localization is solved, thereby improving auditory immersion and user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- GOLDANA TECH CO LTD
- Filing Date
- 2026-01-26
- Publication Date
- 2026-05-19
AI Technical Summary
In existing technologies, sound field reconstruction techniques based on head-related transfer functions suffer from ambiguity in lateral sound image localization, resulting in insufficient auditory immersion.
By acquiring the spatial distribution information of the loudspeakers and the posture information of the listener, and combining the orientation information of the virtual sound source and the head-related transfer function, sound field reconstruction and correction are performed to compensate for the deviation between the physiological structure of the human ear and the actual listening environment, thereby improving the localization clarity of the lateral sound image.
It improves the positioning clarity of lateral sound images, making the sense of sound location highly matched with the spatial orientation of visual content, thus enhancing the user's immersive entertainment experience.
Smart Images

Figure CN122069474A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of audio processing technology, and in particular to an audio processing method, apparatus, electronic device and storage medium. Background Technology
[0002] In entertainment scenarios such as watching movies and playing games, users are not only pursuing highly realistic visual experiences, but also increasingly demanding a more immersive auditory experience.
[0003] Currently, sound field reconstruction technology based on head related transfer function (HRTF) is the mainstream method for realizing sound spatialization. However, because the physiological structure of the human ear is most sensitive to sounds directly in front / behind, and has a weaker ability to perceive sounds from the side, coupled with the interference of room reflection, diffraction and other effects in the actual listening environment, the reconstructed side sound images generally have the problem of blurred positioning in terms of auditory perception.
[0004] In summary, improving the localization clarity of lateral acoustic images has become a pressing technical problem that needs to be solved in this field.
[0005] The above content is only used to help understand the technical solution of this application and does not represent an admission that the above content is prior art. Summary of the Invention
[0006] The main objective of this application is to provide an audio processing method, apparatus, electronic device, and storage medium, which aims to improve the positioning clarity of lateral sound images.
[0007] To achieve the above objectives, this application proposes an audio processing method, which includes: The spatial distribution information of multiple speakers in the target space and the position and posture information of the listener are obtained, as well as the orientation information of the virtual sound source of the audio signal. The position and posture information includes the offset position of the listener relative to each of the speakers and the direction of the head. Based on the orientation information of the virtual sound source, the pose information, and the head-related transfer function acquired in the target space, the audio signal is subjected to sound field reconstruction processing to obtain a binaural audio signal. When the virtual sound source is positioned laterally to the listener, the binaural audio signal is corrected based on the orientation information of the virtual sound source and the spatial distribution information to obtain the target binaural audio signal.
[0008] In one embodiment, before the step of performing sound field reconstruction processing on the audio signal based on the orientation information of the virtual sound source, the pose information, and the head correlation transfer function acquired in the target space to obtain a binaural audio signal, the method further includes: The left surround channel signal and the right surround channel signal in the audio signal are filtered respectively to obtain the left filtered signal and the right filtered signal; Calculate the sum and difference signals of the left filtered signal and the right filtered signal; The absolute value of the difference signal is weighted based on a preset weighting factor, and the weighted result is subtracted from the sum signal to obtain the intermediate signal; The intermediate signal is subjected to half-wave rectification and smoothing filtering to obtain the enhanced signal component. The step of performing sound field reconstruction processing on the audio signal based on the orientation information of the virtual sound source, the pose information, and the head-related transfer function acquired in the target space to obtain binaural audio signals includes: The enhanced signal component is superimposed on the main channel signal in the audio signal to obtain a new audio signal; Based on the orientation information of the virtual sound source, the pose information, and the head-related transfer function acquired in the target space, the new audio signal is subjected to sound field reconstruction processing to obtain a binaural audio signal.
[0009] In one embodiment, the step of performing sound field reconstruction processing on the audio signal based on the orientation information of the virtual sound source, the pose information, and the head correlation transfer function acquired in the target space to obtain a binaural audio signal includes: The head-related transfer function acquired in the target space is subjected to spherical harmonic decomposition to obtain the spherical harmonic coefficients; Based on the orientation information and pose information of the virtual sound source, the position information of the virtual sound source relative to the listener within the target space is determined; Based on the location information and the spherical harmonic coefficients, the audio signal is subjected to sound field reconstruction processing to obtain a binaural audio signal.
[0010] In one embodiment, the step of performing sound field reconstruction processing on the audio signal based on the location information and the spherical harmonic coefficients to obtain a binaural audio signal includes: Calculate the spherical harmonic basis functions based on the location information; The head-related impulse response signal is obtained by multiplying the spherical harmonic coefficients by the spherical harmonic basis functions and then summing the results. The head-related impulse response signal is convolved with the audio signal to obtain the binaural audio signal.
[0011] In one embodiment, the step of correcting the binaural audio signal based on the location information of the virtual sound source and the spatial distribution information to obtain the target binaural audio signal includes: Based on the location information of the virtual sound source and the spatial distribution information, the delay compensation amount of the target loudspeaker corresponding to the main channel in the binaural audio signal is calculated. The binaural time difference is corrected for the binaural audio signal based on the delay compensation amount. The binaural intensity difference adjustment factor is calculated based on the orientation information of the virtual sound source; Based on the binaural intensity difference adjustment factor, the binaural audio signal after binaural time difference correction is subjected to binaural intensity difference correction to obtain the target binaural audio signal.
[0012] In one embodiment, before the step of performing sound field reconstruction processing on the audio signal based on the orientation information of the virtual sound source, the pose information, and the head correlation transfer function acquired in the target space to obtain a binaural audio signal, the method further includes: Control the robotic arm equipped with a measuring microphone to move along a preset grid path within the target listening area; During the movement of the robotic arm, the speakers within the target space are controlled to play test audio. The spatial audio signals received by the measuring microphone at each grid point in the preset grid path are recorded, and the head-related transfer function is calculated based on the spatial audio signals.
[0013] In one embodiment, the step of acquiring the spatial distribution information of multiple speakers in the target space and the positional information of the listener includes: Using a virtual reality headset worn by the listener, the system detects positioning devices installed on multiple speakers in the target space. Based on the detection results of the virtual reality headset, the spatial distribution information of each speaker in the target space is determined; The listener's pose information is determined based on the spatial distribution information and the position tracking sensors configured in the virtual reality headset.
[0014] Furthermore, to achieve the above objectives, this application also proposes an audio processing apparatus, which includes: The information acquisition module is used to acquire spatial distribution information of multiple speakers in the target space and the position and posture information of the listener, as well as the orientation information of the virtual sound source of the audio signal. The position and posture information includes the offset position of the listener relative to each of the speakers and the direction of the head. The binaural audio signal acquisition module is used to perform sound field reconstruction processing on the audio signal based on the orientation information of the virtual sound source, the pose information, and the head correlation transfer function acquired in the target space to obtain binaural audio signals. The signal correction module is used to correct the binaural audio signal based on the orientation information of the virtual sound source and the spatial distribution information when the virtual sound source is located in the lateral position of the listener, so as to obtain the target binaural audio signal.
[0015] In addition, to achieve the above objectives, this application also proposes an electronic device, which includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the audio processing method described above.
[0016] In addition, to achieve the above objectives, this application also proposes a storage medium, which is a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the audio processing method described above.
[0017] One or more technical solutions proposed in this application have at least the following technical effects: In this application, the spatial distribution information of multiple speakers in the target space, the orientation information of the listener, and the orientation information of the virtual sound source of the audio signal to be played are first obtained. Then, based on the head correlation transfer function, the orientation information of the virtual sound source, and the orientation information of the listener pre-collected in the target space, the sound field of the audio signal is reconstructed to ensure that the reconstructed binaural audio signal is adapted to the actual listening environment of the listener. Finally, when the virtual sound source is in the lateral orientation of the listener, the binaural audio signal is corrected according to the orientation information and spatial distribution information of the virtual sound source to obtain the target binaural audio signal. The correction process compensates for the lateral sound image localization deviation caused by the limitations of the human ear's physiological structure and the actual listening environment, thereby improving the localization clarity of the lateral sound image perceived by the listener. This makes the spatial orientation of the sound highly matched with the spatial orientation of the visual content in immersive entertainment scenarios such as watching movies and playing games, thus enhancing the immersive experience of the user. Attached Figure Description
[0018] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0019] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0020] Figure 1 This is a flowchart illustrating an embodiment of the audio processing method of this application. Figure 2This is a schematic diagram of a sound field indication scenario provided in Embodiment 1 of the audio processing method of this application; Figure 3 This is a schematic diagram of a regional positioning scenario provided in Embodiment 1 of the audio processing method of this application; Figure 4 This is a schematic diagram of the audio-visual correction scenario provided in Embodiment 2 of the audio processing method of this application; Figure 5 This is a schematic diagram of the audio processing flow provided in Embodiment 2 of the audio processing method of this application; Figure 6 This is a schematic diagram of the module structure of the audio processing device according to an embodiment of this application; Figure 7 This is a schematic diagram of the device structure of the hardware operating environment involved in the audio processing method in the embodiments of this application.
[0021] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0022] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.
[0023] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.
[0024] In entertainment scenarios such as watching movies and playing games, users are not only pursuing highly realistic visual experiences, but also increasingly demanding a more immersive auditory experience.
[0025] Currently, sound field reconstruction technology based on head-related transfer function is the mainstream method for realizing sound spatialization. However, due to the fact that the physiological structure of the human ear is most sensitive to sounds directly in front / behind, but has a weaker ability to perceive sounds from the side, coupled with the interference of room reflection, diffraction and other effects in the actual listening environment, the reconstructed side sound images generally have the problem of blurred positioning in terms of auditory perception.
[0026] In summary, improving the localization clarity of lateral acoustic images has become a pressing technical problem that needs to be solved in this field.
[0027] To address the aforementioned technical problems, this application's embodiments first acquire the spatial distribution information of multiple speakers within the target space, the listener's pose information, and the orientation information of the virtual sound source of the audio signal to be played. Then, based on the pre-acquired head-related transfer function within the target space, the orientation information of the virtual sound source, and the listener's pose information, the audio signal is reconstructed into a sound field, ensuring that the reconstructed binaural audio signal matches the listener's actual listening environment. Finally, when the virtual sound source is located laterally to the listener, the binaural audio signal is corrected based on the orientation and spatial distribution information of the virtual sound source to obtain the target binaural audio signal. This correction process compensates for the limitations of the human ear's physiological structure and the lateral sound image localization deviation caused by the actual listening environment, thereby improving the clarity of the lateral sound image localization perceived by the listener. This ensures that in immersive entertainment scenarios such as watching movies and playing games, the sense of sound orientation and the spatial direction of visual content can be highly matched, enhancing the immersive user experience.
[0028] It should be noted that the executing entity in this embodiment can be an electronic device with data processing, network communication, and program execution functions.
[0029] The following presents a first embodiment of the audio processing method of this application. (Refer to...) Figure 1 , Figure 1 This is a flowchart illustrating the first embodiment of the audio processing method of this application.
[0030] In this embodiment, the audio processing method includes steps S10 to S30: Step S10: Obtain spatial distribution information of multiple speakers in the target space and the position and pose information of the listener, as well as the orientation information of the virtual sound source of the audio signal. The position and pose information includes the offset position of the listener relative to each speaker and the direction of the head. It should be noted that the target space is the physical space where audio is played, such as a home theater room, VR viewing space, VR game space, etc.; the spatial distribution information includes parameters such as the three-dimensional coordinates, radiation orientation, and relative spacing of each speaker in the target space; the pose information describes the listener's spatial state, where the offset position refers to the spatial position deviation of the listener's head center relative to each speaker, and the head orientation refers to the spatial direction facing the listener's face; the virtual sound source refers to a sound that may not exist in the physical space but is perceived by the listener as emanating from a specific spatial location through audio processing technology, and the orientation information of the virtual sound source includes parameters such as the azimuth angle of the virtual sound source relative to the listener, used to define the virtual direction of sound emission.
[0031] A positioning device can be deployed on each speaker, which can emit detectable positioning signals. Simultaneously, the listener wears a device with positioning signal detection capabilities, such as a virtual reality headset, which receives the signals emitted by the positioning devices on each speaker and analyzes them to obtain the three-dimensional coordinates and relative spacing of each speaker, forming spatial distribution information. At the same time, the device worn by the listener is equipped with a position tracking sensor, which can capture the listener's spatial position and posture in real time. Combined with the acquired speaker spatial distribution information, the listener's offset position relative to each speaker can be obtained through coordinate conversion. Then, the head orientation can be determined through the posture sensor data built into the device. The directional information of the virtual sound source is preset by the audio playback system according to the current playback scene. For example, in a VR movie viewing scene, the directional information of the virtual sound source can correspond to the position of the sound-emitting object in the picture. In a VR game scene, it can be dynamically generated according to the game plot. As the listener moves or rotates in the target space, even if the directional information of the virtual sound source does not change, its sound image in the target space will change with the listener's movement or rotation.
[0032] For example, in a feasible implementation scenario, a schematic diagram of the sound field indication in the target space is shown below. Figure 2 As shown, 1#, 2#, 3#, and 4# are four speakers distributed within the target space. Using the listener's head position as the origin as the coordinate system, when the listener is facing directly forward (i.e., towards Pos1 in the figure), speaker 1# plays the left channel sound, speaker 2# plays the right channel sound, speaker 3# plays the left surround sound, and speaker 4# plays the right surround sound. The four speakers form a speaker system for playing immersive sound. Pos1 represents the current sound image position of the virtual sound source in the target space, and Pos2 represents the current sound image position of the virtual sound source in the target space after the listener turns around.
[0033] like Figure 3 As shown, taking position detection through a virtual reality headset worn by the listener as an example, each of the speakers 1#, 2#, 3#, and 4# in the figure has a positioning device. The offset position of the listener's head relative to each speaker is calculated by the positional offset between the virtual reality headset and the positioning device. When the listener's position moves from position θ1 to position θ2, the angle between each speaker and the listener's head changes.
[0034] Step S20: Based on the orientation information and pose information of the virtual sound source and the head-related transfer function acquired in the target space, the audio signal is processed for sound field reconstruction to obtain the binaural audio signal. It should be noted that the Head Related Transfer Function (HRTF) is essentially a mathematical function, specifically defined as the frequency domain acoustic transfer function of a sound source in a free field to both ears, usually expressed as HRTF(f,θ,φ), where f is the frequency, and θ and φ are the azimuth and elevation angles of the sound source. This function describes the frequency response changes, time delays, and amplitude attenuation caused by physiological structures such as the head and auricles during the propagation of sound waves from a point in space to the tympanic membranes of both ears. Sound field reconstruction processing is the process of converting ordinary audio signals into signals with spatial orientation characteristics. Binaural audio signals are two audio signals adapted to the auditory perception of the left and right ears respectively, which can create the illusion for the listener that the sound comes from a specific spatial direction.
[0035] HRTF data for different locations within the target space is collected beforehand to ensure that the data matches the acoustic environment of the target space. When the speaker plays sound, the actual perceived direction of the virtual sound source relative to the listener is determined by combining the acquired virtual sound source location information and the listener's pose information. Then, the corresponding transfer function parameters are matched from the collected HRTF data. Finally, the input audio signal is processed using the transfer function parameters. By simulating the propagation characteristics of sound in space, the audio signal is converted into binaural audio signals corresponding to the left and right ears respectively, so that the listener can initially perceive the sound from the location of the virtual sound source.
[0036] Step S30: When the virtual sound source is located to the side of the listener, the binaural audio signal is corrected according to the orientation information and spatial distribution information of the virtual sound source to obtain the target binaural audio signal.
[0037] It should be noted that the lateral orientation refers to the azimuth region at an angle of 90° ± 'a' to the listener's direct front. This region is where human hearing is weaker and positioning is prone to ambiguity. The value of 'a' can be based on the speaker distribution in the actual application scenario or can be defined by the user. For example, for the left lateral orientation, the range must be limited to the area between the left channel speaker and the left surround speaker. That is, as long as the sound source is located between the azimuth angles of the left channel speaker and the left surround speaker relative to the listener's direct front, it is considered a left lateral orientation. For the right lateral orientation, the range must be limited to the area between the right channel speaker and the right surround speaker. That is, as long as the sound source is located between the azimuth angles of the right channel speaker and the right surround speaker relative to the listener's direct front, it is considered a right lateral orientation. The correction process is the process of adjusting the time, intensity, and other parameters of the binaural audio signal to address the problem of lateral sound image positioning ambiguity. The target binaural audio signal is the binaural audio signal that, after correction processing, allows the listener to clearly perceive the location of the lateral virtual sound source.
[0038] During the playback of audio by the speaker, it is determined whether the location information of the virtual sound source falls within the lateral location range of the listener. When it is determined that the virtual sound source is in the lateral location, the deviations in the current binaural audio signals in terms of time difference and intensity difference are analyzed by combining the location information of the virtual sound source and the spatial distribution information of the speaker. Then, corresponding correction processing is formulated, such as adjusting the arrival time and signal intensity of the left and right ear audio signals, to obtain the target binaural audio signal.
[0039] In one feasible embodiment, steps A10 to A40 may be included before step S20: Step A10: Filter the left surround channel signal and the right surround channel signal in the audio signal to obtain the left filtered signal and the right filtered signal respectively; It should be noted that the left surround channel signal is the channel signal in the audio signal responsible for the left surround sound output; the right surround channel signal is the channel signal in the audio signal responsible for the right surround sound output; filtering is the process of selecting audio signals in a specific frequency band; the left filtered signal is the signal obtained after filtering the left surround channel signal; and the right filtered signal is the signal obtained after filtering the right surround channel signal.
[0040] The left and right surround channels are distinguished from the audio signal channels. Considering that the vocals and key surround sound components are mainly concentrated in the 200-8kHz frequency band, an FIR linear phase bandpass filter is selected to filter the left and right surround channel signals. This type of filter has linear phase characteristics, which can avoid signal phase distortion and ensure that the filtered signal maintains the original phase relationship. The components in the 200-8kHz frequency band in the left and right surround channel signals are filtered out by this filter, resulting in the left filtered signal (denoted as Ls_filt) and the right filtered signal (denoted as Rs_filt).
[0041] Step A20: Calculate the sum and difference signals of the left and right filtered signals; It should be noted that the sum signal is the signal obtained by directly superimposing the left and right filtered signals (denoted as Sum); the difference signal is the signal obtained by subtracting the left and right filtered signals (denoted as Diff).
[0042] The sum signal (Sum=Ls_filt+Rs_filt) and difference signal (Diff=Ls_filt-Rs_filt) are obtained through signal processing. Sum contains all common components (such as vocals, common noise / reverb) between the left and right surround channel signals, while Diff contains all different components (such as discrete surround effects and most non-vocal accompaniment) between the left and right surround channel signals.
[0043] Step A30: The absolute value of the difference signal is weighted based on a preset weighting factor, and the weighted result is subtracted from the sum signal to obtain the intermediate signal; First, calculate the absolute value of the difference signal, |Diff|, then multiply it by a preset weighting factor α to obtain the weighted difference signal. The preset weighting factor α ranges from 1.5 to 3.0, and its value can be adjusted according to the actual scenario. Subtract this weighted result from the sum signal to obtain the intermediate signal, i.e., intermediate signal = Sum - α. |Diff|. When Sum is strong and Diff is weak (i.e., the left and right signals are highly correlated, such as in a vocal scene), the middle signal is a large positive number; when Sum and Diff are of equal strength or Diff is stronger (i.e., the left and right signals are weakly correlated / uncorrelated, such as in an accompaniment or noise scene), the middle signal is a decimal or negative number.
[0044] Step A40: Perform half-wave rectification and smoothing filtering on the intermediate signal to obtain the enhanced signal component; It should be noted that half-wave rectification is a process that sets all negative values in the signal to zero; smoothing filtering is a process that smooths the signal and eliminates pulse interference; and signal enhancement is a signal obtained after the above processing, mainly composed of human voice components.
[0045] The intermediate signal undergoes half-wave rectification, and its calculation formula can be expressed as: Vocal_Enhanced = Max(0, Sum-α) The |Diff| signal, through half-wave rectification, sets the negative values in the intermediate signal to zero, retaining only the highly correlated positive values. This is precisely the starting and core energy portion of the vocal syllable, forming a pulse train. For example, at the beginning of a vocal syllable, Sum suddenly increases while |Diff| is relatively small, resulting in a positive value and a sudden increase in output; during syllable gaps or consonant segments, Sum decreases or |Diff| is relatively large, resulting in a negative value and a sudden drop in output to zero; in strong accompaniment or noise segments, |Diff| may exceed Sum, resulting in a negative value and a zero output. This alternating pattern forms the pulse train.
[0046] A low-pass filter with a cutoff frequency of 8kHz is used to smooth the pulse train, thereby eliminating the discrete characteristics of the pulse train and restoring the continuous audio signal. This signal is the enhanced signal component, which mainly enhances the human voice component.
[0047] Based on this, step S20 may include steps A50~A60: Step A50: The enhanced signal component is superimposed on the main channel signal in the audio signal to obtain a new audio signal; It should be noted that the main channel signal is the channel signal in the audio signal that is responsible for the main sound output, including the left main channel signal and the right main channel signal.
[0048] First, the left and right channel signals are distinguished from the audio signal. These two channels are the main audio channels, responsible for outputting core content such as vocals and the main melody. Then, a signal superposition algorithm is used to superimpose the previously obtained enhanced signal components with the left and right channel signals respectively, so that the enhanced vocal components are integrated into the main channel signal, making up for the lack of vocal components in the main channel and obtaining a new audio signal.
[0049] Step A60: Based on the orientation information and pose information of the virtual sound source, as well as the head-related transfer function acquired in the target space, the new audio signal is processed for sound field reconstruction to obtain the binaural audio signal.
[0050] HRTF data for different locations within the target space is collected beforehand to ensure that the data matches the acoustic environment of the target space. When the speaker plays sound, the actual perceived direction of the virtual sound source relative to the listener is determined by combining the acquired virtual sound source location information and the listener's pose information. Then, the corresponding transfer function parameters are matched from the collected HRTF data. Finally, the new audio signal after the enhanced signal components are superimposed is processed using the transfer function parameters. By simulating the propagation characteristics of sound in space, the audio signal is converted into binaural audio signals corresponding to the left and right ears respectively, so that the listener can initially perceive the sound from the location of the virtual sound source.
[0051] Therefore, by superimposing the vocal components extracted from the left and right surround channel signals onto the main channel, the clarity of the vocals in the main channel is improved. While ensuring the spatialization effect of the sound field, the layering of the audio experience is further enhanced, allowing listeners to clearly capture vocals while perceiving the spatial surround effect, thus optimizing the auditory experience.
[0052] In one feasible embodiment, step S20 may include steps S201 to S203: Step S201: Perform spherical harmonic decomposition on the head-related transfer function collected in the target space to obtain the spherical harmonic coefficients; The HRTF functions collected within the target space are acquired. Each HRTF function contains transfer functions corresponding to different spatial directions. A spherical harmonic decomposition algorithm is used to process the HRTF in each direction, representing it as a linear combination of spherical harmonic basis functions of different orders. These spherical harmonic basis functions are standardized functions defined based on the spherical coordinate system and can comprehensively cover spatial directions. Through this decomposition process, the spherical harmonic coefficients corresponding to each HRTF are obtained. The spherical harmonic coefficients of all directions are integrated to form a reusable HRTF basic database. Subsequently, HRTF-related data in any direction can be quickly synthesized using this database.
[0053] The expression for spherical harmonic decomposition using HRTF is: ; Where HRTF(θ,φ) is the head-related transfer function corresponding to a certain spatial direction (θ,φ) in the target space, θ is the azimuth angle, φ is the elevation angle, l is the spherical harmonic order, and m is the degree corresponding to the order. For spherical harmonic basis functions, The obtained spherical harmonic coefficients comprehensively characterize the spatial distribution characteristics of HRTF within the target space.
[0054] Step S202: Determine the position information of the virtual sound source relative to the listener within the target space based on the orientation and pose information of the virtual sound source. It should be noted that the location information refers to the spatial relative position parameters of the virtual sound source relative to the listener, including relative angles and relative distances. It's worth mentioning that the azimuth information of the virtual sound source focuses on describing the directional attributes of the virtual sound source relative to the listener, only including parameters such as azimuth and elevation angles to define the virtual direction of sound emission, thus clarifying the sound's orientation. The location information, however, is a spatial relative attribute derived from the azimuth information combined with the listener's posture information. It includes not only the relative angle between the virtual sound source and the listener but also parameters such as relative distance. The relative distance can be set according to the application scenario, such as a preset virtual distance between the sound-emitting object on the screen and the viewer in a VR movie viewing scenario, a preset interactive distance set according to the game's plot in a VR game scenario, or determined based on the maximum distribution range of speakers within the target space. This allows for accurate characterization of the spatial relationship between the virtual sound source and the listener.
[0055] Step S203: Based on the location information and spherical harmonic coefficients, perform sound field reconstruction processing on the audio signal to obtain the binaural audio signal.
[0056] Based on the relative angle between the virtual sound source and the listener's position, correlation coefficients are retrieved from the spherical harmonic coefficient database, and the corresponding spherical harmonic basis function is determined by combining the relative angle. Then, the head-related impulse response (HRIR) corresponding to that direction is synthesized through the operation of the spherical harmonic coefficients and the spherical harmonic basis function. This HRIR can reflect the transmission characteristics of sound from that direction to both ears. The synthesized HRIR is convolved with the input audio signal. This operation simulates the process of sound propagating from the virtual sound source to both ears in space, converting the ordinary audio signal into a binaural audio signal with spatial orientation characteristics, so that the listener can perceive the sound from the direction of the target virtual sound source.
[0057] Therefore, the spatial characteristics of HRTF can be efficiently extracted through spherical harmonic decomposition. The spherical harmonic coefficients obtained by decomposition support the sound field reconstruction of virtual sound sources in any direction without the need to measure HRTF separately for each direction, which simplifies the reconstruction process. The sound field reconstruction is achieved by combining the direction information of the target virtual sound source, which improves the flexibility and versatility of the sound field reconstruction and ensures the accuracy of virtual sound source positioning in panoramic sound scenes.
[0058] In one feasible embodiment, step S203 may include steps S2031 to S2033: Step S2021: Calculate the spherical harmonic basis functions based on the location information; The azimuth angle θ and elevation angle φ of the virtual sound source relative to the listener are extracted from the location information to clarify the order requirement of the spherical harmonic basis function. According to the mathematical definition of the spherical harmonic function, the values of the azimuth angle θ and elevation angle φ are substituted to calculate the spherical harmonic basis function of the corresponding order. This spherical harmonic basis function can accurately describe the spatial characteristics of the relative direction.
[0059] Step S2022: Multiply the spherical harmonic coefficients by the spherical harmonic basis functions and then sum them to obtain the head-related impulse response signal; Calling spherical harmonic coefficients , with spherical harmonic basis functions After multiplying each term and summing them, this summation process is an efficient interpolation operation that can accurately reconstruct the head-related impulse response (HRIR) signal corresponding to the direction of the target virtual sound source. The head-related impulse response signal is a signal that characterizes the impulse response of sound propagation in the direction of the target virtual sound source.
[0060] Step S2023: Convolve the head-related impulse response signal with the audio signal to obtain the binaural audio signal.
[0061] The audio signal used for convolution can be a mono signal (Mono(t)) or a signal after signal enhancement and superposition. The spatial propagation characteristics of HRIR (Head Related Impulse Response, i.e. the temporal response corresponding to HRTF) are imparted to the audio signal through convolution, so that the output binaural audio signal can simulate the effect of sound propagating from the target virtual sound source to both ears.
[0062] The formula for calculating binaural audio signals is: ,in, This represents the binaural audio signal at time t, including the left and right ear channels, used to simulate spatial auditory perception. Let L be the mono audio signal at time t, and L be the highest order of the spherical harmonic function. The spherical harmonic basis functions corresponding to order l, order m, and the azimuth angle θ and elevation angle φ of the virtual sound source.
[0063] Thus, by summing the spherical harmonic basis functions and spherical harmonic coefficients corresponding to the direction of the target virtual sound source, accurate reconstruction of HRIR is achieved; by efficiently fusing audio signals and spatial propagation characteristics through convolution operations, binaural audio signals of virtual sound sources in any direction can be generated without complex on-site measurements, improving the efficiency and accuracy of sound field reconstruction and making sound positioning in panoramic sound scenes more realistic.
[0064] Based on the first embodiment of this application, in the second embodiment of this application, the content that is the same as or similar to that in the first embodiment described above can be referred to the above description and will not be repeated hereafter. Based on this, step S30 may include steps S301 to S304: Step S301: Based on the orientation information and spatial distribution information of the virtual sound source, calculate the delay compensation amount of the target loudspeaker corresponding to the main channel in the binaural audio signal. It should be noted that the delay compensation amount is a compensation value that adjusts the output signal time of the target speaker, and is used to correct the arrival time difference of the audio signals in the left and right ears.
[0065] The location information of the virtual sound source is clearly defined, and its specific lateral angle relative to the listener is determined, such as 90° to the left. Then, based on the spatial distribution information, the position parameters of the target speakers (such as the two speakers located on the left side of the listener) are obtained, the distance between the virtual sound source and each target speaker is calculated, and the delay compensation amount corresponding to each target speaker is calculated by the formula "delay compensation amount = distance difference / speed of sound". This ensures that the audio signals of the left and right ears can reach the listener's ears according to the time pattern of the real sound field.
[0066] Step S302: Perform binaural time difference correction on the binaural audio signal according to the delay compensation amount; It should be noted that when a person sits in the center and faces directly forward to listen to the sound, the listener's sense of the sound image position in the lateral direction will be blurred due to factors such as head diffraction or room effect. Binaural time difference correction is a process of adjusting the time difference between the audio signals of the left and right ears reaching the listener's ears to make it conform to the law of sound propagation in a real sound field.
[0067] Based on the calculated delay compensation amount of each target loudspeaker, the delay time of the output signal of each target loudspeaker is determined. Based on the determined delay time, the signal of the corresponding channel in the binaural audio signal is delayed and adjusted. For example, for a virtual sound source at 90° on the left, the signals output by the two target loudspeakers on the left side of the listener are delayed and adjusted so that the signal received by the left ear and the signal received by the right ear form a time difference that conforms to the real sound field. This time difference can help the listener perceive the position of the lateral sound source more clearly and make up for the time perception deviation caused by head diffraction and room effect.
[0068] Step S303: Calculate the binaural intensity difference adjustment factor based on the orientation information of the virtual sound source; It should be noted that the binaural intensity difference adjustment factor is a coefficient used to adjust the ratio of the intensity of the audio signals in the left and right ears, including the left ear adjustment factor (denoted as a) and the right ear adjustment factor (denoted as b); the binaural intensity difference is the difference in intensity of the same sound source signal received by the left and right ears in a real sound field.
[0069] Based on the location information of the virtual sound source, its lateral position (left or right) is determined. Then, the signal intensities of the virtual sound source received by the left and right ears under this lateral position are obtained, and the energy ratios GL and GR of the left and right ears are calculated. GL = left ear lateral sound source signal intensity / right ear lateral sound source signal intensity, and GR = right ear lateral sound source signal intensity / left ear lateral sound source signal intensity. The binaural intensity difference adjustment factor is then calculated based on this energy ratio. The specific calculation formulas are a = GL / (GL + GR) and b = GR / (GL + GR). This adjustment factor can reflect the intensity perception ratio of the left and right ears of the lateral sound source in the real sound field.
[0070] Step S304: Based on the binaural intensity difference adjustment factor, the binaural audio signal after binaural time difference correction is subjected to binaural intensity difference correction to obtain the target binaural audio signal.
[0071] Acquire the left and right ear audio signals after binaural time difference correction; multiply the left ear adjustment factor b with the corrected left ear audio signal, and multiply the right ear adjustment factor a with the corrected right ear audio signal, i.e., corrected left ear signal = b × time difference corrected left ear signal, corrected right ear signal = a × time difference corrected right ear signal; through this intensity adjustment, the ratio of signal intensity received by the left and right ears conforms to the real sound field law. For example, for a virtual sound source at 90° to the left, the signal intensity received by the left ear after adjustment is stronger than that of the right ear. This intensity difference, combined with the previously corrected time difference, can significantly improve the localization clarity of the lateral sound image, and finally obtain the target binaural audio signal.
[0072] For example, taking the location information of the virtual sound source set at a 90° angle to the left of the listener as an example, such as... Figure 4 As shown, Figure 4 The diagram shows a speaker layout scenario in a target space. Speaker #1 is the left channel speaker, #2 is the right channel speaker, #3 is the left surround speaker, and #4 is the right surround speaker. L, R, Ls, and Rs represent the virtual sound image positions. The X-axis is the vertical axis of the target space, representing the direction the person is facing. The Y-axis is the horizontal central axis of the target space. In this scenario, the virtual sound source is positioned directly to the left of the person, meaning both the virtual sound source and the person are on the Y-axis. θt represents the angle between the virtual sound image positions (L, R, Ls, Rs) and the virtual sound source. φ is the angle between the speaker and the Y-axis. The value of θt ranges from [-φ, +φ]. LE is the length of the speaker distribution in the space. Figure 4 The spacing between speakers #1 and #3, and between speakers #2 and #4, is represented by WI, which is the width of the speaker distribution in the space. Figure 4 The distance between speakers #1 and #2 and between speakers #3 and #4 is defined as follows: d1 is the distance between speaker #3 and the virtual sound image position L, d2 is the distance between speaker #1 and the virtual sound image position L, and the distance between speaker #3 and the virtual sound image position R is equal to d2, and the distance between speaker #1 and the virtual sound image position R is equal to d1.
[0073] At this point, the delay applied to speakers #3 and #1 for the left channel signal is: ; ; For the right channel signal, the delay applied to speakers #3 and #1 is: ; ; The delayed signals are superimposed as follows to obtain the time-difference-corrected binaural audio signals: Binaural(L)=Binaural(L-L3)+Binaural(R-R3); Binaural(R)=Binaural(R-R1)+Binaural(L-L1).
[0074] This process initially corrects the lateral acoustic image localization deviation by compensating for the sound propagation time difference caused by the listener's shifted position.
[0075] Applying personalized frequency equalization to the signal in the lateral direction, for a sound source at a 90° angle to the left: the closer ear receives stronger energy, and the farther ear receives weaker energy. Therefore: Energy ratio between left and right ears: GL=Binaural(L90°) / Binaural(R90°); GR=Binaural(R90°) / Binaural(L90°); Wherein, Binaural(L90°) represents the signal intensity of the sound source at 90° to the left ear, and Binaural(R90°) represents the signal intensity of the sound source at 90° to the right ear.
[0076] Then, based on the energy ratio, the binaural intensity difference adjustment factors a and b are obtained to adapt the lateral acoustic characteristics: a = GL / (GL + GR); b = GR / (GL + GR); The binaural intensity difference adjustment factors a and b can match the human ear's perception pattern in lateral orientation, where sound energy is stronger closer to the ear and weaker farther away.
[0077] The time-difference corrected binaural audio signals Binaural(L) and Binaural(R) are weighted and corrected with binaural intensity difference adjustment factors a and b to match the energy perception characteristics of the lateral orientation. The calculation formula is as follows: Binaural(L)=b Binaural(L); Binaural(R)=a Binaural(R); The corrected Binaural(L) and Binaural(R) are the target binaural audio signals, which can accurately reproduce the sound image perception effect in the lateral direction.
[0078] Thus, by using both time difference and intensity difference correction, a comprehensive lateral acoustic image correction mechanism is constructed, improving the localization clarity of lateral acoustic images.
[0079] In one feasible embodiment, steps B10 to B30 may be included before step S20: Step B10: Control the robotic arm equipped with the measuring microphone to move along a preset grid path within the target listening area; It should be noted that the measuring microphone is a device used to collect spatial audio signals, possessing high sensitivity and wide frequency response characteristics, capable of accurately capturing sound signals; the robotic arm is a multi-degree-of-freedom mechanical device that can move along a preset path, used to carry the measuring microphone; the target listening area is the spatial area where the listener's head may be located, such as the seating area when watching a movie in a home theater; the preset grid path is a high-density grid-like moving path covering the target listening area, with the grid density set according to the measurement accuracy requirements to ensure comprehensive collection of sound field data within the area.
[0080] Based on the size of the target space and the range of the target listening area, a preset grid path is planned. This path needs to cover all possible positions where the listener's head may appear. The grid point spacing is set to a value that can ensure measurement accuracy. The measuring microphone is fixedly installed at the end of the robotic arm, ensuring that the microphone's acquisition direction is facing the speaker array. Through the robotic arm's control system, the robotic arm is driven to move according to the preset grid path. During the movement, the robotic arm's positioning system provides real-time feedback on the current position information to ensure that the movement trajectory is consistent with the preset path.
[0081] Step B20: During the movement of the robotic arm, control each speaker in the target space to play test audio. It should be noted that the movement of the robotic arm is the process of the robotic arm carrying the measuring microphone moving along a preset grid path; the test audio is a specific audio signal used to collect sound field data. Uncorrelated noise can be selected as the test audio. The frequency components of this type of audio signal are evenly distributed, and there is no correlation between the test audio played by different speakers, which can avoid mutual interference.
[0082] As the robotic arm begins to move, all speakers in the target space are controlled to synchronously play test audio in the form of unrelated noise. Throughout the entire movement of the robotic arm, the test audio from each speaker is continuously played, and parameters such as playback volume and frequency characteristics remain stable to ensure that the collected sound field data is consistent and comparable.
[0083] Step B30: Record the spatial audio signals received by the measuring microphone at each grid point in the preset grid path, and calculate the head correlation transfer function based on the spatial audio signals.
[0084] When the robotic arm moves to each grid point of the preset grid path, it simultaneously triggers the measurement microphone to acquire signals and record the spatial audio signal received at that position. The acquisition duration is set according to the signal analysis requirements, such as 1 second. After acquisition, the robotic arm continues to move to the next grid point and repeats the acquisition process until all grid points have acquired signals. After acquisition, the spatial audio signal of each grid point is analyzed. Combined with the original signal of the test audio played by each speaker, the transfer function of each grid point position relative to each speaker is calculated. Then, based on the physiological characteristics of the human ear, the transfer function is converted into a head-related transfer function.
[0085] Thus, by using a robotic arm equipped with a microphone to perform high-density grid measurement, full coverage of the target listening area was achieved, avoiding the errors and limitations of manual measurement. This resulted in higher accuracy for HRTF, ensuring that the audio processing effect was highly adapted to the target space.
[0086] In one feasible embodiment, the step of "acquiring spatial distribution information of multiple speakers in the target space and the pose information of the listener" in step S10 may include steps S101 to S103: Step S101: Using the virtual reality headset worn by the listener, the positioning devices installed on multiple speakers in the target space are detected. Virtual reality headsets (VR devices) are terminal devices worn by listeners and equipped with positioning and detection functions. Each speaker is equipped with a positioning device, such as an infrared positioning module or a UWB (Ultra Wide Band) positioning module. This positioning device can emit specific detection signals, such as infrared signals or UWB signals. The VR device receives these signals to detect the position of the positioning device.
[0087] Step S102: Based on the detection results of the virtual reality headset, determine the spatial distribution information of each speaker in the target space; The VR device calculates the three-dimensional coordinates of each speaker based on parameters such as the signal strength and propagation time detected by the positioning device, combined with its own spatial positioning algorithm. It then integrates the three-dimensional coordinates of all speakers to form spatial distribution information that reflects the relative positional relationship of the speakers and the overall layout.
[0088] Step S103: Determine the listener's pose information based on the spatial distribution information and the position tracking sensors configured in the virtual reality headset.
[0089] It should be noted that the position tracking sensors configured in virtual reality headsets are sensors used to capture the device's own spatial position and attitude, such as gyroscopes, accelerometers, and magnetometers.
[0090] The position tracking sensor in the virtual reality headset collects the device's motion data in real time, including position change data and posture change data. Combined with the device's initial position coordinates, the algorithm calculates the device's real-time three-dimensional coordinates and posture in the target space coordinate system. Since the device is worn by the listener, the device's real-time three-dimensional coordinates can be regarded as the three-dimensional coordinates of the listener's head center, and the device's posture is the listener's head orientation. Combined with the previously determined three-dimensional coordinates of each speaker, the offset position of the listener's head center relative to each speaker is calculated through coordinate conversion. Finally, the calculated offset position and head orientation are integrated to form the listener's pose information, which is updated in real time to ensure that the listener's position movement and posture changes can be reflected in a timely manner.
[0091] Therefore, by using VR devices to detect speaker positioning devices, there is no need to set up additional complex posture detection equipment, which simplifies the system configuration, enables the acquisition of relative positional relationships in real time and accurately, avoids the complex calculations of traditional posture detection methods, and ensures that audio processing dynamically adapts to the listener's head movement.
[0092] For example, to help understand the implementation flow of the audio processing method obtained by combining this embodiment with the first embodiment described above, please refer to... Figure 5 , Figure 5 A simplified flowchart of an audio processing method is provided, specifically: Taking audio processing in a VR movie-watching scenario as an example, the speaker control and distribution process is executed first. Specifically, the positioning devices on each speaker in the target space are detected through the VR headset worn by the listener to determine the spatial distribution information of each speaker. This spatial distribution information is then combined with the sound field area position (i.e., the range of the VR movie-watching target listening area). Without the need for additional posture detection equipment, the relative position of the speakers and the listener can be obtained directly through the VR device, thus simplifying the calculation.
[0093] Then, the spatial data acquisition process is performed. Specifically, the robotic arm equipped with a measuring microphone is controlled to move along a preset grid path within the target listening area of the VR viewing area. At the same time, each speaker is controlled to synchronously play unrelated noise test audio, record the spatial audio signal at each grid point, and calculate the HRTF covering the target area based on the signal. This process generates high-density IR data (i.e., high-density head-related impulse response data), which completely covers all key listening positions within the target area, thereby improving the data density.
[0094] Then, the audio feature extraction process is performed. Specifically, for the VR movie audio signal, the left surround channel and the right surround channel signals are extracted, and then FIR (Finite Impulse Response) linear phase bandpass filtering, intermediate signal calculation, signal half-wave rectification processing based on weight factors, and smoothing filtering processing are performed in sequence to finally obtain the human voice component, that is, the purified and enhanced human voice audio signal, realizing the separation and optimization of human voice audio and surround sound components.
[0095] Then, the sound field reconstruction process is performed. Specifically, the HRTF data is decomposed into spherical harmonics to obtain a database of spherical harmonic coefficients. Combined with the directional information of virtual sound sources in the VR film (such as the direction of the character's voice in the film), the spherical harmonic basis functions of the corresponding directions are calculated. The spherical harmonic coefficients are summed with the basis functions to reconstruct the head-related impulse response (HRIR) of the target direction. Then, the HRIR is convolved with the audio signal superimposed with enhanced human voice components to complete the sound field reconstruction, realizing accurate spatial sound field mapping of any virtual sound source direction.
[0096] Finally, the sound field correction process is performed. Specifically, based on the real-time offset position of the listener relative to each speaker obtained by the VR device, the reconstructed binaural audio signal is subjected to binaural time difference correction. Then, binaural spectral data in the lateral direction is extracted from HRTF, the energy ratio of the left and right ears and the binaural intensity difference adjustment factor are calculated, and the intensity difference correction is performed on the signal after time difference correction. This achieves personalized enhancement of the lateral sound field, thereby correcting the sound image blurring caused by head diffraction and room effect, ensuring accurate lateral sound image positioning, and improving the clarity of the lateral sound field.
[0097] Finally, the processing results of the above steps are integrated through sound field linkage, and the corrected target binaural audio signals are distributed to each speaker in the target space for playback, ultimately achieving an immersive auditory experience in VR movie viewing scenarios with clear and prominent human voices and accurate positioning of sound images on the left and right sides.
[0098] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the audio processing method of this application. Any simple modifications based on this technical concept are within the protection scope of this application.
[0099] This application also provides an audio processing device, please refer to... Figure 6 The audio processing device includes: The information acquisition module 10 is used to acquire the spatial distribution information of multiple speakers in the target space and the position and posture information of the listener, as well as the orientation information of the virtual sound source of the audio signal. The position and posture information includes the offset position of the listener relative to each speaker and the direction of the head. The binaural audio signal acquisition module 20 is used to perform sound field reconstruction processing on the audio signal based on the orientation information and pose information of the virtual sound source and the head correlation transfer function acquired in the target space to obtain the binaural audio signal. The signal correction module 30 is used to correct the binaural audio signal based on the orientation and spatial distribution information of the virtual sound source when the virtual sound source is located to the side of the listener, so as to obtain the target binaural audio signal.
[0100] Optionally, the binaural audio signal acquisition module 20 is also used for: The left surround channel signal and the right surround channel signal in the audio signal are filtered separately to obtain the left filtered signal and the right filtered signal; Calculate the sum and difference signals of the left and right filtered signals; The absolute value of the difference signal is weighted based on a preset weighting factor, and the weighted result is subtracted from the sum signal to obtain the intermediate signal; The intermediate signal is subjected to half-wave rectification and smoothing filtering to obtain the enhanced signal component. The enhanced signal component is superimposed on the main channel signal in the audio signal to obtain a new audio signal; Based on the orientation and pose information of the virtual sound source, as well as the head-related transfer function acquired in the target space, the new audio signal is processed for sound field reconstruction to obtain the binaural audio signal.
[0101] Optionally, the binaural audio signal acquisition module 20 is also used for: The head-related transfer function collected in the target space is subjected to spherical harmonic decomposition to obtain the spherical harmonic coefficients; Based on the orientation and pose information of the virtual sound source, determine the position information of the virtual sound source relative to the listener in the target space; Based on the location information and spherical harmonic coefficients, the audio signal is processed to reconstruct the sound field, thus obtaining the binaural audio signal.
[0102] Optionally, the binaural audio signal acquisition module 20 is also used for: Calculate the spherical harmonic basis functions based on the location information; The head-related impulse response signal is obtained by multiplying the spherical harmonic coefficients by the spherical harmonic basis functions and then summing the results. The head-related impulse response signal is convolved with the audio signal to obtain the binaural audio signal.
[0103] Optionally, the signal correction module 30 is also used for: Based on the location and spatial distribution information of the virtual sound source, the delay compensation amount of the target loudspeaker corresponding to the main channel in the binaural audio signal is calculated. Binaural time difference correction is performed on the binaural audio signals based on the delay compensation amount; The binaural intensity difference adjustment factor is calculated based on the location information of the virtual sound source; Based on the binaural intensity difference adjustment factor, the binaural audio signal after binaural time difference correction is subjected to binaural intensity difference correction to obtain the target binaural audio signal.
[0104] Optionally, the audio processing device further includes a data acquisition module (not shown), which is used for: Control the robotic arm equipped with a measuring microphone to move along a preset grid path within the target listening area; During the movement of the robotic arm, the speakers in the target space are controlled to play test audio. Record the spatial audio signals received by the measurement microphone at each grid point in the preset grid path, and calculate the head correlation transfer function based on the spatial audio signals.
[0105] Optionally, the offset position acquisition module 10 is also used for: Using a virtual reality headset worn by the listener, the system detects positioning devices installed on multiple speakers in the target space. Based on the detection results from the virtual reality headset, the spatial distribution information of each speaker in the target space is determined; Based on spatial distribution information and position tracking sensors configured in the virtual reality headset, the listener's pose information is determined.
[0106] This application provides an electronic device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the status indication method described above.
[0107] The following is for reference. Figure 7 The diagram illustrates an electronic device suitable for implementing embodiments of this application. Figure 7 The electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.
[0108] like Figure 7As shown, the electronic device may include a processing unit 1001 (e.g., a DSP processor), which can perform various appropriate actions and processes according to a program stored in a read-only memory 1002 or a program loaded from a storage device 1003 into a random access memory 1004. The random access memory 1004 also stores various programs and data required for the operation of the electronic device. The processing unit 1001, the read-only memory 1002, and the random access memory 1004 are interconnected via a bus 1005. An input / output interface 1006 is also connected to the bus. Typically, the following systems can be connected to the input / output interface 1006: an input device 1007 including, for example, a microphone, an accelerometer, etc.; an output device 1008 including, for example, a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 allows the electronic device to exchange data wirelessly or via wired communication with other devices. Although electronic devices with various systems are shown in the figures, it should be understood that it is not required to implement or possess all the systems shown. It can be implemented alternatively or with more or fewer systems.
[0109] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from read-only memory 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.
[0110] Compared with the prior art, the beneficial effects of the electronic device provided in this application embodiment are the same as those of the audio processing method provided in the above embodiment, and other technical features in the electronic device are the same as those disclosed in the method of the previous embodiment, and will not be repeated here.
[0111] It should be understood that the various parts disclosed in the embodiments of this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.
[0112] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0113] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, which are used to execute the audio processing method in the above embodiments.
[0114] The computer-readable storage medium provided in this application embodiment may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.
[0115] The aforementioned computer-readable storage medium may be included in an electronic device or may exist independently without being assembled into an electronic device.
[0116] The aforementioned computer-readable storage medium carries one or more programs, which, when executed by an electronic device, cause the electronic device to perform the functions defined in the methods of the embodiments disclosed in this application.
[0117] Computer program code for performing the operations of this application can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, and conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0118] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0119] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.
[0120] The readable storage medium provided in this application embodiment is a computer-readable storage medium, which stores computer-readable program instructions (i.e., computer programs) for performing the above-described audio processing method. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application embodiment are the same as the beneficial effects of the audio processing method provided in the above-described embodiments, and will not be repeated here.
[0121] This application also provides a computer program product, including a computer program, which, when executed by a processor, implements the steps of the status indication method described above. Compared with the prior art, the beneficial effects of the computer program product provided in this application are the same as those of the status indication method provided in the above embodiments, and will not be repeated here.
[0122] The above are only some embodiments of this application and do not limit the patent scope of this application. All equivalent structural transformations made under the technical concept of this application and using the contents of the specification and drawings of this application, or direct / indirect applications in other related technical fields, are included in the patent protection scope of this application.
Claims
1. An audio processing method, characterized in that, The audio processing method includes: The spatial distribution information of multiple speakers in the target space and the position and posture information of the listener are obtained, as well as the orientation information of the virtual sound source of the audio signal. The position and posture information includes the offset position of the listener relative to each of the speakers and the direction of the head. Based on the orientation information of the virtual sound source, the pose information, and the head-related transfer function acquired in the target space, the audio signal is subjected to sound field reconstruction processing to obtain a binaural audio signal. When the virtual sound source is positioned laterally to the listener, the binaural audio signal is corrected based on the orientation information of the virtual sound source and the spatial distribution information to obtain the target binaural audio signal.
2. The audio processing method as described in claim 1, characterized in that, Before the step of performing sound field reconstruction processing on the audio signal based on the orientation information of the virtual sound source, the pose information, and the head-related transfer function acquired in the target space to obtain the binaural audio signal, the method further includes: The left surround channel signal and the right surround channel signal in the audio signal are filtered respectively to obtain the left filtered signal and the right filtered signal; Calculate the sum and difference signals of the left filtered signal and the right filtered signal; The absolute value of the difference signal is weighted based on a preset weighting factor, and the weighted result is subtracted from the sum signal to obtain the intermediate signal; The intermediate signal is subjected to half-wave rectification and smoothing filtering to obtain the enhanced signal component; The step of performing sound field reconstruction processing on the audio signal based on the orientation information of the virtual sound source, the pose information, and the head-related transfer function acquired in the target space to obtain binaural audio signals includes: The enhanced signal component is superimposed on the main channel signal in the audio signal to obtain a new audio signal; Based on the orientation information of the virtual sound source, the pose information, and the head-related transfer function acquired in the target space, the new audio signal is subjected to sound field reconstruction processing to obtain a binaural audio signal.
3. The audio processing method as described in claim 1, characterized in that, The step of performing sound field reconstruction processing on the audio signal based on the orientation information of the virtual sound source, the pose information, and the head-related transfer function acquired in the target space to obtain binaural audio signals includes: The head-related transfer function acquired in the target space is subjected to spherical harmonic decomposition to obtain the spherical harmonic coefficients; Based on the orientation information and pose information of the virtual sound source, the position information of the virtual sound source relative to the listener within the target space is determined; Based on the location information and the spherical harmonic coefficients, the audio signal is subjected to sound field reconstruction processing to obtain a binaural audio signal.
4. The audio processing method as described in claim 3, characterized in that, The step of performing sound field reconstruction processing on the audio signal based on the location information and the spherical harmonic coefficients to obtain a binaural audio signal includes: Calculate the spherical harmonic basis functions based on the location information; The head-related impulse response signal is obtained by multiplying the spherical harmonic coefficients by the spherical harmonic basis functions and then summing the results. The head-related impulse response signal is convolved with the audio signal to obtain the binaural audio signal.
5. The audio processing method as described in claim 1, characterized in that, The step of correcting the binaural audio signal based on the location information of the virtual sound source and the spatial distribution information to obtain the target binaural audio signal includes: Based on the location information of the virtual sound source and the spatial distribution information, the delay compensation amount of the target loudspeaker corresponding to the main channel in the binaural audio signal is calculated. The binaural time difference is corrected for the binaural audio signal based on the delay compensation amount. The binaural intensity difference adjustment factor is calculated based on the orientation information of the virtual sound source; Based on the binaural intensity difference adjustment factor, the binaural audio signal after binaural time difference correction is subjected to binaural intensity difference correction to obtain the target binaural audio signal.
6. The audio processing method as described in claim 1, characterized in that, Before the step of performing sound field reconstruction processing on the audio signal based on the orientation information of the virtual sound source, the pose information, and the head-related transfer function acquired in the target space to obtain the binaural audio signal, the method further includes: Control the robotic arm equipped with a measuring microphone to move along a preset grid path within the target listening area; During the movement of the robotic arm, the speakers within the target space are controlled to play test audio. The spatial audio signals received by the measuring microphone at each grid point in the preset grid path are recorded, and the head-related transfer function is calculated based on the spatial audio signals.
7. The audio processing method according to any one of claims 1 to 6, characterized in that, The steps of acquiring the spatial distribution information of multiple speakers in the target space and the position and pose information of the listener include: Using a virtual reality headset worn by the listener, the system detects positioning devices installed on multiple speakers in the target space. Based on the detection results of the virtual reality headset, the spatial distribution information of each speaker in the target space is determined; The listener's pose information is determined based on the spatial distribution information and the position tracking sensors configured in the virtual reality headset.
8. An audio processing apparatus, characterized in that, The audio processing device includes: The information acquisition module is used to acquire spatial distribution information of multiple speakers in the target space and the position and posture information of the listener, as well as the orientation information of the virtual sound source of the audio signal. The position and posture information includes the offset position of the listener relative to each of the speakers and the direction of the head. The binaural audio signal acquisition module is used to perform sound field reconstruction processing on the audio signal based on the orientation information of the virtual sound source, the pose information, and the head correlation transfer function acquired in the target space to obtain binaural audio signals. The signal correction module is used to correct the binaural audio signal based on the orientation information of the virtual sound source and the spatial distribution information when the virtual sound source is located in the lateral position of the listener, so as to obtain the target binaural audio signal.
9. An electronic device, characterized in that, The electronic device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the audio processing method as described in any one of claims 1 to 7.
10. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, it implements the steps of the audio processing method as described in any one of claims 1 to 7.