A spatial audio processing method and apparatus
Patent Information
- Application Number
- CN202311172643.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-11
- Publication Date
- 2026-08-28
- Estimated Expiration
- 2043-09-11
AI Technical Summary
[0002]空间音频处理是虚拟、增强和混合现实应用中的一种重要的音频信号处理技术,空间音频技术可以重放音频的定位信息,合成逼真的声学环境,创造沉浸式的音频体验,空间听觉研究表明,人对声音的空间感知主要受到双耳因素影响,基于双耳声信号的虚拟重放技术,主要利用头相关传输函数(Head-Related Transfer Function,HRTF)或双耳房间脉冲响应(Binaural Room Impulse Response,BRIR)进行信号处理,从而得到双耳声信号,通过耳机重放双耳声信号让用户感知到相应虚拟声源,受限于消费者或者个人用户的消费成本,基于HRTF的虚拟听觉重放拥有天然的优势,现有技术中基于HRTF的虚拟听觉重放技术主要应用于静态自由场声源场景,而对混响环境下的动态声源的处理不足
[0047] This invention utilizes a dynamically updated virtual sound source rendering algorithm to process the audio signal input from the virtual sound source and obtain a virtual sound source rendering signal. It then uses a preset environment rendering algorithm to process the audio signal input from the virtual sound source in each direction and obtain an environment rendering signal. Based on the superposition signal of the environment rendering signal and the virtual sound source rendering signal after cross-transition processing, an audio signal with spatial sound effects is determined. This method can render natural dynamic virtual sound sources with natural spatial direction transitions, appropriate reverberation, and a good sense of immersion.
Smart Images

Figure CN117061985B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of audio processing technology, and more specifically to a spatial audio processing method and apparatus. Background Technology
[0002] Spatial audio processing is an important audio signal processing technology in virtual, augmented, and mixed reality applications. Spatial audio technology can reproduce the location information of audio, synthesize realistic acoustic environments, and create immersive audio experiences. Spatial hearing research shows that human spatial perception of sound is mainly influenced by binaural factors. Virtual playback technology based on binaural sound signals mainly uses the Head-Related Transfer Function (HRTF) or Binaural Room Impulse Response (BRIR) for signal processing to obtain binaural sound signals. By reproducing the binaural sound signals through headphones, users can perceive the corresponding virtual sound sources. Due to the limited cost for consumers or individual users, virtual auditory playback based on HRTF has a natural advantage. In existing technologies, virtual auditory playback technology based on HRTF is mainly applied to static free-field sound source scenarios, while it is insufficient for processing dynamic sound sources in reverberant environments. Summary of the Invention
[0003] The present invention aims to at least solve one of the technical problems existing in the prior art. To this end, a first aspect of the present invention proposes a spatial audio processing method, comprising:
[0004] Acquire dynamic positioning information of virtual sound sources from multiple directions. The positioning information includes the virtual sound source's orientation relative to the user, distance information, and movement speed information.
[0005] The virtual sound source rendering algorithm is used to process the audio signal input from the virtual sound source to obtain the virtual sound source rendering signal. The virtual sound source rendering algorithm is determined based on the preset filter construction algorithm and the preset dynamic effect processing algorithm. The parameters of the virtual sound source filtering algorithm are updated according to the dynamic directional information of the virtual sound source from multiple directions, and the parameters of the dynamic effect processing algorithm are updated according to the dynamic distance information and motion speed information of the virtual sound source from multiple directions.
[0006] Using a preset environment rendering algorithm, the audio signal input from the virtual sound source in each direction is processed to obtain the environment rendering signal. The environment rendering algorithm is determined based on the first-order sound field coding algorithm and the first-order sound field decoding algorithm.
[0007] Based on the superposition of the environmental rendering signal and the virtual sound source rendering signal after cross-transition processing, the audio signal with spatial sound effects is determined.
[0008] Optionally, the steps of processing the audio signal input to the virtual sound source using a dynamically updated virtual sound source rendering algorithm to obtain the virtual sound source rendering signal include:
[0009] Using a preset filter construction algorithm, the system dynamically processes audio signals input from virtual sound sources in multiple directions and outputs multi-channel filtered signals.
[0010] Using a preset dynamic effects processing algorithm, the multi-channel filtered output signals are processed to obtain virtual sound source rendering signals.
[0011] Optionally, the steps for obtaining dynamic localization information of virtual sound sources from multiple directions include:
[0012] The source of the audio signal input from the virtual sound source is obtained. The source of the audio signal includes pre-set by the program, dynamic input by the user, or on-site acquisition by the sensor.
[0013] Based on the source of the audio signal, determine the location information of the virtual sound source relative to the user at any given moment.
[0014] Optionally, the step of dynamically processing audio signals input from multiple virtual sound sources using a preset filter construction algorithm and outputting multi-channel filtered signals includes:
[0015] Based on the preprocessed head-related transfer function database, the basis vector, left and right ear weighting coefficients, and head-related delay parameters are obtained respectively.
[0016] Based on the location information of the virtual sound source relative to the user at any given moment, determine the spatial angle information of the virtual sound source relative to the user at any given moment.
[0017] Based on the spatial angle information of the virtual sound source relative to the user at any given moment, the control coefficients for the multi-directional aspects at any given moment are determined. The control coefficients include the left and right ear weighting coefficients and the head-related delay parameters.
[0018] Based on the basis vectors and the multi-directional control coefficients at any given time, update the parameters of the preset filter construction algorithm;
[0019] By using a preset filter construction algorithm with updated parameters, audio signals input from virtual sound sources from multiple directions are processed to obtain multi-channel filtered signals.
[0020] Optionally, before updating the parameters of the preset filter construction algorithm based on the basis vectors and the control coefficients in multiple orientations at any given time, the algorithm further includes processing the control coefficients using a spatial smoothing algorithm. The spatial smoothing algorithm includes a bilinear interpolation algorithm, and the step of processing the control coefficients using the bilinear interpolation algorithm includes:
[0021] Acquire the orientation information of each virtual sound source relative to the user at any given time. The orientation information includes the first three-dimensional coordinate information of the virtual sound source with the user as the center point.
[0022] Based on the four second three-dimensional coordinates of each first three-dimensional coordinate information in the adjacent directions, determine the linear expression of each first three-dimensional coordinate information;
[0023] Using a preset interpolation coefficient calculation formula, the interpolation coefficients of the linear expression of each first three-dimensional coordinate information are calculated, thereby obtaining the control coefficients of each virtual sound source relative to the user after smoothing.
[0024] Optionally, the steps of processing the output multi-channel filtered signals using a preset dynamic effects processing algorithm to obtain the virtual sound source rendering signal include:
[0025] Based on the distance and speed information of the virtual sound source relative to the user at any given moment, the corresponding distance gain coefficient and Doppler effect coefficient are obtained respectively.
[0026] The parameters of the preset dynamic effect processing algorithm are updated using the distance gain coefficient and Doppler effect coefficient at any given time.
[0027] By using the dynamic effect processing algorithm after parameter update, the multi-channel filtered output signal is processed to obtain the virtual sound source rendering signal.
[0028] Optionally, the steps of processing the audio signals input from virtual sound sources in each direction using a preset environment rendering algorithm to obtain the environment rendering signal include:
[0029] The first-order sound field coding algorithm is used to process the audio signals input from virtual sound sources in each direction to obtain the encoded audio signals. The expression of the first-order sound field coding algorithm includes:
[0030]
[0031]
[0032]
[0033]
[0034] Among them, s i It is the i-th input audio signal among N virtual sound sources, θ i φ i Let X, Y, and Z represent the horizontal and vertical angles of the i-th virtual sound source, respectively, and let W, X, Y, and Z represent the first, second, third, and fourth output data of the first-order sound field coding algorithm, respectively.
[0035] The encoded audio signal is decoded using a matrix constructed from the binaural room impulse response signal and the spherical harmonic function matrix to obtain the environmental rendering signal. The matrix constructed from the binaural room impulse response signal and the spherical harmonic function matrix includes:
[0036] D = Y H (Y * Y H ) -1 B
[0037]
[0038] Y represents Ω in different directions S The spherical harmonic function matrix of the virtual sound source, where B represents the binaural room impulse response signal of the audio signal input from the virtual sound source in each direction.
[0039] In another aspect, the present invention provides a spatial audio processing apparatus, comprising:
[0040] The information acquisition module is used to acquire dynamic positioning information of virtual sound sources from multiple directions. The positioning information includes the orientation information, distance information, and movement speed information of the virtual sound sources relative to the user.
[0041] The first processing module is used to process the audio signal input from the virtual sound source using a dynamically updated virtual sound source rendering algorithm to obtain the virtual sound source rendering signal. The virtual sound source rendering algorithm is determined based on a preset filter construction algorithm and a preset dynamic effect processing algorithm. The parameters of the virtual sound source filtering algorithm are updated according to the dynamic directional information of the virtual sound source from multiple directions, and the parameters of the dynamic effect processing algorithm are updated according to the dynamic distance information and motion speed information of the virtual sound source from multiple directions.
[0042] The second processing module is used to process the audio signals input from the virtual sound source in each direction using a preset environment rendering algorithm to obtain the environment rendering signal. The environment rendering algorithm is determined based on the first-order sound field coding algorithm and the first-order sound field decoding algorithm.
[0043] The output module is used to determine the audio signal with spatial sound effects based on the superposition signal of the environmental rendering signal and the virtual sound source rendering signal after cross-processing.
[0044] In another aspect, the present invention provides an electronic device comprising a processor and a memory, wherein the memory stores at least one instruction, at least one program, code set or instruction set, and the at least one instruction, at least one program, code set or instruction set is loaded and executed by the processor to implement the spatial audio processing method as described in any one of the first aspects 1-7.
[0045] In another aspect, the present invention provides a computer-readable storage medium storing at least one instruction, at least one program, code set, or instruction set, wherein the at least one instruction, at least one program, code set, or instruction set is loaded and executed by a processor to implement the spatial audio processing method as described in any one of the first aspects 1-7.
[0046] This invention provides a spatial audio processing method and apparatus, which have the following advantages compared with the prior art:
[0047] This invention utilizes a dynamically updated virtual sound source rendering algorithm to process the audio signal input from the virtual sound source and obtain a virtual sound source rendering signal. It then uses a preset environment rendering algorithm to process the audio signal input from the virtual sound source in each direction and obtain an environment rendering signal. Based on the superposition signal of the environment rendering signal and the virtual sound source rendering signal after cross-transition processing, an audio signal with spatial sound effects is determined. This method can render natural dynamic virtual sound sources with natural spatial direction transitions, appropriate reverberation, and a good sense of immersion. Attached Figure Description
[0048] To more clearly illustrate the technical solutions of the present invention, the accompanying drawings used in the description of the embodiments or prior art will be briefly introduced below. Obviously, the drawings described below are merely some embodiments of the present invention, and those skilled in the art can obtain other drawings based on these drawings without any creative effort.
[0049] Figure 1 A flowchart of a spatial audio processing method provided in an embodiment of the present invention;
[0050] Figure 2 A flowchart illustrating the steps of a spatial audio processing method for obtaining dynamic positioning information of virtual sound sources from multiple directions, as provided in an embodiment of the present invention.
[0051] Figure 3 A flowchart illustrating the steps of obtaining a virtual sound source rendering signal in a spatial audio processing method according to an embodiment of the present invention;
[0052] Figure 4 A flowchart illustrating the steps of a spatial audio processing method for outputting a multi-channel filtered signal, as provided in an embodiment of the present invention;
[0053] Figure 5 A flowchart illustrating the steps of obtaining a virtual sound source rendering signal in a spatial audio processing method according to an embodiment of the present invention;
[0054] Figure 6 A step diagram illustrating the steps of obtaining an environmental rendering signal in a spatial audio processing method provided in an embodiment of the present invention;
[0055] Figure 7 A schematic diagram of the encoding and decoding of environmental rendering signals for a spatial audio processing method provided in an embodiment of the present invention;
[0056] Figure 8 This is a schematic diagram of a spatial audio processing method provided in an embodiment of the present invention. Detailed Implementation
[0057] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0058] This specification provides the operational steps for the methods described in the embodiments or flowcharts, but may include more or fewer operational steps based on conventional or non-inventive labor. In actual system or server product execution, the methods shown in the embodiments or drawings may be executed sequentially or in parallel (e.g., in a parallel processor or multi-threaded processing environment).
[0059] Figure 1 A flowchart of a spatial audio processing method provided in an embodiment of the present invention is shown below. Figure 1 As shown, the processing methods include:
[0060] Step 101: Obtain dynamic positioning information of virtual sound sources from multiple directions. The positioning information includes the orientation information, distance information, and movement speed information of the virtual sound sources relative to the user.
[0061] It should be noted that this invention addresses the input of audio information from virtual sound sources when the user is in motion. The number of virtual sound sources is at least one, and they are located in different directions. As the user moves, the location information changes, thus the direction information of the sound source transmission to the user is constantly changing. Simultaneously, the user's distance information and movement speed information are also constantly changing. When the user is stationary, the location information remains unchanged. This invention primarily addresses how to achieve a realistic and immersive experience in reproducing dynamic sound sources when the user's location information changes. Of course, this invention can also be used to achieve better processing speed and effects when the user is stationary.
[0062] Virtual sound sources can be used in virtual environments that emit audio signals, such as in Virtual Reality (VR) or Augmented Reality (AR).
[0063] Specifically, such as Figure 2As shown, step 101 of obtaining dynamic localization information of virtual sound sources from multiple directions includes:
[0064] Step 1011: Obtain the source of the audio signal input from the virtual sound source. The source of the audio signal includes pre-set by the program, dynamic input by the user, or on-site acquisition by the sensor.
[0065] Among them, the program pre-sets dynamic sound effects that do not require human intervention in the program's specified scene. For example, when a train scene is designed to run from one end to the other in a VR program, the sound effects change dynamically.
[0066] User dynamic input is applied to action sound effects that require human intervention, such as the dynamic process of footsteps in a VR game moving from far to near and from near to far.
[0067] Sensor-based on-site data acquisition utilizes sensors at different locations to transmit sound effects from different locations.
[0068] Step 1012: Determine the location information of the virtual sound source relative to the user at any given moment, based on the source of the audio signal.
[0069] Specifically, different types of location information are determined from audio signals from different sources, such as location information from the screen to the user, location information from the operation action to the user, and distance information from the sensor to the user. At least one of these sources is selected to obtain the location information based on actual use.
[0070] Step 102: Using a dynamically updated virtual sound source rendering algorithm, process the audio signal input from the virtual sound source to obtain the virtual sound source rendering signal. The virtual sound source rendering algorithm is determined based on a preset filter construction algorithm and a preset dynamic effect processing algorithm. The parameters of the virtual sound source filtering algorithm are updated according to the dynamic directional information of the virtual sound source from multiple directions, and the parameters of the dynamic effect processing algorithm are updated according to the dynamic distance information and motion speed information of the virtual sound source from multiple directions.
[0071] Specifically, such as Figure 3 As shown, step 102, which uses a dynamically updated virtual sound source rendering algorithm to process the audio signal input from the virtual sound source and obtain the virtual sound source rendering signal, includes:
[0072] Step 1021: Using a preset filter construction algorithm, dynamically process the audio signals input from virtual sound sources in multiple directions and output multi-channel filtered signals.
[0073] Among them, such as Figure 4 As shown, step 1021, which uses a preset filter construction algorithm to dynamically process audio signals input from virtual sound sources in multiple directions and outputs multi-channel filtered signals, includes:
[0074] Step 10211: Based on the preprocessed head-related transfer function database, obtain the basis vector, left and right ear weight coefficients, and head-related delay parameters respectively;
[0075] The preprocessing steps include:
[0076] The head-related transfer function database is processed using a phase transformation algorithm to obtain a head-related transfer function database with minimum phase characteristics.
[0077] Principal component analysis (PCA) is used to process the head correlation transfer function (HRF) database with minimum phase characteristics, resulting in a low-volume HRF database. The processed data volume is approximately 4% of the original data volume. Therefore, the above steps significantly reduce the data volume and improve processing speed.
[0078] Step 10212: Based on the orientation information of the virtual sound source relative to the user at any given time, determine the spatial angle information of the virtual sound source relative to the user at any given time.
[0079] Among them, the location information of the virtual sound source relative to the user at any given time includes the three-dimensional coordinate information of the virtual sound source at any given time. Using the coordinate system transformation method of three-dimensional coordinates and spatial angles, the three-dimensional coordinate information of the virtual sound source is transformed into the spatial angle information of the virtual sound source.
[0080] Step 10213: Based on the spatial angle information of the virtual sound source relative to the user at any given time, determine the multi-directional control coefficients at any given time. The control coefficients include the left and right ear weighting coefficients and the head-related delay parameters.
[0081] It should be noted that the left and right ear weighting coefficients and head-related delay parameters stored in the HRTF database are determined based on spatial angle information. Since the spatial angle of the virtual sound source relative to the user changes at any given moment, the spatial angle information can be used to match with the HRTF database. The left and right ear weighting coefficients and head-related delay parameters corresponding to the spatial angle information of the virtual sound source stored in the HRTF database can be selected. Its function is to determine different left and right ear weighting coefficients and head-related delay parameters for different virtual sound source links.
[0082] Step 10214: Update the parameters of the preset filter construction algorithm based on the basis vectors and the control coefficients of the multi-directional direction at any time.
[0083] The formula for the preset filter construction algorithm is as follows:
[0084]
[0085] In the formula,
[0086] Let w represent the i-th filter.qi The left and right ear weight coefficients of the i-th filter, d q Let τ represent the basis vector of the filter, τ be the head-related delay parameter, and t-τ represent the delay of the basis vector.
[0087] Step 10215: Using the updated filter construction algorithm, process the audio signals input from virtual sound sources from multiple directions to obtain multi-channel filtered signals.
[0088] The filter construction algorithm constructs filters that correspond to virtual sound sources. Each virtual sound source in a given direction corresponds to a filter. After filtering the virtual sound sources, a filtered signal is obtained.
[0089] Step 1022: Using a preset dynamic effect processing algorithm, process the output multi-channel filtered signal to obtain the virtual sound source rendering signal.
[0090] Among them, such as Figure 5 As shown, step 1022, which uses a preset dynamic effects processing algorithm to process the output multi-channel filtered signals to obtain the virtual sound source rendering signal, includes:
[0091] Step 10221: Based on the distance information and motion speed information of the virtual sound source relative to the user at any given time, obtain the corresponding distance gain coefficient and Doppler effect coefficient respectively;
[0092] The dynamic effects processing algorithm calculates the distance gain coefficient and Doppler effect coefficient based on the relative distance and velocity of the virtual sound source. These two coefficients constitute the control coefficients. The formula for calculating the distance gain coefficient is as follows:
[0093]
[0094] d ref It is a reference distance, A ref This is the attenuation coefficient for the reference distance. Typically, the attenuation coefficient A for the reference distance is set. ref =0.5, meaning that as the distance doubles, the gain decreases by 6dB.
[0095] Doppler coefficient The calculation formula is:
[0096]
[0097] v ls It is the speed at which the user leaves the sound source, v sl It refers to the speed at which the sound source moves away from the user, who is the listener.
[0098] Step 10222: Update the parameters of the preset dynamic effect processing algorithm using the distance gain coefficient and Doppler effect coefficient at any given time.
[0099] The preset dynamic effects processing algorithm can be selected from commonly used dynamic effects processors in this field. The parameters of commonly used dynamic effects processors are determined according to the distance gain coefficient and Doppler effect coefficient of the virtual sound source at any time according to this invention.
[0100] Step 10223: Using the updated dynamic effect processing algorithm, process the output multi-channel filtered signal to obtain the virtual sound source rendering signal.
[0101] After processing, the input mono sound source is output as a binaural sound signal; if it is a dynamic virtual sound source, the output is a binaural sound signal of the dynamic virtual sound source.
[0102] Before updating the parameters of the preset filter construction algorithm based on the basis vectors and the control coefficients at any given time in multiple directions, the algorithm further includes processing the control coefficients using a spatial smoothing algorithm. The spatial smoothing algorithm includes a bilinear interpolation algorithm, and the step of processing the control coefficients using the bilinear interpolation algorithm includes:
[0103] Acquire the orientation information of each virtual sound source relative to the user at any given time. The orientation information includes the first three-dimensional coordinate information of the virtual sound source with the user as the center point.
[0104] Based on the four second three-dimensional coordinates of each first three-dimensional coordinate information in the adjacent directions, determine the linear expression of each first three-dimensional coordinate information;
[0105] Using a preset interpolation coefficient calculation formula, the interpolation coefficients of the linear expression of each first three-dimensional coordinate information are calculated, thereby obtaining the control coefficients of each virtual sound source relative to the user after smoothing.
[0106] w′=a1w1+a2w2+a3w3+a4w4
[0107]
[0108]
[0109] Where w' is the coefficient vector of the filter after smoothing the target orientation (θ0,φ0), and w1, w2, w3, w4 are the four filter coefficient vectors of the target spatial orientation's neighboring spatial orientations (θ1,φ1), (θ2,φ1), (θ1,φ2), and (θ2,φ2), respectively, representing the four second three-dimensional coordinates of the adjacent orientations of the first three-dimensional coordinates.
[0110] Step 103: Using a preset environment rendering algorithm, process the audio signal input from the virtual sound source in each direction to obtain the environment rendering signal. The environment rendering algorithm is determined based on the first-order sound field encoding algorithm and the first-order sound field decoding algorithm.
[0111] Specifically, such as Figure 6 and 7 As shown, step 103, which uses a preset environment rendering algorithm to process the audio signals input from virtual sound sources in each direction to obtain the environment rendering signal, includes:
[0112] Step 1031: Process the audio signals input from the virtual sound sources in each direction using a first-order sound field coding algorithm to obtain the encoded audio signals. The expression for the first-order sound field coding algorithm includes:
[0113]
[0114]
[0115]
[0116]
[0117] Among them, s i It is the i-th input audio signal among N virtual sound sources, θ i φ i Let X, Y, and Z represent the horizontal and vertical angles of the i-th virtual sound source, respectively, and let W, X, Y, and Z represent the first, second, third, and fourth output data of the first-order sound field coding algorithm, respectively.
[0118] Step 1032: Decode the encoded audio signal using a matrix constructed from the binaural room impulse response signal and the spherical harmonic function matrix to obtain the environmental rendering signal. The matrix constructed from the binaural room impulse response signal and the spherical harmonic function matrix includes:
[0119] D = Y H (Y * Y H ) -1 B
[0120]
[0121] Y represents Ω in different directions S The spherical harmonic function matrix of the virtual sound source, where B represents the binaural room impulse response signal of the audio signal input from the virtual sound source in each direction.
[0122] Step 104: Determine the audio signal with spatial sound effects based on the superposition signal of the environmental rendering signal and the virtual sound source rendering signal after cross-transition processing.
[0123] This invention enhances the realism and immersion of virtual spatial sound sources by superimposing the ambient reverberation sound rendered by the ambient rendering algorithm onto the virtual sound source rendering signal after cross-transition processing, i.e., the binaural sound signal.
[0124] In another aspect, the present invention also provides a spatial audio processing device 200, such as... Figure 8 As shown, the device includes:
[0125] The information acquisition module 201 is used to acquire dynamic positioning information of virtual sound sources from multiple directions. The positioning information includes the orientation information, distance information and movement speed information of the virtual sound sources relative to the user.
[0126] The first processing module 202 is used to process the audio signal input from the virtual sound source using a dynamically updated virtual sound source rendering algorithm to obtain a virtual sound source rendering signal. The virtual sound source rendering algorithm is determined based on a preset filter construction algorithm and a preset dynamic effect processing algorithm. The parameters of the virtual sound source filtering algorithm are updated according to the dynamic directional information of the virtual sound source from multiple directions, and the parameters of the dynamic effect processing algorithm are updated according to the dynamic distance information and motion speed information of the virtual sound source from multiple directions.
[0127] The second processing module 203 is used to process the audio signal input from the virtual sound source in each direction using a preset environment rendering algorithm to obtain the environment rendering signal. The environment rendering algorithm is determined based on the first-order sound field coding algorithm and the first-order sound field decoding algorithm.
[0128] Output module 204 is used to determine an audio signal with spatial sound effects based on the superposition signal of the environmental rendering signal and the virtual sound source rendering signal after cross-transition processing.
[0129] In another embodiment of the present invention, a device is also provided, the device including a processor and a memory, the memory storing at least one instruction, at least one program, code set or instruction set, the at least one instruction, the at least one program, the code set or instruction set being loaded and executed by the processor to implement the spatial audio processing method described in the embodiment of the present invention.
[0130] In another embodiment of the present invention, a computer-readable storage medium is also provided, wherein at least one instruction, at least one program, code set or instruction set is stored in the storage medium, and the at least one instruction, the at least one program, the code set or instruction set is loaded and executed by a processor to implement the spatial audio processing method described in the embodiment of the present invention.
[0131] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes multiple computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates multiple available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state disk (SSD)).
[0132] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0133] The various embodiments in this specification are described in a related manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.
[0134] The above description is merely a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention are included within the scope of protection of the present invention.
Claims
1. A spatial audio processing method, characterized in that, include: Acquire dynamic positioning information of virtual sound sources from multiple directions. The positioning information includes the virtual sound source's orientation relative to the user, distance information, and movement speed information. The virtual sound source rendering algorithm is used to process the audio signal input from the virtual sound source to obtain the virtual sound source rendering signal. The virtual sound source rendering algorithm is determined based on the preset filter construction algorithm and the preset dynamic effect processing algorithm. The parameters of the virtual sound source filtering algorithm are updated according to the dynamic directional information of the virtual sound source from multiple directions, and the parameters of the dynamic effect processing algorithm are updated according to the dynamic distance information and motion speed information of the virtual sound source from multiple directions. Using a preset environment rendering algorithm, the audio signal input from the virtual sound source in each direction is processed to obtain the environment rendering signal. The environment rendering algorithm is determined based on the first-order sound field coding algorithm and the first-order sound field decoding algorithm. Based on the superposition of the environmental rendering signal and the virtual sound source rendering signal after cross-transition processing, the audio signal with spatial sound effects is determined. The step of processing the audio signal input from the virtual sound source using a dynamically updated virtual sound source rendering algorithm to obtain the virtual sound source rendering signal includes: Using a preset filter construction algorithm, the system dynamically processes audio signals input from virtual sound sources in multiple directions and outputs multi-channel filtered signals. Using a preset dynamic effects processing algorithm, the multi-channel filtered output signals are processed to obtain the virtual sound source rendering signal. The step of dynamically processing audio signals input from multiple virtual sound sources using a preset filter construction algorithm and outputting multi-channel filtered signals includes: Based on the preprocessed head-related transfer function database, the basis vector, left and right ear weight coefficients, and head-related delay parameters are obtained respectively. The preprocessing steps include: processing the head-related transfer function database using a phase transformation algorithm to obtain a head-related transfer function database with minimum phase characteristics; and processing the head-related transfer function database with minimum phase characteristics using principal component analysis to obtain a head-related transfer function database with low data volume. Based on the location information of the virtual sound source relative to the user at any given moment, determine the spatial angle information of the virtual sound source relative to the user at any given moment. Based on the spatial angle information of the virtual sound source relative to the user at any given moment, multi-directional control coefficients are determined at any given moment. The control coefficients include left and right ear weighting coefficients and head-related delay parameters. The left and right ear weighting coefficients and head-related delay parameters stored in the low-data-volume head-related transfer function database are determined based on the spatial angle information. When the spatial angle of the virtual sound source relative to the user changes at any given moment, the spatial angle information can be used to match with the low-data-volume head-related transfer function database, thereby selecting the left and right ear weighting coefficients and head-related delay parameters stored in the low-data-volume head-related transfer function database that correspond to the spatial angle information of the virtual sound source. Different left and right ear weighting coefficients and head-related delay parameters are determined for different virtual sound source links. Based on the basis vectors and the multi-directional control coefficients at any given time, update the parameters of the preset filter construction algorithm; By using a preset filter construction algorithm with updated parameters, audio signals input from virtual sound sources from multiple directions are processed to obtain multi-channel filtered signals.
2. The spatial audio processing method as described in claim 1, characterized in that, The steps for obtaining dynamic positioning information of virtual sound sources from multiple directions include: The source of the audio signal input from the virtual sound source is obtained. The source of the audio signal includes pre-set by the program, dynamic input by the user, or on-site acquisition by the sensor. Based on the source of the audio signal, determine the location information of the virtual sound source at any given moment.
3. The spatial audio processing method as described in claim 1, characterized in that, Before updating the parameters of the preset filter construction algorithm based on the basis vectors and the control coefficients at any given time in multiple directions, the algorithm further includes processing the control coefficients using a spatial smoothing algorithm. The spatial smoothing algorithm includes a bilinear interpolation algorithm, and the step of processing the control coefficients using the bilinear interpolation algorithm includes: Acquire the orientation information of each virtual sound source relative to the user at any given time. The orientation information includes the first three-dimensional coordinate information of the virtual sound source with the user as the center point. Based on the four second three-dimensional coordinates of each first three-dimensional coordinate information in the adjacent directions, determine the linear expression of each first three-dimensional coordinate information; Using a preset interpolation coefficient calculation formula, the interpolation coefficients of the linear expression of each first three-dimensional coordinate information are calculated, thereby obtaining the control coefficients of each virtual sound source relative to the user after smoothing.
4. The spatial audio processing method as described in claim 1, characterized in that, The step of processing the output multi-channel filtered signals using a preset dynamic effects processing algorithm to obtain the virtual sound source rendering signal includes: Based on the distance and speed information of the virtual sound source relative to the user at any given moment, the corresponding distance gain coefficient and Doppler effect coefficient are obtained respectively. The parameters of the preset dynamic effect processing algorithm are updated using the distance gain coefficient and Doppler effect coefficient at any given time. By using the dynamic effect processing algorithm after parameter update, the multi-channel filtered output signal is processed to obtain the virtual sound source rendering signal.
5. A spatial audio processing device, characterized in that, include: The information acquisition module is used to acquire dynamic positioning information of virtual sound sources from multiple directions. The positioning information includes the orientation information, distance information, and movement speed information of the virtual sound sources relative to the user. The first processing module is used to process the audio signal input from the virtual sound source using a dynamically updated virtual sound source rendering algorithm to obtain the virtual sound source rendering signal. The virtual sound source rendering algorithm is determined based on a preset filter construction algorithm and a preset dynamic effect processing algorithm. The parameters of the virtual sound source filtering algorithm are updated according to the dynamic directional information of the virtual sound source from multiple directions, and the parameters of the dynamic effect processing algorithm are updated according to the dynamic distance information and motion speed information of the virtual sound source from multiple directions. The second processing module is used to process the audio signals input from the virtual sound source in each direction using a preset environment rendering algorithm to obtain the environment rendering signal. The environment rendering algorithm is determined based on the first-order sound field coding algorithm and the first-order sound field decoding algorithm. The output module is used to determine the audio signal with spatial sound effects based on the superposition signal of the environmental rendering signal and the virtual sound source rendering signal after cross-processing.
6. An electronic device, characterized in that, The electronic device includes a processor and a memory, wherein the memory stores at least one instruction, at least one program, a code set, or an instruction set, and the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by the processor to implement the spatial audio processing method as described in any one of claims 1-4.
7. A computer-readable storage medium, characterized in that, The storage medium stores at least one instruction, at least one program, code set, or instruction set, wherein the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by a processor to implement the spatial audio processing method as described in any one of claims 1-4.
Citation Information
Patent Citations
Spatial audio for interactive audio environments
CN112567768A
Immersive adaptive rendering method, processor and system for audio file
CN115604644A