A panoramic audio encoding method based on deep learning
Through the deep learning-based panoramic audio encoding method, the conflict between panoramic audio functions and appearance and structural design in consumer electronic products is solved by using microphone arrays and deep neural network encoders, and the rapid and low-cost panoramic audio functions are achieved, reducing the complexity of hardware design and computing costs.
Patent Information
- Application Number
- CN202310424297.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-20
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2043-04-20
AI Technical Summary
The prior art is difficult to quickly and at low cost to realize panoramic audio functions in consumer electronic products, and the product development threshold is high, making it difficult to be compatible with appearance and structural design.
Using a panoramic audio encoding method based on deep learning, the microphone array output is directly encoded into a panoramic audio signal format through a microphone array and a deep neural network encoder, reducing the cost of traditional mathematical and physical analysis and debugging.
It realizes the rapid and low-cost implementation of panoramic audio functions in consumer electronic products, reduces hardware design complexity and computing costs, and improves R&D and production efficiency.
Smart Images

Figure FDA0004187670820000011 
Figure FDA0004187670820000012 
Figure FDA0004187670820000021
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of panoramic audio coding. Background Art
[0002] Panoramic audio technology uses microphone arrays to record the spatial pressure and velocity information of sound signals, recreating the original sound field as accurately as possible. To achieve this goal, various solutions have been proposed and implemented. Among them, the Ambisonics system, a three-dimensional sound field reconstruction technology proposed by Michael Gerzon of Oxford University in the 1970s, is currently primarily used in commercial applications such as mobile audio and video devices and electronic games due to its relative ease of implementation, diverse formats, and ability to reproduce panoramic sound fields without relying on large and complex speaker systems.
[0003] Initially, Gerzon and others used omnidirectional and figure-8 microphones to collect the zero-order and first-order information of the sound field in three orthogonal directions, obtaining four audio signals (called W, X, Y, and Z components), which were then reproduced using speakers in the corresponding directions. This system is called First Order Ambisonics (FOA) and is now widely used in VR games, 360° videos, and more. During playback, it is assumed that the listener is at the center of a 360° sphere, with both ears receiving sounds from all directions of the sphere. If the spatial sound field recorded during the recording process can be transmitted to both ears in this way, rather than just to the two speakers in front, it will give us a more credible and immersive experience.
[0004] Ambisonics technology can encode the sound field in all spatial dimensions. Therefore, this audio spatial encoding method has a particularly obvious advantage: it can create spatial sound mixing and transmit spatial distance information, lateral and vertical positioning information and height information to the listener.
[0005] First-order Ambisonics can only reconstruct the spatial sound field in a very small area, and the spatial resolution is relatively low. Daniel et al. proposed higher-order ambisonics (HOA), which is based on the spherical harmonic decomposition of the spatial sound field. This method uses a combination of basis vectors defined in a spherical coordinate system to represent the spatial sound field information. This can be compared to the Taylor expansion of a function in a rectangular coordinate system or the Fourier series of a periodic function. It's easy to see that higher-order expansion coefficients increase the spatial resolution, resulting in superior spatial resolution and more detailed information when capturing or reproducing the sound field.
[0006] To achieve relatively high spatial resolution, an Ambisonics system requires a large number of channels, as traditional first-order Ambisonics has limited spatial fitting capabilities. Third-order Ambisonics (with 16 audio channels) or higher is required to truly achieve good spatial resolution. If listeners desire real-world recordings comparable to the object-based computational audio rendering used in video games, even higher-order recording systems (such as sixth-order Ambisonics with 49 audio channels) are required. For such systems, a spherical three-dimensional microphone array is the ideal solution, as spherical arrays offer simpler and more efficient signal processing and are relatively compact. However, spherical array-based Ambisonics recording systems are difficult to deploy in miniaturized, portable audio and video electronics, particularly those targeting consumer electronics. This is because the layout and mechanical design of spherical arrays, while simultaneously meeting the requirements of a spherical form factor, far-field free space, and hollow or rigid housings, are difficult to integrate into the device itself, limiting their application areas.
[0007] In addition, microphone array beamforming algorithms are theoretically capable of processing signals recorded by any array. However, these processing algorithms usually require the array to meet certain shape requirements in order to operate relatively simply and efficiently, such as a flat ring, flat polygon, linear array, or cross array. This type of technology has already been put into practical use in fields such as smart speakers and in-vehicle voice systems. However, this type of beamforming technology usually exhibits the following characteristics: as the number of microphones increases, the main lobe width becomes narrower and the number of side lobes increases. For the same number of microphones, the larger the spacing, the greater the number of side lobes and the higher the peak gain of the side lobes. In order to ensure that the performance indicators meet the requirements, careful array parameter debugging is still required, and a variety of hardware measurement methods are needed to detect whether the spatial response of the device meets the design requirements.
[0008] Beamforming filters can only accurately restore the beam shape in one or a limited number of directions at a time, allowing the microphone to "listen" to signals from a specific direction. Considering a relatively simple case, for a uniform linear array, simply adding all signals together can achieve a directional beam in the 0-degree direction. For complex arrays of arbitrary shapes, this "perfect" mathematical property does not exist. Signal processing for complex arrays requires high computational costs, using devices such as multi-core DSPs and programmable logic. This will also increase the cost of the overall solution. At the same time, the more complex electronic design, greater power consumption, and heat dissipation will also limit its application in consumer electronics and portable audio and video equipment.
[0009] Even more complex, for a three-dimensional arbitrary microphone array designed for installation inside an electronic device, not only does it require analysis of how external sound waves propagate through the device, but it also requires acoustic measurements and theoretical derivation of the electronic product's internal structure to create a series of acoustic models that account for non-ideal propagation conditions such as complex reflection, scattering, and attenuation. Based on this acoustic model and the array design, the internal parameters of the beamforming filter are then calculated. Often, these calculations cannot directly meet the requirements, requiring experienced acousticians to perform extensive measurements and fine-tuning for each product design case and specific application scenario. This debugging is not only time-consuming and labor-intensive, but also requires real-time computer simulation and professional anechoic chamber experiments, resulting in a high threshold and long product development cycle. Furthermore, the aforementioned solutions are generally sufficient only for simpler applications, such as identifying the direction of speech. Achieving high-fidelity panoramic sound field playback would be technically more challenging and even less feasible than traditional methods. Summary of the Invention
[0010] The technical problem to be solved by the present invention is: how to overcome the problems existing in the background technology and provide a panoramic audio encoding method based on deep learning, which can quickly, effectively and cost-effectively solve the conflict between the appearance and structural design requirements of consumer electronic products and the requirements for the realization of panoramic audio functions, accelerate the research and development, debugging and production speed of corresponding products, and make the panoramic audio media form more likely to be widely used.
[0011] The technical solution adopted by the present invention is: a panoramic audio encoding method based on deep learning, in which a reference sound source S composed of multiple sound sources and a microphone array A composed of multiple array elements exist in the same space, and a three-dimensional rectangular coordinate system is established with the first array element position of the microphone array A or the geometric center of the microphone array A as the coordinate origin. The x-axis and y-axis are any two perpendicular lines on the horizontal plane where the coordinate origin is located, and the z-axis is a line perpendicular to the horizontal plane where the coordinate origin is located. The directions of the x-axis, y-axis, and z-axis are arbitrarily specified, including the following steps:
[0012] Step 1: Input the driving signal of the audio signal of L frames generated by each sound source, the azimuth angle of each sound source relative to the coordinate origin, the altitude angle of each sound source relative to the coordinate origin, and the spatial straight-line distance of each sound source relative to the coordinate origin into the reference signal generator R of the feedback module F. The reference signal generator R of the feedback module F generates a matrix REF with m rows and L columns. Each row of the matrix REF is the same reference signal spherical harmonic domain component arranged in sequence along the time, and each column of the matrix REF is a different reference signal spherical harmonic domain component arranged in sequence at the same time.
[0013] Step 2: When each sound source of the reference sound source S is driven by a driving signal and the sound waves emitted are broadcast into space, the microphone array A receives the sound waves and records the audio signals as L frames, which are transmitted to the deep neural network panoramic sound encoding module N. The number of array elements is the same as or different from the number of sound sources. The input of the deep neural network panoramic sound encoding module N is a matrix I with the number of rows equal to the number of array elements and the number of columns L. Each row of the matrix I is the signal amplitude received by the same array element of the microphone array A arranged in sequence along the time, and each column of the matrix I is the signal amplitude received by different array elements arranged in order at the same time. The output of the deep neural network panoramic sound encoding module N is a matrix O with m rows and L columns, where m is the number of channels of the spherical harmonic decomposition of the signal. Each row of the matrix O is the same output signal spherical harmonic domain component arranged in sequence along the time, and each column of the matrix O is the different output signal spherical harmonic domain components arranged in order at the same time.
[0014] Step 3. Input the matrix O and the matrix REF into the evaluator E of the feedback module F. The evaluator E obtains the difference evaluation error loss of the matrix O and the matrix REF based on statistical indicators. If the difference evaluation error loss is less than the set value ε or the step is returned to step 1 u times continuously and the decrease in the difference evaluation error loss each time is less than the set value σ, the matrix O is the panoramic audio encoding of the current spatial sound field signal of the microphone array A. Otherwise, after changing the position of the reference sound source S itself, return to step 1.
[0015] In microphone array A, each microphone constitutes an array element, the number of array elements is n_array, the array elements are sorted, and the output signal of the i-th array element is A i , the coordinates of the ith element are (x i ,y i ,z i ); In the reference sound source S, there are n_source sound sources in total. Sort the sound sources and record the wth sound source as S w , the driving signal of the wth sound source is f w (t), that is, in an ideal state, the vibration of the w-th sound source satisfies the amplitude-time relationship f w (t), the azimuth angle of the w-th sound source relative to the coordinate origin is The altitude angle is θ w , spatial straight-line distance r w , m = 2*order+1, order is the order of the panoramic audio system, n_array>(order+1) 2 .
[0016] The deep neural network panoramic sound encoding module N is a deep neural network inside, and the matrix I can be expressed as The matrix O can be expressed as Where t is the number of frames, Ai , t Indicates the signal amplitude of the t-th frame time received by the i-th array element, i is a natural number less than or equal to n_array, t is an integer from 0 to L, C v,t The spherical harmonic component of the vth channel component of the spherical harmonic decomposition of the reference signal at the tth frame time, where v is a natural number less than or equal to m.
[0017] The reference signal generator R of the feedback module F is an Ambisonics encoder based on the theory of sound wave propagation under ideal conditions. The sound field signal generated by all reference sound sources at the coordinate origin has a spherical harmonic expansion in, is a spherical harmonic function, p is a positive integer, q is an integer whose absolute value is not greater than p, j is an imaginary unit, k is the wave number, h is the Hankel function on the sphere, C′ v,t Represents the spherical harmonic component of the vth channel component of the spherical harmonic decomposition of the output signal at the tth frame time, where v is a natural number less than or equal to m.
[0018] The difference evaluation error loss = g(O, REF), where g(O, REF) is a scalar-valued function with the matrix REF and the matrix O as variables. The scalar-valued function is one or a combination of square difference function, MSE function, and MAE function.
[0019] The beneficial effects of the present invention are as follows: by designing appropriate deep neural network input and output and training methods, the present invention can directly obtain a microphone array panoramic audio encoder, and encode the output signal of any microphone array into a panoramic audio signal format commonly used in the industry, saving the huge theoretical research workload and experimental costs such as debugging and measurement required to process the special mechanical structure, special sound propagation characteristics, and special non-ideal factors of any array through traditional mathematical and physical analysis methods.
[0020] The panoramic audio encoder provided by the present invention has a model training process that can be completed in advance. When encoding the audio signal, it is only necessary to load the parameters and complete the inference process. This process is composed of a series of general neural network operators. Most current mobile device system-level solutions (SoC chips) already include neural network processor modules (NPU modules). Only driver adaptation and model adaptation are required to accelerate such operations on existing hardware platforms, effectively utilizing existing conditions and occupying fewer other hardware resources to achieve panoramic audio functions. At the same time, it also avoids increasing the complexity of hardware design and the use of additional devices. DETAILED DESCRIPTION
[0021] A panoramic audio coding method based on deep learning includes a reference sound source S consisting of two sound sources and a microphone array A consisting of four array elements in the same space. A three-dimensional rectangular coordinate system is established with the geometric center of the microphone array A as the coordinate origin. The x-axis and y-axis are any two perpendicular lines on the horizontal plane where the coordinate origin is located. The z-axis is a line perpendicular to the horizontal plane where the coordinate origin is located and points from the coordinate origin to the first array element. The directions of the x-axis, y-axis, and z-axis can be arbitrarily specified. The method includes the following steps:
[0022] Step 1. The driving signal of the audio signal generated by each sound source for L frames, the azimuth angle of each sound source relative to the coordinate origin, the altitude angle of each sound source relative to the coordinate origin, and the spatial straight-line distance of each sound source relative to the coordinate origin are input into the reference signal generator R of the feedback module F. The reference signal generator R of the feedback module F generates a matrix REF with m=4 rows and L=1024 columns. Each row of the matrix REF is the same reference signal spherical harmonic domain component arranged in sequence along the time, and each column of the matrix REF is a different reference signal spherical harmonic component arranged in order at the same time.
[0023] In microphone array A, each microphone constitutes an array element, the number of array elements is n_array=4, the array elements are sorted (they can be sorted by personnel, such as the arrangement method in patent number ZL202120871777.0), the output signal of the i-th array element is A i , the coordinates of the ith element are (x i ,y i ,z i ); In the reference sound source S, there are n_source = 2 sound sources. Sort the sound sources and record the wth sound source as S w , the driving signal of the wth sound source is f w (t), that is, in an ideal state, the vibration of the w-th sound source satisfies the amplitude-time relationship f w (t). Here, f1(t) represents the sound of an instrument and f2(t) represents the sound of an explosion, both of which are obtained from pre-recorded waveform files. However, regardless of the actual sound, the generality remains the same. Generally, the choice of sound source waveforms should be diverse, encompassing as many common sound features as possible to ensure the generalizability of the trained model.
[0024] The azimuth angle of the wth sound source relative to the coordinate origin is The altitude angle is θ w , spatial straight-line distance r w , m = 2*order+1, order is the order of the panoramic audio system (order is manually specified and is currently set to 1). In order to achieve good spatial resolution at this order, it should satisfy n_array>(order+1) 2.
[0025] When the microphone array distribution is close to spherical, it is generally sufficient to meet the lower limit of the condition. Otherwise, the number of array elements n_array should be appropriately increased in order to achieve better coding effect.
[0026] Step 2: When each sound source of the reference sound source S is driven by a driving signal and emits a sound wave that is broadcast into space, the microphone array A receives the sound wave and records it as L=1024 frames of audio signals, which are transmitted to the deep neural network panoramic sound encoding module N. The number of array elements is the same as or different from the number of sound sources. The input of the deep neural network panoramic sound encoding module N is a matrix I with the number of rows equal to the number of array elements and the number of columns L=1024. Each row of the matrix I is the signal amplitude received by the same array element of the microphone array A arranged in sequence along the time, and each column of the matrix I is the signal amplitude received by different array elements arranged in sequence at the same time. The output of the deep neural network panoramic sound encoding module N is a matrix O with m=4 rows and L=1024 columns, where m is the number of channels of the signal spherical harmonic decomposition. Each row of the matrix O is the same output signal spherical harmonic domain component arranged in sequence along the time, and each column of the matrix O is the different output signal spherical harmonic domain components arranged in sequence at the same time.
[0027] The deep neural network panoramic sound encoding module N is a deep neural network inside, and the matrix I can be expressed as The matrix O can be expressed as Where t is the number of frames, A i , t Indicates the signal amplitude of the t-th frame time received by the i-th array element, i is a natural number less than or equal to n_array, t is an integer from 0 to L, C v,t The spherical harmonic component of the vth channel component of the spherical harmonic decomposition of the reference signal at the tth frame time, where v is a natural number less than or equal to m.
[0028] L is the frame interval The product of the sampling rate s_rate = 44100, that is, L = t int *s_rate.
[0029] Step three, input the matrix O and the matrix REF into the evaluator E of the feedback module F. The evaluator E obtains the difference evaluation error loss of the matrix O and the matrix REF based on statistical indicators. If the difference evaluation error loss is less than the set value ε or the difference evaluation error loss decreases less than the set value σ each time the step is returned to step one u times, the matrix O is the panoramic audio encoding of the current spatial sound field signal of the microphone array A. Otherwise, after changing the position of the reference sound source S itself, return to step one. In this example, random changes are used. Whenever returning to step one, all reference sound sources are within a certain distance range, such as a radius of 0-5m and all angles of the global surface, that is, the azimuth angle. Take 0-360°, altitude angle θ i Randomly update its own position within ±90°) and send out waveform f i (t).
[0030] The evaluator E collects the Atmos reference signal matrix REF and the Atmos coded signal output matrix O from the network, derives a difference evaluation loss for the two signals based on statistical indicators, and outputs the result to a numerical optimization solver outside of the present invention. (For artificial neural network training, iterative adjustment of network internal parameters using numerical optimization methods is common knowledge.) This solver iteratively updates the internal variable parameters of the deep neural network Atmos coding module N (the number and organization of the variable parameters depends on the scale and structure of the neural network model used. That is, regardless of the number of variable parameters and their organization, as long as the model falls within the scope of artificial neural networks, it is considered without loss of generality).
[0031] Training ends when the evaluation error loss meets the termination criteria, such as falling below a sufficiently small value ε = 1e-2 or failing to decrease for several consecutive steps (iterations u), meaning the decrease per step is less than σ = 1e-5. (After manually setting the termination criteria, whether to terminate after this iteration is determined by the optimization solver's process control.) At this point, all variable parameters in the deep neural network's panoramic sound encoding module N should be saved. Otherwise, they are randomly modified and the training process is repeated in step 1.
[0032] The reference signal generator R of the feedback module F is an Ambisonics encoder based on the theory of sound wave propagation under ideal conditions. The sound field signal generated by all reference sound sources at the coordinate origin has a spherical harmonic expansion in, is a spherical harmonic function, p is a positive integer, q is an integer whose absolute value is not greater than p, j is the imaginary unit, k is the wave number, and h is the Hankel function on the sphere.
[0033] When the system order is given, p increases from 0 and q increases from -p, which are the spherical harmonic components C1 to C mThat is, once the reference sound source position, output characteristics and system parameters are determined, R will output the reference signal column vector at any time t = t0 as In a discrete time from 0 to the frame length L, in units of 1 / s_rate, the column vector is formed as follows It is particularly important to note that the REF and O matrices have the same physical meaning, both representing panoramic sound field signals in the spherical harmonic domain. For the same system implemented according to the present invention, the length and width of these two matrices should be equal, and in this example, both have 4 rows and 1024 columns.
[0034] C′ v,t Represents the spherical harmonic component of the vth channel component of the spherical harmonic decomposition of the output signal at the tth frame time, where v is a natural number less than or equal to m.
[0035] The difference evaluation error loss = g(O, REF), where g(O, REF) is a scalar-valued function with the matrix REF and the matrix O as variables. The scalar-valued function is one or a combination of square difference function, MSE function, and MAE function.
[0036] When performing model inference to process an input signal using the panoramic audio encoder described in the present invention:
[0037] In this patent, only microphone array A and deep neural network panoramic sound encoding module N are involved in the work. The internal variable parameters of N are set according to the parameter values saved after training is completed. Then the real sound field signal is collected by microphone array A. The microphone signal is organized into matrix I and sent to deep neural network panoramic sound encoding module N. Deep neural network panoramic sound encoding module N obtains the encoding of the current spatial sound field signal, namely matrix O. The present invention includes but is not limited to the following three implementation methods:
[0038] This implementation is based on computer software simulation. Its characteristics are that the reference sound source S, the spatial propagation path and propagation characteristics of the sound, the azimuth measurement module M1, the elevation measurement module M2, the distance measurement module M3, the spatial placement of the microphone array, and the frequency response characteristics are all implemented using physical simulation tools or programming algorithms based on computer software. Position information and waveform information are generated by the software or pre-provided according to the program. The subsequent neural network algorithm is also processed by the computer program. The internal modules of the entire implementation do not need to exist in a real hardware form, and there is no need for actual connection and actual physical interaction between modules. Only the corresponding data needs to be transmitted or read. This implementation method is mainly used to verify principles, collect data, complete the early stages of model training, or conduct quantitative evaluation of trained models.
[0039] The implementation method is based on audio equipment and embedded hardware. Its characteristics are that actual sound-generating equipment, sound-receiving equipment, and angle and distance measurement equipment are used to implement the present invention. The spatial propagation of sound adopts a real space with air as the medium or an environment with some artificially controlled conditions. It is worth noting that the reference signal generator R is not only provided by an algorithm, but can also be a set of standard structures, designed based on ideal sound wave propagation theory, with a microphone array and its supporting signal processing equipment with verified performance. The panoramic sound signal used for reference is obtained by actual recording rather than pure theoretical calculation. The azimuth measurement module M1, altitude measurement module M2, and distance measurement module M3 required for spatial position measurement can be integrated and implemented using the same principle, or they can be constructed separately or implemented using different principles. The modules themselves can perform coordinate measurement of array elements based on some physical phenomenon (such as emitting laser or electromagnetic wave pulses, or using a machine vision system to identify array elements and then estimate them), or the array elements can autonomously measure coordinates of a common reference point (for example, using a UWB indoor positioning system, with the located terminal installed on the array element and supporting facilities that can determine the reference point installed in the room, so that each array element obtains its own position based on the reference point), and then report its position to the center via communication. Neural network models can run on conventional computers, devices with such computing capabilities, or dedicated processors.
[0040] Hybrid and distributed implementations. The implementation cases in which the real implementation and software simulation implementation of the various parts of the present invention mentioned above are used together to meet specific needs belong to the hybrid implementation of the present invention. The various parts of the present invention do not have to be arranged in the same device, set of equipment, the same computer or some artificially defined adjacent locations such as the same room, the same local area network, etc., that is, the distributed implementation of the present invention (for example, for some specific needs, the signal of a remote microphone array is transmitted as input through the Internet, and the microphone array mentioned in the present invention is constructed using multiple sub-array components with independent structures, which transmit and process signals respectively, and measure or report relative positions respectively, and the model training process uses multi-node parallel computing, and all or part of the results obtained from the model training and the corresponding processing steps are solidified into a certain dedicated circuit, so as to deviate from the common form of the neural network algorithm but achieve the same function, etc.) only needs to logically meet its working principle and connection relationship.
Claims
1. A deep learning-based panoramic audio encoding method. A reference sound source S consisting of multiple sound sources and a microphone array A consisting of multiple elements exist in the same space. A three-dimensional rectangular coordinate system is established with the first element position of microphone array A or the geometric center of microphone array A as the coordinate origin. The x-axis and y-axis are any two perpendicular lines on the horizontal plane where the coordinate origin is located, and the z-axis is a line perpendicular to the horizontal plane where the coordinate origin is located. The directions of the x-axis, y-axis, and z-axis can be arbitrarily specified. The method is characterized by: The steps include: Step 1: Input the driving signal of the audio signal of L frames generated by each sound source, the azimuth angle of each sound source relative to the coordinate origin, the altitude angle of each sound source relative to the coordinate origin, and the spatial straight-line distance of each sound source relative to the coordinate origin into the reference signal generator R of the feedback module F. The reference signal generator R of the feedback module F generates a matrix REF with m rows and L columns. Each row of the matrix REF is the same reference signal spherical harmonic domain component arranged in sequence along the time, and each column of the matrix REF is a different reference signal spherical harmonic domain component arranged in sequence at the same time. Step 2: When each sound source of the reference sound source S is driven by a driving signal and the sound waves emitted are broadcast into space, the microphone array A receives the sound waves and records the audio signals as L frames, which are transmitted to the deep neural network panoramic sound encoding module N. The number of array elements is the same as or different from the number of sound sources. The input of the deep neural network panoramic sound encoding module N is a matrix I with the number of rows equal to the number of array elements and the number of columns L. Each row of the matrix I is the signal amplitude received by the same array element of the microphone array A arranged in sequence along the time, and each column of the matrix I is the signal amplitude received by different array elements arranged in order at the same time. The output of the deep neural network panoramic sound encoding module N is a matrix O with m rows and L columns, where m is the number of channels of the spherical harmonic decomposition of the signal. Each row of the matrix O is the same output signal spherical harmonic domain component arranged in sequence along the time, and each column of the matrix O is the different output signal spherical harmonic domain components arranged in order at the same time. Step 3. Input the matrix O and the matrix REF into the evaluator E of the feedback module F. The evaluator E obtains the difference evaluation error loss of the matrix O and the matrix REF based on statistical indicators. If the difference evaluation error loss is less than the set value ε or the step is returned to step 1 u times continuously and the decrease in the difference evaluation error loss each time is less than the set value σ, the matrix O is the panoramic audio encoding of the current spatial sound field signal of the microphone array A. Otherwise, after changing the position of the reference sound source S itself, return to step 1.
2. The deep learning-based panoramic audio encoding method according to claim 1, wherein: In the microphone array A, each microphone constitutes an array element, the number of array elements is n_array, the array elements are sorted, and the output signal of the i-th array element is A i , the coordinates of the ith element are (x i ,y i ,z i ); In the reference sound source S, there are n_source sound sources in total. Sort the sound sources and record the wth sound source as S w , the driving signal of the wth sound source is f w (t), that is, in an ideal state, the vibration of the w-th sound source satisfies the amplitude-time relationship f w (t), the azimuth angle of the w-th sound source relative to the coordinate origin is The altitude angle is θ w , spatial straight-line distance r w , m = 2*order+1, order is the order of the panoramic audio system, n_array>(order+1) 2 .
3. The method for encoding panoramic audio based on deep learning according to claim 2, wherein: The deep neural network panoramic sound encoding module N is a deep neural network inside, and the matrix I can be expressed as The matrix O can be expressed as Where t is the number of frames, A i,t Indicates the signal amplitude of the t-th frame time received by the i-th array element, i is a natural number less than or equal to n_array, t is an integer from 0 to L, C v,t The spherical harmonic component of the vth channel component of the spherical harmonic decomposition of the reference signal at the tth frame time, where v is a natural number less than or equal to m.
4. The method for encoding panoramic audio based on deep learning according to claim 2, wherein: The reference signal generator R of the feedback module F is an Ambisonics encoder based on the theory of sound wave propagation under ideal conditions. The sound field signal generated by all reference sound sources at the coordinate origin has a spherical harmonic expansion in, is a spherical harmonic function, p is a positive integer, q is an integer whose absolute value is not greater than p, j is an imaginary unit, k is the wave number, h is the Hankel function on the sphere, C′ v,t Represents the spherical harmonic component of the vth channel component of the spherical harmonic decomposition of the output signal at the tth frame time, where v is a natural number less than or equal to m.
5. The method for encoding panoramic audio based on deep learning according to claim 2, wherein: The difference evaluation error loss = g(O, REF), where g(O, REF) is a scalar-valued function with the matrix REF and the matrix O as variables. The scalar-valued function is one or a combination of square difference function, MSE function, and MAE function.
Citation Information
Patent Citations
Microphone array for panoramic audio
CN215529286U
Audio acquisition device and Dolby Atmos encoding scheme based on device
CN108877817A
Panoramic audio processing method for panoramic camera
CN113347530A