Panoramic sound audio generation method and system based on artificial intelligence, and storage medium

By combining multimodal feature fusion and spherical harmonic function modeling with adaptive rendering algorithms, the problems of low efficiency and insufficient accuracy in traditional panoramic audio generation are solved, realizing efficient and automated panoramic audio generation, and improving spatial positioning accuracy and device adaptability.

CN121547723AActive Publication Date: 2026-02-17SHANGHAI RUIHEFENG ELECTRONIC TECHNOLOGY CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202511911870.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-17
Publication Date
2026-02-17
Estimated Expiration
2045-12-17

AI Technical Summary

Technical Problem

Traditional panoramic sound audio generation relies on manual intervention, resulting in lengthy production cycles, high costs, and unstable spatial positioning accuracy. Existing AI technologies lack accurate modeling of three-dimensional spatial characteristics and device adaptability, making it impossible to effectively reproduce sound fields in complex acoustic environments.

Method used

By fusing time-domain and frequency-domain features and spatial scene features through a multimodal feature extraction algorithm, a 3D spatial sound field model is constructed based on the spherical harmonic function, and an adaptive rendering algorithm is used to generate multi-channel panoramic audio based on the target device parameters.

Benefits of technology

It achieves efficient and automated panoramic audio generation, improves spatial positioning accuracy and device compatibility, and provides an immersive auditory experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121547723A_ABST
    Figure CN121547723A_ABST
Patent Text Reader

Abstract

The invention discloses a panoramic sound audio generation method and system based on artificial intelligence, and a storage medium, and the method comprises the following steps: S1, inputting an original audio signal and scene parameters, obtaining the time domain-frequency domain features and space scene features of the original audio signal through a multi-modal feature extraction algorithm, carrying out weighted fusion on the time domain-frequency domain features and the space scene features to obtain fusion features; s2, constructing a mathematical expression basis of a 3D space sound field model based on a spherical harmonic function, and mapping fusion features to three-dimensional space coordinates and the 3D space sound field model through a sound field modeling algorithm; and S3, converting the 3D space sound field model into multi-channel panoramic sound audio output according to the parameters of the target playing equipment by using an adaptive rendering algorithm. According to the method, feature extraction, fusion and sound field modeling are carried out by inputting the original audio signals and scene parameters, automatic panoramic sound audio generation is realized based on adaptive rendering conversion output, and the method has the advantages of automation, high efficiency, high precision and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to audio signal processing technology, and more specifically, to an artificial intelligence-based panoramic audio generation method, system, and computer-readable storage medium. Background Technology

[0002] Immersive sound technology provides users with a highly immersive auditory experience through multi-channel layouts or three-dimensional spatial sound field restoration mechanisms, and is widely used in film and television production, home theater systems, in-vehicle entertainment devices, video games, and virtual reality and augmented reality applications. However, the traditional immersive sound audio generation process heavily relies on the manual intervention of professional mixing engineers. These engineers must manually adjust spatial parameters for each sound source, including key elements such as azimuth, elevation, distance attenuation, and reverberation time. This approach not only results in lengthy production cycles and high labor costs, but also makes it difficult to maintain stable spatial positioning accuracy due to subjective experience differences, especially in dynamic scenes or complex acoustic environments. For example, in non-ideal acoustic spaces such as car cabins, manual adjustments cannot effectively compensate for environmental reflections and noise interference, causing distortion in sound field restoration.

[0003] Current AI-driven audio generation technologies primarily focus on speech synthesis and music composition, and their model architectures generally lack targeted optimization for the spatial characteristics of immersive audio. Existing solutions often only process one-dimensional time-domain or frequency-domain signals, failing to fully integrate three-dimensional spatial coordinate information. This results in an inability to accurately model the sound field's orientation, elevation angle changes, and distance attenuation behavior in a spherical coordinate system. Some simplified methods use channel duplication or fixed-delay superposition strategies to simulate spatial effects, but these techniques inherently cannot reproduce the dynamic characteristics of a real physical sound field, such as directional attenuation in sound propagation, the time-varying characteristics of environmental reverberation, and multipath interference effects. Furthermore, these methods have significant shortcomings when adapting to playback devices with different channel configurations. When the number of channels on the target device changes, the sound field mapping logic cannot be automatically adjusted, resulting in a loss of spatial layering in the output audio. Users experience blurred sound positioning or a break in immersion during actual playback.

[0004] To address the aforementioned issues, existing technologies urgently need improvement. Summary of the Invention

[0005] The purpose of this invention is to provide a leak detection device and method for fuel cell systems to solve the problems existing in the prior art.

[0006] The above-mentioned technical objective of the present invention is achieved through the following technical solution:

[0007] An artificial intelligence-based panoramic sound audio generation method includes the following steps:

[0008] S1. Input the original audio signal and scene parameters, obtain the time-domain-frequency domain features and spatial scene features of the original audio signal through a multimodal feature extraction algorithm, and perform weighted fusion of the time-domain-frequency domain features and spatial scene features to obtain the fused features;

[0009] S2. The mathematical expression basis for constructing a 3D spatial sound field model based on spherical harmonic functions is to map the fused features to three-dimensional spatial coordinates through a sound field modeling algorithm, thus creating a 3D spatial sound field model.

[0010] S3. Using an adaptive rendering algorithm, the 3D spatial sound field model is converted into multi-channel panoramic audio output based on the parameters of the target playback device.

[0011] Furthermore, the implementation method of step S1 is as follows:

[0012] S11. Perform a short-time Fourier transform on the original audio signal to obtain time-frequency domain characteristics;

[0013]

[0014] in, Represents the spectral characteristics of the original audio signal; For frequency point index, For frame index, Represents the time-domain sampled values ​​of the original audio signal; N represents the number of FFT points; R represents the frame shift length; Represents the Hanning window function; Represents a complex exponential kernel function;

[0015] S12. Encode the scene parameters into standardized feature vectors to obtain spatial scene features;

[0016]

[0017] in, This represents the standardized feature vector, where θ is the azimuth angle, φ is the elevation angle, and d is the distance. For the maximum effective distance, T 60 For reverberation time, The maximum reference reverberation time; T is the row and column interchange operator for a vector or matrix;

[0018] S13. The time-domain-frequency domain features and spatial scene features are weighted and fused using an attention-weighted fusion formula to obtain the fused features;

[0019] The formula for calculating attention weights is as follows:

[0020]

[0021] The formula for generating fused features is as follows:

[0022] in, Indicates fusion features, Let M(·) represent the attention weights, and M(·) be the Mel-spectral mapping function. Represents the spectral characteristics of the original audio signal; represents the standardized feature vector; Softmax(·) is the normalization exponential function; MLP is a multilayer perceptron.

[0023] Furthermore, the implementation method of step S2 is as follows:

[0024] S21. Constructing spherical harmonic functions

[0025] Define a spherical harmonic function of order L and order index m (-L≤m≤L);

[0026]

[0027] in, These are spherical harmonic functions used to map the azimuth angle θ and elevation angle φ to spherical basis functions, forming the mathematical expression of the spatial sound field; To relate Legendre polynomials, It is a complex exponential term;

[0028] S22. Utilize fusion features to dynamically adjust the spherical harmonic coefficients and construct a 3D spatial sound field model;

[0029]

[0030]

[0031] in, For spherical harmonic coefficients, As a feature of fusion, This is the frequency weight matrix; It is a 3D spatial sound field model. For the largest order, It is a spherical harmonic function.

[0032] Furthermore, the implementation method of step S3 is as follows:

[0033] S31. For the M channels of the target playback device, construct the spatial response matrix for each channel;

[0034]

[0035] in, Here, M represents the space-frequency response matrix, where M is the total number of channels in the target playback device. Let m be the spatial coordinates of the m-th speaker;

[0036] S32. Perform frequency domain convolution between the 3D spatial sound field model and the spatial response matrix to generate a multi-channel frequency domain output signal;

[0037]

[0038] in, It is a multi-channel frequency domain output signal used to preserve all frequency domain information of the 3D spatial sound field model;

[0039] S33. Perform an inverse short-time Fourier transform on the frequency domain output of each channel to output a time-domain panoramic audio signal;

[0040]

[0041] in, The time-domain panoramic audio signal is used to construct a multi-channel panoramic audio output; ISTFT(·) is the inverse short-time Fourier transform.

[0042] An artificial intelligence-based panoramic sound audio generation system includes:

[0043] The feature extraction module receives the raw audio signal and scene parameters, processes them using a multimodal feature extraction algorithm, and generates fused features.

[0044] The spatial modeling module, connected to the feature extraction module, establishes the mathematical expression basis of the sound field model based on the spherical harmonic function, and maps the fused features to three-dimensional spatial coordinates to generate a 3D spatial sound field model.

[0045] An adaptive rendering module, connected to the spatial modeling module, converts the 3D spatial sound field model into multi-channel panoramic audio output using an adaptive rendering algorithm based on the target playback device parameters.

[0046] The control module is connected to the feature extraction module, spatial modeling module, and adaptive rendering module respectively, and is used to receive scene parameters and device parameters and schedule the above modules to work together.

[0047] A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-described AI-based panoramic audio generation method.

[0048] In summary, the present invention has the following beneficial effects:

[0049] By extracting, fusing, and modeling features from the input raw audio signal and scene parameters, and then converting the output based on adaptive rendering, automated panoramic audio generation is achieved, which has the advantages of automation, high efficiency, and high precision. Attached Figure Description

[0050] Figure 1 This is a schematic diagram of the sensor unit described in this invention. Detailed Implementation

[0051] To make the technical means, creative features, objectives and effects of this invention easier to understand, the invention will be further described below with reference to the figures and specific embodiments.

[0052] like Figure 1 As shown, the present invention proposes an artificial intelligence-based panoramic audio generation method, which includes the following steps:

[0053] S1. Input the original audio signal and scene parameters, obtain the time-domain-frequency domain features and spatial scene features of the original audio signal through a multimodal feature extraction algorithm, and perform weighted fusion of the time-domain-frequency domain features and spatial scene features to obtain the fused features;

[0054] S2. The mathematical expression basis for constructing a 3D spatial sound field model based on spherical harmonic functions is to map the fused features to three-dimensional spatial coordinates through a sound field modeling algorithm, thus creating a 3D spatial sound field model.

[0055] S3. Using an adaptive rendering algorithm, the 3D spatial sound field model is converted into multi-channel panoramic audio output based on the parameters of the target playback device.

[0056] Therefore, this invention can achieve efficient, high-precision, and adaptive generation of panoramic audio, effectively solving the problems of low efficiency, high cost, and insufficient spatial positioning accuracy of traditional methods, as well as the lack of accurate spatial characteristic modeling and device adaptability of existing AI technologies.

[0057] For ease of understanding, the following explains some key terms in this embodiment:

[0058] Raw audio signals refer to unprocessed, directly recorded or generated audio data, which usually exists in waveform form and carries the original information of the sound.

[0059] Scene parameters describe various attributes of the sound environment, such as the location of the sound source (azimuth, elevation, distance), the room's reverberation time, and the ambient noise level. These parameters are crucial for creating a realistic spatial auditory experience.

[0060] Multimodal feature extraction algorithms are algorithms that can extract and integrate useful features from different types of data (such as time-domain and frequency-domain information of audio and spatial information of a scene). Their purpose is to comprehensively capture the complex information of sound events and their environment.

[0061] Time-frequency domain characteristics describe the properties of audio signals in the time and frequency dimensions, such as spectrum, energy distribution, pitch, and timbre. These characteristics are fundamental to understanding the content of sound.

[0062] Spatial scene features describe the distribution and perceived characteristics of sound in three-dimensional space, such as the direction of the sound source, the sense of distance, and spatial diffusion. These features are crucial for constructing immersive panoramic sound.

[0063] Weighted fusion is a method that combines features from different sources or of different importance according to preset or learned weights. It aims to highlight key information and suppress redundant or unimportant information, thereby forming more representative fused features.

[0064] The fusion feature is a comprehensive feature representation after weighted fusion. It contains time-domain and frequency-domain information and spatial scene information of the original audio signal, providing comprehensive input for subsequent sound field modeling.

[0065] Spherical harmonic functions are a set of orthogonal basis functions defined in spherical coordinates, often used to describe the decomposition of any function on a sphere. In acoustics, they are widely used to represent sound fields in three-dimensional space, accurately capturing the directionality and spatial distribution of the sound field.

[0066] A 3D spatial sound field model is a mathematical or physical representation of the sound distribution in three-dimensional space. It includes the location, intensity, and direction of the sound source, as well as the influence of the environment on sound propagation, and aims to accurately simulate the acoustic environment felt by the listener in different locations.

[0067] Sound field modeling algorithms are algorithms that transform abstract blending features into concrete 3D spatial sound field models. They utilize mathematical tools (such as spherical harmonic functions) to map the characteristic information of sound to three-dimensional spatial coordinates, thereby constructing a sound field that can be rendered.

[0068] Adaptive rendering algorithms are algorithms that can convert 3D spatial sound field models into corresponding channel outputs based on the characteristics and layout of different target playback devices (such as headphones, multi-channel speaker systems). It can dynamically adjust rendering strategies to ensure the best immersive sound experience across various devices.

[0069] The target playback device parameters describe the specific configuration information of the target audio playback device, such as the number, location, and type of speakers, or the HRTF (Head-Related Transfer Function) of the headphones. These parameters are the basis for the adaptive rendering algorithm to perform optimizations.

[0070] Multi-channel surround sound audio output refers to multi-channel audio signals generated for a specific playback device. These signals are carefully processed to reproduce a surround sound effect with a sense of direction, distance, and immersion during playback.

[0071] In step S1, the original audio signal and scene parameters need to be input. The original audio signal can be uploaded manually by the user, for example, by importing a WAV or MP3 audio file through the file selection interface. Scene parameters can be input manually by the user, for example, by specifying the azimuth, elevation, distance, and reverberation time of the sound source through text boxes or drop-down menus. The time-domain-frequency domain features and spatial scene features of the original audio signal are obtained through a multimodal feature extraction algorithm. The time-domain-frequency domain features can be obtained by performing a simple Fourier transform or wavelet transform on the original audio signal to extract its spectral envelope, fundamental frequency, and other information. Spatial scene features can be obtained by directly using the input scene parameters as feature vectors, for example, by directly combining the values ​​of azimuth, elevation, and distance. Finally, the time-domain-frequency domain features and spatial scene features are weighted and fused to obtain the fused features. Weighted fusion can be achieved through simple linear weighting, for example, by multiplying the time-domain-frequency domain features and spatial scene features by a fixed weight (e.g., 0.5 and 0.5 respectively), and then summing them.

[0072] In step S2, the mathematical expression basis of the 3D spatial sound field model is constructed based on the spherical harmonic function. The spherical harmonic function can be pre-calculated and stored as a set of basis functions, and its order and order index can be fixed according to the required spatial resolution, for example, choosing a lower order to simplify the calculation.

[0073] Therefore, a 3D spatial sound field model is generated by mapping the fused features to three-dimensional spatial coordinates through a sound field modeling algorithm. The fused features can be directly used to adjust the preset sound field model parameters. For example, some components of the fused features can be directly used as the intensity or position offset of the sound source, thereby roughly locating the sound source in three-dimensional space.

[0074] In step S3, an adaptive rendering algorithm is used to convert the 3D spatial sound field model into a multi-channel panoramic audio output based on the target playback device parameters. The adaptive rendering algorithm can employ a simple distance-based attenuation model, adjusting the volume according to the distance between the sound source and each speaker. The target playback device parameters can simply include the number of speakers and their approximate layout on a two-dimensional plane.

[0075] This invention utilizes artificial intelligence technology to achieve multimodal fusion of raw audio signals and scene parameters, constructing an accurate 3D spatial sound field model that can adaptively render according to the target playback device. This effectively solves the problems of low efficiency, high cost, and limited spatial positioning accuracy in traditional panoramic sound generation, and overcomes the shortcomings of existing AI technologies in spatial characteristic modeling and device adaptability, providing a high-precision, immersive auditory experience.

[0076] In some of the embodiments of the present invention, a multimodal feature extraction algorithm is proposed to obtain the time-frequency domain features and spatial scene features of the original audio signal and perform weighted fusion. However, in its implementation, the specific algorithms for feature extraction and fusion are not described in detail, which may lead to inaccurate feature representation and inability to effectively fuse multi-source information, thereby reducing the spatial positioning accuracy and adaptive rendering effect of panoramic audio.

[0077] To address this, the present invention further proposes the following implementation method:

[0078] S11. Perform a short-time Fourier transform on the original audio signal to obtain time-frequency domain characteristics;

[0079]

[0080] in, Represents the spectral characteristics of the original audio signal; For frequency point index, For frame index, Represents the time-domain sampled values ​​of the original audio signal; N represents the number of FFT points; R represents the frame shift length; Represents the Hanning window function; Represents a complex exponential kernel function;

[0081] S12. Encode the scene parameters into standardized feature vectors to obtain spatial scene features;

[0082]

[0083] in, This represents the standardized feature vector, where θ is the azimuth angle, φ is the elevation angle, and d is the distance. For the maximum effective distance, T 60 For reverberation time, The maximum reference reverberation time; T is the row and column interchange operator for a vector or matrix;

[0084] S13. The time-domain-frequency domain features and spatial scene features are weighted and fused using an attention-weighted fusion formula to obtain the fused features;

[0085] The formula for calculating attention weights is as follows:

[0086]

[0087] The formula for generating fused features is as follows:

[0088] in, Indicates fusion features, Let M(·) represent the attention weights, and M(·) be the Mel-spectral mapping function. Represents the spectral characteristics of the original audio signal; represents the standardized feature vector; Softmax(·) is the normalization exponential function; MLP is a multilayer perceptron.

[0089] Specifically, in step S11, a Short-Time Fourier Transform (STFT) is performed on the original audio signal to convert the time-domain audio signal into a time-frequency domain representation, thereby revealing the energy distribution of the signal at different times and frequencies. This transformation provides fine-grained frequency-dimensional features for subsequent spatial sound field modeling and is a key step in capturing the dynamic characteristics of the audio signal. In implementation, a fixed window length and frame shift STFT can be used, for example, a 2048-point FFT with a window length of 2048 sampling points and a frame shift of 512 sampling points, to balance time and frequency resolution. Alternatively, a multi-resolution STFT can be used, dynamically adjusting the window length according to the frequency range; for example, a long window is used for the low-frequency portion to improve frequency resolution, and a short window is used for the high-frequency portion to improve time resolution, thereby better capturing the transient and steady-state characteristics of the audio signal. The formula for calculating the spectral characteristics of the original audio signal... The FFT represents the time-domain sampled value of the original audio signal; N represents the number of FFT points, which determines the frequency resolution; R represents the frame shift length, which determines the time resolution. This represents the Hanning window function, used to reduce spectral leakage; This represents a complex exponential kernel function used to project a signal into the frequency domain. The precise definition of these parameters ensures the accuracy of the time-frequency domain characteristics.

[0090] In step S12, scene parameters are encoded into standardized feature vectors. This aims to unify the scene parameters describing the location of the sound source in three-dimensional space and the acoustic characteristics of the environment. This standardization helps eliminate dimensional differences between different parameters, preventing certain parameters from having an excessive impact on model training, thereby improving the stability and convergence speed of model training. This allows the model to be effectively processed by machine learning models and fused with the time-frequency features of the audio. The standardized feature vector includes azimuth, elevation, distance, maximum effective distance, reverberation time, and maximum reference reverberation time. These parameters collectively define the location of the sound source in space and the reverberation characteristics of the environment. By normalizing the distance and reverberation time, the numerical comparability of different scene parameters is ensured, laying the foundation for subsequent feature fusion. T represents the row and column interchange operator for vectors or matrices, used to ensure the correct dimensions of the feature vector.

[0091] In step S13, the temporal-frequency domain features and spatial scene features are weighted and fused using an attention-weighted fusion formula. This aims to dynamically adjust the contributions of temporal-frequency domain features and spatial scene features to the final fused features. This mechanism allows the model to focus more on the feature information most relevant to the current task (panoramic audio generation), thereby improving the expressiveness and discriminative power of the fused features. In implementation, a self-attention mechanism can be used, where attention weights are dynamically generated by calculating the similarity between features (e.g., dot product, additive attention), allowing the model to automatically learn which feature combinations are more important during fusion. Alternatively, a cross-attention mechanism can be used, where one feature (e.g., spatial scene features) serves as the query, and another feature (e.g., temporal-frequency domain features) serves as the key and value, thus achieving guided feature fusion from one modality to another. In the formula for calculating attention weights, an MLP (Multilayer Perceptron) is used to learn the complex nonlinear relationship between temporal-frequency domain features and standardized feature vectors, generating the original attention score. Softmax(·) is a normalization exponential function that transforms these scores into a probability distribution between 0 and 1, ensuring that the sum of all weights is 1, thus achieving weighting. In the formula for generating the fused features, M(·) is the Mel-spectral mapping function, which maps spectral features to the Mel frequency scale, which better matches the auditory perception characteristics of the human ear. Then, the final fused features are generated by weighting and combining the Mel-spectral features with attention weights and normalized feature vectors.

[0092] Through the above technical solutions, this invention ensures the accurate capture of the time-frequency domain features of the original audio signal by precisely defining the parameters of the short-time Fourier transform, avoiding the loss of frequency information caused by simple time-domain processing. Simultaneously, by encoding scene parameters into standardized feature vectors, the spatial representation of different scenes is unified, effectively solving the spatial positioning deviation caused by parameter inconsistencies. Furthermore, by dynamically integrating time-frequency domain features and spatial scene features using an attention-weighted fusion formula, the low fusion efficiency caused by traditional simple averaging or splicing is overcome, significantly enhancing the discriminative power of the fused features. This refined feature extraction and fusion method enables the fused features to more accurately reflect the audio content and its position and environmental information in three-dimensional space, thus providing high-quality input for subsequent 3D spatial sound field model construction, ultimately improving the spatial positioning accuracy and adaptive rendering effect of panoramic audio.

[0093] In some of the embodiments of the present invention described above, a 3D spatial sound field model based on spherical harmonic functions is proposed to map the fusion features to three-dimensional spatial coordinates. However, in its implementation, there is a lack of specific spherical harmonic function definition and dynamic adjustment mechanism, resulting in insufficient accuracy and adaptability of spatial sound field modeling, and an inability to accurately reproduce the sense of direction, distance attenuation characteristics, and adapt to the needs of different playback devices in the real sound field.

[0094] In response, the present invention further proposes a method for implementing the above step S2, specifically including:

[0095] S21. Constructing spherical harmonic functions

[0096] Define a spherical harmonic function of order L and order index m (-L≤m≤L);

[0097]

[0098] in, These are spherical harmonic functions used to map the azimuth angle θ and elevation angle φ to spherical basis functions, forming the mathematical expression of the spatial sound field; To relate Legendre polynomials, It is a complex exponential term;

[0099] S22. Utilize fusion features to dynamically adjust the spherical harmonic coefficients and construct a 3D spatial sound field model;

[0100]

[0101]

[0102] in, For spherical harmonic coefficients, As a feature of fusion, This is the frequency weight matrix; It is a 3D spatial sound field model. For the largest order, It is a spherical harmonic function.

[0103] Spherical harmonic functions (SHFCs) are mathematical tools used to describe the distribution of a three-dimensional sound field in spherical coordinates. Their core function is to accurately map the azimuth angle θ and elevation angle φ of a sound source onto a series of spherical basis functions, thus laying the foundation for the mathematical expression of the spatial sound field. This mapping mechanism effectively captures the propagation characteristics of sound waves in different directions, providing a precise mathematical framework for subsequent sound field modeling. Specifically, SHFCs with order L and order index m (-L≤m≤L) can be defined. This definition is a widely adopted standard method in acoustics, providing a complete orthogonal basis to represent any sound field distribution on a sphere. In practical applications, the maximum order L can be flexibly chosen based on the required spatial resolution and computational resources. For example, when capturing more macroscopic sound field direction information, a lower L value (e.g., L=1 or L=2) can be selected; while when achieving finer spatial details and higher positioning accuracy, a higher L value (e.g., L=4 or L=7) can be used. The combination of order L and index m determines the number and shape of basis functions, which in turn affects the precision of sound field modeling and the ability to express complex sound field structures.

[0104] Building upon the constructed spherical harmonic function, this invention further utilizes fusion features to dynamically adjust the spherical harmonic coefficients to construct a 3D spatial sound field model. This step is crucial for achieving the adaptability and accuracy of the sound field model. The fusion features encompass the time-frequency domain features and spatial scene features of the original audio signal, providing rich and meaningful input for dynamic adjustment. One implementation approach is to employ deep learning models, such as multilayer perceptrons (MLPs), recurrent neural networks (RNNs), or convolutional neural networks (CNNs), to learn the complex mapping relationship between the fusion features and the spherical harmonic coefficients. The fusion features serve as input to the network, and the trained network outputs predicted spherical harmonic coefficients. This data-driven approach captures complex nonlinear relationships, enabling high-precision dynamic adjustment. Another implementation approach is to determine the spherical harmonic coefficients through optimization algorithms. For example, a loss function can be designed to quantify the difference between the current sound field model and the desired sound field implied by the fusion features. Subsequently, gradient descent or other iterative optimization methods can be used to progressively adjust the spherical harmonic coefficients to minimize this loss function. In this process, the frequency weight matrix serves as an important constraint or weighting factor, emphasizing the importance of specific frequency ranges in sound field modeling. This allows the constructed 3D spatial sound field model to more accurately reflect the frequency-spatial characteristics of the sound field. The introduction of the frequency weight matrix enables differentiated processing of the spatial distribution of different frequency components when constructing a 3D spatial sound field model. For example, this matrix can be designed based on the sensitivity of the human ear to different frequencies, the spectral characteristics of the audio content, or the frequency response characteristics of the target playback device, thereby ensuring that the sound field model provides high-quality spatial reproduction across the entire frequency range.

[0105] Through the above technical solution, this invention first constructs a well-defined spherical harmonic function in step S21, providing a rigorous mathematical foundation for the 3D spatial sound field model. This mapping based on azimuth angle θ and elevation angle φ ensures the accuracy and universality of the sound field model in three-dimensional space, effectively solving the problem of positioning errors caused by the lack of mathematical rigor in sound field modeling in traditional methods. It allows the sound field to be decomposed into a series of orthogonal basis functions, thereby accurately representing and reconstructing complex spatial sound fields. On this basis, step S22 uses fusion features to dynamically adjust the spherical harmonic coefficients, enabling the sound field model to adaptively adjust according to real-time audio content and scene parameters. Combined with the frequency weight matrix, the spatial distribution of different frequency components can be finely controlled, thus solving the problems of traditional methods failing to accurately reproduce the azimuth, distance attenuation characteristics, and adaptability to different playback devices in a real sound field. This dynamic adjustment mechanism makes the generated 3D spatial sound field model not only highly accurate but also highly flexible and adaptable, better simulating the acoustic environment of the real world. Overall, by combining the fusion features extracted in step S1 with the adaptive rendering in step S3, this solution ensures the accurate transmission and restoration of spatial information throughout the entire chain from the original audio signal to the final multi-channel panoramic audio output through precise modeling in steps S21 and S22. The fusion features provide rich and meaningful input for dynamic adjustments, while the accurate 3D spatial sound field model provides high-quality source data for subsequent adaptive rendering, resulting in excellent performance in spatial positioning, immersion, and device compatibility for the final panoramic audio output.

[0106] In some of the above-mentioned solutions of the present invention, an adaptive rendering algorithm is proposed to convert the 3D spatial sound field model into a multi-channel output. However, in its implementation, it is necessary to ensure spatial positioning accuracy and efficient adaptation to different playback devices, and avoid sound field distortion and positioning error caused by simple copying or delay superposition.

[0107] In response, the present invention further proposes a method for implementing step S3, specifically including:

[0108] S31. For the M channels of the target playback device, construct the spatial response matrix for each channel;

[0109]

[0110] in, Here, M represents the space-frequency response matrix, where M is the total number of channels in the target playback device. Let m be the spatial coordinates of the m-th speaker;

[0111] S32. Perform frequency domain convolution between the 3D spatial sound field model and the spatial response matrix to generate a multi-channel frequency domain output signal;

[0112]

[0113] in, It is a multi-channel frequency domain output signal used to preserve all frequency domain information of the 3D spatial sound field model;

[0114] S33. Perform an inverse short-time Fourier transform on the frequency domain output of each channel to output a time-domain panoramic audio signal;

[0115]

[0116] in, The time-domain panoramic audio signal is used to construct a multi-channel panoramic audio output; ISTFT(·) is the inverse short-time Fourier transform.

[0117] Specifically, for the M channels of the target playback device, a spatial response matrix is ​​constructed for each channel. This spatial response matrix describes the contribution or reception characteristics of a specific speaker to the sound field in three-dimensional space. This matrix captures the speaker's position, orientation, frequency response, and the geometric and acoustic characteristics relative to the listening position. The purpose of constructing this matrix is ​​to accurately map the abstract 3D spatial sound field model onto the specific physical playback device, ensuring that the audio output from each channel can accurately reproduce the spatial information in the sound field. Specifically, this can be achieved, for example, through measurement or simulation. Given the speaker and listening positions, the distance and angle from each speaker to the listening point are calculated using a geometric acoustic model (such as ray tracing or mirror source method), and combined with the speaker's own frequency response characteristics, a frequency-space response function is constructed. This function, after discretization, forms the spatial response matrix. Alternatively, parametric modeling can be performed using preset speaker layout templates and room acoustic parameters. For example, for common 5.1, 7.1.4, or 22.2 channel layouts, the standard position of each speaker can be predefined, and combined with parameters such as room reverberation time and sound absorption coefficient, the sound pressure contribution of each speaker to the listening area at different frequencies can be calculated using acoustic propagation models (such as those based on wave equations or finite element methods), thereby forming a spatial response matrix.

[0118] Furthermore, the 3D spatial sound field model is convolved with the spatial response matrix in the frequency domain to generate a multi-channel frequency domain output signal. Frequency domain convolution is the core step in realizing the conversion from spatial sound field model to multi-channel output. This operation "mixes" abstract, device-independent 3D spatial sound field information with the acoustic characteristics of a specific playback device (represented by the spatial response matrix), thereby calculating the signal content that each physical channel should play. Convolution in the frequency domain can efficiently handle complex frequency responses and time delays while preserving the spatial details of the sound field. Specifically, this can be achieved, for example, by performing a frequency-wise multiplication operation between the 3D spatial sound field model (typically represented as spherical harmonic coefficients or point source distribution) and the spatial response matrix of each channel in the frequency domain. For each frequency component, the spectral representation of the sound field model is multiplied by the spectral representation of the spatial response matrix of the corresponding channel to obtain the output spectrum of that channel at that frequency. Alternatively, the 3D spatial sound field model can be decomposed into a series of virtual sound sources, and then each virtual sound source can be rendered using the spatial response matrix. In the frequency domain, this means convolving (or multiplying) the spectrum of each virtual sound source with the response in the spatial response matrix corresponding to the location of that virtual sound source, and then summing up the contributions of all virtual sound sources to each channel to form the final multi-channel frequency domain output signal.

[0119] Based on this, an inverse short-time Fourier transform (ISTFT) is performed on the frequency domain output of each channel to output a time-domain immersive audio signal. The inverse short-time Fourier transform (ISTFT) is a necessary step to convert the frequency-domain processed multi-channel signal back to the time domain. In S32, we obtained the frequency domain output signal for each channel, which contains precise frequency and phase information, but these are abstract mathematical representations. ISTFT restores this frequency domain data to a continuous time-domain waveform, which is an analog or digitally sampled signal that the loudspeaker can directly play, thus completing the generation of immersive audio. Specifically, this can be implemented, for example, using the standard ISTFT algorithm, which typically involves performing an inverse Fourier transform on each frequency domain frame and then stitching these time-domain frames together using windowing and overlap-add methods to restore the continuity of the original signal and eliminate artifacts introduced by the window function. Alternatively, a filter bank-based synthesis method can be used, treating the frequency domain output signal as the output of a filter bank, and then reconstructing these frequency domain components into a time-domain signal using a synthesis filter bank.

[0120] Through the above technical solution, this invention can accurately convert a 3D spatial sound field model into a multi-channel panoramic audio output adapted to different playback devices. Specifically, by constructing a spatial response matrix for each channel, this invention can fully consider the specific spatial layout and acoustic characteristics of the M channels of the target playback device, thereby providing an accurate physical basis for subsequent sound field rendering and effectively avoiding sound field distortion and positioning offset caused by device differences. Based on this, frequency domain convolution is performed between the 3D spatial sound field model and the spatial response matrix, which can efficiently and accurately map abstract sound field information to specific physical channels, ensuring that the complete frequency domain information of the sound field model is preserved, avoiding information loss or distortion, and achieving accurate spatial mapping. Finally, by performing an inverse short-time Fourier transform on the frequency domain output of each channel, the processed frequency domain signal is restored to a playable time-domain panoramic audio signal, ensuring the real-time performance and compatibility of the final audio. Through rigorous mathematical operations, the entire process enhances the panoramic sound audio's adaptability to different playback devices and the fidelity of the sound field, thereby solving the problem of accurately adapting to different playback devices and maintaining the realism of the sound field during the rendering process, and significantly improving the user's immersive listening experience.

[0121] This invention also proposes an AI-based panoramic sound audio generation system. This system achieves end-to-end automated processing through the collaborative work of a feature extraction module, a spatial modeling module, an adaptive rendering module, and a control module. The feature extraction module receives the raw audio signal and scene parameters, and processes them using a multimodal feature extraction algorithm to generate fused features. Specifically, the multimodal feature extraction algorithm analyzes the raw audio signal to extract time-frequency features, such as the spectral envelope and fundamental frequency information, and combines these with scene parameters such as the sound source azimuth, elevation angle, and distance as spatial features, performing weighted fusion to form fused features. This fusion method comprehensively captures the complex information of sound events and their environment, avoiding the inefficiency of manual adjustments and ensuring the automation and accuracy of feature extraction.

[0122] The spatial modeling module is connected to the feature extraction module, establishing the mathematical foundation for the sound field model based on spherical harmonic functions. Spherical harmonic functions, as a set of orthogonal basis functions defined in spherical coordinates, can accurately represent the directionality and spatial distribution of the sound field. This module maps the fused features to three-dimensional spatial coordinates, generating a 3D spatial sound field model. This accurately reproduces the directional perception and distance attenuation characteristics of the sound source, overcoming the shortcomings of simple channel replication or delay superposition, and achieving high-precision spatial modeling.

[0123] The adaptive rendering module connects to the spatial modeling module. Based on the parameters of the target playback device, it uses an adaptive rendering algorithm to convert the 3D spatial sound field model into multi-channel panoramic audio output. The target playback device parameters include the number of speakers, their placement, etc. The adaptive rendering algorithm dynamically adjusts the rendering strategy, such as adjusting the volume attenuation based on the distance between the sound source and the speakers, to ensure a consistent panoramic sound experience on playback devices with different numbers of channels, effectively solving the device compatibility problem.

[0124] The control module is connected to the feature extraction module, spatial modeling module, and adaptive rendering module, respectively. It receives scene parameters and device parameters and schedules these modules to work collaboratively. This module acts as a central coordination mechanism, ensuring efficient linkage between modules and achieving automated operation of the entire system, thus avoiding the problem of module disconnection in traditional methods.

[0125] Through the above technical solutions, this invention achieves efficient, high-precision, and adaptive generation of panoramic sound audio. The multimodal fusion of the feature extraction module ensures the comprehensiveness of the input features; the spatial modeling module accurately constructs a sound field model based on spherical harmonic functions, significantly improving spatial positioning accuracy; and the adaptive rendering module dynamically adapts to different playback devices, guaranteeing output quality. Overall, the system effectively solves the problems of low efficiency, high cost, and limited spatial positioning accuracy in traditional panoramic sound generation, and overcomes the shortcomings of existing AI technologies in spatial characteristic modeling and device adaptability, providing users with an immersive auditory experience.

[0126] The present invention further proposes a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-described artificial intelligence-based panoramic audio generation method.

[0127] Specifically, the computer-readable storage medium refers to a physical medium capable of storing digital data and allowing computers or other digital devices to read and write information. This medium serves as a carrier for software programs, enabling the programs to be persistently stored, transmitted, and deployed. For example, the medium can be non-volatile memory such as hard disk drives, solid-state drives, or flash memory devices, which retain data even after power is lost; or, during program execution, it can be random access memory (RAM) used to temporarily store running program code and data.

[0128] The computer program is a series of instructions arranged in a specific order, designed to guide the computer to perform specific tasks or operations. In this invention, the program encapsulates the logic and algorithm of an artificial intelligence-based panoramic audio generation method, enabling the method to be executed automatically. The program can exist in the form of compiled code (such as machine code or bytecode compiled from C++ or Java), which can be directly executed by the processor or run on a virtual machine; or it can be interpreted code (such as Python or JavaScript scripts), which is translated and executed line by line by an interpreter at runtime.

[0129] Processor execution refers to the process by which a computing unit, such as a central processing unit (CPU) or graphics processing unit (GPU), performs data processing, logical operations, and control operations according to the instruction sequence defined in a computer program. Through processor execution, the computer program stored on a computer-readable storage medium is transformed into actual computational behavior, thereby realizing the panoramic sound audio generation method. This can be accomplished by a general-purpose CPU performing instruction decoding and arithmetic and logical operations, or by using a GPU for accelerated execution, which is particularly suitable for the large amount of parallel computation involved in panoramic sound audio generation methods, such as Fourier transforms, matrix operations, and neural network inference.

[0130] By executing a computer program through a processor, all steps of the aforementioned AI-based panoramic audio generation method can be fully performed, encompassing the entire process from inputting raw audio signals and scene parameters to outputting multi-channel panoramic audio. This can be achieved by encapsulating steps such as feature extraction, spatial modeling, and adaptive rendering into independent software modules or functions, which are then called and coordinated by the main program; alternatively, existing deep learning frameworks or audio processing libraries can be used to build and execute the various algorithmic components within the method.

[0131] The following example will provide a more detailed explanation of the above technical solution:

[0132] Suppose a content creator needs to generate immersive panoramic audio for a virtual reality (VR) scene. The VR scene is set up so that user A is in a virtual forest, which includes birdsong from a specific location and distance, as well as the sound of a flowing stream. The target playback device is a 7.1 channel surround sound system.

[0133] First, the system receives raw audio signals, such as pre-recorded birdsong and stream sounds, along with scene parameters. These parameters include the azimuth, elevation, and distance of the birdsong source, as well as environmental information such as the relative position and reverberation time of the stream source. In the feature extraction module, a short-time Fourier transform is performed on the raw audio signal, converting it into time-domain-frequency domain features. This transform represents the spectral characteristics of the raw audio signal in both time and frequency dimensions, revealing its frequency components that change over time. Simultaneously, the scene parameters are encoded into standardized feature vectors to obtain spatial scene features. For example, parameters such as the azimuth θ, elevation φ, and distance d of the birdsong source are quantized and normalized, forming a vector describing the spatial location of the sound source and the acoustic characteristics of the environment. Subsequently, the system uses an attention-weighted fusion formula to weightedly fuse the time-domain-frequency domain features and spatial scene features, generating fused features. This fusion method allows the system to simultaneously consider the spectral characteristics of the audio content itself and its position and environmental attributes in three-dimensional space. Unlike traditional methods that simply separate audio from spatial parameters, this method uses an attention mechanism to intelligently identify and emphasize multimodal information that is crucial for panoramic sound generation. This provides richer and more accurate input for subsequent sound field modeling, solving the problem of limited spatial positioning accuracy in traditional methods.

[0134] Next, the spatial modeling module constructs the mathematical foundation for the 3D spatial sound field model based on spherical harmonic functions. Spherical harmonic functions map the azimuth and elevation angles of a sound source onto a set of orthogonal spherical basis functions, providing a mathematical framework for describing the sound field distribution in three-dimensional space. Using the fusion features obtained in the preceding steps, the system dynamically adjusts the spherical harmonic coefficients. The audio content and spatial information contained in the fusion features are used to precisely modulate these coefficients, thereby constructing a 3D spatial sound field model. This model mathematically and accurately describes the sound pressure distribution and propagation characteristics of birdsong and stream sounds in a virtual forest in three-dimensional space, including their orientation, distance attenuation, and reverberation effects. This method of dynamically adjusting spherical harmonic coefficients based on fusion features overcomes the inefficiency and accuracy limitations of manually adjusting parameters in traditional methods, achieving high-precision sound field modeling. It solves the problems of low efficiency and high cost in traditional panoramic sound generation and compensates for the lack of accurate modeling of panoramic sound spatial characteristics in existing AI audio generation technologies.

[0135] Finally, the adaptive rendering module converts the 3D spatial sound field model into multi-channel panoramic audio output based on the target playback device parameters. For the target 7.1-channel surround sound system, the system first constructs a spatial response matrix for each of its M channels (i.e., 8 channels). This matrix describes the response characteristics of each speaker to the sound field at a specific spatial location. Subsequently, the 3D spatial sound field model is convolved with these spatial response matrices in the frequency domain to generate a multi-channel frequency domain output signal. This convolution operation "projects" the abstract 3D sound field model onto the specific speaker layout, preserving all the frequency domain information of the 3D spatial sound field model. Finally, an inverse short-time Fourier transform is performed on the frequency domain output of each channel to output a time-domain panoramic audio signal. These time-domain signals constitute the multi-channel panoramic audio output, which can be played through a 7.1-channel system, allowing user A to experience an immersive auditory experience in a VR scene, such as birdsong coming from a specific direction and the sound of a stream flowing around them. Unlike existing technologies that simply copy or delay audio channels, this method uses an adaptive rendering algorithm to accurately convert the 3D sound field model into a suitable output based on the playback device parameters with different numbers of channels. This ensures the directional sense, distance attenuation characteristics, and device compatibility of the panoramic sound, solving the problems of existing technologies being unable to reproduce the true sound field characteristics and being difficult to adapt to different playback devices. This achieves high-precision, adaptive panoramic sound audio generation.

[0136] In this document, the terms "upper," "lower," "front," "back," "left," "right," "top," "bottom," "inner," "outer," "vertical," and "horizontal," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are only used for the clarity of expressing the technical solution and for the convenience of description, and therefore should not be construed as limiting the present invention.

[0137] In this document, the terms “comprising,” “including,” or any other variations thereof are intended to cover non-exclusive inclusion, which includes not only the elements listed but also other elements not expressly listed.

[0138] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of this invention is defined by the appended claims and their equivalents.

Claims

1. An artificial intelligence-based ambisonics audio generation method, characterized by, The method comprises the following steps: S1. inputting an original audio signal and a scene parameter, obtaining time-frequency domain features and spatial scene features of the original audio signal through a multi-modal feature extraction algorithm, and performing weighted fusion on the time-frequency domain features and the spatial scene features to obtain fused features; S2. constructing a mathematical expression basis of a 3D spatial sound field model based on a spherical harmonic function, mapping the fused features to a three-dimensional spatial coordinate through a sound field modeling algorithm, and obtaining the 3D spatial sound field model; S3. converting the 3D spatial sound field model into a multi-channel ambisonic audio output by using an adaptive rendering algorithm according to a target playback device parameter. 2.The AI-based ambisonics audio generation method of claim 1, wherein, The implementation method of the step S1 is as follows: S11. performing a short-time Fourier transform on the original audio signal to obtain time-frequency domain features; ; wherein, represents a spectral feature of the original audio signal; is a frequency point index, is a frame index, represents a time domain sample value of the original audio signal; N represents a FFT point number; R represents a frame shift length; represents a Hanning window function; represents a complex exponential kernel function; S12. encoding the scene parameter into a standardized feature vector to obtain spatial scene features; ; wherein, denotes a normalized eigenvector, θ is an azimuth angle, φ is an elevation angle, d is a distance, is a maximum effective distance, T 60 is a reverberation time, is a maximum reference reverberation time; T is a row and column transposing operator of a vector or a matrix; S13. performing weighted fusion on the time-frequency domain features and the spatial scene features through an attention weighted fusion formula to obtain fused features; The attention weight calculation formula is as follows: ; The fusion feature generation public key is generated as follows: ; wherein, denotes a fused feature, denotes an attention weight, M(·) is a mel-spectrogram mapping function, denotes a spectral feature of the original audio signal; denotes a normalized feature vector; Softmax(·) is a normalized exponential function, and MLP is a multi-layer perceptron. 3.The AI-based ambisonics audio generation method of claim 1, wherein, The implementation method of the step S2 is as follows: S21. constructing a spherical harmonic function defining a spherical harmonic function with an order L and an order index m (-L≤m≤L); ; wherein is a spherical harmonic function for mapping the azimuth angle θ and the elevation angle φ to a spherical basis function forming the mathematical basis for the spatial sound field; is an associated Legendre polynomial, is a complex exponential term; S22. dynamically adjusting spherical harmonic coefficients by using the fused features to construct a 3D spatial sound field model; ; ; wherein, is a spherical harmonic coefficient, is a fused feature, is a frequency weight matrix; is a 3D spatial sound field model, is a maximum order, is a spherical harmonic function. 4.The AI-based ambisonics audio generation method of claim 1, wherein, The implementation method of the step S3 is as follows: S31. constructing a spatial response matrix of each channel for M channels of the target playback device; ; wherein, is a spatial-frequency response matrix, M is the total number of channels of the target playback device, is the spatial coordinate of the mth loudspeaker; S32. performing frequency domain convolution on the 3D spatial sound field model and the spatial response matrix to generate a multi-channel frequency domain output signal; ; wherein is a multi-channel frequency domain output signal for preserving all frequency domain information of the 3D spatial sound field model; S33. performing inverse short-time Fourier transform on the frequency domain output of each channel to output a time-domain ambisonic audio signal; ; wherein is a time-domain ambisonic audio signal for constituting a multi-channel ambisonic audio output; ISTFT(·) is an inverse short-time Fourier transform.

5. An artificial intelligence-based ambisonics audio generation system, characterized by, The method comprises the following steps: A feature extraction module is configured to receive an original audio signal and a scene parameter, process the original audio signal and the scene parameter through a multi-modal feature extraction algorithm, and generate fused features. A spatial modeling module is connected to the feature extraction module, and is configured to establish a mathematical expression basis of a sound field model based on a spherical harmonic function, map the fused features to a three-dimensional spatial coordinate to generate a 3D spatial sound field model. An adaptive rendering module is connected to the spatial modeling module, and is configured to convert the 3D spatial sound field model into a multi-channel ambisonic audio output by using an adaptive rendering algorithm according to a target playback device parameter. A control module is connected to the feature extraction module, the spatial modeling module, and the adaptive rendering module, and is configured to receive a scene parameter and a device parameter, and schedule the modules to work cooperatively.

6. A computer readable storage medium having a computer program stored thereon, wherein the computer program is executed by a processor to implement the ambisonic audio generation method based on artificial intelligence according to any one of claims 4.

Citation Information

Patent Citations

  • 3D sound field establishment method

    CN106211017A

  • Audio control method of multiple projection control hosts, medium and system

    CN118214992A

  • Head-related sound field transfer function generation method and coefficient generation model training method

    CN119012077A

  • Intelligent sound equipment control method and system based on sound field adaptive adjustment

    CN120848190A