Audio processing method and system, storage medium and electronic equipment

By rendering user-generated dry audio from top, surround, and near-ear sources, and combining acoustic and physical models, the problem of ambiguous sound source localization in online mode was solved, achieving accurate restoration of the immersive atmosphere of a concert space and enhancing the audio experience.

CN121547722APending Publication Date: 2026-02-17TENCENT MUSIC ENTERTAINMENT TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511903095.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-16
Publication Date
2026-02-17

AI Technical Summary

Technical Problem

Existing technologies cannot recreate the immersive spatial experience of a concert online, mainly because the concert sound field is a large open/semi-open space with acoustic characteristics such as long sound propagation distance, significant air attenuation, and a complex sound field caused by the superposition of multiple speaker systems, resulting in blurred sound source localization.

Method used

By performing top sound source spatial rendering, surround sound source spatial rendering, and near-ear detail rendering on the user's dry audio, top simulated sound source, surround simulated sound source, and near-ear simulated sound source that match the target scene are simulated respectively. Spatial fusion processing is then performed, and by combining acoustic theory and physical models, the sound field layers of the target scene are accurately restored.

Benefits of technology

It achieves the spatial immersion of a real concert in an online format, enhancing the authenticity and immersion of the music listening experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121547722A_ABST
    Figure CN121547722A_ABST
Patent Text Reader

Abstract

The invention discloses an audio processing method and system, a storage medium and electronic equipment, relates to the technical field of acoustic simulation and spatial audio rendering, and aims to simulate a top simulation sound source conforming to a target scene by performing top sound source spatial rendering processing on user dry sound audio. The surrounding sound source space rendering processing is carried out on the basic accompaniment audio to simulate a surrounding simulation sound source of the target scene, and the near-ear detail rendering processing is carried out on the basic accompaniment audio to simulate a near-ear simulation sound source conforming to the target scene. The scene surrounding sound source has a surrounding dispersion feeling around the stadium, and the ear-attached accompaniment has a detail fitting feeling, so that the sound field levels of long-distance propagation, multi-direction surrounding and near-field details of the target scene can be accurately restored, which is closer to the physical reality. Therefore, the spatial immersion of a real concert can be restored when the music is listened online.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of acoustic simulation and spatial audio rendering, and more specifically, to an audio processing method, system, storage medium, and electronic device. Background Technology

[0002] With the development of technology, users can listen to music online through relevant music software, logging into specific platforms, and other online methods, which brings great convenience to users.

[0003] Because the sound field of a concert is a large open / semi-open space, its acoustic characteristics include long sound propagation distance, significant air attenuation, and complex sound field superimposed by multiple speaker systems. This results in the current rendering of acoustic characteristics and the fuzzy location of sound sources. Therefore, when users listen to music online, it is difficult to reproduce the spatial immersion of a real concert.

[0004] Therefore, how to recreate the spatial immersion of a real concert when listening to music online is a problem that this application urgently needs to solve. Summary of the Invention

[0005] In view of this, this application discloses an audio processing method, system, storage medium, and electronic device, aiming to recreate the spatial immersion of a real concert when listening to music online.

[0006] To achieve the above objectives, the disclosed technical solution is as follows:

[0007] The first aspect of this application discloses an audio processing method, the method comprising:

[0008] Obtain the user's dry audio and basic accompaniment audio;

[0009] Based on the user's dry audio, a top sound source space rendering process is performed to obtain a top simulated sound source that conforms to the target scene.

[0010] Based on the basic accompaniment audio, surround sound source space rendering processing is performed to obtain a surround simulated sound source that conforms to the target scene;

[0011] Based on the aforementioned basic accompaniment audio, near-ear detail rendering is performed to obtain a near-ear simulated sound source that matches the target scene;

[0012] The top simulated sound source, the surround simulated sound source, and the near-ear simulated sound source are spatially fused to obtain target spatial audio that matches the target scene.

[0013] Preferably, the step of performing top sound source spatial rendering processing based on the user's dry audio to obtain a top simulated sound source that conforms to the target scene includes:

[0014] Acquire room impulse response data for the target scene;

[0015] The user's dry audio is sidechained to obtain the compressed accompaniment signal;

[0016] Spatial reverberation simulation is performed using the room impulse response data and the compressed accompaniment signal to obtain a simulated top sound source that matches the target scene.

[0017] Preferably, the step of performing surround sound source spatial rendering processing based on the basic accompaniment audio to obtain a surround simulated sound source that conforms to the target scene includes:

[0018] Acquire surround sound room impulse response data for the target scene;

[0019] Spatial reverb enhancement is applied to the base accompaniment audio;

[0020] The difference in the arrival time of a sound source at both ears in a target scene is simulated based on the head-related transfer function; wherein, the difference includes at least a time difference and an intensity difference;

[0021] Based on the surround sound room impulse response data, the basic accompaniment audio after spatial reverberation enhancement, the time difference, and the intensity difference, a surround simulated sound source that matches the target scene is determined.

[0022] Preferably, the near-ear detail rendering processing based on the basic accompaniment audio to obtain a near-ear simulated sound source that conforms to the target scene includes:

[0023] Extract the detailed components of the base accompaniment audio;

[0024] The width and sense of depth of the detailed components are enhanced by a preset rendering algorithm;

[0025] Adaptive gain adjustment is performed on the enhanced detail components;

[0026] A near-ear simulated sound source that matches the target scene is simulated based on the detailed components after adaptive gain adjustment.

[0027] Preferably, the step of spatially fusing the top simulated sound source, the surround simulated sound source, and the near-ear simulated sound source to obtain target spatial audio that conforms to the target scene includes:

[0028] The top simulated sound source, the surrounding simulated sound source, and the near-ear simulated sound source are accelerated by convolution;

[0029] The top simulated sound source, the surround simulated sound source, and the near-ear simulated sound source, which are accelerated by convolution, are shared and merged to obtain the target spatial audio that matches the target scene.

[0030] A second aspect of this application discloses an audio processing system, the system comprising:

[0031] The acquisition unit is used to acquire the user's dry audio and the basic accompaniment audio.

[0032] The first rendering unit is used to perform top sound source space rendering processing based on the user's dry audio to obtain a top simulated sound source that conforms to the target scene.

[0033] The second rendering unit is used to perform surround sound source space rendering processing based on the basic accompaniment audio to obtain a surround simulated sound source that conforms to the target scene.

[0034] The third rendering unit is used to perform near-ear detail rendering based on the basic accompaniment audio to obtain a near-ear simulated sound source that conforms to the target scene.

[0035] The fusion unit is used to spatially fuse the top simulated sound source, the surround simulated sound source, and the near-ear simulated sound source to obtain target spatial audio that conforms to the target scene.

[0036] Preferably, the first rendering unit includes:

[0037] The first acquisition module is used to acquire room impulse response data of the target scene;

[0038] The sidechain compression module is used to perform sidechain compression on the user's dry audio to obtain a compressed accompaniment signal;

[0039] The spatial reverberation simulation module is used to perform spatial reverberation simulation using the room impulse response data and the compressed accompaniment signal to obtain a simulated top sound source that matches the target scene.

[0040] Preferably, the second rendering unit includes:

[0041] The second acquisition module is used to acquire the surround sound room impulse response data of the target scene;

[0042] A spatial reverb enhancement module is used to enhance the spatial reverb of the basic accompaniment audio.

[0043] A simulation module is used to simulate the difference in the arrival time of a sound source at both ears in a target scene based on a head-related transfer function; wherein the difference includes at least a time difference and an intensity difference;

[0044] The determination module is used to determine the surround sound simulation source that matches the target scene based on the surround sound room impulse response data, the basic accompaniment audio after spatial reverberation enhancement, the time difference, and the intensity difference.

[0045] A third aspect of this application discloses a storage medium comprising stored instructions, wherein, when the instructions are executed, the device in which the storage medium is located is controlled to perform an audio processing method as described in any one of the first aspects.

[0046] The fourth aspect of this application discloses an electronic device including a memory and one or more instructions, wherein one or more instructions are stored in the memory and configured to be executed by one or more processors using the audio processing method as described in any of the first aspects.

[0047] As can be seen from the above technical solution, this application discloses an audio processing method, system, storage medium, and electronic device, which acquires user dry audio and basic accompaniment audio, performs top sound source spatial rendering processing based on user dry audio to obtain a top simulated sound source that conforms to the target scene, performs surround sound source spatial rendering processing based on basic accompaniment audio to obtain a surround simulated sound source that conforms to the target scene, performs near-ear detail rendering processing based on basic accompaniment audio to obtain a near-ear simulated sound source that conforms to the target scene, and performs spatial fusion processing on the top simulated sound source, surround simulated sound source, and near-ear simulated sound source to obtain target spatial audio that conforms to the target scene. Based on the above, top sound source spatial rendering processing is performed on the user's dry audio to simulate a top simulated sound source that matches the target scene. Surround sound source spatial rendering processing is performed on the basic accompaniment audio to simulate a surround simulated sound source in the target scene. Near-ear detail rendering processing is performed on the basic accompaniment audio to simulate a near-ear simulated sound source that matches the target scene. Since the top sound source in the scene has the open height of the main stage sound, the surround sound source in the scene has the surround diffusion of the venue, and the near-ear accompaniment has the detail fit, the sound field layer of the target scene with long-distance propagation, multi-directional surround and near-field details can be accurately reproduced. This is closer to physical reality, so as to realize the spatial immersion of a real concert when listening to music online. Attached Figure Description

[0048] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0049] Figure 1 This is a schematic flowchart of an audio processing method disclosed in an embodiment of this application;

[0050] Figure 2 This is a schematic diagram of the product side disclosed in an embodiment of this application;

[0051] Figure 3 This is a schematic diagram of the structure of an audio processing system disclosed in an embodiment of this application;

[0052] Figure 4 This is a schematic diagram of the structure of the electronic device disclosed in the embodiments of this application. Detailed Implementation

[0053] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.

[0054] In this application, the terms "comprising," "including," or any other variations thereof are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0055] As the background technology shows, when users listen to live concerts online, the sound field of the concert is a large open / semi-open space. Its acoustic characteristics include long sound propagation distance, significant air attenuation, and complex sound field superimposed by multiple speaker systems. This makes it difficult to locate the sound source in the current acoustic feature rendering, and thus it is difficult to reproduce the spatial immersion of a real concert.

[0056] To address the aforementioned issues, this application discloses an audio processing method, system, storage medium, and electronic device. It performs top-sound source spatial rendering processing on the user's dry audio to simulate a top-sound source matching the target scene; performs surround sound source spatial rendering processing on the basic accompaniment audio to simulate a surround sound source matching the target scene; and performs near-ear detail rendering processing on the basic accompaniment audio to simulate a near-ear sound source matching the target scene. Because the top-sound source in the scene possesses the expansive height of the main stage sound, the surround sound source in the scene possesses the surround diffusion of the venue's surroundings, and the near-ear accompaniment possesses detailed fit, it can accurately reproduce the sound field layers of the target scene, including long-distance propagation, multi-directional surround sound, and near-field details. This is closer to physical reality, thereby achieving a spatial immersion similar to a real concert when listening to music online. The specific implementation is described in detail through the following embodiments.

[0057] It should be noted that the audio processing method, system, storage medium, and electronic device provided in this application can be used in the fields of audio signal processing, acoustic simulation, and spatial audio rendering. The above is only an example and does not limit the application field of the audio processing method, system, storage medium, and electronic device provided in this application.

[0058] refer to Figure 1 The diagram shown is a flowchart of an audio processing method disclosed in an embodiment of this application. The audio processing method mainly includes the following steps:

[0059] S101: Obtain the user's dry audio and basic accompaniment audio.

[0060] User-generated dry audio refers to the original human voice captured directly by microphones or other acquisition devices during the recording process, without any post-processing or effects modification.

[0061] Basic accompaniment audio refers to the background music created in music production to highlight the main melody (such as vocals or lead instruments). Basic accompaniment audio provides support through elements such as harmony, rhythm, and bass lines, making the overall work fuller and more three-dimensional.

[0062] S102: Perform top sound source space rendering processing based on user dry audio to obtain a top simulated sound source that matches the target scene.

[0063] The types of target scenarios include, but are not limited to, concert venues, stadiums, and concert halls.

[0064] The user's dry audio is rendered using a top-sound-source spatial rendering process based on the first accompaniment layer and a physical model. This simulates the top-level sound source of the target scene, i.e., the sound source above the stage (such as the main stage speaker group), to recreate the sense of spaciousness and height experienced by the listener as the sound travels from a distance to the product side. The first accompaniment layer is the accompaniment layer for the simulated top-level sound source of the scene. (The product side is shown in the image.) Figure 2 As shown.

[0065] The physical model includes, but is not limited to, the distance decay model. The distance decay model is preferred in this application.

[0066] The distance attenuation model follows the spherical diffusion law in terms of sound propagation attenuation, and the formula for the distance attenuation model is shown in formula (1). Formula (1) represents the sound pressure level attenuation of the distance attenuation model.

[0067] (1)

[0068] in, The sound pressure level at a distance r from the sound source; For reference distance The sound pressure level at that location, Usually, 1 meter (m) is used. The air absorption coefficient, It is related to frequency, humidity, temperature, etc., and can be calculated by referring to tables or through empirical formulas.

[0069] The specific process of obtaining a top-mounted simulated sound source that matches the target scenario is shown in A1-A3.

[0070] A1: Collect room impulse response (IR) data for the target scene.

[0071] IR data can be synthesized using acoustic software and other methods.

[0072] A2: Perform sidechain compression on the user's dry audio to obtain the compressed accompaniment signal.

[0073] It should be noted that, in order to ensure the fusion of the human voice signal and the accompaniment, the user's dry audio obtained in the target scene needs to be side-chained compressed using dynamic range to ensure the dynamic adaptation of the accompaniment when the human voice signal is dominant. The specific formula for side-chain compression is shown in formula (2).

[0074] (2)

[0075] in, The input is dry audio from the user; R is the compressed signal after sidechain compression, i.e., the compressed accompaniment signal; T is the compression ratio; and T is the threshold.

[0076] A3: Spatial reverberation simulation is performed using the room impulse response IR data of the target scene and the compressed accompaniment signal to obtain a top simulated sound source that matches the target scene.

[0077] In A3, convolutional reverberation is used in the spatial reverberation simulation process. Specifically, the measured or synthesized IR data is convolved with the user's dry sound signal to obtain the simulated top sound source of the target scene, so as to simulate the sense of openness of long-distance propagation.

[0078] The specific formula for spatial reverberation simulation is shown in formula (3).

[0079] (3)

[0080] in, Simulate a sound source at the top of the target scene; For users' dry audio signals; For IR signals; * indicates convolution operation.

[0081] The simulated sound source at the top of the target scene is output to the corresponding channel to simulate the feeling of "sound coming from a high place in front".

[0082] In this embodiment, during the process of simulating the top sound source of the scene, the top sound source of the target scene is simulated through the first accompaniment layer and the physical model, thereby restoring the sense of openness and height of the sound traveling from a distance to the listener's ears, so that when listening to live concerts and other scenarios online, the spatial immersion of a real concert can be restored.

[0083] S103: Based on the basic accompaniment audio, perform surround sound source spatial rendering processing to obtain a surround simulated sound source that matches the target scene.

[0084] The second accompaniment layer and physical model are used to render the surrounding sound source space to simulate the diffuse sound sources (such as side field fill speakers and environmental reflections) around the venue / distant from the target scene, that is, the surround simulated sound source of the target scene, so as to restore the surround feeling of "sound entering the ears from a distance and from multiple directions".

[0085] The second accompaniment layer is an accompaniment layer that simulates the surround sound source of the scene.

[0086] The specific process of obtaining a surround sound source that matches the target scene is shown in B1-B4.

[0087] B1: Acquire surround sound room impulse response (IR) data for the target scene.

[0088] Among them, surround sound IR data can be generated through sound field simulation software and other methods.

[0089] B2: Spatial reverb enhancement for the basic accompaniment audio.

[0090] It should be noted that spatial reverb enhancement is applied to the basic accompaniment audio to increase the density of reflected sound and delay.

[0091] In B2, the reverberation model of the accompaniment layer surrounding the sound source in the scene is extended to increase the simulation of the later reflected sound, and the energy attenuation of the reflected sound conforms to formula (4).

[0092] (4)

[0093] in, The reflected sound energy at time t; As initial energy, The instantaneous acoustic energy of the original signal from the surrounding sound source in the corresponding scene, without propagation or reflection processing, at the output end of the sound source (such as the diaphragm of the side field supplementary sound box unit), does not include the energy lost through subsequent air propagation attenuation, interface reflection loss, and other energy loss processes. The reflection attenuation coefficient is... Related to the sound-absorbing materials used in the venue.

[0094] B3: Simulate the difference in sound source arrival at both ears in the target scene using the head-related transfer function (HRTF); where the difference includes at least the interaural time difference (ITD) and the interaural level difference (ILD).

[0095] In B3, the time difference and intensity difference of the sound source in the target scene reaching both ears (i.e., the binaural signal) are adjusted by the HRTF function to simulate a multi-directional surround sound effect.

[0096] The ITD is determined by the geometric dimensions of the human head; let the azimuth angle of the sound source be... The distance between the two ears is d, which is the distance between the two ears. The calculation formula of the ITD model is shown in formula (5).

[0097] (5)

[0098] in, d is the adjusted time difference between the two ears; d is the distance between the two ears, which is approximately 0.2m; c is the speed of sound, which is approximately 343 meters per second (m / s).

[0099] ILD is determined by the head occlusion effect and is related to frequency f. ILD can be obtained by fitting HRTF measurement data.

[0100] B4: Based on the surround sound room impulse response data, the basic accompaniment audio after spatial reverberation enhancement, the time difference, and the intensity difference, determine the surround simulated sound source that matches the target scene.

[0101] The surround sound room impulse response data, the basic accompaniment audio after spatial reverberation enhancement, the time difference of the adjusted binaural signals, and the intensity difference of the adjusted binaural signals are used to determine the surround sound simulation source that matches the target scene.

[0102] The surround sound room impulse response data, the basic accompaniment audio enhanced with spatial reverberation, the time difference of the adjusted binaural signals, and the intensity difference of the adjusted binaural signals are output to the side / rear channels to create an immersive feeling of "sound spreading from a distance".

[0103] In this embodiment, the accompaniment layer and physical model of the scene surround sound source are used to simulate the diffuse sound source around / far away from the venue that matches the target scene, i.e., the scene surround sound source, so as to restore the "surround feeling of sound entering the ears from a distance and from multiple directions".

[0104] S104: Perform near-ear detail rendering based on the basic accompaniment audio to obtain a near-ear simulated sound source that matches the target scene.

[0105] By using a third accompaniment layer (i.e., the ear-to-ear accompaniment layer) and a physical model, the basic accompaniment audio is rendered with near-ear detail to simulate the accompaniment close to the singer's / listener's ear. The accompaniment close to the singer's / listener's ear, such as the in-ear signal and near-field fill, restores the "detail and fit of sound propagation at close range".

[0106] The specific process of obtaining a near-ear simulated sound source that matches the target scenario is shown in C1-C4.

[0107] C1: Extract detailed components from the basic accompaniment audio.

[0108] Among them are detailed elements such as high-frequency instruments and vocal reverberation.

[0109] C2: Enhances the width and depth of detail components through preset rendering algorithms.

[0110] The preset rendering algorithm includes, but is not limited to, a near-field stereo rendering algorithm. The preset rendering algorithm of this application preferably uses a near-field stereo rendering algorithm.

[0111] In enhancing the breadth and layering of detail in the basic accompaniment audio through a near-field stereo rendering algorithm, a stereo expansion algorithm is employed to further enhance the width and layering of the sound. Let the left channel signal... The right channel signal is The extended signal is shown in formulas (6) and (7).

[0112] (6)

[0113] (7)

[0114] in, This is the left channel signal after being expanded using the stereo expansion algorithm. This is the left channel signal; The cross-mixing coefficient, Used to control the extension strength; For binaural delay, Used to simulate differences in head size; This is the right channel signal after being expanded using the stereo expansion algorithm. This is the right channel signal.

[0115] C3: Adaptive gain adjustment for enhanced detail components.

[0116] In C3, to highlight details, adaptive gain adjustment is performed on the enhanced detail components to emphasize near-field details.

[0117] The formula for adaptive gain adjustment is shown in Equation (8).

[0118] (8)

[0119] Where G is the gain and E is the signal energy; This is the gain coefficient. Used to balance detail and distortion.

[0120] C4: Simulate a near-ear simulated sound source that matches the target scene based on the detailed components after adaptive gain adjustment.

[0121] The adaptive gain-adjusted detail components are output to the corresponding channels, that is, the adaptive gain-adjusted detail components are output to the near-field channels to complete the process of simulating close-up accompaniment.

[0122] Near-field channels include, but are not limited to, in-ear monitoring simulation channels and front-row fill speakers, in order to simulate the feeling of "sound surrounding the ears".

[0123] S105: Spatial fusion processing is performed on the top simulated sound source, the surround simulated sound source, and the near-ear simulated sound source to obtain target spatial audio that matches the target scene.

[0124] The specific process of spatially fusing the top simulated sound source, the surround simulated sound source, and the near-ear simulated sound source is shown in D1-D2.

[0125] D1: Convolve and accelerate the top simulated sound source, the surround simulated sound source, and the near-ear simulated sound source.

[0126] The top simulated sound source, the surround simulated sound source, and the near-ear simulated sound source are convolved in blocks using the overlapping addition method. The convolved sound sources are then processed according to the fast Fourier transform to complete the process of optimizing the convolution calculation of the top simulated sound source, the surround simulated sound source, and the near-ear simulated sound source.

[0127] The long signals from the top analog sound source, the surround analog sound source, and the near-ear analog sound source are combined using the overlap-addition method. and of Perform block convolution processing. Let the length be N, then the calculation formula for the overlapping addition method is shown in formula (9).

[0128] (9)

[0129] in, It is a long signal (i.e., a signal with a length greater than IR). It refers to the input signal of the convolution operation, which corresponds to different input sources in different scenarios; The IR used is the intraocular impulse response (BRIR). for The kth signal.

[0130] Compared to existing pure temporal convolutions (complexity) Block convolution combined with Fast Fourier Transform (FFT) can reduce the complexity to Furthermore, block convolution combined with FFT can support real-time rendering and optimize computational efficiency.

[0131] D2: The top simulated sound source, the surround simulated sound source, and the near-ear simulated sound source after convolution acceleration are shared and merged to obtain the target spatial audio that matches the target scene.

[0132] It should be noted that since the top-mounted, surround-mounted, and near-ear-mounted analog sound sources are sound sources of the same data type but located in different positions within the same physical space, they meet the conditions for sharing and merging. The top-mounted, surround-mounted, and near-ear-mounted analog sound sources satisfy the conditions of having the same direction, type, and propagation path within the same physical space, such as multiple speakers on the left side of the stage playing the same accompaniment.

[0133] The IR values ​​of the top simulated sound source, the surround simulated sound source, and the near-ear simulated sound source can be combined into an equivalent value. . The calculation formula is shown in formula (10).

[0134] (10)

[0135] in, This is the result of convolving the IR and sound sources after merging, i.e., the result of merging multiple sound sources. Let be the IR of the i-th sound source; M is the number of sound sources.

[0136] For sound sources of the same type and on the same side of the stage, such as the left-side supplementary sound box group, their IR is merged and then convolved with the audio signal. According to relevant experiments, the number of convolutions can be reduced from M to 1, thereby reducing the amount of computation and supporting real-time processing of multi-sound source superposition scenes, thus optimizing computational efficiency.

[0137] This application uses layered rendering and acoustic models to accurately reproduce the sound field layers of the target scene, including "long-distance propagation, multi-directional surround sound, and near-field details." Listeners can perceive the spaciousness and height of the main stage sound source, the surround diffusion around the venue, and the close-fitting details of the accompaniment, thereby achieving a sense of spatial immersion similar to a real concert when listening to music online.

[0138] Based on acoustic theory formulas, such as spherical attenuation and HRTF models, and without relying on empirical fitting, it is closer to physical reality and can be adapted to different scenarios, such as concerts, stadiums, and concert halls, to achieve a realistic experience for the audience.

[0139] This application relates to the technical fields of audio signal processing, acoustic simulation and spatial audio rendering, focusing on large open space scenarios such as concerts. By accurately simulating the acoustic propagation characteristics, it achieves immersive audio rendering and enhances the recording and listening experience.

[0140] This application utilizes layered rendering based on acoustic theory, simulating a concert sound field through a "three-layer accompaniment layer." By combining convolutional computation optimization and IR applications, it achieves a balance between realism and efficiency. It enables immersive audio rendering of the target scene, accurately reproducing the long-distance propagation characteristics of stage sound sources, the complex sound field layers of multiple superimposed speakers, and the attenuation and reflection patterns of sound in open spaces. This allows for a truly immersive spatial experience when listening to music online. Furthermore, it supports real-time processing of multi-source superimposed scenes, optimizing computational efficiency.

[0141] This application embodiment performs top sound source spatial rendering processing on the user's dry audio to simulate a top simulated sound source that matches the target scene, performs surround sound source spatial rendering processing on the basic accompaniment audio to simulate a surround simulated sound source in the target scene, and performs near-ear detail rendering processing on the basic accompaniment audio to simulate a near-ear simulated sound source that matches the target scene. Since the top sound source in the scene has the open height of the main stage sound, the surround sound source in the scene has the surround diffusion of the venue, and the near-ear accompaniment has the detail fit, it can accurately reproduce the sound field layers of the target scene with long-distance propagation, multi-directional surround and near-field details. This is closer to physical reality, thereby realizing the spatial immersion of a real concert when listening to music online. Furthermore, by optimizing the convolution calculation of the top simulated sound source, the surround simulated sound source, and the near-ear simulated sound source, and by performing spatial fusion processing on the top simulated sound source, the surround simulated sound source, and the near-ear simulated sound source, i.e., multi-sound source sharing and merging, the computational complexity of real-time rendering sound sources is reduced, and real-time processing of multi-sound source superposition scenes is supported, thereby improving the computational efficiency of real-time rendering sound sources.

[0142] Based on the above embodiments Figure 1 The disclosed audio processing method, in addition to the corresponding audio processing system, also includes an audio processing method described in this application. Figure 3 As shown, the audio processing system includes:

[0143] Acquisition unit 301 is used to acquire user dry audio and basic accompaniment audio;

[0144] The first rendering unit 302 is used to perform top sound source space rendering processing based on the user's dry audio to obtain a top simulated sound source that conforms to the target scene.

[0145] The second rendering unit 303 is used to perform surround sound source space rendering processing based on the basic accompaniment audio to obtain a surround simulated sound source that conforms to the target scene.

[0146] The third rendering unit 304 is used to perform near-ear detail rendering based on the basic accompaniment audio to obtain a near-ear simulated sound source that matches the target scene.

[0147] The fusion unit 305 is used to perform spatial fusion processing on the top simulated sound source, the surround simulated sound source and the near-ear simulated sound source to obtain target spatial audio that conforms to the target scene.

[0148] Furthermore, the first rendering unit 302 includes:

[0149] The first acquisition module is used to acquire room impulse response data of the target scene;

[0150] The sidechain compression module is used to perform sidechain compression on the user's dry audio to obtain the compressed accompaniment signal;

[0151] The spatial reverberation simulation module is used to simulate spatial reverberation using room impulse response data and compressed accompaniment signals to obtain a simulated top sound source that matches the target scene.

[0152] Furthermore, the second rendering unit 303 includes:

[0153] The second acquisition module is used to acquire the surround sound room impulse response data of the target scene;

[0154] The spatial reverb enhancement module is used to enhance the spatial reverb of the basic accompaniment audio.

[0155] The simulation module is used to simulate the difference in the arrival time of sound sources at both ears in a target scene based on the head-related transfer function; wherein the difference includes at least the time difference and the intensity difference.

[0156] The determination module is used to determine the surround sound simulation source that matches the target scene based on the surround sound room impulse response data, the basic accompaniment audio after spatial reverberation enhancement, the time difference, and the intensity difference.

[0157] Furthermore, the third rendering unit 304 includes:

[0158] The extraction module is used to extract detailed components from the basic accompaniment audio.

[0159] The enhancement module is used to enhance the width and depth of detail components through preset rendering algorithms;

[0160] The adjustment module is used to adaptively adjust the gain of the enhanced detail components;

[0161] The simulation module is used to simulate a near-ear simulated sound source that matches the target scene based on the detailed components after adaptive gain adjustment.

[0162] Furthermore, the fusion unit 305 includes:

[0163] The convolution acceleration module is used to accelerate the top simulated sound source, the surround simulated sound source, and the near-ear simulated sound source through convolution.

[0164] The shared merging module is used to share and merge the top simulated sound source, the surround simulated sound source, and the near-ear simulated sound source after convolution acceleration to obtain target spatial audio that matches the target scene.

[0165] This application embodiment performs top sound source spatial rendering processing on the user's dry audio to simulate a top simulated sound source that matches the target scene, performs surround sound source spatial rendering processing on the basic accompaniment audio to simulate a surround simulated sound source in the target scene, and performs near-ear detail rendering processing on the basic accompaniment audio to simulate a near-ear simulated sound source that matches the target scene. Since the top sound source in the scene has the open height of the main stage sound, the surround sound source in the scene has the surround diffusion of the venue, and the near-ear accompaniment has the detail fit, it can accurately reproduce the sound field layers of the target scene with long-distance propagation, multi-directional surround and near-field details. This is closer to physical reality, thereby realizing the spatial immersion of a real concert when listening to music online. Furthermore, by optimizing the convolution calculation of the top simulated sound source, the surround simulated sound source, and the near-ear simulated sound source, and by performing spatial fusion processing on the top simulated sound source, the surround simulated sound source, and the near-ear simulated sound source, i.e., multi-sound source sharing and merging, the computational complexity of real-time rendering sound sources is reduced, and real-time processing of multi-sound source superposition scenes is supported, thereby improving the computational efficiency of real-time rendering sound sources.

[0166] This application embodiment also provides a storage medium, the storage medium including stored instructions, wherein, when the instructions are executed, the device where the storage medium is located is controlled to perform the audio processing method described above.

[0167] This application also provides an electronic device, the structural schematic diagram of which is shown below. Figure 4 As shown, it specifically includes a memory 401 and one or more instructions 402, wherein one or more instructions 402 are stored in the memory 401 and configured to be executed by one or more processors 403 to perform the above-mentioned audio processing method.

[0168] For the foregoing method embodiments, in order to simplify the description, they are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, because according to this application, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.

[0169] It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For apparatus embodiments, since they are basically similar to method embodiments, the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments.

[0170] The steps in the methods of the various embodiments of this application can be adjusted, combined, or deleted according to actual needs.

[0171] Finally, it should be noted that in this paper, relational terms such as first and second are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations.

[0172] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

[0173] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.

Claims

1. An audio processing method, characterized in that, The method includes: Obtain the user's dry audio and basic accompaniment audio; Based on the user's dry audio, a top sound source space rendering process is performed to obtain a top simulated sound source that conforms to the target scene. Based on the basic accompaniment audio, surround sound source space rendering processing is performed to obtain a surround simulated sound source that conforms to the target scene; Based on the aforementioned basic accompaniment audio, near-ear detail rendering is performed to obtain a near-ear simulated sound source that matches the target scene; The top simulated sound source, the surround simulated sound source, and the near-ear simulated sound source are spatially fused to obtain target spatial audio that matches the target scene.

2. The method according to claim 1, characterized in that, The step of performing top sound source spatial rendering processing based on the user's dry audio to obtain a top simulated sound source that conforms to the target scene includes: Acquire room impulse response data for the target scene; The user's dry audio is sidechained to obtain the compressed accompaniment signal; Spatial reverberation simulation is performed using the room impulse response data and the compressed accompaniment signal to obtain a simulated top sound source that matches the target scene.

3. The method according to claim 1, characterized in that, The process of rendering the surround sound source space based on the basic accompaniment audio to obtain a surround simulated sound source that conforms to the target scene includes: Acquire surround sound room impulse response data for the target scene; Spatial reverb enhancement is applied to the base accompaniment audio; The difference in the arrival time of a sound source at both ears in a target scene is simulated based on the head-related transfer function; wherein, the difference includes at least a time difference and an intensity difference; Based on the surround sound room impulse response data, the basic accompaniment audio after spatial reverberation enhancement, the time difference, and the intensity difference, a surround simulated sound source that matches the target scene is determined.

4. The method according to claim 1, characterized in that, The near-ear detail rendering process based on the basic accompaniment audio to obtain a near-ear simulated sound source that conforms to the target scene includes: Extract the detailed components of the base accompaniment audio; The width and sense of depth of the detailed components are enhanced by a preset rendering algorithm; Adaptive gain adjustment is performed on the enhanced detail components; A near-ear simulated sound source that matches the target scene is simulated based on the detailed components after adaptive gain adjustment.

5. The method according to claim 1, characterized in that, The step of spatially fusing the top simulated sound source, the surround simulated sound source, and the near-ear simulated sound source to obtain target spatial audio that matches the target scene includes: The top simulated sound source, the surrounding simulated sound source, and the near-ear simulated sound source are accelerated by convolution; The top simulated sound source, the surround simulated sound source, and the near-ear simulated sound source, which are accelerated by convolution, are shared and merged to obtain the target spatial audio that matches the target scene.

6. An audio processing system, characterized in that, The system includes: The acquisition unit is used to acquire the user's dry audio and the basic accompaniment audio. The first rendering unit is used to perform top sound source space rendering processing based on the user's dry audio to obtain a top simulated sound source that conforms to the target scene. The second rendering unit is used to perform surround sound source space rendering processing based on the basic accompaniment audio to obtain a surround simulated sound source that conforms to the target scene. The third rendering unit is used to perform near-ear detail rendering based on the basic accompaniment audio to obtain a near-ear simulated sound source that conforms to the target scene. The fusion unit is used to spatially fuse the top simulated sound source, the surround simulated sound source, and the near-ear simulated sound source to obtain target spatial audio that conforms to the target scene.

7. The system according to claim 6, characterized in that, The first rendering unit includes: The first acquisition module is used to acquire room impulse response data of the target scene; The sidechain compression module is used to perform sidechain compression on the user's dry audio to obtain a compressed accompaniment signal; The spatial reverberation simulation module is used to perform spatial reverberation simulation using the room impulse response data and the compressed accompaniment signal to obtain a simulated top sound source that matches the target scene.

8. The system according to claim 6, characterized in that, The second rendering unit includes: The second acquisition module is used to acquire the surround sound room impulse response data of the target scene; A spatial reverb enhancement module is used to enhance the spatial reverb of the basic accompaniment audio. A simulation module is used to simulate the difference in the arrival time of a sound source at both ears in a target scene based on a head-related transfer function; wherein the difference includes at least a time difference and an intensity difference; The determination module is used to determine the surround sound simulation source that matches the target scene based on the surround sound room impulse response data, the basic accompaniment audio after spatial reverberation enhancement, the time difference, and the intensity difference.

9. A storage medium, characterized in that, The storage medium includes stored instructions, wherein, when the instructions are executed, the device containing the storage medium is controlled to perform the audio processing method as described in any one of claims 1 to 5.

10. An electronic device, characterized in that, It includes a memory, and one or more instructions, wherein one or more instructions are stored in the memory and configured to be executed by one or more processors as described in any one of claims 1 to 5.