Schrodinger bridge-based spatial audio synthesis method and system

By adopting a Schrödinger bridge-based method in spatial audio synthesis, using the iterative sampling path of the random differential equation and the parameterization goal of the two-channel Schrödinger bridge, the problems of slow and poor quality of spatial audio synthesis in the existing technology are solved, and high-quality and fast spatial audio synthesis is achieved.

CN119964544AActive Publication Date: 2025-05-09TSINGHUA UNIVERSITY

Patent Information

Application Number
CN202411954295.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-27
Publication Date
2025-05-09
Estimated Expiration
2044-12-27

AI Technical Summary

Technical Problem

Existing spatial audio synthesis technology has problems of slow speed and poor quality.

Method used

Using the Spatial Audio Synthesis Method based on Schrödinger Bridge, two-channel spatial audio is generated through the pre-trained spatial audio synthesis model and the iterative sampling path of the stochastic differential equation. This method trains the neural network model through parameterized targets and boundary-assisted supervision of the two-channel Schrödinger bridge.

Benefits of technology

It significantly improves the quality of spatial audio synthesis, surpasses the performance of previous methods on five indicators, and greatly improves the synthesis speed by compressing the sampling steps, making it comparable to the single-step inference speed of end-to-end neural networks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119964544A_ABST
    Figure CN119964544A_ABST
Patent Text Reader

Abstract

The invention provides a Schrodinger bridge-based spatial audio synthesis method and system, and the method comprises the steps: obtaining a monaural sound source signal, and constructing a prior signal and a noisy representation based on the monaural sound source signal; inputting the prior signal and the noisy representation into a pre-trained spatial audio synthesis model, and generating a final dual-track spatial audio through the spatial audio synthesis model based on a stochastic differential equation iterative sampling path; wherein the spatial audio synthesis model is obtained by training a preset neural network model through a predefined dual-channel Schrodinger bridge parameterization target and boundary auxiliary supervision. According to the invention, the problems of slow synthesis speed and poor quality of the existing spatial audio are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of audio synthesis, and in particular to a spatial audio synthesis method and system based on Schrödinger bridge. Background Art

[0002] Binaural Audio Synthesis (BAS) refers to the process of synthesizing a monophonic audio signal at the sound source into a binaural audio signal based on relative position, head orientation and other information. In the real world, the process of the audio signal at the sound source being transmitted to the left and right ears of the receiver is affected by many factors, such as the indoor environment, the receiver's head model, and environmental noise. How to synthesize spatial audio signals with high quality and speed is a key issue in building a BAS system.

[0003] Compared with mono audio signals, dual-channel audio enables users to sense spatial position information from audio signals. In virtual reality, augmented reality and other scenarios such as online conference rooms, movies, and games, spatial audio synthesis technology can greatly enhance the user's sense of immersion and experience. Therefore, this technology has a wide range of applications in these scenarios.

[0004] In the past methods, spatial audio synthesis technology can be divided into three main technical routes. The first route is based on traditional signal processing methods (Digital Signal Processing Methods, DSP). The advantage of this type of method is that there is no need to train neural networks, but it requires the collection of many physical parameters to represent indoor environmental information (such as Room Impulsive Response, RIR), user head information (such as Head-related Transfer Function, HRTF), and even environmental noise (Environmental Noise) and other information. Because the measurement of this information is very expensive and has high requirements for the measurement device, this type of method is usually unable to obtain accurate physical parameters in actual application scenarios, and can only use approximate values, which inevitably brings modeling errors. In addition, this type of method is also unable to model the nonlinear process in spatial audio synthesis. These two defects seriously limit the quality of spatial audio synthesis of DSP methods.

[0005] The second technical route is based on deterministic regression neural networks (Mapping Based-Networks). This type of method uses neural networks such as convolutional neural networks to fit the synthesis process of mono audio to binaural audio end-to-end. Compared with traditional signal processing methods, this type of method no longer requires expensive parameter measurement processes, and can simultaneously model linear and nonlinear changes in spatial audio synthesis, thereby improving the quality of both the main and guest synthesis. However, because these methods usually use objective indicators of the waveform space (Waveform), amplitude spectrum space (Amplitude), or phase spectrum space (Phase) of binaural signals, such as mean square error (MSE), as neural network training, and sample deterministic single-step regression for binaural audio synthesis, the synthesis quality is still relatively limited and cannot effectively capture the detailed information of high-sampling rate spatial audio. The third technical route is based on diffusion models. This type of work no longer directly models the propagation process from mono to binaural, but transforms this process into a conditional generation process, using conditional diffusion models to synthesize spatial audio. A two-stage conditional diffusion model was developed. In the first stage, a single-channel diffusion model models the mean part of the binaural signal from the mono signal at the sound source to the left and right ears of the receiver, that is, the "common information" of the binaural signal is modeled. In the second stage, a dual-channel diffusion model models the ground-truth binaural audio signals from the mean of the left and right ear signals to the left and right ear signals, that is, the "specific information" of the left and right ears is modeled from the "common information". The overall design of BinauralGrad realizes the generation process from commonality to specificity, from coarseness to fine-grainedness in spatial audio tasks, and achieves subjective and objective synthesis quality that exceeds previous work.

[0006] Among the three technical routes mentioned above, the traditional signal processing DSP method uses manually constructed filters to model various factors that affect sound wave transmission. The physical parameters of these filters, such as RIR and HRTF, usually require expensive or time-consuming measurement processes. In actual application scenarios, approximate values ​​are usually used because accurate physical parameters cannot be obtained. On the other hand, the complex spatial audio transmission process cannot be modeled using only simple linear filters, and the nonlinear change process cannot be modeled using DSP methods. Therefore, the quality of spatial audio synthesis based on DSP methods is always limited.

[0007] The end-to-end neural network method based on deterministic regression cannot accurately capture the detailed information in the audio waveform due to its optimization goal and single-step regression synthesis method. Although the quality is improved compared to the DSP method, the synthesis quality is still limited.

[0008] The spatial audio synthesis method based on the diffusion model uses an iterative sampling synthesis method to gradually model the spatial audio, which has greatly improved the quality compared with previous methods. However, its synthesis process requires 12 steps of iterative sampling, and each sampling step needs to be calculated in the audio waveform space with a sampling rate of 48kHz. Therefore, its synthesis speed is very slow, much slower than previous methods. Summary of the invention

[0009] The present invention provides a spatial audio synthesis method and system based on Schrödinger bridge, which are used to solve the problems of slow speed and poor quality of existing spatial audio synthesis.

[0010] The present invention provides a spatial audio synthesis method based on Schrödinger bridge, comprising: Acquire a monophonic sound source signal, and construct a priori signal and a noisy representation based on the monophonic sound source signal; Inputting the prior signal and the noisy representation into a pre-trained spatial audio synthesis model, and generating final binaural spatial audio through the spatial audio synthesis model based on a stochastic differential equation iterative sampling path; The spatial audio synthesis model is obtained by training a preset neural network model through a predefined binaural Schrödinger bridge parameterized target and boundary auxiliary supervision.

[0011] According to a spatial audio synthesis method based on Schrödinger bridge provided by the present invention, the spatial audio synthesis model is obtained by training a preset neural network model through a predefined binaural Schrödinger bridge parameterization target and boundary auxiliary supervision, specifically comprising: Obtain the quaternion of the monophonic sound source signal, the relative spatial position coordinates of the sound source and the receiver, and the receiver's head orientation information; Building a binaural Schrodinger bridge between the monophonic sound source signal and the audio signals received by the left and right ears of the receiver; Obtaining a noisy representation at each time step in the two-channel Schrödinger bridge, and parameterizing the two-channel Schrödinger bridge by shallowly adding noise based on the noisy representation; Based on the binaural Schrödinger bridge parameterization target, the preset spatial audio synthesis model is trained through the relative spatial position coordinates of the sound source and the receiver, the quaternion of the receiver's head orientation information, and boundary-assisted supervision.

[0012] According to a spatial audio synthesis method based on Schrödinger bridge provided by the present invention, a binaural Schrödinger bridge is built between the monophonic sound source signal and the audio signals received by the left and right ears of the receiver, specifically comprising: Duplicating the monophonic sound source signal to form a dual-channel speech signal; The audio signals received by the left and right ears of the receiver are used as the generation target, and the binaural speech signal is used as the prior signal of the audio signals received by the left and right ears of the receiver; A binaural Schrödinger bridge is built between the prior signal and the generated target.

[0013] According to a spatial audio synthesis method based on Schrödinger bridge provided by the present invention, the step of obtaining a noisy representation at each time step in the binaural Schrödinger bridge and parameterizing the binaural Schrödinger bridge by shallowly adding noise based on the noisy representation specifically includes: Obtaining a noisy representation at each time step in the binaural Schrödinger bridge; Based on the noisy characterization, the noise content is first increased to the peak of the noise content, and then the noise content is gradually reduced until the noise disappears, thereby completing the parameterization of the binaural Schrodinger bridge.

[0014] According to a spatial audio synthesis method based on Schrödinger bridge provided by the present invention, The training of the preset spatial audio synthesis model is completed based on the binaural Schrödinger bridge parameterization target through the relative spatial position coordinates of the sound source and the receiver, the quaternion of the receiver's head orientation information, and boundary auxiliary supervision, specifically including: The quaternion of the relative spatial position coordinates between the sound source and the receiver and the receiver's head orientation information is used as conditional information to guide the training of a preset spatial audio synthesis model; Based on the binaural Schrödinger bridge parameterization target, the quaternion of the relative spatial position coordinates between the sound source and the receiver and the receiver's head orientation information is used as conditional information to guide the training of the preset spatial audio synthesis model. The loss function is added through boundary auxiliary supervision to supervise the training of the preset spatial audio synthesis model. The loss function includes: multi-scale amplitude loss and phase spectrum loss.

[0015] According to a spatial audio synthesis method based on Schrödinger bridge provided by the present invention, the prior signal and the noisy representation are input into a pre-trained spatial audio synthesis model, and the final two-channel spatial audio is generated through the spatial audio synthesis model based on a random differential equation iterative sampling path, which specifically includes: Inputting the prior signal and the noisy representation into a spatial audio synthesis model to generate a single-step prediction target; Based on the single-step prediction target, sampling is performed through a stochastic differential equation iterative sampling path, and the final two-channel spatial audio is synthesized according to the sampling results.

[0016] The present invention also provides a spatial audio synthesis system based on Schrödinger bridge, the system comprising: An information acquisition module, used to acquire a monophonic sound source signal, and construct a priori signal and a noisy representation based on the monophonic sound source signal; An audio synthesis module, configured to input the prior signal and the noisy representation into a pre-trained spatial audio synthesis model, and generate final binaural spatial audio through the spatial audio synthesis model based on an iterative sampling path of a stochastic differential equation; The spatial audio synthesis model is obtained by training a preset neural network model through a predefined binaural Schrödinger bridge parameterized target and boundary auxiliary supervision.

[0017] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the spatial audio synthesis method based on the Schrödinger bridge as described above is implemented.

[0018] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the spatial audio synthesis method based on the Schrödinger bridge as described above is implemented.

[0019] The present invention also provides a computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the spatial audio synthesis method based on the Schrödinger bridge as described above is implemented.

[0020] The present invention provides a Schrödinger bridge-based spatial audio synthesis method and system, which improve the quality of spatial audio synthesis through a new Schrödinger bridge-based spatial audio synthesis model; significantly surpass the previous methods in five indicators for measuring synthesis quality; and can compress the sampling steps based on the random differential equation iterative sampling path, greatly improving the inference speed of the spatial audio synthesis system based on iterative sampling. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0022] Figure 1It is a flow chart of a spatial audio synthesis method based on Schrödinger bridge provided by the present invention.

[0023] Figure 2 It is a noise scheduling comparison schematic diagram provided by the present invention.

[0024] Figure 3 The present invention provides a schematic diagram of module connection of a spatial audio synthesis system based on a Schrödinger bridge.

[0025] Figure 4 It is a structural schematic diagram of the electronic device provided by the present invention.

[0026] Reference numerals: 110: information acquisition module; 120: audio synthesis module; 410: processor; 420: communication interface; 430: memory; 440: communication bus. DETAILED DESCRIPTION

[0027] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0028] Previous spatial audio synthesis methods, such as DSP and end-to-end single-step regression based on neural networks, have limited stereo audio synthesis quality. The synthesis quality of the diffusion model-based system is acceptable, but the sampling speed is slow, requiring 12 steps of iterative sampling in the stereo 48kHz waveform space.

[0029] The purpose of this invention is to propose a new spatial audio synthesis framework, which surpasses all previous methods in synthesis quality while being comparable to the single-step reasoning speed of an end-to-end neural network in terms of inference speed, thus realizing for the first time a high-quality, fast and efficient spatial audio synthesis system.

[0030] Combine the following Figure 1 A spatial audio synthesis method based on Schrödinger bridge of the present invention is described, comprising: Step 100: Obtain a monophonic sound source signal, and construct a priori signal and a noisy representation based on the monophonic sound source signal.

[0031] The present invention adopts a generative model based on Schrödinger bridge to realize the spatial audio synthesis from mono to binaural by data-to-data generation. That is, the sound wave transmission process from mono sound source to binaural spatial audio is implicitly modeled by using an iterative sampling path based on stochastic differential equations (SDE).

[0032] In the spatial audio synthesis task, given a monophonic sound source signal , represents the coordinates of the relative spatial position of the sound source and the receiver , and the quaternion representing the receiver's head orientation information , the spatial audio synthesis system aims to synthesize the two-channel audio signals received by the left and right ears of the receiver , and Represents the signals received by the left and right ears respectively, and the sound source signal have the same dimensions.

[0033] Under this task definition, the performance of the spatial audio synthesis system mainly includes (1) synthesis quality, (2) synthesis speed, and (3) the number of model parameters. Higher synthesis quality, faster synthesis speed, and fewer model parameters are the main goals of establishing a spatial audio synthesis system.

[0034] Step 200: input the prior signal and the noisy representation into a pre-trained spatial audio synthesis model, and generate a final two-channel spatial audio through the spatial audio synthesis model based on an iterative sampling path of a random differential equation; wherein the spatial audio synthesis model is obtained by training a preset neural network model through a predefined two-channel Schrödinger bridge parameterized target and boundary-assisted supervision.

[0035] Specifically including: obtaining the quaternion of the monophonic sound source signal, the relative spatial position coordinates between the sound source and the receiver, and the receiver's head orientation information; Building a binaural Schrodinger bridge between the monophonic sound source signal and the audio signals received by the left and right ears of the receiver; Obtaining a noisy representation at each time step in the two-channel Schrödinger bridge, and parameterizing the two-channel Schrödinger bridge by shallowly adding noise based on the noisy representation; Based on the binaural Schrödinger bridge parameterization target, the preset spatial audio synthesis model is trained through the relative spatial position coordinates of the sound source and the receiver, the quaternion of the receiver's head orientation information, and boundary-assisted supervision.

[0036] The Schrödinger bridge-based generative model used in the present invention is defined by the following forward and reverse stochastic differential equations: .

[0037] in and They represent the two endpoint distributions of the "data to data" synthesis process of the Schrödinger bridge, namely the prior distribution and the target data distribution. and are the drift term and diffusion term of the stochastic differential equation, and Represents the Brownian motion in forward and reverse time in Schrödinger. and When both boundary distributions are Dirac-delta distributions, the nonlinear drift components of the above forward and inverse stochastic differential equations are and The definition is as follows: .

[0038] .

[0039] Specifically, a binaural Schrödinger bridge is built between the monophonic sound source signal and the audio signals received by the left and right ears of the receiver, including: Duplicating the monophonic sound source signal to form a dual-channel speech signal; The audio signals received by the left and right ears of the receiver are used as the generation target, and the monophonic sound source signal is used as the prior signal of the audio signals received by the left and right ears of the receiver; A binaural Schrödinger bridge is built between the prior signal and the generated target.

[0040] In the present invention, a spatial audio synthesis system based on Schrödinger bridge is proposed. In the BAS task, a monophonic signal at the sound source is used. At the same time, the left and right ears of the receiver receive the signal and Specifically, the monophonic speech signal at the sound source is copied to form a binaural speech signal as a priori signal. , taking the audio signals received by the left and right ears of the receiver as the generation target , a two-channel Schrödinger bridge is built between the information of the two, where the nonlinear drift terms in the forward and reverse SDE become: .

[0041] Specifically, obtaining a noisy representation at each time step in the binaural Schrödinger bridge; Based on the noisy characterization, the noise content is first increased to the peak of the noise content, and then the noise content is gradually reduced until the noise disappears, thereby completing the parameterization of the binaural Schrodinger bridge.

[0042] In the present invention, given the prior signal and the generated target, the noisy representation at each time t in the Schrödinger bridge built on the BAS task is calculated as: .

[0043] In the parameterization target of Schrödinger bridge, "noise prediction" is used as the parameterization target of binaural Schrödinger bridge, and the training target of spatial audio synthesis model is: .

[0044] Among them, the coordinates representing the relative position information , and the quaternion representing the receiver's head orientation information They are also used as neural network inputs to guide the spatial audio synthesis process as conditional information. They have been omitted in the above formula for simplicity of expression.

[0045] In previous Schrödinger bridge research, only symmetric noise scheduling and asymmetric noise scheduling were discussed. However, previous research has not yet studied the appropriate degree of noise addition in the Schrödinger forward process. In the present invention, the intensity of the noise part in the Schrödinger bridge is studied for the first time.

[0046] Compared with the "noise to data" generation method of the diffusion model, the "data to data" synthesis process of the Schrödinger bridge makes its noise content not gradually increase from t=0 to t=1 like the diffusion model, but is divided into two stages: the noise content first increases to the peak of the noise content, and then decreases to almost nothing.

[0047] At the peak of the noise content, if the noise intensity is too large, the prior signal will be covered by the noise to a large extent, that is, the prompt information provided by the prior signal to the generated target will be seriously damaged by the noise. Compared with the "noise to data" synthesis paradigm of the diffusion model, the "data to data" paradigm of the Schrödinger bridge will partially lose its advantage of utilizing the prior signal due to the excessive noise content.

[0048] refer to Figure 2 , in the spatial audio task, two noise schedules are shown, with DeepNoise on the left and ShallowNoise on the right. It can be seen that at the peak of the noise content of the Schrödinger bridge, the structure of the prior signal (Prior) has been destroyed in the representation of the DeepNoise schedule. In the ShallowNoise schedule, the noise amplitude is extremely small compared to the speech signal amplitude. At the peak of the noise content, the dual-channel Schrödinger bridge representation still maintains the contour information of the speech signal.

[0049] Comparing the left and right figures, it can be seen that when using shallow noise technology, the dual-channel Schrödinger bridge is easier to maintain voice signal information and generate target data (Data), i.e., dual-channel audio, with less burden, thereby improving the sample synthesis quality and speed compared to the deep noise design.

[0050] Based on the binaural Schrödinger bridge parameterized objective, the preset spatial audio synthesis model is trained through boundary auxiliary supervision and an additional loss function is added as auxiliary supervision; The loss function includes multi-scale including: multi-scale amplitude loss and phase spectrum loss.

[0051] In the present invention, according to the training objective of the spatial audio synthesis model, after about 1M steps of model iteration, the neural network is characterized by the noise at time t and at time t. As input, you will get the output: .

[0052] In the past work related to Schrödinger bridge, different model parameterization methods such as noise, data, speed, score, etc. have been explored. However, relying solely on these equivalent parameterization methods may not necessarily achieve the best sample synthesis quality and synthesis speed. Therefore, this invention is the first to directly consider the sample quality as the optimization target of Schrödinger bridge. Specifically, in addition to the original Schrödinger bridge optimization target In addition, an additional loss function is added as auxiliary supervision: .

[0053] in, and represents the additional loss function at t=0 and t=1 of the Schrödinger bridge, which is defined as follows: .

[0054] That is, the clean data at t=0 and t=1 are used as auxiliary supervision targets. In the spatial audio synthesis task, the mono source audio and binaural spatial audio are used as auxiliary supervision targets to strengthen the training of the neural network. and Represents two loss functions commonly used to characterize the quality of audio data: Multi-Scale Short-Time Fourier Transform and Phase loss. to Represents the weight of the corresponding loss function, which is an adjustable parameter during model training.

[0055] In actual calculation, at each step of neural network training, for generating target The single-step calculation of has been shown above, which can be calculated Using the output of the spatial audio synthesis model, the t=1 prior signal can also be calculated in a single step as follows: .

[0056] You can calculate ,Finally, as shown above, when using auxiliary supervision technology, the loss function of the spatial audio synthesis system based on Schrödinger bridge consists of three parts, taking into account the optimization of Schrödinger bridge and the optimization of sample quality at each sampling step.

[0057] Step 500: Based on the spatial audio synthesis model, the iterative sampling path of the stochastic differential equation is used to implicitly model the sound wave transmission process from the monophonic sound source signal to the binaural spatial audio, and the final binaural spatial audio is iteratively generated.

[0058] Specifically, it includes: using the noisy representation and prior signal acquired in real time as the input of the spatial audio synthesis model, and outputting a single-step prediction generation target; The target is generated based on a single-step prediction sampling method based on stochastic differential equations, and the final two-channel spatial audio is iteratively generated.

[0059] In the present invention, at each step of sampling, according to the output of the spatial audio synthesis model, the target can be first predicted and generated in a single step: .

[0060] The final two-channel spatial audio is iteratively generated according to the sampling method based on ordinary differential equations (ODE): .

[0061] By building a dual-channel Schrödinger bridge on the spatial audio task, unlike the three technical paradigms on the three previous BAS tasks, the spatial audio synthesis process is implicitly modeled using the "data to data" synthesis process based on the probabilistic generative model for the first time, directly building a generative model between the mono audio at the sound source and the binaural audio at the receiver, and no longer limited to the "noise to data" synthesis paradigm of the diffusion model. Compared with methods such as single-step regression and DSP, the probabilistic generative model has stronger expressive power in fitting data distribution, and can usually capture more fine-grained information and achieve higher accuracy in audio waveform modeling.

[0062] In a specific embodiment, the spatial audio synthesis quality of the present invention is compared with the previous three technical routes. In Table 1, It is a spatial audio synthesis system built. Auxiliary Supervision (AS) is added to the system. Other methods include represents the DSP method implemented by BinauralGrad, represents the experimental results achieved on WarpNet proposed by DopplerBAS, and BG represents the experimental results of the traditional diffusion model. It can be seen that in five objective indicators such as waveform, spectrum amplitude, phase amplitude, etc., it has surpassed the previous three technical routes at the same time.

[0063] Table 1 .

[0064] See Table 2 for a comparison of the computational efficiency of the spatial audio synthesis system with the three previous representative methods.

[0065] Table 2 .

[0066] As shown in Table 2, the traditional signal processing method DSP only needs to be calculated on the CPU, does not need to use GPU for inference, and does not require model training, but its synthesis quality is very limited. WarpNet only needs a single-stage neural network, end-to-end one-step prediction and synthesis of spatial audio, and performs best in the real-time factor (RTF) that measures the synthesis speed. The model size is 8.59M.

[0067] A spatial audio synthesis system based on a diffusion model released by a research institute requires two stages of neural network training and 12 steps of sampling during synthesis. Measured in terms of RTF, the speed is more than 6 times slower than WarpNet, which greatly increases the latency of the synthesis stage.

[0068] This paper maintains the advantages of single-stage, single neural network training of WarpNet, and does not require the two-stage training of the diffusion model. At the same time, a smaller neural network (6.91M) is used, and the 12-step sampling of the diffusion model is compressed to 2 steps during sampling, which significantly improves the sampling speed. RTF is comparable to the non-iterative model WarpNet (RTF 0.076 vs RTF 0.063).

[0069] refer to Figure 3 The present invention also discloses a spatial audio synthesis system based on Schrödinger bridge, the system comprising: An information acquisition module 110, configured to acquire a monophonic sound source signal, and construct a priori signal and a noisy representation based on the monophonic sound source signal; An audio synthesis module 120, configured to input the prior signal and the noisy representation into a pre-trained spatial audio synthesis model, and generate a final two-channel spatial audio through the spatial audio synthesis model based on a random differential equation iterative sampling path; The spatial audio synthesis model is obtained by training a preset neural network model through a predefined binaural Schrödinger bridge parameterized target and boundary auxiliary supervision.

[0070] The spatial audio synthesis model is obtained by training the preset neural network model through the predefined binaural Schrödinger bridge parameterized objective and boundary auxiliary supervision, including: Obtain the quaternion of the monophonic sound source signal, the relative spatial position coordinates of the sound source and the receiver, and the receiver's head orientation information; Building a binaural Schrodinger bridge between the monophonic sound source signal and the audio signals received by the left and right ears of the receiver; Obtaining a noisy representation at each time step in the two-channel Schrödinger bridge, and parameterizing the two-channel Schrödinger bridge by shallowly adding noise based on the noisy representation; Based on the binaural Schrödinger bridge parameterization target, the preset spatial audio synthesis model is trained through the relative spatial position coordinates of the sound source and the receiver, the quaternion of the receiver's head orientation information, and boundary-assisted supervision.

[0071] A binaural Schrödinger bridge is built between the monophonic sound source signal and the audio signals received by the left and right ears of the receiver, specifically including: Duplicating the monophonic sound source signal to form a dual-channel speech signal; The audio signals received by the left and right ears of the receiver are used as the generation target, and the binaural speech signal is used as the prior information of the audio signals received by the left and right ears of the receiver; A binaural Schrödinger bridge is built between the prior information and the generated target.

[0072] Obtaining a noisy representation at each time step in the two-channel Schrödinger bridge, and parameterizing the two-channel Schrödinger bridge by shallowly adding noise based on the noisy representation, specifically including: Obtaining a noisy representation at each time step in the binaural Schrödinger bridge; Based on the noisy characterization, the noise content is first increased to the peak of the noise content, and then the noise content is gradually reduced until the noise disappears, thereby completing the parameterization of the binaural Schrodinger bridge.

[0073] Based on the binaural Schrödinger bridge parameterization target, the preset spatial audio synthesis model is trained through the relative spatial position coordinates of the sound source and the receiver, the quaternion of the receiver's head orientation information, and boundary auxiliary supervision, including: The quaternion of the relative spatial position coordinates between the sound source and the receiver and the receiver's head orientation information is used as conditional information to guide the training of a preset spatial audio synthesis model; Based on the binaural Schrödinger bridge parameterization target, the quaternion of the relative spatial position coordinates between the sound source and the receiver and the receiver's head orientation information is used as conditional information to guide the training of the preset spatial audio synthesis model. The loss function is added through boundary auxiliary supervision to supervise the training of the preset spatial audio synthesis model. The loss function includes: multi-scale amplitude loss and phase spectrum loss.

[0074] The prior signal and the noisy representation are input into a pre-trained spatial audio synthesis model, and the final two-channel spatial audio is generated through the spatial audio synthesis model based on a random differential equation iterative sampling path, specifically including: Inputting the prior signal and the noisy representation into a spatial audio synthesis model to generate a single-step prediction target; Based on the single-step prediction target, sampling is performed through a stochastic differential equation iterative sampling path, and the final two-channel spatial audio is synthesized according to the sampling results.

[0075] A spatial audio synthesis system based on Schrödinger bridge provided by the present invention improves the quality of spatial audio synthesis through a new spatial audio synthesis model based on Schrödinger bridge; significantly surpasses previous methods in five indicators for measuring synthesis quality; and the sampling steps can be compressed based on the iterative sampling path of random differential equations, which greatly improves the inference speed of the spatial audio synthesis system based on iterative sampling.

[0076] Figure 4 An example of a physical structure diagram of an electronic device is shown in FIG. Figure 4 As shown, the electronic device may include: a processor 410, a communication interface 420, a memory 430 and a communication bus 440, wherein the processor 410, the communication interface 420 and the memory 430 communicate with each other through the communication bus 440. The processor 410 may call the logic instructions in the memory 430 to execute a spatial audio synthesis method based on a Schrödinger bridge, the method comprising: obtaining a monophonic sound source signal, constructing a priori signal and a noisy representation based on the monophonic sound source signal; inputting the priori signal and the noisy representation into a pre-trained spatial audio synthesis model, and generating a final binaural spatial audio through the spatial audio synthesis model based on a random differential equation iterative sampling path; wherein the spatial audio synthesis model is obtained by training a preset neural network model through a predefined binaural Schrödinger bridge parameterized target and boundary auxiliary supervision.

[0077] In addition, the logic instructions in the above-mentioned memory 430 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc. Various media that can store program codes.

[0078] On the other hand, the present invention also provides a computer program product, which includes a computer program, which can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute a spatial audio synthesis method based on Schrödinger bridge provided by the above methods, the method including: obtaining a monophonic sound source signal, and constructing a priori signal and noisy representation based on the monophonic sound source signal; inputting the priori signal and noisy representation into a pre-trained spatial audio synthesis model, and generating a final dual-channel spatial audio through the spatial audio synthesis model based on an iterative sampling path of random differential equations; wherein the spatial audio synthesis model is obtained by training a preset neural network model through a predefined dual-channel Schrödinger bridge parameterized target and boundary-assisted supervision.

[0079] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute a spatial audio synthesis method based on a Schrödinger bridge provided by the above-mentioned methods, the method comprising: obtaining a monophonic sound source signal, and constructing a priori signal and a noisy representation based on the monophonic sound source signal; inputting the priori signal and the noisy representation into a pre-trained spatial audio synthesis model, and generating a final dual-channel spatial audio through the spatial audio synthesis model based on an iterative sampling path of a random differential equation; wherein the spatial audio synthesis model is obtained by training a preset neural network model through a predefined dual-channel Schrödinger bridge parameterized target and boundary-assisted supervision.

[0080] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.

[0081] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0082] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A spatial audio synthesis method based on Schrödinger bridge, characterized in that: include: Acquire a monophonic sound source signal, and construct a priori signal and a noisy representation based on the monophonic sound source signal; Inputting the prior signal and the noisy representation into a pre-trained spatial audio synthesis model, and generating final binaural spatial audio through the spatial audio synthesis model based on a stochastic differential equation iterative sampling path; The spatial audio synthesis model is obtained by training a preset neural network model through a predefined binaural Schrödinger bridge parameterized target and boundary auxiliary supervision.

2. The spatial audio synthesis method based on Schrödinger bridge according to claim 1, characterized in that: The spatial audio synthesis model is obtained by training a preset neural network model through a predefined binaural Schrödinger bridge parameterized target and boundary auxiliary supervision, specifically including: Obtain the quaternion of the monophonic sound source signal, the relative spatial position coordinates of the sound source and the receiver, and the receiver's head orientation information; Building a binaural Schrodinger bridge between the monophonic sound source signal and the audio signals received by the left and right ears of the receiver; Obtaining a noisy representation at each time step in the two-channel Schrödinger bridge, and parameterizing the two-channel Schrödinger bridge by shallowly adding noise based on the noisy representation; Based on the binaural Schrödinger bridge parameterization target, the preset spatial audio synthesis model is trained through the relative spatial position coordinates of the sound source and the receiver, the quaternion of the receiver's head orientation information, and boundary-assisted supervision.

3. The spatial audio synthesis method based on Schrödinger bridge according to claim 2, characterized in that: The step of building a binaural Schrödinger bridge between the monophonic sound source signal and the audio signals received by the left and right ears of the receiver specifically includes: Duplicating the monophonic sound source signal to form a dual-channel speech signal; The audio signals received by the left and right ears of the receiver are used as the generation target, and the binaural speech signal is used as the prior signal of the audio signals received by the left and right ears of the receiver; A binaural Schrödinger bridge is built between the prior signal and the generated target.

4. The spatial audio synthesis method based on Schrödinger bridge according to claim 2, characterized in that: The step of obtaining a noisy representation at each time step in the two-channel Schrödinger bridge and parameterizing the two-channel Schrödinger bridge by shallowly adding noise based on the noisy representation specifically includes: Obtaining a noisy representation at each time step in the binaural Schrödinger bridge; Based on the noisy characterization, the noise content is first increased to the peak of the noise content, and then the noise content is gradually reduced until the noise disappears, thereby completing the parameterization of the binaural Schrodinger bridge.

5. The spatial audio synthesis method based on Schrödinger bridge according to claim 2, characterized in that: The training of the preset spatial audio synthesis model is completed based on the binaural Schrödinger bridge parameterization target through the relative spatial position coordinates of the sound source and the receiver, the quaternion of the receiver's head orientation information, and boundary auxiliary supervision, specifically including: The quaternion of the relative spatial position coordinates between the sound source and the receiver and the receiver's head orientation information is used as conditional information to guide the training of a preset spatial audio synthesis model; Based on the binaural Schrödinger bridge parameterization target, the quaternion of the relative spatial position coordinates between the sound source and the receiver and the receiver's head orientation information is used as conditional information to guide the training of the preset spatial audio synthesis model. The loss function is added through boundary auxiliary supervision to supervise the training of the preset spatial audio synthesis model. The loss function includes: multi-scale amplitude loss and phase spectrum loss.

6. The spatial audio synthesis method based on Schrödinger bridge according to claim 1, characterized in that: The step of inputting the priori signal and the noisy representation into a pre-trained spatial audio synthesis model, and generating a final two-channel spatial audio through the spatial audio synthesis model based on a stochastic differential equation iterative sampling path, specifically includes: Inputting the prior signal and the noisy representation into a spatial audio synthesis model to generate a single-step prediction target; Based on the single-step prediction target, sampling is performed through a stochastic differential equation iterative sampling path, and the final two-channel spatial audio is synthesized according to the sampling results.

7. A spatial audio synthesis system based on Schrödinger bridge, characterized in that: The system comprises: An information acquisition module, used to acquire a monophonic sound source signal, and construct a priori signal and a noisy representation based on the monophonic sound source signal; An audio synthesis module, configured to input the prior signal and the noisy representation into a pre-trained spatial audio synthesis model, and generate final binaural spatial audio through the spatial audio synthesis model based on an iterative sampling path of a stochastic differential equation; The spatial audio synthesis model is obtained by training a preset neural network model through a predefined binaural Schrödinger bridge parameterized target and boundary auxiliary supervision.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the spatial audio synthesis method based on Schrödinger bridge as claimed in any one of claims 1 to 6 is implemented.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the spatial audio synthesis method based on Schrödinger bridge as claimed in any one of claims 1 to 6 is implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the spatial audio synthesis method based on Schrödinger bridge as claimed in any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Speech synthesis method and device, electronic equipment and readable storage medium

    CN117854470A

  • Signal generation and road reconstruction method and system based on Riemannian diffusion Schrodinger bridge

    CN118089702A

  • CONDITIONAL DIFFUSION MODEL FOR DATA-TO-DATA TRANSLATION

    DE102024103309A1

Cited By

  • Audio super-division model training method and device, audio super-division processing method and device and electronic equipment

    CN120319255A