A method and system for spatial audio synthesis based on a schrodinger bridge
Through Schrödinger bridge parameterization and boundary-assisted supervised neural network training, combined with stochastic differential equation iterative sampling, the problems of slow speed and poor quality of spatial audio synthesis are solved, and efficient spatial audio synthesis is achieved.
Patent Information
- Application Number
- CN202411954295.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-27
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2044-12-27
AI Technical Summary
Existing spatial audio synthesis technology has problems of slow speed and poor quality. Traditional signal processing methods require expensive physical parameter measurements and cannot model nonlinear processes. Neural networks based on deterministic regression cannot capture detailed information, and the synthesis speed of diffusion models is too slow.
A spatial audio synthesis method based on Schrödinger bridge is adopted. The preset neural network model is trained with predefined binaural Schrödinger bridge parameterized objectives and boundary-assisted supervision, and the stochastic differential equation iterative sampling path is used to generate binaural spatial audio.
It significantly surpasses existing methods in synthesis quality and speed, achieving high-quality and fast spatial audio synthesis, improving synthesis quality and compressing sampling steps.
Smart Images

Figure CN119964544B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of audio synthesis, and in particular to a spatial audio synthesis method and system based on a Schrödinger bridge. BACKGROUND
[0002] Binaural Audio Synthesis (BAS) refers to a process of synthesizing Binaural Audio Signals from Monaural Audio Signals at a sound source according to relative positions, head orientations, and other information. In the real world, the process of propagating audio signals at a sound source to the left and right ears of a receiver is affected by many factors, such as indoor environment, receiver head model, environmental noise, and the like. How to synthesize spatial audio signals with high quality and quickly is a key problem in building a BAS system.
[0003] Compared with monaural audio signals, binaural audio signals can enable users to perceive spatial position information from the audio signals. In virtual reality, augmented reality, and other scenarios such as online conference rooms, movies, and games, spatial audio synthesis technology can greatly enhance the sense of immersion and experience of users. Therefore, this technology has a wide range of applications in these scenarios.
[0004] In the past, spatial audio synthesis technology can be divided into three main technical routes. The first route is based on traditional Digital Signal Processing Methods (DSP). The advantage of this method is that it does not require training of a neural network, but it needs to collect a large number of physical parameters to represent indoor environment information (such as Room Impulsive Response, RIR), user head information (such as Head-related Transfer Function, HRTF), and even environmental noise (Environmental Noise) and other information. Because the measurement of these information is very expensive and requires high requirements for the measuring device, this method cannot obtain accurate physical parameters in actual application scenarios, and can only use approximate values, which inevitably brings modeling errors. In addition, this method cannot model the nonlinear process in spatial audio synthesis. These two defects seriously limit the quality of spatial audio synthesis of the DSP method.
[0005] The second technical approach involves mapping-based networks (BANs) based on deterministic regression. These methods use neural networks, such as convolutional neural networks (CNNs), to perform an end-to-end fit for the synthesis process from mono to binaural audio. Compared to traditional signal processing methods, these methods eliminate the need for expensive parameter measurements and can simultaneously model both linear and nonlinear variations in spatial audio synthesis, improving both the subjective and objective synthesis quality. However, because these methods typically use objective metrics such as mean square error (MSE) in the waveform, amplitude, or phase space of the binaural signal to train the neural network and employ a deterministic single-step regression approach for binaural audio synthesis, their synthesis quality remains limited and they are unable to effectively capture the detailed information of high-sampled spatial audio. The third technical approach is based on diffusion models. This type of work no longer directly models the propagation process from mono to binaural. Instead, it transforms this process into a conditional generative process, utilizing conditional diffusion models to synthesize spatial audio. This approach develops a two-stage conditional diffusion model. In the first stage, a single-channel diffusion model models the mean of the binaural signals from the source to the left and right ears, thereby modeling the "common information" of the binaural signals. In the second stage, a two-channel diffusion model models the ground-truth binaural audio signals from the mean of the left and right ear signals. This modeling shifts from "common information" to "specific information" of the left and right ears. BinauralGrad's overall design achieves a generative process from commonality to specificity, from coarseness to fine-grainedness, for spatial audio tasks, achieving subjective and objective synthesis quality that surpasses previous work.
[0006] Among the three technical approaches mentioned above, traditional signal processing (DSP) methods utilize manually constructed filters to model various factors affecting sound wave transmission. The physical parameters of these filters, such as those in RIR and HRTF, often require expensive or time-consuming measurement processes. In practical applications, approximate values are often used due to the unavailability of accurate physical parameters. Furthermore, the complex spatial audio transmission process cannot be modeled using simple linear filters alone, and the nonlinear changes involved cannot be modeled using DSP methods. Consequently, the quality of spatial audio synthesis based on DSP methods is always limited.
[0007] The end-to-end neural network method based on deterministic regression is unable to accurately capture the detailed information in the audio waveform due to its optimization goal and single-step regression synthesis method. Although the quality is improved compared to the DSP method, the synthesis quality is still limited.
[0008] Spatial audio synthesis methods based on diffusion models use iterative sampling to gradually model spatial audio, significantly improving quality compared to previous methods. However, the synthesis process requires 12 iterative sampling steps, each of which requires computation in the audio waveform space at a 48kHz sampling rate. As a result, the synthesis speed is very slow, far slower than previous methods. Summary of the Invention
[0009] The present invention provides a spatial audio synthesis method and system based on Schrödinger bridge, which are used to solve the problems of slow speed and poor quality of existing spatial audio synthesis.
[0010] The present invention provides a spatial audio synthesis method based on Schrödinger bridge, comprising:
[0011] Acquire a monophonic sound source signal, and construct a priori signal and a noisy representation based on the monophonic sound source signal;
[0012] Inputting the prior signal and the noisy representation into a pre-trained spatial audio synthesis model, and generating final binaural spatial audio through the spatial audio synthesis model based on an iterative sampling path of stochastic differential equations;
[0013] The spatial audio synthesis model is obtained by training a preset neural network model through a predefined binaural Schrödinger bridge parameterized target and boundary auxiliary supervision.
[0014] According to the present invention, a spatial audio synthesis method based on Schrödinger bridge is provided. The spatial audio synthesis model is obtained by training a preset neural network model using a predefined binaural Schrödinger bridge parameterized target and boundary auxiliary supervision, specifically comprising:
[0015] Obtain the quaternion of the monophonic sound source signal, the relative spatial position coordinates of the sound source and the receiver, and the receiver's head orientation information;
[0016] Building a binaural Schrödinger bridge between the monophonic sound source signal and the audio signals received by the left and right ears of the receiver;
[0017] Obtaining a noisy representation of the binaural Schrödinger bridge at each time step, and parameterizing the binaural Schrödinger bridge by shallowly adding noise based on the noisy representation;
[0018] Based on the binaural Schrödinger bridge parameterization target, the preset spatial audio synthesis model is trained through the relative spatial position coordinates of the sound source and the receiver, the quaternion of the receiver's head orientation information, and boundary-assisted supervision.
[0019] According to a spatial audio synthesis method based on a Schrödinger bridge provided by the present invention, a binaural Schrödinger bridge is constructed between the monophonic sound source signal and the audio signals received by the left and right ears of the receiver, specifically comprising:
[0020] Duplicating the monophonic sound source signal to form a dual-channel speech signal;
[0021] The audio signals received by the left and right ears of the receiver are used as the generation target, and the binaural speech signal is used as the prior signal of the audio signals received by the left and right ears of the receiver;
[0022] A binaural Schrödinger bridge is built between the prior signal and the generated target.
[0023] According to a spatial audio synthesis method based on a Schrödinger bridge provided by the present invention, obtaining a noisy representation at each time step in the binaural Schrödinger bridge and parameterizing the binaural Schrödinger bridge by shallowly adding noise based on the noisy representation specifically includes:
[0024] Obtaining a noisy representation of each time step in the binaural Schrödinger bridge;
[0025] Based on the noisy characterization, the noise content is first increased to a peak of the noise content, and then the noise content is gradually reduced until the noise disappears, thereby completing the parameterization of the binaural Schrödinger bridge.
[0026] According to a spatial audio synthesis method based on Schrödinger bridge provided by the present invention,
[0027] The training of the preset spatial audio synthesis model based on the binaural Schrödinger bridge parameterized target is completed through the relative spatial position coordinates of the sound source and the receiver, the quaternion of the receiver's head orientation information, and boundary auxiliary supervision, specifically including:
[0028] The quaternion of the relative spatial position coordinates of the sound source and the receiver, and the receiver's head orientation information is used as conditional information to guide the training of a preset spatial audio synthesis model;
[0029] Based on the binaural Schrödinger bridge parameterization target, the quaternion of the relative spatial position coordinates of the sound source and the receiver and the receiver's head orientation information is used as conditional information to guide the training of the preset spatial audio synthesis model. On this basis, the loss function is added through boundary-assisted supervision to supervise the training of the preset spatial audio synthesis model.
[0030] The loss function includes: multi-scale amplitude loss and phase spectrum loss.
[0031] According to a Schrödinger bridge-based spatial audio synthesis method provided by the present invention, the prior signal and the noisy representation are input into a pre-trained spatial audio synthesis model, and the final two-channel spatial audio is generated by the spatial audio synthesis model based on an iterative sampling path of a stochastic differential equation. Specifically, the method includes:
[0032] Inputting the prior signal and the noisy representation into a spatial audio synthesis model to generate a single-step prediction target;
[0033] Sampling is performed based on the single-step prediction target through a stochastic differential equation iterative sampling path, and final two-channel spatial audio is synthesized according to the sampling results.
[0034] The present invention also provides a spatial audio synthesis system based on Schrödinger bridge, the system comprising:
[0035] An information acquisition module, configured to acquire a monophonic sound source signal and construct a priori signal and a noisy representation based on the monophonic sound source signal;
[0036] An audio synthesis module, configured to input the prior signal and the noisy representation into a pre-trained spatial audio synthesis model, and generate final binaural spatial audio through the spatial audio synthesis model based on an iterative sampling path of stochastic differential equations;
[0037] The spatial audio synthesis model is obtained by training a preset neural network model through a predefined binaural Schrödinger bridge parameterized target and boundary auxiliary supervision.
[0038] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the spatial audio synthesis method based on the Schrödinger bridge as described above is implemented.
[0039] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the spatial audio synthesis method based on Schrödinger bridge as described above is implemented.
[0040] The present invention also provides a computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements any of the above-mentioned Schrödinger bridge-based spatial audio synthesis methods.
[0041] The present invention provides a Schrödinger bridge-based spatial audio synthesis method and system, which improves the quality of spatial audio synthesis through a novel Schrödinger bridge-based spatial audio synthesis model. The method and system significantly surpass previous methods in five indicators for measuring synthesis quality. The iterative sampling path based on stochastic differential equations can compress the number of sampling steps, significantly improving the inference speed of the iterative sampling-based spatial audio synthesis system. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction is given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0043] Figure 1 It is a flow chart of a spatial audio synthesis method based on Schrödinger bridge provided by the present invention.
[0044] Figure 2 2 is a noise scheduling comparison diagram provided by the present invention.
[0045] Figure 3 This is a schematic diagram of module connections of a spatial audio synthesis system based on a Schrödinger bridge provided by the present invention.
[0046] Figure 4 It is a structural schematic diagram of the electronic device provided by the present invention.
[0047] Reference numerals: 110 : information acquisition module; 120 : audio synthesis module; 410 : processor; 420 : communication interface; 430 : memory; 440 : communication bus. DETAILED DESCRIPTION
[0048] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0049] Previous spatial audio synthesis methods, such as DSP and neural network-based end-to-end single-step regression, have limited the quality of binaural audio synthesis. Systems based on diffusion models have acceptable synthesis quality, but the sampling speed is slow, requiring 12 iterative sampling steps in the binaural 48kHz waveform space.
[0050] The purpose of the present application is to propose a brand-new spatial audio synthesis framework, which has a synthesis quality beyond all previous methods while having a reasoning speed comparable to that of a single-step reasoning of an end-to-end neural network, and which realizes a high-quality and fast efficient spatial audio synthesis system for the first time.
[0051] The present application is described below Figure 1 A spatial audio synthesis method based on a Schrodinger bridge of the present application comprises:
[0052] Step 100, obtaining a monaural sound source signal, constructing a prior signal and a noisy representation based on the monaural sound source signal.
[0053] The present application adopts a generation model based on a Schrodinger bridge to realize spatial audio synthesis from monaural to binaural in a data-to-data generation process. That is, an iterative sampling path based on Stochastic Differential Equations (SDE) is used to implicitly model the sound wave transmission process from a monaural sound source to binaural spatial audio.
[0054] In the spatial audio synthesis task, given a monaural sound source signal , a coordinate representing the relative spatial position of the sound source and the receiver , and a quaternion representing the head orientation information of the receiver , the spatial audio synthesis system aims to synthesize the binaural audio signals received by the left and right ears of the receiver , and represent the signals received by the left and right ears, respectively, and have the same dimension as the sound source signal .
[0055] Under this task definition, the performance of the spatial audio synthesis system mainly includes (1) synthesis quality, (2) synthesis speed, (3) model parameter amount, etc. Higher synthesis quality, faster synthesis speed, and fewer model parameter amounts are the main goals of establishing a spatial audio synthesis system.
[0056] Step 200, inputting the prior signal and noisy representation into a pre-trained spatial audio synthesis model, generating the final binaural spatial audio based on the iterative sampling path of the random differential equation through the spatial audio synthesis model; wherein the spatial audio synthesis model is obtained by training a preset neural network model through a pre-defined binaural Schrodinger bridge parameterization target and boundary auxiliary supervision.
[0057] Specifically, it comprises: obtaining a monaural sound source signal, a coordinate of the relative spatial position of the sound source and the receiver, and a quaternion of the head orientation information of the receiver.
[0058] Building a binaural Schrödinger bridge between the monophonic sound source signal and the audio signals received by the left and right ears of the receiver;
[0059] Obtaining a noisy representation of the binaural Schrödinger bridge at each time step, and parameterizing the binaural Schrödinger bridge by shallowly adding noise based on the noisy representation;
[0060] Based on the binaural Schrödinger bridge parameterization target, the preset spatial audio synthesis model is trained through the relative spatial position coordinates of the sound source and the receiver, the quaternion of the receiver's head orientation information, and boundary-assisted supervision.
[0061] The Schrödinger bridge-based generative model used in the present invention is defined by the following forward and reverse stochastic differential equations:
[0062] .
[0063] in and They represent the two endpoint distributions of the "data to data" synthesis process of the Schrödinger bridge, namely the prior distribution and the target data distribution. and are the drift term and diffusion term of the stochastic differential equation, and Represents the Brownian motion in forward and reverse time in Schrödinger. and When both boundary distributions are Dirac-delta distributions, the nonlinear drift components of the forward and inverse stochastic differential equations are and The definition is as follows:
[0064] .
[0065] .
[0066] Specifically, a binaural Schrödinger bridge is built between the monophonic sound source signal and the audio signals received by the left and right ears of the receiver, including:
[0067] Duplicating the monophonic sound source signal to form a dual-channel speech signal;
[0068] The audio signals received by the left and right ears of the receiver are used as the generation target, and the mono sound source signal is used as the prior signal of the audio signals received by the left and right ears of the receiver;
[0069] A binaural Schrödinger bridge is built between the prior signal and the generated target.
[0070] In this invention, a spatial audio synthesis system based on Schrödinger bridge is proposed. In the BAS task, the monophonic signal at the sound source is used to generate the sound. At the same time, the left and right ears of the receiver receive the signal and Specifically, the monophonic speech signal at the sound source is copied to form a binaural speech signal as a priori signal. , taking the audio signals received by the left and right ears of the receiver as the generation target , a two-channel Schrödinger bridge is built between the two information, where the nonlinear drift terms in the forward and reverse SDE become:
[0071] .
[0072] Specifically, obtaining a noisy representation at each time step in the binaural Schrödinger bridge;
[0073] Based on the noisy characterization, the noise content is first increased to a peak of the noise content, and then the noise content is gradually reduced until the noise disappears, thereby completing the parameterization of the binaural Schrödinger bridge.
[0074] In the present invention, given the prior signal and the generated target, the noisy representation at each time t in the Schrödinger bridge built on the BAS task is calculated as:
[0075] .
[0076] In the parameterization objective of the Schrödinger bridge, "noise prediction" is used as the parameterization objective of the binaural Schrödinger bridge. The training objective of the spatial audio synthesis model is:
[0077] .
[0078] Among them, the coordinates representing the relative position information , and the quaternion representing the receiver's head orientation information They are also used as neural network inputs and serve as conditional information to guide the spatial audio synthesis process. They have been omitted in the above formula for simplicity.
[0079] Previous research on Schrödinger bridges has focused on noise scheduling, focusing on symmetric and asymmetric noise scheduling. However, the appropriate level of noise addition during the Schrödinger forward process has not been investigated. This paper, for the first time, investigates the intensity of the noise component in the Schrödinger bridge.
[0080] Compared with the "noise to data" generation method of the diffusion model, the "data to data" synthesis process of the Schrödinger bridge makes its noise content not gradually increase from t=0 to t=1 like the diffusion model, but is divided into two stages: the noise content first increases to the peak of the noise content and then decreases to almost nothing.
[0081] At the peak of the noise content, if the noise intensity is too high, the prior signal will be largely overwhelmed by the noise. This means that the information provided by the prior signal for generating the target will be severely corrupted by the noise. Therefore, compared to the "noise-to-data" synthesis paradigm of the diffusion model, the "data-to-data" paradigm of the Schrödinger bridge loses some of its advantages in utilizing the prior signal due to the excessive noise content.
[0082] refer to Figure 2 In the spatial audio task, two noise configurations are shown: Deep Noise on the left and Shallow Noise on the right. It can be seen that at the peak of the Schrödinger bridge noise content, the structure of the prior signal is destroyed in the Deep Noise representation. In contrast, with Shallow Noise, the noise amplitude is extremely small compared to the speech signal amplitude. At the peak of the noise content, the dual-channel Schrödinger bridge representation still retains the contour information of the speech signal.
[0083] Comparing the left and right figures, it can be seen that when using shallow noisy technology, the dual-channel Schrödinger bridge is more likely to preserve voice signal information and generate target data (Data), i.e., two-channel audio, with less burden, thereby improving the sample synthesis quality and speed compared to the deep noisy design.
[0084] Based on the binaural Schrödinger bridge parameterized objective, the preset spatial audio synthesis model is trained through boundary auxiliary supervision and an additional loss function as auxiliary supervision;
[0085] The loss function includes multi-scale including: multi-scale amplitude loss and phase spectrum loss.
[0086] In the present invention, according to the training goal of the spatial audio synthesis model, after about 1M steps of model iteration, the neural network is characterized by the noise at time t and at time t. As input, you will get the output:
[0087] .
[0088] In previous work on Schrödinger bridges, different model parameterization methods such as noise, data, speed, and score have been explored. However, relying solely on these equivalent parameterization methods does not necessarily achieve the best sample synthesis quality and synthesis speed. Therefore, this invention is the first to directly incorporate sample quality into the optimization objective of the Schrödinger bridge. Specifically, in addition to the original Schrödinger bridge optimization objective In addition, an additional loss function is added as auxiliary supervision:
[0089] .
[0090] in, and represents the additional loss function at t=0 and t=1 of the Schrödinger bridge, which is defined as follows:
[0091] .
[0092] That is, the clean data at t=0 and t=1 are used as auxiliary supervision targets. In the spatial audio synthesis task, both the monophonic source audio and the binaural spatial audio are used as auxiliary supervision targets to strengthen the training of the neural network. and Represents two loss functions commonly used to characterize the quality of audio data: Multi-Scale Short-Time Fourier Transform loss and Phase loss. to Represents the weight of the corresponding loss function, which is an adjustable parameter during model training.
[0093] In actual calculation, at each step of neural network training, the generated target The single-step calculation of has been shown above, which can be calculated Using the output of the spatial audio synthesis model, for the t=1 prior signal, the single-step calculation can also be performed as follows:
[0094] .
[0095] You can calculate Finally, as shown above, when using auxiliary supervision technology, the loss function of the spatial audio synthesis system based on Schrödinger bridge consists of three parts, taking into account the optimization of Schrödinger bridge and the optimization of sample quality at each sampling step.
[0096] Step 500: Based on the spatial audio synthesis model, an iterative sampling path of a stochastic differential equation is used to implicitly model the sound wave transmission process from a monophonic sound source signal to a binaural spatial audio, and the final binaural spatial audio is iteratively generated.
[0097] Specifically, the method uses the real-time acquired noisy representation and prior signal as input to the spatial audio synthesis model, and outputs a single-step prediction generation target.
[0098] The target is generated based on a single-step prediction sampling method based on stochastic differential equations, and the final two-channel spatial audio is iteratively generated.
[0099] In the present invention, at each step of sampling, based on the output of the spatial audio synthesis model, a single-step prediction generation target can be first performed:
[0100] .
[0101] The final two-channel spatial audio is iteratively generated according to the sampling method based on Ordinary Differential Equations (ODE):
[0102] .
[0103] By building a two-channel Schrödinger bridge for spatial audio tasks, unlike the three previous technical paradigms for the three BAS tasks, this approach implicitly models the spatial audio synthesis process by utilizing a "data-to-data" synthesis process based on a probabilistic generative model. This model directly builds a generative model between the monophonic audio at the source and the binaural audio at the receiver, rather than being constrained by the "noise-to-data" synthesis paradigm of the diffusion model. Compared to methods such as single-step regression and DSP, probabilistic generative models demonstrate greater expressive power in fitting data distributions and are typically able to capture more fine-grained information and achieve higher accuracy in audio waveform modeling.
[0104] In a specific embodiment, the spatial audio synthesis quality of the present invention is compared with the three previous technical routes. In Table 1, It is a spatial audio synthesis system built. Auxiliary Supervision (AS) is added to the system. Other methods include Represents the DSP method implemented by BinauralGrad, The represents the experimental results achieved on the WarpNet proposed by DopplerBAS, while the BG represents the experimental results of the traditional diffusion model. It can be seen that in five objective indicators such as waveform, spectral amplitude, and phase amplitude, it surpasses the three previous technical approaches.
[0105] Table 1
[0106] .
[0107] The computing efficiency of the spatial audio synthesis system is compared with the previous representative three methods in Table 2.
[0108] Table 2
[0109]
[0110] As shown in Table 2, the traditional signal processing method DSP only needs to be calculated on the CPU, does not need to use the GPU for inference, and does not need model training, but the synthesis quality is very limited. WarpNet only needs a single-stage neural network, and can synthesize spatial audio in one step in an end-to-end manner, and has the best performance in the real-time factor (RTF) for measuring synthesis speed. The model size is 8.59M.
[0111] The spatial audio synthesis system based on the diffusion model published by an institute needs a two-stage neural network to be trained respectively, and needs 12 steps of sampling during synthesis. The speed measured by RTF is more than 6 times slower than WarpNet, greatly increasing the time delay in the synthesis stage.
[0112] The present application maintains the advantages of single-stage and single neural network training of WarpNet, and does not need the two-stage training of the diffusion model. At the same time, a smaller neural network (6.91M) is used, and the 12-step sampling of the diffusion model is compressed to 2 steps during sampling, which significantly improves the sampling speed, and the RTF is comparable to the non-iterative model WarpNet (RTF 0.076 vs RTF 0.063).
[0113] Reference Figure 3 The present application also discloses a spatial audio synthesis system based on a Schrodinger bridge, which comprises:
[0114] An information acquisition module 110 is configured to acquire a monophonic sound source signal, and construct a prior signal and a noisy representation based on the monophonic sound source signal;
[0115] An audio synthesis module 120 is configured to input the prior signal and the noisy representation into a pre-trained spatial audio synthesis model, and generate a final binaural spatial audio based on a stochastic differential equation iterative sampling path of the spatial audio synthesis model.
[0116] The spatial audio synthesis model is obtained by training a preset neural network model based on a pre-defined binaural Schrodinger bridge parameterization target and boundary auxiliary supervision.
[0117] The spatial audio synthesis model is obtained by training a preset neural network model based on a pre-defined binaural Schrodinger bridge parameterization target and boundary auxiliary supervision, and specifically includes:
[0118] Obtain the quaternion of the monophonic sound source signal, the relative spatial position coordinates of the sound source and the receiver, and the receiver's head orientation information;
[0119] Building a binaural Schrödinger bridge between the monophonic sound source signal and the audio signals received by the left and right ears of the receiver;
[0120] Obtaining a noisy representation of the binaural Schrödinger bridge at each time step, and parameterizing the binaural Schrödinger bridge by shallowly adding noise based on the noisy representation;
[0121] Based on the binaural Schrödinger bridge parameterization target, the preset spatial audio synthesis model is trained through the relative spatial position coordinates of the sound source and the receiver, the quaternion of the receiver's head orientation information, and boundary-assisted supervision.
[0122] Building a binaural Schrödinger bridge between the monophonic sound source signal and the audio signals received by the left and right ears of the receiver specifically includes:
[0123] Duplicating the monophonic sound source signal to form a dual-channel speech signal;
[0124] The audio signals received by the left and right ears of the receiver are used as the generation target, and the binaural speech signal is used as the prior information of the audio signals received by the left and right ears of the receiver;
[0125] A binaural Schrödinger bridge is built between the prior information and the generated target.
[0126] Obtaining a noisy representation at each time step in the two-channel Schrödinger bridge, and parameterizing the two-channel Schrödinger bridge by shallowly adding noise based on the noisy representation, specifically including:
[0127] Obtaining a noisy representation of each time step in the binaural Schrödinger bridge;
[0128] Based on the noisy characterization, the noise content is first increased to a peak of the noise content, and then the noise content is gradually reduced until the noise disappears, thereby completing the parameterization of the binaural Schrödinger bridge.
[0129] Based on the binaural Schrödinger bridge parameterization objective, the preset spatial audio synthesis model is trained using the relative spatial position coordinates of the sound source and receiver, the quaternion of the receiver's head orientation information, and boundary-assisted supervision. Specifically, the training includes:
[0130] The quaternion of the relative spatial position coordinates of the sound source and the receiver, and the receiver's head orientation information is used as conditional information to guide the training of a preset spatial audio synthesis model;
[0131] Based on the binaural Schrödinger bridge parameterization target, the quaternion of the relative spatial position coordinates of the sound source and the receiver and the receiver's head orientation information is used as conditional information to guide the training of the preset spatial audio synthesis model. On this basis, the loss function is added through boundary-assisted supervision to supervise the training of the preset spatial audio synthesis model.
[0132] The loss function includes: multi-scale amplitude loss and phase spectrum loss.
[0133] Inputting the prior signal and the noisy representation into a pre-trained spatial audio synthesis model, and generating the final two-channel spatial audio through the spatial audio synthesis model based on an iterative sampling path of a stochastic differential equation, specifically including:
[0134] Inputting the prior signal and the noisy representation into a spatial audio synthesis model to generate a single-step prediction target;
[0135] Sampling is performed based on the single-step prediction target through a stochastic differential equation iterative sampling path, and final two-channel spatial audio is synthesized according to the sampling results.
[0136] The spatial audio synthesis system based on Schrödinger bridge provided by the present invention improves the quality of spatial audio synthesis through a new Schrödinger bridge-based spatial audio synthesis model; significantly surpasses previous methods in five indicators for measuring synthesis quality; and the iterative sampling path based on stochastic differential equations can compress the number of sampling steps, greatly improving the inference speed of the spatial audio synthesis system based on iterative sampling.
[0137] Figure 4 An example of a physical structure diagram of an electronic device is shown below. Figure 4 As shown, the electronic device may include: a processor 410, a communications interface 420, a memory 430, and a communications bus 440, wherein the processor 410, the communications interface 420, and the memory 430 communicate with each other via the communications bus 440. The processor 410 may call logic instructions in the memory 430 to execute a spatial audio synthesis method based on a Schrödinger bridge, the method comprising: obtaining a monophonic sound source signal, constructing a priori signal and a noisy representation based on the monophonic sound source signal; inputting the priori signal and the noisy representation into a pre-trained spatial audio synthesis model, and generating final binaural spatial audio through the spatial audio synthesis model based on an iterative sampling path of a stochastic differential equation; wherein the spatial audio synthesis model is obtained by training a preset neural network model using a predefined binaural Schrödinger bridge parameterized target and boundary-assisted supervision.
[0138] Furthermore, the logic instructions in the aforementioned memory 430 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product, stored in a storage medium, includes instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0139] On the other hand, the present invention also provides a computer program product, which includes a computer program, which can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute a spatial audio synthesis method based on Schrödinger bridge provided by the above methods, the method including: obtaining a monophonic sound source signal, constructing a priori signal and noisy representation based on the monophonic sound source signal; inputting the priori signal and noisy representation into a pre-trained spatial audio synthesis model, and generating the final binaural spatial audio through the spatial audio synthesis model based on an iterative sampling path of a random differential equation; wherein the spatial audio synthesis model is obtained by training a preset neural network model through a predefined binaural Schrödinger bridge parameterized target and boundary-assisted supervision.
[0140] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute a spatial audio synthesis method based on a Schrödinger bridge provided by the above-mentioned methods, the method comprising: obtaining a monophonic sound source signal, constructing a priori signal and a noisy representation based on the monophonic sound source signal; inputting the priori signal and the noisy representation into a pre-trained spatial audio synthesis model, and generating the final binaural spatial audio through the spatial audio synthesis model based on an iterative sampling path of a random differential equation; wherein the spatial audio synthesis model is obtained by training a preset neural network model through a predefined binaural Schrödinger bridge parameterized target and boundary-assisted supervision.
[0141] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.
[0142] Through the above description of the embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a necessary general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods described in each embodiment or certain portions of the embodiments.
[0143] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A spatial audio synthesis method based on Schrödinger bridge, characterized in that: include: Acquire a monophonic sound source signal, and construct a priori signal and a noisy representation based on the monophonic sound source signal; Inputting the prior signal and the noisy representation into a pre-trained spatial audio synthesis model, and generating final binaural spatial audio through the spatial audio synthesis model based on an iterative sampling path of stochastic differential equations; The spatial audio synthesis model is obtained by training a preset neural network model through a predefined binaural Schrödinger bridge parameterized target and boundary auxiliary supervision.
2. The spatial audio synthesis method based on Schrödinger bridge according to claim 1, characterized in that: The spatial audio synthesis model is obtained by training a preset neural network model using a predefined binaural Schrödinger bridge parameterized objective and boundary-assisted supervision, specifically including: Obtain the quaternion of the monophonic sound source signal, the relative spatial position coordinates of the sound source and the receiver, and the receiver's head orientation information; Building a binaural Schrödinger bridge between the monophonic sound source signal and the audio signals received by the left and right ears of the receiver; Obtaining a noisy representation of the binaural Schrödinger bridge at each time step, and parameterizing the binaural Schrödinger bridge by shallowly adding noise based on the noisy representation; Based on the binaural Schrödinger bridge parameterization target, the preset spatial audio synthesis model is trained through the relative spatial position coordinates of the sound source and the receiver, the quaternion of the receiver's head orientation information, and boundary-assisted supervision.
3. The spatial audio synthesis method based on Schrödinger bridge according to claim 2, characterized in that: The step of building a binaural Schrödinger bridge between the monophonic sound source signal and the audio signals received by the left and right ears of the receiver specifically includes: Duplicating the monophonic sound source signal to form a dual-channel speech signal; The audio signals received by the left and right ears of the receiver are used as the generation target, and the binaural speech signal is used as the prior signal of the audio signals received by the left and right ears of the receiver; A binaural Schrödinger bridge is built between the prior signal and the generated target.
4. The spatial audio synthesis method based on Schrödinger bridge according to claim 2, characterized in that: The obtaining of a noisy representation at each time step in the binaural Schrödinger bridge and parameterizing the binaural Schrödinger bridge by shallowly adding noise based on the noisy representation specifically includes: Obtaining a noisy representation of each time step in the binaural Schrödinger bridge; Based on the noisy characterization, the noise content is first increased to a peak of the noise content, and then the noise content is gradually reduced until the noise disappears, thereby completing the parameterization of the binaural Schrödinger bridge.
5. The spatial audio synthesis method based on Schrödinger bridge according to claim 2, characterized in that: The training of the preset spatial audio synthesis model based on the binaural Schrödinger bridge parameterized target is completed through the relative spatial position coordinates of the sound source and the receiver, the quaternion of the receiver's head orientation information, and boundary auxiliary supervision, specifically including: The quaternion of the relative spatial position coordinates of the sound source and the receiver, and the receiver's head orientation information is used as conditional information to guide the training of a preset spatial audio synthesis model; Based on the binaural Schrödinger bridge parameterization target, the quaternion of the relative spatial position coordinates of the sound source and the receiver and the receiver's head orientation information is used as conditional information to guide the training of the preset spatial audio synthesis model. On this basis, the loss function is added through boundary-assisted supervision to supervise the training of the preset spatial audio synthesis model. The loss function includes: multi-scale amplitude loss and phase spectrum loss.
6. The spatial audio synthesis method based on Schrödinger bridge according to claim 1, characterized in that: Inputting the prior signal and the noisy representation into a pre-trained spatial audio synthesis model, and generating final binaural spatial audio through the spatial audio synthesis model based on an iterative sampling path of a stochastic differential equation, specifically includes: Inputting the prior signal and the noisy representation into a spatial audio synthesis model to generate a single-step prediction target; Sampling is performed based on the single-step prediction target through a stochastic differential equation iterative sampling path, and final two-channel spatial audio is synthesized according to the sampling results.
7. A spatial audio synthesis system based on Schrödinger bridge, characterized in that: The system comprises: An information acquisition module, configured to acquire a monophonic sound source signal and construct a priori signal and a noisy representation based on the monophonic sound source signal; An audio synthesis module, configured to input the prior signal and the noisy representation into a pre-trained spatial audio synthesis model, and generate final binaural spatial audio through the spatial audio synthesis model based on an iterative sampling path of stochastic differential equations; The spatial audio synthesis model is obtained by training a preset neural network model through a predefined binaural Schrödinger bridge parameterized target and boundary auxiliary supervision.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the spatial audio synthesis method based on Schrödinger bridge according to any one of claims 1 to 6 is implemented.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the spatial audio synthesis method based on Schrödinger bridge according to any one of claims 1 to 6 is implemented.
10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the spatial audio synthesis method based on Schrödinger bridge according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Speech synthesis method and device, electronic equipment and readable storage medium
CN117854470A
Signal generation and road reconstruction method and system based on Riemannian diffusion Schrodinger bridge
CN118089702A