Method and device for generating binaural spatial audio with multiple sound sources according to text

Through the cascading method, the large language model and diffusion model are used to generate binaural spatial audio, which solves the problems of inaccurate sound source position and low audio quality in the prior art, and achieves higher quality and more realistic binaural audio generation.

CN120199227APending Publication Date: 2025-06-24WUHAN UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510413478.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

Existing audio generation technologies are difficult to accurately generate binaural audio, resulting in inaccurate sound source position and low audio quality.

Method used

The cascading method is adopted, first preprocessing the text through a large language model to generate structural information; then using the diffusion model to generate single-channel audio; then using the binaural rendering model to render single-channel audio as binaural audio; finally synthesize the target binaural audio based on the timing information in the text.

Benefits of technology

The quality of binaural audio and the accuracy of sound source position are improved, and the generated audio is more realistic in human ear perception, solving the problems of high noise and inaccurate sound source position in the prior art.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120199227A_ABST
    Figure CN120199227A_ABST
Patent Text Reader

Abstract

The invention discloses a method and a device for generating a binaural spatial audio with multiple sound sources according to a text. The method comprises the following steps of: inputting a description type text or a parameter type text of the audio; preprocessing the description type text or the parameter type text by adopting a big language model to generate structural information including sound events, sound duration, sound source position information and time sequence information; generating a plurality of single-channel audios corresponding to sound events and sound durations in the input text by using a diffusion model; adopting a binaural rendering model to render all the single-channel audios into binaural audios conforming to the sound source position information in the input text; and synthesizing each binaural audio obtained by rendering into a target binaural audio according to the time sequence information of each sound source in the input text. According to the method, the reasonable sound source orientation can be given according to the physical law when the sound source position is missing, and the accuracy of converting the text into the binaural space audio is greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of audio signal generation, mainly to text-to-audio methods and methods for rendering single-channel audio to binaural spatial audio (TTA, text to audio). More specifically, it relates to a method and apparatus for generating binaural spatial audio with multiple sound sources based on text. Background Art

[0002] Audio generation technology has been widely applied in various educational systems and entertainment systems. For example, in virtual classrooms, text-to-audio can help teachers better complete teaching and improve teaching efficiency; in virtual reality, the generation of special effect audio can make virtual reality more realistic. In these applications of audio generation, a crucial part is to generate the spatial sense and orientation sense of the audio. However, in existing audio generation technologies, most still only generate single-channel audio, and a small part can generate spatial audio. Spatial audio is different from binaural audio. Spatial audio does not consider the influence of the human torso and ear canal on sound propagation and reception. When wearing headphones, spatial audio cannot very accurately reflect the position of the sound source. Therefore, generating binaural spatial audio based on text is a very useful research topic, and binaural audio with accurate sound source positions will bring a more realistic experience to users.

[0003] Regarding the problem of rendering binaural spatial audio, the head-related transfer function (HRTF) has shown effectiveness in practical applications. After the audio is transformed by the fast Fourier transform and multiplied by the HRTF, the spectrum of binaural spatial audio can be obtained. However, the recorded HRTF data is discontinuous and cannot meet the requirement of generating binaural spatial audio at arbitrary angles. At the same time, recording the HRTF data also requires a large amount of time and manpower. Although there are some existing technologies that predict the HRTF at arbitrary angles through a network, when using the predicted HRTF to render binaural spatial audio, the resulting audio has disadvantages such as high noise and inaccurate sound source positions. Summary of the Invention

[0004] Aiming at the technical problems existing in the prior art, the present invention provides a cascaded method for generating binaural spatial audio with controllable duration based on text. First, the generation of text-to-single-channel audio is completed, then the single-channel audio is directly rendered into binaural audio by a network, and finally, the individual binaural audios are synthesized into one binaural audio, reducing the network training difficulty while improving the quality and azimuth accuracy of the generated binaural audio.

[0005] To achieve the above object, the first aspect of the present invention provides a method for generating binaural spatial audio with multiple sound sources based on text, including:

[0006] Input descriptive text or parametric text of the audio;

[0007] Preprocess descriptive text or parametric text using a large language model to generate structural information including sound events, sound durations, sound source location information, and temporal information;

[0008] Use a diffusion model to generate a number of single-channel audio corresponding to the sound events and sound durations in the input text;

[0009] Use a binaural rendering model to render all single-channel audio into binaural audio that matches the sound source location information in the input text;

[0010] Synthesize the rendered binaural audio into target binaural audio according to the temporal information of each sound source in the input text.

[0011] In one implementation, preprocess descriptive text or parametric text using a large language model to generate structural information including sound events, sound durations, sound source location information, and temporal information, including:

[0012] Preprocess the input text using a large language model, extract the speech information and sound source features in the input text, and generate structural information including sound events, sound durations, sound source location information, and temporal information. Among them, the descriptive text or parametric text of the audio is text containing sound events, sound source locations, and times. When the sound source location information is not explicitly given, reasonable sound source location information is inferred according to the physical laws of the objective world and the context semantics.

[0013] In one implementation, use a diffusion model to generate a number of single-channel audio corresponding to the sound events and sound durations in the input text, including:

[0014] Use a pre-trained text-to-speech model to generate a number of single-channel audio for each sound event on the condition of the sound events and sound durations output by the large language model;

[0015] Sort the generated single-channel audio after CLAP scoring, and retain the single-channel audio that best matches the sound event information and sound duration.

[0016] In one implementation, the pre-trained text-to-speech model first encodes the audio using the variational encoder of the audio generation tool during the training phase, then scores the generated audio in combination with the CLAP model, selects the optimal audio, and the loss function is obtained by combining the flow matching loss and the preference optimization loss.

[0017] In one implementation, the sound source location information includes the sound source azimuth and the sound source distance. Use a binaural rendering model to render all single-channel audio into binaural audio that matches the sound source location information in the input text, including:

[0018] Adopt a binaural rendering network model based on Fourier transform. Conditional on the azimuth and distance of each sound source output by the large language model, for each single-channel audio, after frame division and windowing, predict the sound attenuation and time delay of its left and right ear channels respectively;

[0019] Through a preset signal processing method, stack each frame of the left and right channels into complete left and right channels to obtain a binaural audio that conforms to the sound source position information in the input text.

[0020] In one implementation, the loss function of the binaural rendering network model based on Fourier transform is obtained based on l2 loss, phase loss, binaural intensity difference loss, and multi-resolution short-time Fourier transform loss.

[0021] In one implementation, the timing information includes the start time of the sound. Synthesize the rendered binaural audios into a target binaural audio according to the timing information of each sound source in the input text, including:

[0022] According to the start time of each sound output by the large language model, perform corresponding time shift operations on the binaural audios generated by the binaural rendering model, and splice all the binaural audios into a target binaural audio.

[0023] Based on the same inventive concept, the second aspect of the present invention provides an apparatus for generating binaural spatial audio with multiple sound sources based on text, including:

[0024] An input module for inputting descriptive text or parametric text of the audio;

[0025] A preprocessing module for preprocessing the descriptive text or parametric text using a large language model to generate structural information including sound events, sound durations, sound source position information, and timing information;

[0026] A single-channel audio generation module for generating a number of single-channel audios corresponding to the sound events and sound durations in the input text using a diffusion model;

[0027] A rendering module that uses a binaural rendering model to render all single-channel audios into binaural audios that conform to the sound source position information in the input text;

[0028] A target binaural audio synthesis module for synthesizing the rendered binaural audios into a target binaural audio according to the timing information of each sound source in the input text

[0029] Based on the same inventive concept, the third aspect of the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the method for generating binaural spatial audio with multiple sound sources based on text described in the first aspect.

[0030] Based on the same inventive concept, a fourth aspect of the present invention provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the method for generating binaural spatial audio with multiple sound sources based on text described in the first aspect.

[0031] Compared with the prior art, the advantages and beneficial technical effects of the present invention are as follows:

[0032] The present invention provides a method for generating binaural spatial audio with multiple sound sources based on text. First, a large language model is used to preprocess the input descriptive or parametric text to extract the voice information and sound source features in the text. Then, a text-to-audio model is used to generate several single-channel audio corresponding to the text and duration, ensuring that the duration and content of each audio are consistent with the text description. Next, a binaural rendering model is used to render all single-channel audio into binaural audio that conforms to the position information in the text, particularly considering the influence of the human torso and ear canal structure on sound propagation, making the generated audio have higher spatial realism. Finally, according to the timing information of each sound source in the text, the individual binaural audio are synthesized into a binaural audio, ensuring the accuracy of the time sequence and spatial relationship. Text-to-binaural spatial audio is an audio generation technology that can ensure that the generated audio has the same direction as expected when perceived by the human ear. Different from the existing text-to-spatial audio technology, binaural spatial audio takes into account the influence of the human torso, ear canal, etc. on sound propagation in the sound propagation path from the sound source to the human ear, so it can generate audio that is more consistent with the described sound source orientation. At the same time, due to the strong information extraction ability of the large language model, this method can give a reasonable sound source orientation according to physical laws when the sound source position is missing, greatly improving the accuracy and robustness of text-to-binaural spatial audio. Description of the Drawings

[0033] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0034] Figure 1 It is a flowchart of the method for generating binaural spatial audio with multiple sound sources and controllable duration based on text in cascade type in the embodiment of the present invention.

[0035] Figure 2 It is a detailed flowchart of the whole process in the embodiment of the present invention.

[0036] Figure 3 It is a structural diagram of the text-to-audio model in the embodiment of the present invention.

[0037] Figure 4 It is a structural diagram of the binaural audio rendering network according to an embodiment of the present invention.

[0038] Figure 5 It is a schematic diagram of the influence of factors such as the human torso on the sound propagation process according to an embodiment of the present invention.

[0039] Figure 6 It is a comparison table of various parameters of the single-channel audio generated in the text-to-audio stage according to an embodiment of the present invention.

[0040] Figure 7 It is a comparison table of various parameters of the binaural audio generated in the binaural audio rendering stage according to an embodiment of the present invention.

[0041] Figure 8 It is a comparison chart of the distribution and average score of the subjective position perception score and the accuracy chart of the subjective perceived position and the actual sound source position according to an embodiment of the present invention.

[0042] Figure 9 It is a module diagram of a cascaded device for generating binaural spatial audio with controllable duration based on text according to an embodiment of the present invention. Specific Embodiments

[0043] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0044] Embodiment 1

[0045] This embodiment discloses a method for generating binaural spatial audio with multiple sound sources based on text. Please refer to Figure 1 , including:

[0046] S101: Input a descriptive text or parametric text of the audio;

[0047] S102: Use a large language model to preprocess the descriptive text or parametric text to generate structural information including sound events, sound durations, sound source position information, and timing information;

[0048] S103: Use a diffusion model to generate a number of single-channel audios corresponding to the sound events and sound durations in the input text;

[0049] S104: Use a binaural rendering model to render all single-channel audios into binaural audios that match the sound source position information in the input text;

[0050] S105: Synthesize the rendered binaural audios into a target binaural audio according to the timing information of each sound source in the input text.

[0051] Specifically, the content of the input text may include descriptions of sound events, spatial position information of sound sources, the time sequence of sounds, etc., providing basic information for subsequent processing.

[0052] In one implementation, S102 is implemented in the following manner:

[0053] Use a large language model to preprocess the input text, extract the speech information and sound source features in the input text, and generate structural information including sound events, sound durations, sound source position information, and timing information. Among them, the descriptive text or parametric text of the audio is text containing sound events, sound source positions, and times. When the sound source position information is not explicitly given, reasonable sound source position information is inferred according to the physical laws of the objective world and the context semantics.

[0054] Specifically, using the large language model, convert the input of the descriptive or parametric text into several sets of processing conditions for subsequent steps, including sound event information, sound duration, sound source azimuth, sound source distance, and sound start time. If the sound source position information is not explicitly given, then infer a reasonable sound source azimuth according to the physical laws of the objective world and the context semantics to ensure the accuracy and rationality of the spatial positioning of the generated audio.

[0055] In one implementation, S103 is implemented in the following manner:

[0056] Use a pre-trained text-to-speech model to generate several single-channel audios for each sound event with the sound events and sound durations output by the large language model as conditions;

[0057] Sort the generated single-channel audios after CLAP scoring, and retain the single-channel audio that best matches the sound event information and sound duration.

[0058] Specifically, use the pre-trained text-to-speech model, with the sound event information and sound duration in step S102 as conditions. For each sound event information, use the diffusion model to generate several single-channel audios. The generated single-channel audios are sorted after CLAP scoring, and the single-channel audio that best fits the event information is left until corresponding single-channel audios are generated for all sound event information, ensuring that the content and duration of each audio are highly consistent with the text description.

[0059] CLAP (Contrastive Language-Assembly Pre-training) is a model evaluation method based on natural language supervision. Its core lies in achieving cross-modal semantic alignment through contrastive learning, and it is applicable to scoring tasks in scenarios such as code representation and image generation.

[0060] In one implementation, S104 is achieved in the following way:

[0061] Adopt a binaural rendering network model based on the Fourier transform. Conditional on the azimuth and distance of each sound source output by the large language model, after windowing each single-channel audio frame by frame, predict the sound attenuation and time delay of its left and right ear channels respectively;

[0062] Overlay each frame of the left and right channels into complete left and right channels through a preset signal processing method to obtain binaural audio that matches the sound source position information in the input text.

[0063] Specifically, use a binaural rendering network model based on the Fourier transform. Conditional on the azimuth and distance of each sound source in step S102, after windowing each single-channel audio frame by frame, predict the sound attenuation and time delay of its left and right ear channels respectively, and finally overlay each frame of the left and right channels into complete left and right channels through the WOLA method; repeat the above operations until each single-channel audio is rendered into binaural audio according to the corresponding sound source azimuth and distance, making the generated audio have higher spatial realism and immersion.

[0064] The WOLA method processes the input signal in blocks, then processes each block using a weighting function, and finally overlaps and adds them to obtain the output signal.

[0065] In one implementation, S105 is achieved in the following way:

[0066] According to the start time of each sound output by the large language model, perform corresponding time shift operations on each binaural audio generated by the binaural rendering model, and splice all binaural audios into the target binaural audio.

[0067] Specifically, according to the start time of each sound given in step S102, perform corresponding time shift operations on each binaural audio generated in step S104 and then splice all binaural audios into a binaural audio, ensuring that the finally synthesized audio is completely consistent with the text description in terms of time sequence and spatial relationship.

[0068] Based on the actual application scenarios, considering the convenience of generating binaural spatial audio and the impact of the sound quality and spatial sense of the generated audio in applications such as education and virtual reality that require real - space audio, the present invention proposes a method for generating binaural spatial audio with multiple sound sources and controllable duration based on text. Through a cascaded overall structure, first, a large - language model is used to pre - process the information in the text. Subsequently, a text - to - audio model is used to generate several single - channel audio. Then, all single - channel audio is rendered in corresponding spatial orientations according to the text conditions. Finally, the binaural spatial audio obtained from each rendering is time - shifted according to the conditions and merged into a large binaural spatial audio. In this way, the difficulty of model training and the computational complexity can be greatly reduced while realizing the generation of multi - channel binaural audio, and at the same time, the inference time when generating binaural audio can be reduced, and the quality and semantic consistency of the generated audio can be improved. The method provided by the present invention can be implemented by computer software technology. For the specific flowchart, see Figure 1 and Figure 2 。

[0069] The large - language model (LLM) is used to process descriptive or parametric text, converting the text input into several sets of structural outputs containing sound event information, duration, sound source azimuth distance, and start time. The output structure of each set is <sound event>@<duration>@<sound source azimuth angle, sound source elevation angle>@<sound source distance>@<sound event start time>.

[0070] Subsequently, the output of the large - language model is fed into the text - to - audio model. In this stage, a pre - trained model is used, and the model structure is as Figure 3 shown. This model generally uses a diffusion model. In the training stage, first, Stable Audio VAE (Variational Auto - Encoder for Audio Generation) is used to encode the audio. Subsequently, the CLAP model is combined to score the generated audio, and the optimal audio is selected. Among them, x1 is the latent representation of the audio sample, encoded by the Variational Auto - Encoder (VAE), x0 is the noise output, following the standard normal distribution x0 ∼ N(0, I), where I is the identity matrix, v t is defined as representing the change direction of x t towards the target x1, where t is a scalar sampled from the uniform distribution U(0, 1), representing the normalized time in the diffusion process, u(x t , t; θ) is the speed predicted by the neural network model with parameters θ, used to predict the speed v t , and represent the speed fields of the preferred audio and the inferior - selected audio respectively. The loss function for intermediate - state training is as follows:

[0071]

[0072] Among them is the preference optimization loss function, where σ is the sigmoid function and β is a hyperparameter used to control the intensity of preference optimization. Among them is the flow matching loss function, and the goal is to make the velocity u(x t , t; θ) close to the true velocity v t , where π0 is the initial base model obtained through pre-training, and π n is the model checkpoint after the nth iteration.

[0073] At the same time, CLAP is used as a proxy reward model to evaluate the matching degree between the generated audio and the input text. and respectively represent the intermediate states of the preferred audio and the inferior audio evaluated by CLAP, and θ ref represents the parameters of the current base model.

[0074] is the total loss function, which combines the flow matching loss and the preference optimization loss to stabilize the training and avoid over-optimization.

[0075] The model obtained through the above training can complete several single-channel audio generation tasks. Subsequently, this method trains a binaural rendering model for single-channel to binaural audio. The overall framework of this model is as shown in Figure 4 . The main goal of this model is to predict the frequency-domain outputs of the left ear and the right ear, which are obtained by applying amplitude attenuation and phase shift to the mono-spectrum X(k) =. The formula is as follows:

[0076]

[0077] where X(k) is calculated through the discrete Fourier transform (DFT), k is the frequency index, is the angular frequency, K is the number of frequency points, is the amplitude attenuation coefficient, indicating the energy attenuation of the sound during propagation. C is the number of channels (left ear L, right ear R). represents the phase shift, indicating the time delay of the sound during propagation.

[0078] In the figure is the spatial position of the sound source, represented as three-dimensional coordinates. is the direction of the sound source, represented as a quaternion, and g is the geometric delay, calculated based on the direct path distance from the sound source to the listener. is a strictly positive amplitude coefficient, obtained through a non-linear activation function. is the phase shift, obtained through a non-linear activation function, and its value is restricted to be within half of the frame length. is the final binaural audio output, through It is obtained by performing inverse discrete Fourier transform (IDFT) and weighted overlap and add (WOLA) synthesis. respectively represent the frequency-domain outputs of the predicted left ear and the predicted right ear, and σ L (c,k), σ R (c,k) represent the left-ear channel amplitude attenuation coefficient and the right-ear channel amplitude attenuation coefficient respectively. The training loss function is as follows:

[0079]

[0080] Among them, these four items are the l2 loss, which measures the mean square error of the time-domain signal, is the predicted binaural audio, y is the real binaural audio, which is used for training, the phase loss, which measures the norm of the frequency-domain phase difference, is the phase angle of the predicted binaural audio, ∠Y is the phase angle of the real binaural audio for training, and the interaural intensity difference loss measures the difference in interaural intensity difference, is the predicted interaural intensity difference, IID(y) is the real interaural intensity difference for training, and there is also the multi-resolution short-time Fourier transform loss which measures the difference in multi-resolution short-time Fourier transform, is the difference between the short-time Fourier transform of the predicted binaural audio and the real binaural audio.

[0081] According to the above structure and loss function, a binaural-rendered audio with good effect and high stability can be trained. After each single-channel audio is rendered, it is time-shifted according to the start time of each sound event and then spliced into a binaural spatial audio to complete the whole process of generating the target audio.

[0082] Figure 5 shows a schematic diagram of the influence of factors such as the human torso on the sound propagation process in the present invention. When there is no influence of the human torso, the sound propagation is only affected by the audio reflected by the wall, and there is no occlusion in the propagation path. However, when considering the human torso, it is necessary to consider the blocking effect of the human torso on sound propagation, and at the same time consider the influence of the ear canal on the human body when receiving sound. Therefore, the obtained binaural spatial audio will be more realistic.

[0083] Figure 6Shows the comparison table of various parameters of the single-channel audio generated in the text-to-audio stage of the present invention. The performance indicators of four audio generation methods are compared in the table. The up and down arrows indicate that the better the effect, the larger or smaller this item is. FD (Frechet Distance) is used to measure the distribution difference between the generated audio and the real audio. The smaller the value, the closer the generated result is to the real data. KL (Kullback-Leibler Divergence) is used to evaluate the similarity between the generated distribution and the real distribution. The smaller the value, the closer the two are. IS (Inception Score) is used to comprehensively evaluate the quality and diversity of the generated audio. The larger the value, the better the generation effect. CLAP (Audio-Text Alignment Score) is used to measure the matching degree between the text and the audio. The higher the value, the more matching. MOS-Q (Mean Opinion Score - Quality) is used to represent the audio quality evaluated manually. The higher the score, the higher the audio quality. MOS-F (Mean Opinion Score - Fidelity) is used to represent the fidelity evaluated manually. The higher the score, the more realistic the generated audio. Inference Time represents the time required to generate a single audio, and the shorter the better. The optimal indicator is marked with *, and the sub-optimal indicator is marked with The various parameters show that the audio generated by the method in the present invention has the characteristics of high quality, fast speed, and semantic fit. In the table, AudioLDM, Make-An-Audio2, TangoFlux, and TangoFlux-NFS-woNI are all audio generation methods. Among them, the first three methods generate single-channel audio, and the last method generates binaural audio because it passes through the NFS-woNI binaural rendering network. We take one of the channels for comparison.

[0084] Figure 7 Shows the comparison table of various parameters of the binaural audio generated in the binaural audio rendering stage of the present invention. The performance indicators of four audio generation methods are compared in the table. l2 (L2 error) is used to measure the waveform reconstruction error, and the smaller the value, the better. (Magnitude spectrum loss) is used to evaluate the accuracy of the spectrum magnitude, and the smaller the value, the better. (Phase spectrum loss) is used to evaluate the accuracy of phase reconstruction, and the smaller the value, the better. (Short-time Fourier transform loss) is used to evaluate the comprehensive spectrum error, and the smaller the value, the better. PESQ (Perceptual Evaluation of Speech Quality) is used to evaluate the speech quality, and the larger the value, the better. MOS-Q (Mean Opinion Score - Quality) represents the audio quality evaluated manually. The higher the score, the higher the audio quality. MOS-P (Mean Opinion Score - Position) is used to evaluate the accuracy of the artificial perception of the binaural audio position, and the higher the value, the better. The optimal indicator is marked with *, and the sub-optimal indicator is marked with The various parameters show that the method in the present invention generates binaural spatial audio with high azimuth accuracy and high audio quality, which conforms to the characteristics of binaural audio in reality. The table shows four methods for generating binaural audio. The generation method in the first stage is controlled as TangoFlux, and four methods are used for transformation in the second stage to compare multiple indicators, and finally the table is obtained.

[0085] Figure 8 It shows a comparison chart of the distribution and average score of subjective position perception scores in the present invention, as well as an accuracy chart of the subjective perceived position and the actual sound source position. Among them, Correct is marked for correct and Incorrect is marked for incorrect. MOS-P (Mean Opinion Score - Position) is used to evaluate the accuracy of artificially perceiving the binaural audio position, and the higher the value, the better. The results show that the accuracy rate of the subjective evaluators in selecting the azimuth reaches 86.25%, which indicates that the overall method of the present invention is feasible, and the generated multi-source binaural audio has a good sense of spatial azimuth.

[0086] Embodiment 2

[0087] Based on the same inventive concept, this embodiment discloses a device for generating binaural spatial audio of multi-sources according to text. Please refer to Figure 9 , including:

[0088] An input module 201, configured to input a descriptive text or a parametric text of the audio;

[0089] A preprocessing module 202, configured to preprocess the descriptive text or the parametric text by using a large language model to generate structural information including sound events, sound durations, sound source position information, and timing information;

[0090] A single-channel audio generation module 203, configured to generate a plurality of single-channel audios corresponding to the sound events and sound durations in the input text by using a diffusion model;

[0091] A rendering module 204, configured to render all single-channel audios into binaural audios that conform to the sound source position information in the input text by using a binaural rendering model;

[0092] A target binaural audio synthesis module 205, configured to synthesize the rendered binaural audios into a target binaural audio according to the timing information of each sound source in the input text.

[0093] Since the device introduced in Embodiment 2 of the present invention is the device adopted for implementing the method for generating binaural spatial audio of multi-sources according to text in Embodiment 1 of the present invention, based on the method introduced in Embodiment 1 of the present invention, those skilled in the art can understand the specific structure and variations of the device, so it will not be elaborated here. All devices adopted for the method in Embodiment 1 of the present invention fall within the scope of protection of the present invention.

[0094] Embodiment 3

[0095] Based on the same inventive concept, the present invention further provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the method described in Embodiment 1 is implemented.

[0096] Since the computer-readable storage medium introduced in Embodiment 3 of the present invention is the computer-readable storage medium used for implementing the method of generating binaural spatial audio with multiple sound sources based on text in Embodiment 1 of the present invention, based on the method introduced in Embodiment 1 of the present invention, those skilled in the art can understand the specific structure and variations of the computer-readable storage medium, so it will not be elaborated here. Any computer-readable storage medium used in the method of Embodiment 1 of the present invention falls within the scope of protection of the present invention.

[0097] Embodiment 4

[0098] The present invention further provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, the method described in Embodiment 1 is implemented.

[0099] Since the computer device introduced in Embodiment 4 of the present invention is the computer device used for implementing the method of generating binaural spatial audio with multiple sound sources based on text in Embodiment 1 of the present invention, based on the method introduced in Embodiment 1 of the present invention, those skilled in the art can understand the specific structure and variations of the computer device, so it will not be elaborated here. Any computer device used in the method of Embodiment 1 of the present invention falls within the scope of protection of the present invention.

[0100] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0101] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and combinations of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general purpose computer, special purpose computer, embedded processor, or other programmable data processing device to produce a machine, such that the instructions executed by the processor of the computer or other programmable data processing device create means for implementing the functions specified in one flow Figure 1 one flow or more flows and / or blocks Figure 1 or in one block or more blocks.

[0102] Although the preferred embodiments of the present invention have been described, additional changes and modifications can be made by those skilled in the art once they learn of the basic creative concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications that fall within the scope of the present invention. Obviously, those skilled in the art can make various changes and variations to the embodiments of the present invention without departing from the spirit and scope of the embodiments of the present invention. Thus, if these modifications and variations of the embodiments of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.

Claims

1. A method for generating binaural spatial audio with multiple sound sources based on text, characterized in that: include: Enter descriptive text or parameter text for the audio; Use a large language model to preprocess descriptive text or parameter text to generate structural information including sound events, sound duration, sound source location information and timing information; Generate a number of single-channel audios corresponding to the sound events and sound durations in the input text using a diffusion model; A binaural rendering model is used to render all single-channel audio into binaural audio that matches the sound source location information in the input text. The rendered binaural audios are synthesized into target binaural audio according to the timing information of each sound source in the input text.

2. The method for generating binaural spatial audio with multiple sound sources based on text according to claim 1, characterized in that: Use a large language model to preprocess descriptive text or parameter text to generate structural information including sound events, sound duration, sound source location information and timing information, including: A large language model is used to preprocess the input text, extract the voice information and sound source features in the input text, and generate structural information including sound events, sound duration, sound source location information and timing information. The descriptive text or parameter text of the audio is the text containing the sound event, sound source location and time. When the sound source location information is not clearly given, reasonable sound source location information is inferred based on the physical laws of the objective world and the contextual semantics.

3. The method for generating binaural spatial audio with multiple sound sources based on text according to claim 1, characterized in that: Use the diffusion model to generate several single-channel audios corresponding to the sound events and sound durations in the input text, including: Using the pre-trained text-to-speech model, we generate several single-channel audios for each sound event based on the sound events and sound duration output by the large language model. The generated single-channel audio is sorted after CLAP scoring, and the single-channel audio that best matches the sound event information and sound duration is retained.

4. The method for generating binaural spatial audio with multiple sound sources based on text according to claim 3, characterized in that: During the training phase, the pre-trained text-to-speech model first encodes the audio using the variational encoder of the audio generation tool, and then scores the generated audio in combination with the CLAP model to select the optimal audio. The loss function is obtained by combining the stream matching loss and the preference optimization loss.

5. The method for generating binaural spatial audio with multiple sound sources based on text according to claim 1, characterized in that: The sound source location information includes the sound source direction and the sound source distance. The binaural rendering model is used to render all single-channel audio into binaural audio that matches the sound source location information in the input text, including: A binaural rendering network model based on Fourier transform is used. Based on the sound source orientation and distance of each group output by the large language model, each single-channel audio is framed and windowed to predict the sound attenuation and time delay of the left and right ear channels respectively. The frames of the left and right channels are superimposed into complete left and right channels through a preset signal processing method to obtain binaural audio that matches the sound source position information in the input text.

6. The method for generating binaural spatial audio with multiple sound sources based on text according to claim 5, characterized in that: The loss function of the Fourier transform-based binaural rendering network model is obtained based on l2 loss, phase loss, binaural intensity difference loss and multi-resolution short-time Fourier transform loss.

7. The method for generating binaural spatial audio with multiple sound sources based on text according to claim 1, characterized in that: The timing information includes the start time of the sound. The rendered binaural audios are synthesized into the target binaural audio according to the timing information of each sound source in the input text, including: According to the start time of each sound output by the large language model, a corresponding time shift operation is performed on each binaural audio generated by the binaural rendering model, and all binaural audios are spliced ​​into the target binaural audio.

8. A device for generating binaural spatial audio with multiple sound sources based on text, characterized in that: include: An input module, used to input descriptive text or parameter text for audio; A preprocessing module is used to preprocess the descriptive text or parameter text using a large language model to generate structural information including sound events, sound duration, sound source location information and timing information; A single-channel audio generation module is used to generate a number of single-channel audios corresponding to the sound events and sound durations in the input text using a diffusion model; The rendering module uses a binaural rendering model to render all single-channel audio into binaural audio that matches the sound source location information in the input text; The target binaural audio synthesis module is used to synthesize the rendered binaural audios into target binaural audios according to the timing information of each sound source in the input text.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method for generating binaural spatial audio of multiple sound sources based on text as claimed in any one of claims 1 to 7 is implemented.

10. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the method for generating binaural spatial audio of multiple sound sources based on text is implemented as claimed in any one of claims 1 to 7.

Citation Information

Cited By

  • Audio generation method and device, equipment, storage medium and program product

    CN120998175A