Sound generation method based on multiple modes

Through the multimodal input and feature alignment network method, the problem of insufficient cross-modal unified capability in the audio generation method is solved, and high-quality audio generation is realized, which is suitable for a variety of application scenarios.

CN120452412AInactive Publication Date: 2025-08-08DATA TRANSMISSION GRP
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510571515.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-06
Publication Date
2025-08-08
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing audio and music generation methods lack cross-modal unified capabilities, making it difficult to effectively integrate diversified inputs, resulting in difficulty in model convergence and degradation of generation quality, and are subject to the scarcity of high-quality multimodal training data.

Method used

Using multimodal input, feature extraction, alignment, splicing and large-model training methods, high-quality audio or music is generated through feature alignment networks and dynamic mask training strategies, combining diffusion models and autoregressive models.

Benefits of technology

It realizes cross-modal fusion capability, improves the correlation and robustness of generated audio and input scenes, and is suitable for scenes such as film and television soundtracks, game sound effects, etc., reduces labeling costs and improves the diversity and stability of generated results, and supports multi-task transfer learning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120452412A_ABST
    Figure CN120452412A_ABST
Patent Text Reader

Abstract

The invention discloses a sound generation method based on multiple modes, and belongs to the technical field of artificial intelligence and multimedia generation, and the method comprises the following steps: S1, multi-mode input: inputting multi-mode contents, including texts, videos, images, music and audios; s2, feature extraction; s3, feature alignment: additionally adding three corresponding small networks for the three extracted features, aligning the three extracted features in dimensions, and generating three aligned features; s4, feature splicing: splicing the three aligned features front and back, and inputting the spliced features together to generate a large model; s5, training a large model; s6, calculating a loss function; and S7, outputting audio or music. According to the method, multiple modes including texts, videos, images, music and audios serve as input, high-quality audios or music is generated in combination with the generation capacity of the large model, and due to the fact that the input modes are multiple modes, corresponding music or audios can be generated only by inputting one or more modes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence and multimedia generation technology, and in particular to a multimodal sound generation method. Background Art

[0002] With the development of deep learning and artificial intelligence technologies over the past decade, the audio field has experienced rapid growth. Many high-performance algorithms have emerged, including those for voice cloning, speech recognition, timbre recognition, and sound generation.

[0003] For various audio tasks, the most commonly used method is to first convert the audio input into a series of discrete features, and then use the learning ability of the neural network to learn a specific data set to achieve good task results.

[0004] Audio and music generation has become a key task in numerous applications, but existing methods have significant limitations: they operate independently, lack cross-modal unification capabilities, are constrained by the scarcity of high-quality multimodal training data, and have difficulty effectively integrating diverse inputs. The feature dimensions and semantic spaces of different modalities differ significantly, and direct splicing can easily lead to model convergence difficulties and reduced generation quality. Therefore, we propose a multimodal sound generation method to address this problem. Summary of the Invention

[0005] The object of the present invention is to provide a multimodal sound generation method to solve the problems raised in the above background technology.

[0006] In order to achieve the above object, the present invention adopts the following technical solutions:

[0007] A multimodal sound generation method comprises the following steps:

[0008] S1. Multimodal input: Input multimodal content, including text, video, image, music, and audio;

[0009] S2. Feature extraction: Input audio or music into a speech-based model to extract a text description of the music. Input audio or music into an audio model to generate the main content of the audio. For images, this can be expanded into a video along the time dimension. For a video sampling model, video features are extracted.

[0010] S3, feature alignment: For the three extracted features, three corresponding small networks are added to align the three extracted features in dimension to generate three aligned features;

[0011] S4, feature stitching: stitch the three aligned features together and input them together to generate a large model;

[0012] S5. Large model training: Train the generated large model. During training, a certain mask is randomly applied to the three features. The generalization ability of the model is improved through learning of the large model.

[0013] S6. Loss function calculation: Compare the input audio or music with the generated audio or music, calculate the loss function using a diffusion model, autoregressive model, or flow model, and train to generate a large model.

[0014] S7. Audio or music output: One or more of text, video, image, music, and audio are input into a large model to generate corresponding music or audio.

[0015] Preferably, in S2, a 3D Vae model or a Clip-Vit model is used to extract features of the video clip;

[0016] For music or audio, use open source music encoder to extract features;

[0017] For text description, T5 or CLIP is used to extract features.

[0018] Preferably, in S2, the audio content includes but is not limited to the type of music, instrument, emotion, and rhythm.

[0019] Preferably, in S3, the small network includes a Transformer layer, and the Transformer layer aligns the three extracted features in dimension.

[0020] Preferably, in said S5, for the video clip, a portion of the area is randomly blocked;

[0021] For music clips, randomly delete a certain section in the middle;

[0022] For text descriptions, randomly delete a certain paragraph of the text description.

[0023] Preferably, the network framework for generating the large model is Unet or Dit.

[0024] Preferably, in said S3, the specific steps are:

[0025] S301, feature standardization preprocessing

[0026] The features of the three modalities, video, audio, and text, are standardized separately. For video features, LayerNorm is used to eliminate dimensional differences; for audio features, mean-variance normalization is used; for text features, Softmax weight normalization is applied. If the dimensions of the modal features are inconsistent, they are mapped to a unified dimension through a linear projection layer.

[0027] S302, Multimodal Alignment Network Design: Video features are input into the Transformer layer to capture spatiotemporal correlations; audio features enhance temporal sequence coherence through a self-attention mechanism; text features fuse visual and auditory semantics through a cross-attention layer; attention weight matrices are calculated for video and text features, and weighted summation is used to generate a joint representation; InfoNCE loss is used to maximize the cosine similarity of video-audio, video-text, and audio-text pairs, and a domain discriminator is introduced to distinguish modal sources and force feature distribution alignment; initially, only high-confidence modal pairs are aligned, and low-confidence modal pairs are gradually added;

[0028] S303, alignment result verification and screening: Calculate the cosine similarity matrix of video-audio, video-text, and audio-text, screen feature pairs with similarity below the threshold, and trigger the realignment process;

[0029] S304, abnormal feature removal: use the isolation forest algorithm to detect outlier feature points and remove noise interference features.

[0030] Preferably, in said S4, the specific steps are as follows:

[0031] S401, splicing structure design: The aligned video, audio, and text features are all 512-dimensional, and are spliced into a 1536-dimensional vector along the channel dimension; a learnable modal identification vector is inserted to enhance the model's perception of the modal source;

[0032] S402, position coding injection: add temporal position coding to video features; use BERT-style position coding for text features; apply frequency domain position coding to audio features;

[0033] S403, feature dimensionality reduction and fusion: Use linear layer for feature dimensionality reduction, and the activation function is GELU;

[0034] Introduce residual connections to retain original feature information;

[0035] S404, multi-scale feature fusion: Dynamically adjust the weights of each modal feature through the SENet module, and select the dominant modal feature based on the gated recurrent unit;

[0036] S405, post-splicing feature verification: input the spliced features to the discriminator network, predict the input modality combination type, and when the accuracy is lower than the preset value, trigger the feature return realignment process;

[0037] S406, Noise robustness enhancement: Add Gaussian noise and random mask to train the model's anti-interference ability.

[0038] The beneficial effects of the present invention are:

[0039] In the present invention, the multimodal sound generation method described effectively solves the semantic fragmentation problem between different modal data in traditional methods through multimodal feature alignment technology, significantly improving the relevance of generated audio to the input scene. For example, when a video clip is input, the model can automatically parse the picture content and generate background music that matches its rhythm, while accurately capturing the emotional tone of the picture, so that the generated audio is highly synchronized with the visual content in terms of style and dynamic changes. This cross-modal fusion capability breaks through the limitations of single-modal generation and provides a more natural interactive experience for scenes such as film and television soundtracks and game sound effects.

[0040] In the present invention, a multimodal sound generation method and a dynamic mask training strategy enhance the model's robustness to incomplete input. Even if only a single modal information is provided, the model can still generate complete and semantically consistent audio content through a cross-modal completion mechanism. By introducing an adversarial feature alignment network, the model continuously optimizes the distribution consistency of multimodal features during training, reducing dependence on specific data distributions, thereby significantly reducing annotation costs and improving the diversity of generated results. This technology is particularly suitable for real-time interactive scenarios, such as in a virtual reality environment, where users can instantly generate sound effects adapted to the environment through gestures or voice commands.

[0041] The present invention describes a multimodal sound generation method. Its hybrid generation architecture combines the high fidelity of a diffusion model with the temporal coherence of an autoregressive model. This ensures rich audio detail while optimizing the stability of long-term generation. Through end-to-end joint training, the model can flexibly adapt to different task requirements, such as accurately restoring instrument timbre and chord progressions in music creation or simulating the spatial sense of complex sound fields in environmental sound synthesis. This architectural design not only improves generation efficiency but also supports multi-task transfer learning, providing a scalable solution for vertical fields such as music production and smart cockpit audio effects.

[0042] In the present invention, the multimodal sound generation method described above has an open design of a multimodal input interface, which enables a single model to cover diverse scenarios such as music creation, sound effect synthesis, and interactive media generation. This avoids the compatibility issues caused by modal switching in traditional cascade systems. Through a dynamic feature fusion mechanism, the model can automatically identify the primary and secondary relationships of input modalities and weightedly integrate features. For example, in video dubbing tasks, visual synchronization is prioritized, and in poetry recitation generation, text semantic analysis is emphasized. This flexibility significantly reduces deployment costs and provides a technical foundation for personalized audio generation.

[0043] In the present invention, the multimodal sound generation method described takes multiple modalities including text, video, image, music and audio as input, combines the generation capability of a large model, and generates high-quality audio or music. Since the input of the present invention is multimodal, the present invention only needs to input one or more of them in the inference stage to generate the corresponding music or audio. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 This is a flowchart of a multimodal sound generation method proposed by the present invention. DETAILED DESCRIPTION

[0045] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments.

[0046] Reference Figure 1 , a multimodal sound generation method, comprising the following steps:

[0047] S1. Multimodal input: Input multimodal content, including text, video, image, music, and audio;

[0048] S2. Feature extraction: Input audio or music into a speech-based model to extract a text description of the music. Input audio or music into an audio model to generate the main content of the audio. For images, this can be expanded into a video along the time dimension. For a video sampling model, video features are extracted.

[0049] S3, feature alignment: For the three extracted features, three corresponding small networks are added to align the three extracted features in dimension to generate three aligned features;

[0050] S4, feature stitching: stitch the three aligned features together and input them together to generate a large model;

[0051] S5. Large model training: Train the generated large model. During training, a certain mask is randomly applied to the three features. The generalization ability of the model is improved through learning of the large model.

[0052] S6. Loss function calculation: Compare the input audio or music with the generated audio or music, calculate the loss function using a diffusion model, autoregressive model, or flow model, and train to generate a large model.

[0053] S7. Audio or music output: One or more of text, video, image, music, and audio are input into a large model to generate corresponding music or audio.

[0054] In this embodiment, in S2, a 3D Vae model or a Clip-Vit model is used to extract features of the video clip;

[0055] For music or audio, use open source music encoder to extract features;

[0056] For text description, T5 or CLIP is used to extract features.

[0057] In this embodiment, in S2, the content of the audio includes but is not limited to the type of music, instrument, emotion, and rhythm.

[0058] In this embodiment, in S3, the small network includes a Transformer layer, which aligns the three extracted features in dimension.

[0059] In this embodiment, in S5, for the video clip, a portion of the area is randomly blocked;

[0060] For music clips, randomly delete a certain section in the middle;

[0061] For text descriptions, randomly delete a certain paragraph of the text description.

[0062] In this embodiment, the network framework for generating the large model is Unet or Dit.

[0063] In this embodiment, in S3, the specific steps are:

[0064] S301, feature standardization preprocessing

[0065] The features of the three modalities, video, audio, and text, are standardized separately. For video features, LayerNorm is used to eliminate dimensional differences; for audio features, mean-variance normalization is used; for text features, Softmax weight normalization is applied. If the dimensions of the modal features are inconsistent, they are mapped to a unified dimension through a linear projection layer.

[0066] S302, Multimodal Alignment Network Design: Video features are input into the Transformer layer to capture spatiotemporal correlations; audio features enhance temporal sequence coherence through a self-attention mechanism; text features fuse visual and auditory semantics through a cross-attention layer; attention weight matrices are calculated for video and text features, and weighted summation is used to generate a joint representation; InfoNCE loss is used to maximize the cosine similarity of video-audio, video-text, and audio-text pairs, and a domain discriminator is introduced to distinguish modal sources and force feature distribution alignment; initially, only high-confidence modal pairs are aligned, and low-confidence modal pairs are gradually added;

[0067] S303, alignment result verification and screening: Calculate the cosine similarity matrix of video-audio, video-text, and audio-text, screen feature pairs with similarity below the threshold, and trigger the realignment process;

[0068] S304, abnormal feature removal: use the isolation forest algorithm to detect outlier feature points and remove noise interference features.

[0069] In this embodiment, in S4, the specific steps are as follows:

[0070] S401, splicing structure design: The aligned video, audio, and text features are all 512-dimensional, and are spliced into a 1536-dimensional vector along the channel dimension; a learnable modal identification vector is inserted to enhance the model's perception of the modal source;

[0071] S402, position coding injection: add temporal position coding to video features; use BERT-style position coding for text features; apply frequency domain position coding to audio features;

[0072] S403, feature dimensionality reduction and fusion: Use linear layer for feature dimensionality reduction, and the activation function is GELU;

[0073] Introduce residual connections to retain original feature information;

[0074] S404, multi-scale feature fusion: Dynamically adjust the weights of each modal feature through the SENet module, and select the dominant modal feature based on the gated recurrent unit;

[0075] S405, post-splicing feature verification: input the spliced features to the discriminator network, predict the input modality combination type, and when the accuracy is lower than the preset value, trigger the feature return realignment process;

[0076] S406, Noise robustness enhancement: Add Gaussian noise and random mask to train the model's anti-interference ability.

[0077] The above is a detailed introduction to a multimodal sound generation method provided by the present invention. Specific embodiments are used herein to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present invention, the present invention can also be improved and modified in several ways, and these improvements and modifications also fall within the scope of protection of the claims of the present invention.

Claims

1. A multimodal sound generation method, characterized in that: The following steps are involved: S1. Multimodal input: Input multimodal content, including text, video, image, music, and audio; S2. Feature extraction: Input audio or music into a speech-based model to extract a text description of the music. Input audio or music into an audio model to generate the main content of the audio. For images, this can be expanded into a video along the time dimension. For a video sampling model, video features are extracted. S3, feature alignment: For the three extracted features, three corresponding small networks are added to align the three extracted features in dimension to generate three aligned features; S4, feature stitching: stitch the three aligned features together and input them together to generate a large model; S5. Large model training: Train the generated large model. During training, a certain mask is randomly applied to the three features. The generalization ability of the model is improved through learning of the large model. S6. Loss function calculation: Compare the input audio or music with the generated audio or music, calculate the loss function using a diffusion model, autoregressive model, or flow model, and train to generate a large model. S7. Audio or music output: One or more of text, video, image, music, and audio are input into a large model to generate corresponding music or audio.

2. The multimodal sound generation method according to claim 1, wherein: In S2, the 3DVae model or the Clip-Vit model is used to extract features of the video clip; For music or audio, use open source music encoder to extract features; For text description, T5 or CLIP is used to extract features.

3. The multimodal sound generation method according to claim 1, wherein: In S2, the content of the audio includes but is not limited to the type of music, instrument, emotion, and rhythm.

4. The multimodal sound generation method according to claim 1, wherein: In S3, the small network includes a Transformer layer, which aligns the three extracted features in dimension.

5. The multimodal sound generation method according to claim 1, wherein: In said S5, for the video clip, a portion of the area is randomly blocked; For music clips, randomly delete a certain section in the middle; For text descriptions, randomly delete a certain paragraph of the text description.

6. The multimodal sound generation method according to claim 1, wherein: The network framework for generating large models is Unet or Dit.

7. The multimodal sound generation method according to claim 1, wherein: In S3, the specific steps are: S301, feature standardization preprocessing The features of the three modalities, video, audio, and text, are standardized separately. For video features, LayerNorm is used to eliminate dimensional differences; for audio features, mean-variance normalization is used; for text features, Softmax weight normalization is applied. If the dimensions of the modal features are inconsistent, they are mapped to a unified dimension through a linear projection layer. S302, Multimodal Alignment Network Design: Video features are input to the Transformer layer to capture spatiotemporal correlations; Audio features enhance temporal coherence through a self-attention mechanism; text features fuse visual and auditory semantics through a cross-attention layer; an attention weight matrix is calculated for video and text features, and a weighted sum is used to generate a joint representation; the InfoNCE loss is used to maximize the cosine similarity of video-audio, video-text, and audio-text pairs, and a domain discriminator is introduced to distinguish modal sources and force feature distribution alignment; initially, only high-confidence modal pairs are aligned, and low-confidence modal pairs are gradually added; S303, alignment result verification and screening: Calculate the cosine similarity matrix of video-audio, video-text, and audio-text, screen feature pairs with similarity below the threshold, and trigger the realignment process; S304, abnormal feature removal: use the isolation forest algorithm to detect outlier feature points and remove noise interference features.

8. The multimodal sound generation method according to claim 1, wherein: In said S4, the specific steps are as follows: S401, splicing structure design: The aligned video, audio, and text features are all 512-dimensional, and are spliced into a 1536-dimensional vector along the channel dimension; a learnable modal identification vector is inserted to enhance the model's perception of the modal source; S402, position coding injection: add temporal position coding to video features; use BERT-style position coding for text features; apply frequency domain position coding to audio features; S403, feature dimensionality reduction and fusion: Use linear layer for feature dimensionality reduction, and the activation function is GELU; Introduce residual connections to retain original feature information; S404, multi-scale feature fusion: Dynamically adjust the weights of each modal feature through the SENet module, and select the dominant modal feature based on the gated recurrent unit; S405, post-splicing feature verification: input the spliced features to the discriminator network, predict the input modality combination type, and when the accuracy is lower than the preset value, trigger the feature return realignment process; S406, Noise robustness enhancement: Add Gaussian noise and random mask to train the model's anti-interference ability.

Citation Information

Cited By

  • Digital human speech synthesis method and system based on multi-modal speech feature fusion

    CN120833777A

  • Digital human speech synthesis method and system based on multi-modal speech feature fusion

    CN120833777B

  • Large language model training method and system for multi-modal content output

    CN122065956A