Multi-parameter driven aigc personalized song generation method

By using a multi-parameter-driven AIGC song generation method, combined with various interaction methods and a deep generation model, the problems of insufficient emotional expression and scene adaptation in existing systems are solved, achieving high-quality, personalized song generation and improving generation efficiency and user experience.

CN121543739BActive Publication Date: 2026-04-21XIANGJIANG LAB
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
XIANGJIANG LAB
Filing Date
2026-01-15
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing AIGC song generation systems are unable to deeply explore users' emotional needs, resulting in generated works that lack emotional tension and personalized expression. They also suffer from limited interaction methods, weak scene adaptability, poor generation efficiency and sound quality, and a lack of exception handling mechanisms.

Method used

It adopts a multi-parameter driven approach, supports interactive methods such as text, voice description, and reference audio upload, generates emotion tags by combining a voice representation model and a large language model, creates lyrics and melodies through Transformer and diffusion model, dynamically adjusts melody complexity and harmony to achieve multi-scene adaptation, and introduces a neural vocoder for high-fidelity synthesis, equipped with user feedback iteration and anomaly handling mechanisms.

Benefits of technology

It achieves high accuracy in emotional matching, good coordination of melody and harmony, a balance between generation speed and sound quality, strong scene adaptability, and optimized user experience, ensuring the originality and stability of the generated works.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121543739B_ABST
    Figure CN121543739B_ABST
Patent Text Reader

Abstract

This invention discloses a multi-parameter driven AIGC personalized song generation method, relating to the fields of intelligent audio generation and emotional state adaptation. It supports multiple interaction methods, collecting personalized parameters and emotional reference information such as music style, lyrics, theme, and singing style. A speech representation model generates emotional tags, and a large language model expands the adjective set and constructs an emotional-theme semantic network. Multiple input types are transformed into multimodal feature vectors, which are then processed through semantic modeling and normalization to generate unified input data. Lyrics, melody, and harmony are collaboratively generated based on a combined architecture. Melody, rhythm, and timbre features are optimized as needed. A neural vocoder synthesizes the vocals and performs multi-track mixing to output an original song. This invention achieves accurate emotional adaptation, with lyrics, melody, harmony, and emotional theme highly consistent. It exhibits excellent multimodal feature fusion, generates high-quality, stylistically unified songs with strong scene adaptability, supports rapid iterative optimization, balances generation efficiency and sound quality, and enhances the personalized creation experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent audio generation and emotional state adaptation technology, and in particular to a multi-parameter driven AIGC personalized song generation method. Background Technology

[0002] With the rapid penetration of AIGC (Artificial Intelligence Generated Content) technology into the content creation field, personalized song generation has become a research hotspot in audio processing. However, existing technologies still have many key shortcomings. Most current AIGC song generation systems focus only on simple matching of basic parameters such as musical style and tempo, failing to deeply explore the intrinsic connection between user emotional needs and song creation. The generated works often lack emotional tension and personalized expression. In different scenarios such as campus activities, advertising soundtracks, and game audio, users not only need music of specific styles but also hope that songs can accurately convey core emotions such as joy, excitement, and relaxation. However, existing systems cannot translate emotional inclinations into specific creative parameters, resulting in a significant deviation between the generated songs and user emotional expectations. At the same time, user input methods are limited, mostly confined to text commands, lacking diversified interactive channels such as voice descriptions and uploading reference audio, making it difficult to fully capture users' implicit needs.

[0003] The imperfection of the emotional feature extraction and fusion mechanism is a core bottleneck restricting the quality of generated content. In existing technologies, speech emotion recognition and song creation are disconnected. The emotional tags generated by the speech representation model lack detailed processing and cannot be transformed into specific guidelines for lyric writing and melody arrangement. Large language models, during lyric generation, do not fully integrate emotional tags for semantic expansion and optimization, resulting in a disconnect between lyric content and emotional tone, and sentence structure and vocabulary selection failing to meet the needs of emotional expression. Furthermore, the multimodal feature fusion effect is poor; natural language descriptions, structured parameters, and emotional features are not effectively mapped to a unified feature space, making it difficult for deep generation models to accurately understand the user's comprehensive needs. This, in turn, affects the collaborative creation effect of lyrics, melody, and harmony, often resulting in monotonous melodies, disharmonious harmonies, and mismatches between singing style and emotion.

[0004] Existing technologies exhibit significant shortcomings in terms of generation efficiency, scene adaptation, and quality assurance. Most systems employ a single model architecture, failing to achieve a dynamic balance between generation speed and sound quality. Generating simple songs takes excessive time, while generating complex songs results in poor sound quality. Scene adaptation capabilities are weak, lacking dedicated parameter configurations for different scenarios such as campus events, advertising soundtracks, and game audio. Generated songs often fail to meet the specific needs of these scenarios, such as the strong memorability required for advertising soundtracks or the energetic rhythm needed for game battle scenes. User interaction and feedback mechanisms are lacking; once generated results are output, they are difficult to adjust and cannot be iteratively optimized based on user preferences, leading to high costs for repeated user operations. Furthermore, copyright risk detection and anomaly handling mechanisms are inadequate, potentially resulting in generated works with high similarity to existing works. There is also a lack of effective solutions for issues such as parameter conflicts and model lag, impacting the system's stability and usability. Summary of the Invention

[0005] This invention proposes a multi-parameter driven AIGC personalized song generation method to solve the problems mentioned in the prior art.

[0006] To achieve the above objectives, the present invention adopts the following technical solution:

[0007] A multi-parameter-driven AIGC personalized song generation method includes the following steps:

[0008] The user needs collection process supports three interaction methods: text, input voice description, and reference audio upload. It collects personalized needs parameters specified by the user and synchronously receives emotional reference information from the reference audio file or natural language description uploaded by the user.

[0009] In the emotional feature extraction step, the speech representation model uses a pre-trained audio emotion recognition architecture to analyze audio files and generate emotion tags; the large language model expands and generates an adjective set based on the emotion tags; based on the emotion tags and the generated adjective set, an emotion-theme association semantic network is constructed in conjunction with the lyrics theme to clarify the core emotional tone and expression direction of the song.

[0010] The feature parsing and parameter encoding steps transform the user's input natural language description and structured parameters into multimodal feature vectors. Semantic modeling and style labeling are performed on the emotion tag, style type, rhythm parameter, and vocal range. The vectors are mapped to a high-dimensional feature space through an embedding layer, and feature normalization is used to generate model input data in a unified format.

[0011] The deep generative model-driven process, based on an architecture combining Transformer and diffusion models, first uses an encoder to perform hierarchical feature fusion and semantic understanding of multimodal feature vectors, and then uses a decoder to automatically generate lyrics, simultaneously completing melody creation and harmony arrangement, generating single-track melody audio and harmony track audio and lyrics text that meet the requirements of emotional tone and style.

[0012] The generation parameter dynamic adjustment step adopts a multi-task joint training strategy and a controllable generation mechanism. It adjusts the melody complexity, rhythm density and timbre characteristics in real time according to user needs, and dynamically optimizes the melody fluctuation amplitude and harmony richness by combining the intensity of emotional tags, so that the generated content is highly consistent with the user's emotional needs.

[0013] The multi-track audio synthesis process involves high-fidelity vocal synthesis using a neural vocoder, mixing the generated vocal track with the accompaniment track and harmonic track effects audio, and using an adaptive mixing algorithm to adjust the volume balance, frequency distribution, and spatial effects of each track to output a high-quality original song audio file.

[0014] Furthermore, it also includes a dynamic adjustment step for melody complexity. During the deep generative model-driven process, it combines the intensity of sentiment tags with the user's demand for melody complexity, using a formula... Calculate the optimal melody complexity coefficient, where The melody complexity coefficient is set to a value between 0.1 and 1.0. The emotional complexity coefficient is determined by the intensity of the emotion; the closer the value is to 1.0. The emotional complexity weight is set to a value between 0.6 and 0.8. The base value for rhythm complexity is set based on the rhythm speed. The rhythmic complexity weight is set to 0.2-0.4. Based on the calculation results, the number of melodic notes, interval span, and proportion of ornaments are adjusted. When the emotion is strong, the interval span and the number of ornaments are increased to enhance expressiveness. When the rhythm is slow, the melodic structure is simplified to enhance smoothness.

[0015] Furthermore, the process includes lyrics emotional adaptation and polishing steps. After the large language model generates the initial draft of the lyrics, the lyrics are semantically optimized by combining emotional tags and adjective sets, using formulas... Calculate the emotional fit of the lyrics, among which The emotional fit of the lyrics is set to a value between 0 and 1. The semantic similarity between the lyrics text and the emotion tags. The semantic similarity weight is set to a value between 0.7 and 0.8. The fit coefficient between lyric structure and emotional expression is calculated by assigning a higher percentage of exclamatory sentences, with the value approaching 1. The sentence structure is weighted between 0.2 and 0.3. When the fit is lower than the preset threshold, words with higher emotional relevance are automatically replaced to adjust the sentence structure, optimize the rhythm and rhyme, and strengthen the consistency between the lyrics and the core emotional tone.

[0016] Furthermore, it also includes a step for precise adaptation of harmonic style. During the harmonic arrangement process, based on the user-specified music style and emotional tags, a pre-built style-harmonic mapping database is invoked to match the corresponding style's harmonic paradigm chord type and modulation rules. At the same time, the chord connection method is dynamically adjusted according to the melody direction and emotional changes. Ascending harmonics are used when the emotion progresses, smooth harmonic connections are used when the emotion is relaxed, and transitional chords are added when the emotion changes, so that the harmony and melody emotion form an organic whole.

[0017] Furthermore, it includes steps for customizing singing style and timbre. It collects user-specified singing style reference features, including parameters such as articulation clarity, breath control intensity, vibrato frequency, transitions, and tail note processing. Combined with emotional tags, it adjusts timbre features: a bright and clear timbre corresponds to a cheerful emotion, a deep and resonant timbre corresponds to a sad emotion, and a high-pitched and powerful timbre corresponds to an excited emotion. It simulates the vocalization methods of different singing techniques through a speech synthesis model, adding details such as breath sounds, glissando, and vibrato. It also supports users uploading their own voice samples, extracting core timbre features through a timbre extraction algorithm and embedding them into the generation process to achieve a personalized singing effect with the user's individual timbre characteristics.

[0018] Furthermore, it includes a dynamic rhythm and tempo optimization step, receiving user-defined basic rhythm and tempo parameters, and adjusting rhythm density and beat intensity in conjunction with emotional tags and music style. A rhythm template library is constructed by analyzing the rhythm patterns of classic songs of similar styles; upbeat emotions correspond to higher rhythm density to emphasize strong beats, soothing emotions reduce rhythm density to weaken accents and highlight melodic smoothness, and exciting emotions use syncopated rhythms to enhance the sense of rhythm; at the same time, a rhythm gradient effect is set at the transitions between song sections, gradually increasing rhythm density and tempo intensity from verse to chorus, and appropriately adjusting the rhythm pattern from chorus to bridge to enhance the song's layering and expressiveness.

[0019] Furthermore, it also includes multi-scene adaptation and adjustment steps, preset parameter configuration templates for multiple application scenarios, each template corresponding to specific emotional tendencies, style characteristics, rhythm, speed, audio mixing standards, and duration limits; users can directly select scene templates, and the system will automatically match the corresponding generation parameters, while also supporting users to fine-tune personalized needs based on the templates.

[0020] Furthermore, it also includes steps to optimize the balance between generation speed and sound quality. It employs model quantization and inference acceleration technology to improve generation efficiency while maintaining sound quality. The balance is achieved by dynamically adjusting the model inference accuracy. Low-precision inference is used to speed up the generation of simple songs, while high-precision inference is used to ensure sound quality for complex songs. A multi-level sound quality evaluation standard is established, which monitors the sound quality in real time from five dimensions: frequency response, signal-to-noise ratio, vocal clarity, acoustic coordination, and accompaniment blending. If the sound quality is found to be substandard during the generation process, the model parameters are automatically adjusted to re-optimize the generation results of the corresponding track.

[0021] Furthermore, it also includes user interaction feedback iteration steps. After generating the initial song audio, it provides interactive modification entry points such as lyrics modification, melody adjustment, rhythm speed change, timbre change, and harmony optimization, and receives user feedback and adjustment instructions. The user feedback is converted into quantitative parameters, the semantic modeling results of the feature analysis and parameter encoding module are updated, and the data is re-input into the deep generation model for secondary generation. At the same time, it records the user's modification preferences and aesthetic inclinations, establishes a personalized user profile, stores user preference parameters and generation history data, and automatically adapts to user preferences in subsequent generation processes, reducing repeated modification operations and gradually improving the fit between the generated results and user needs.

[0022] Furthermore, it includes anomaly handling and quality assurance steps. During the user input stage, it checks the completeness and rationality of parameters, automatically pushes default configuration suggestions when core parameters are missing, and prompts users to adjust contradictory items when parameters conflict. During the feature parsing stage, it resolves ambiguity in vague semantic descriptions and determines the optimal interpretation through contextual analysis. During the generation process, it monitors the model's running status and automatically switches to a backup model architecture to continue generation when generation stalls or abnormal results occur. After generation, it performs copyright risk detection by combining melody feature hash comparison with lyric semantic similarity analysis to check the similarity between the lyrics melody and existing works. When the similarity exceeds a threshold, it automatically adjusts the core melody towards the semantic expression of the lyrics and the harmonic arrangement.

[0023] Compared with existing technologies, the beneficial effects of this invention are:

[0024] This invention presents a multi-parameter-driven AIGC personalized song generation method, with emotional state adaptation at its core. It optimizes the entire process from demand collection to song output, bringing significant advantages across multiple dimensions. Regarding the accuracy of emotional adaptation, it generates emotional tags through a speech representation model, combined with the expansion of the adjective set from a large language model and the construction of an emotion-theme related semantic network, ensuring that song creation revolves around the core emotional tone throughout. The lyrics are refined through emotional adaptation, with sentence structure and vocabulary precisely matching the emotional expression; the complexity and fluctuation of the melody are dynamically adjusted according to the intensity of the emotion; and the harmonic arrangement is deeply linked to emotional changes, using ascending harmonics for emotional progression and adding transitional chords for emotional transitions, making the emotional expression of the entire song coherent and full of tension.

[0025] Multi-parameter fusion and generation quality are both significantly improved. User needs collection supports diverse methods such as text input, voice description, and reference audio uploads, comprehensively capturing various parameters including music style, tempo, and emotional tone. The feature parsing and parameter encoding stages utilize semantic modeling, style labeling, and feature normalization to transform multimodal data into a unified model input format, ensuring the semantic understanding and feature fusion effects of the deep generation model. Based on a combined architecture of Transformer and diffusion models, coupled with a multi-task joint training strategy, lyrics, melody, and harmony are collaboratively generated. After processing by a neural vocoder and adaptive mixing algorithm, the generated work boasts high-fidelity sound quality, balanced and harmonious tracks, smooth and natural melodies, and rich and full harmonies.

[0026] The system boasts significantly optimized scene adaptability and user experience, offering pre-set parameter configuration templates for multiple application scenarios, covering various usage needs such as campus activities, advertising soundtracks, and game audio. Users can directly select templates or make personalized adjustments, obtaining works that meet specific scenario requirements without complex operations. A dynamic balance mechanism between generation speed and sound quality adjusts the model's inference accuracy based on the complexity of the song's style, balancing efficiency and quality. A user interaction feedback iteration mechanism supports multiple operations such as lyrics modification and melody adjustment, converting user preferences into quantifiable parameters to build personalized user profiles. Subsequent generation automatically adapts to preferences, reducing redundant modifications. Anomaly handling and copyright risk detection mechanisms ensure stable system operation and the originality of works, comprehensively meeting users' core needs in personalized creation and commercial applications, and driving the development of AIGC song generation technology towards precision, personalization, and practicality. Attached Figure Description

[0027] Figure 1 This is a schematic block diagram of a multi-parameter driven AIGC personalized song generation method proposed in this invention;

[0028] Figure 2 A comparison chart showing the degree of relevance between emotions and themes in different scenarios;

[0029] Figure 3 Generate speed versus sound quality balance comparison charts for songs with different levels of complexity;

[0030] Figure 4 A comparison chart of the emotional fit of lyrics under different emotional types;

[0031] Figure 5 A comparison chart of multimodal feature fusion effects;

[0032] Figure 6 A comparison chart of iteration efficiency based on user feedback. Detailed Implementation

[0033] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0034] In the description of this invention, it should be understood that the terms "center," "longitudinal," "lateral," "length," "width," "thickness," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," "outer," "clockwise," and "counterclockwise," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are only for the convenience of describing this invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on this invention.

[0035] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of the stated features. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified. Furthermore, the terms "installed," "connected," and "linked" should be interpreted broadly; for example, they may refer to a fixed connection, a detachable connection, or an integral connection; they may refer to a mechanical connection or an electrical connection; they may refer to a direct connection or an indirect connection through an intermediate medium; and they may refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances. The invention will now be described in further detail with reference to the accompanying drawings.

[0036] Reference Figures 1 to 6A multi-parameter-driven AIGC personalized song generation method includes the following steps:

[0037] The user requirement collection process supports three interactive methods: text input, voice description, and reference audio upload. It collects personalized requirement parameters specified by the user, such as music style, lyric theme, singing style, rhythm speed, vocal range, and emotional tendency, and simultaneously receives reference audio files or emotional reference information described in natural language uploaded by the user.

[0038] In the emotional feature extraction step, the speech representation model uses a pre-trained audio emotion recognition architecture to analyze audio files and generate emotion tags; the large language model expands and generates an adjective set based on the emotion tags; the large language model constructs an emotion-theme association semantic network based on the emotion tags and the generated adjective set, combined with the lyrics theme, to clarify the core emotional tone and expression direction of the song.

[0039] The feature parsing and parameter encoding steps transform the user's input natural language description and structured parameters into multimodal feature vectors. Semantic modeling and style labeling are performed on sentiment tags, style types, rhythm parameters, and vocal range. The vectors are mapped to a high-dimensional feature space through an embedding layer, and feature normalization is used to generate model input data in a unified format.

[0040] The deep generative model-driven process, based on an architecture combining Transformer and diffusion models, first uses an encoder to perform hierarchical feature fusion and semantic understanding of multimodal feature vectors, and then uses a decoder to automatically generate lyrics, simultaneously completing melody creation and harmony arrangement, generating single-track melody audio and harmony track audio and lyrics text that meet the requirements of emotional tone and style.

[0041] The generation parameter dynamic adjustment step adopts a multi-task joint training strategy and a controllable generation mechanism to adjust the melody complexity, rhythm density and timbre characteristics in real time according to user needs. Combined with the intensity of emotional tags, the melody fluctuation amplitude and harmony richness are dynamically optimized so that the generated content is highly consistent with the user's emotional needs.

[0042] The multi-track audio synthesis process involves high-fidelity vocal synthesis using a neural vocoder, followed by mixing the generated vocal track with the accompaniment track, harmony track, and special effects audio. An adaptive mixing algorithm is then used to adjust the volume balance, frequency distribution, and spatial sound effects of each track, resulting in a high-quality original song audio file.

[0043] This invention also includes a dynamic adjustment step for melody complexity. During the deep generative model-driven process, the intensity of the sentiment label and the user's demand for melody complexity are combined using a formula.

[0044] Calculate the optimal melody complexity coefficient, where The melody complexity coefficient is set to a value between 0.1 and 1.0. The emotional complexity coefficient is determined by the intensity of the emotion; the closer the value is to 1.0. The emotional complexity weight is set to a value between 0.6 and 0.8. The base value for rhythm complexity is set based on the rhythm speed. The rhythmic complexity weight is set to 0.2-0.4. Based on the calculation results, the number of melodic notes, interval span, and proportion of ornaments are adjusted. When the emotion is strong, the interval span and the number of ornaments are increased to enhance expressiveness. When the rhythm is slow, the melodic structure is simplified to enhance smoothness.

[0045] This invention also includes a lyrics emotion adaptation and polishing step. After the large language model generates the initial lyrics draft, the lyrics are semantically optimized by combining emotion tags and adjective sets, using a formula.

[0046] Calculate the emotional fit of the lyrics, among which The emotional fit of the lyrics is set to a value between 0 and 1. The semantic similarity between the lyrics text and the emotion tags. The semantic similarity weight is set to a value between 0.7 and 0.8. The fit coefficient between lyric structure and emotional expression is calculated by assigning a higher percentage of exclamatory sentences, with the value approaching 1. The sentence structure matching weight is set to 0.2-0.3. When the matching degree is lower than the preset threshold, words with higher emotional relevance are automatically replaced, sentence structure is adjusted, and rhythm is optimized to strengthen the consistency between the lyrics and the core emotional tone.

[0047] This invention also includes a step for precise harmonic style adaptation. During the harmonic arrangement process, based on the user-specified music style and emotional tags, a pre-built style-harmonic mapping database is invoked to match the corresponding style's harmonic progression paradigm, chord type, and modulation rules. For classical styles, triads and seventh chords are prioritized to enhance harmonic fullness; for pop styles, suspended chords and added ninth chords are added to enhance the modern feel; and for rock styles, the use of fifth chords is strengthened to enhance the sense of power. At the same time, the chord connection method is dynamically adjusted according to the melody direction and emotional changes. Ascending harmonic progressions are used when the emotion progresses, smooth harmonic connections are used when the emotion is relaxed, and transitional chords are added when the emotion changes, so that the harmony and melody emotion form an organic whole.

[0048] This invention also includes a singing style and timbre customization step. It collects user-specified singing style reference features, including parameters such as articulation clarity, breath control intensity, vibrato frequency, transition method, and tail note processing. Combined with emotional tags, it adjusts timbre features: a bright and clear timbre corresponds to a cheerful emotion, a deep and resonant timbre corresponds to a sad emotion, and a high-pitched and powerful timbre corresponds to an excited emotion. It simulates the vocalization methods of different singing techniques through a speech synthesis model, adding details such as breath sounds, glissando, and vibrato. At the same time, it supports users to upload their own voice samples. The timbre extraction algorithm extracts the core timbre features and embeds them into the generation process to achieve a personalized singing effect with the user's personal timbre characteristics.

[0049] This invention also includes a dynamic rhythm and speed optimization step, which receives basic rhythm and speed parameters set by the user, and adjusts the rhythm density and beat intensity in combination with emotional tags and music styles. A rhythm template library is constructed by analyzing the rhythm patterns of classic songs of the same style. Upbeat emotions correspond to higher rhythm density and emphasize strong beats, while soothing emotions reduce rhythm density, weaken accents, and highlight the smoothness of the melody. Exhilarating emotions use syncopated rhythms to enhance the sense of rhythm. At the same time, a rhythm gradient effect is set at the transition between song sections, gradually increasing the rhythm density and speed intensity from the verse to the chorus, and appropriately adjusting the rhythm pattern from the chorus to the bridge to enhance the song's layering and expressiveness.

[0050] This invention also includes multi-scene adaptation and adjustment steps, with preset parameter configuration templates for multiple application scenarios such as campus activities, personalized creations, advertising music, game audio, wedding celebrations, and film and television soundtracks. Each template corresponds to specific emotional tendencies, style characteristics, rhythm and speed audio mixing standards, and time limits. The campus activity scenario emphasizes concise and lively melodies, positive lyrics, and a lively rhythm. The advertising music scenario highlights memorable melodies and controls the song length to 15-60 seconds to enhance brand compatibility. The game audio scenario adjusts the emotional atmosphere according to the game type, with battle scenarios using exciting rhythms and strong harmonies, and leisure scenarios using soothing melodies and gentle accompaniment. Users can directly select a scene template, and the system automatically matches the corresponding generation parameters. At the same time, users can fine-tune their personalized needs based on the template.

[0051] This invention also includes a step to optimize the balance between generation speed and sound quality. It employs model quantization and inference acceleration technology to improve generation efficiency while maintaining sound quality. The balance is achieved by dynamically adjusting the model inference accuracy. Low-precision inference is used to speed up the generation of simple songs, while high-precision inference is used to ensure sound quality for complex songs. A multi-level sound quality evaluation standard is established, which monitors the sound quality in real time from five dimensions: frequency response, signal-to-noise ratio, vocal clarity, acoustic coordination, and accompaniment blending. If the sound quality is found to be substandard during the generation process, the model parameters are automatically adjusted to re-optimize the generation results of the corresponding track, so that the output song meets both the generation efficiency requirements and has a high-fidelity sound quality effect.

[0052] This invention also includes a user interaction feedback iteration step. After generating the initial song audio, it provides interactive modification entry points such as lyrics modification, melody adjustment, rhythm speed change, timbre replacement, and harmony optimization, and receives user feedback and adjustment instructions. The user feedback is converted into quantitative parameters, and the semantic modeling results of the feature analysis and parameter encoding module are updated and re-input into the deep generation model for secondary generation. At the same time, the user's modification preferences and aesthetic inclinations are recorded to establish a personalized user profile. User preference parameters and generation history data are stored, and the system automatically adapts to user preferences in subsequent generation processes, reducing repeated modification operations and gradually improving the fit between the generated results and user needs.

[0053] This invention also includes anomaly handling and quality assurance steps. During the user input phase, the completeness and rationality of parameters are checked; if core parameters are missing, default configuration suggestions are automatically pushed; and if parameters conflict, the user is prompted to adjust contradictory items. During the feature parsing phase, ambiguous semantic descriptions are resolved by resolving polysemy and determining the optimal interpretation through contextual analysis. During the generation process, the model's running status is monitored; if generation stalls or abnormal results occur, a backup model architecture is automatically switched to continue generation. After generation, copyright risk detection is performed, using a combination of melody feature hash comparison and lyric semantic similarity analysis to check the similarity between the lyrics melody and existing works. If the similarity exceeds a threshold, the core melody is automatically adjusted to better express the semantics of the lyrics and harmonic arrangement, ensuring the originality and legality of the output song.

[0054] The present invention will be further illustrated below through two embodiments:

[0055] Example 1

[0056] Application of generating theme songs for school anniversary celebrations in campus event scenarios

[0057] This example is applied to the customization of theme songs for university anniversary galas. The user is the person in charge of student union activities and needs to generate a youthful, uplifting, and positive pop-rock style song suitable for choral performances, while incorporating campus elements and emotional resonance. The system executes the entire process according to a multi-parameter driven flow, as follows:

[0058] I. Execution of Core Processes and Key Steps

[0059] User Request Collection: Users submit their requests through two methods supported by the system: text input and upload of reference audio. Text input specifies the music style as pop-rock, lyrics themes revolving around school anniversary, youth, unity, and pursuing dreams, a singing style suitable for choral performance, a tempo of 120 BPM, and a vocal range covering both male and female mid-range voices. Users also upload three classic campus choral songs as reference audio, supplementing with a natural language description expressing a desire for songs with a strong, memorable chorus, a soothing and narrative verse, and a powerful and energetic chorus. The system simultaneously collects potential user requests, providing interactive prompts to supplement relevant information such as the number of chorus members being 30 and the performance venue being an indoor gymnasium.

[0060] Emotional Feature Extraction: The speech representation model employs a pre-trained audio emotion recognition architecture, performing spectral analysis, rhythmic pattern extraction, and emotional feature capture on reference audio to generate three core emotional tags: joy, excitement, and passion. Based on these emotional tags, the large language model expands to generate a set of adjectives such as bright, uplifting, steadfast, and surging. Combining the lyric theme and emotional tags, the large language model constructs an emotion-theme-related semantic network, using youth as the core node and linking it to sub-themes such as campus life, friendships, and pursuing dreams. Each sub-theme corresponds to 2-3 emotional adjectives, clarifying the expressive direction of the song's narrative and lyrical verses and the emotional outburst in the chorus.

[0061] Feature parsing and parameter encoding: The system transforms the user-input natural language description and structured parameters into multimodal feature vectors. The music style vector includes rhythmic patterns and chord progressions characteristic of pop and rock music; the sentiment vector integrates the intensity values ​​of three core sentiment tags; and the rhythm parameter vector corresponds to a 120 BPM beat. Semantic modeling and style labeling are performed on sentiment tags, style types, rhythm parameters, and vocal range. These are mapped to a high-dimensional feature space through an embedding layer. Min-Max normalization is used to unify the feature values ​​of each dimension to the 0-1 range, generating 384-dimensional unified format model input data.

[0062] Deep Generative Model-Driven and Parameter Tuning: Based on an architecture combining Transformer and diffusion models, the encoder performs hierarchical feature fusion on multimodal feature vectors. First, it captures the correlation features between style and emotion through a self-attention mechanism, and then fuses thematic semantic information through a cross-attention mechanism. Based on the fused features, the decoder first generates the lyrics text. The verses revolve around the confusion of entering school and the time spent with classmates, while the chorus focuses on the joy of the school anniversary celebration and the determination to pursue dreams together. Simultaneously, the melody is composed. The verses use a gentle pentatonic scale, while the chorus adds an ascending leap to enhance tension. The harmonic arrangement calls on a style-harmonic mapping database, with pop-rock style emphasizing the use of fifth chords, and ascending harmonics used for emotional progression.

[0063] In the dynamic parameter adjustment phase, the intensity of the emotional label and the user's demand for melody complexity are combined using formulas. Calculate the optimal melody complexity coefficient. Among these, the emotional complexity coefficient... =0.9, the value approaches 1.0 when the emotion is strong. =0.7, the base value of pacing complexity =0.6, set according to a tempo of 120 BPM. =0.3, substituting into the calculation yields =0.9×0.7+0.6×0.3=0.63+0.18=0.81. Based on this result, the number of notes in the chorus is increased by 30%, the interval span is expanded to a sixth, and ornaments and glissandos are added to enhance expressiveness; the verses have a simplified melodic structure, a reduced proportion of ornaments, and enhanced fluency.

[0064] Customized vocal style and multi-track audio synthesis: Reference vocal characteristics of pop-rock style are collected, with medium articulation clarity, a vibrato frequency of 5Hz, and natural glissando transitions. Combined with joyful and passionate emotional tags, the timbre is adjusted to be bright and clear. A voice synthesis model simulates choral vocal production, adding group harmony effects. Soft breathy tones are incorporated into the verses, while the power of the voice is enhanced in the chorus. In the multi-track audio synthesis stage, a neural vocoder generates a high-fidelity choral vocal track, which is mixed with an accompaniment track (guitar, bass, drums), a harmony track, and applause effects audio. An adaptive mixing algorithm is used to adjust the volume ratio of the vocal track to 40%, the accompaniment track to 50%, and the effects track to 10%. Frequency allocation is optimized to avoid frequency band conflicts, spatial sound effects are added to enhance the sense of presence, and lossless audio files are output.

[0065] User feedback iteration and quality assurance: After generating the initial song, users suggested adjustments such as making the chorus lyrics more prominently feature the anniversary theme and slightly speeding up the verse tempo. The system converted this feedback into quantitative parameters, increased the semantic weight of the anniversary theme, adjusted the verse tempo to 125 BPM, updated the semantic modeling results of the feature analysis and parameter encoding modules, and re-input the data into the deep generation model for secondary generation. Simultaneously, user preferences were recorded to create a personalized profile of the campus chorus's energetic style. In the anomaly handling phase, the system automatically pushed a default suitable vocal range suggestion when the user's initial input vocal range was too narrow and conflicted with choral fit. After generation, copyright risk detection showed that the similarity between the lyrics and melody was 12%, below the 30% threshold, ensuring originality.

[0066] II. Data Representation and Interpretation

[0067] Table 1 Comparison of Song Generation Performance for Campus Activity Scenarios

[0068]

[0069] Table 1 clearly demonstrates the advantages of this invention in campus event scenarios. Traditional generation methods lack deep emotional integration and scene-specific configuration, with an emotional-theme fit of only 65% ​​and melody-harmonic coordination of less than 70%. This fails to accurately adapt to the needs of choral performances, resulting in inefficient iterations requiring 25 minutes per user feedback iteration, and originality is difficult to guarantee. This invention achieves 95% emotional-theme fit and 93% melody-harmonic coordination through emotional-theme semantic network construction and precise style-harmony matching. The combination of scene-adaptive templates and personalized adjustments precisely meets the needs of choral performances at school anniversary galas, improving user feedback iteration efficiency to 5 minutes per iteration. 96% originality fully complies with campus event copyright requirements, delivering high-quality theme songs from all angles.

[0070] Example 2

[0071] Application of product promotion music generation in advertising soundtrack scenarios

[0072] This example is applied to the customization of background music for new product promotion advertisements of fast-moving consumer goods brands. The user is a brand marketer who needs to generate a catchy, upbeat, and memorable pop song, 30 seconds long, highlighting the product's core selling points of convenience and style. The system executes the process according to a multi-parameter driven flow, as follows:

[0073] I. Execution of Core Processes and Key Steps

[0074] User Request Collection: Users submit their requests via text input and voice description. Text input specifies the music style as upbeat and pop, lyrics focusing on product convenience, fashion, and youthfulness, a sweet and lively vocal style, a tempo of 130 BPM, and a vocal range targeting the female mid-range. Voice descriptions add a memorable melody at the beginning of the song, repeating the core selling point in the chorus, with the overall duration kept under 30 seconds. If no reference audio is uploaded, the system automatically recommends three similar advertising soundtracks for users to confirm their emotional inclination.

[0075] Emotional Feature Extraction: The speech representation model analyzes the reference advertising background music confirmed by users, generating three core emotional tags: upbeat, cheerful, and lively. Based on these emotional tags, the large language model expands to generate a set of adjectives: bright, simple, vivid, and playful. Combining the lyric theme and emotional tags, the large language model constructs an emotion-theme-related semantic network. With product convenience as the core node, it connects to sub-themes such as daily use, efficiency, and fashionable styling. Each sub-theme corresponds to two emotional adjectives, clearly defining the song's overall upbeat and smooth flow, and the repeated reinforcement of core selling points.

[0076] Feature parsing and parameter encoding: The system transforms user-input natural language descriptions, structured parameters, and sentiment tendencies into multimodal feature vectors. The music style vector includes upbeat and popular rhythms and instrumentation features. The sentiment vector integrates the intensity values ​​of three core sentiment tags. The rhythm parameter vector corresponds to a 130 BPM beat characteristic and a 30-second duration limit. Semantic modeling and style labeling are performed on sentiment tags, style types, rhythm parameters, and vocal ranges. These are mapped to a high-dimensional feature space through an embedding layer, and Z-score normalization is used to generate 256-dimensional unified format model input data.

[0077] Deep generative model-driven lyrics optimization: Based on an architecture combining Transformer and diffusion models, the encoder performs hierarchical feature fusion on multimodal feature vectors, focusing on capturing short-duration, memorable scene-specific features; the decoder first generates a draft of the lyrics, with the verse concisely describing the product's usage scenario, and the chorus repeating convenient and stylish core selling points, all within a 30-second timeframe. In the emotional adaptation and polishing stage, semantic optimization is performed by combining emotional tags and adjective sets, using formulas... Calculate the emotional fit of the lyrics. The semantic similarity between the lyrics text and the sentiment tag is 0.85. =0.75, The fit coefficient between lyric structure and emotional expression is 0.9, with a high proportion of exclamatory sentences. =0.25, substituting into the calculation yields =0.85×0.75+0.9×0.25=0.6375+0.225=0.8625, which is higher than the preset threshold of 0.8. There is no need to replace words; just adjust the sentence structure and rhythm to better match the upbeat melody.

[0078] In the melody and harmony generation stage, a multi-task joint training strategy is adopted to dynamically adjust the melody complexity, highlight the memorability of the first four measures, and use a combination of repetitive rhythmic patterns and leap intervals. The harmony arrangement calls up the style-harmony mapping database, adding suspended chords and added ninth chords to the upbeat pop style, using smooth harmony connections when the emotion is stable, and adding upbeat syncopated chord rhythms to the chorus to enhance the sense of rhythm.

[0079] Rhythm optimization and multi-track audio synthesis: In the dynamic optimization stage of rhythm and speed, combining the upbeat emotional label and pop style, the rhythm density was adjusted to 8 notes per measure, emphasizing strong beats and adding drum emphasis to the core selling points of the chorus. A rhythmic gradient effect was set at the transitions between song sections, with a quick cut from the intro to the verse and maintaining rhythmic consistency from the verse to the chorus to enhance memorability. In the customized singing style stage, sweet and lively singing reference characteristics were collected, with clear and crisp enunciation, gentle breath control, a vibrato frequency of 3Hz, and a light and cheerful ending. The timbre was adjusted to a bright and sweet type based on the emotional label, and a slight breathy tone was added to enhance intimacy.

[0080] In the multi-track audio synthesis stage, the neural vocoder generates a high-fidelity female vocal track, which is mixed with the accompaniment track composed of an electronic keyboard, drum kit, and bass, the harmony track, and product-specific sound effects including switch sounds and water flow sounds. An adaptive mixing algorithm is used to adjust the volume ratio of the vocal track to 45%, the accompaniment track to 50%, and the effects track to 5%, controlling the total song length to 30 seconds, and outputting a compressed audio file adapted for advertising.

[0081] Scene Adaptation and Quality Assurance: The system calls upon the advertising background music scene parameter configuration template, automatically matching the generation standards of short duration, strong memorability, and brand suitability, requiring no additional adjustments from the user. During the generation speed and sound quality balance optimization phase, due to the simple song style, low-precision inference is used to accelerate the generation speed, with a generation time of only 3 minutes. Simultaneously, sound quality monitoring shows a frequency domain response deviation of 0.3dB, a signal-to-noise ratio of 45dB, and vocal clarity of 92%, all dimensions meeting standards. During the user feedback iteration phase, users suggested that the core selling point of the chorus could be more prominent; the system increased the number of chorus repetitions, and the regenerated version met the requirements. In the anomaly handling phase, the system automatically fine-tuned the tempo to 128 BPM, as the initial user input tempo was too fast and potentially conflicted with the clarity of the lyrics; copyright risk detection showed a 10% similarity between the lyrics and melody, not exceeding the 30% threshold, ensuring originality and legality.

[0082] II. Data Representation and Interpretation

[0083] Table 2 Comparison of Song Generation Performance for Advertising Background Music Scenes

[0084]

[0085] Table 2 data highlights the advantages of this invention in advertising soundtrack scenarios. Traditional generation methods cannot simultaneously meet the commercial demands of short duration and strong memorability. Memorability is only 60%, emotional and product selling point alignment is 68%, duration control accuracy is 55%, generation time is 15 minutes, and the signal-to-noise ratio is only 38dB, making it difficult to balance speed and sound quality, resulting in insufficient commercial applicability. This invention, through scene template adaptation and melody memorability enhancement, achieves a memorability of 94%, quickly conveying the core selling points of the product; deep emotional-thematic integration achieves 92% alignment; a perfect balance is achieved between 3-minute generation time and a 45dB signal-to-noise ratio; 98% duration adaptation accuracy and 95% commercial application adaptability fully meet the commercial needs of advertising.

[0086] Figure 2 This chart visually demonstrates the core advantages of this invention in terms of emotion and theme adaptation. Traditional generation methods lack in-depth extraction of emotional features and construction of semantic networks, simply matching style parameters, resulting in an emotion-theme fit rate generally below 70% across various scenarios, making it difficult to meet the emotional expression needs of specific scenarios. This invention generates accurate emotional tags through a speech representation model and constructs an emotion-theme related semantic network using a large language model, ensuring that lyrics, melody, and harmony revolve around the core emotion and theme throughout. The fit rate in various scenarios exceeds 90%, reaching 95% in campus activity scenarios, perfectly achieving a deep integration of emotion and theme, significantly enhancing the song's appeal and scene adaptability.

[0087] Figure 3 This figure highlights the technological breakthrough of this invention in balancing generation speed and sound quality. Traditional generation methods use a single model architecture without dynamic adjustment of inference accuracy. This leads to a significant increase in generation time and a continuous decline in sound quality as song complexity increases. Extremely complex songs can take up to 40 minutes with a signal-to-noise ratio of only 32dB, failing to balance efficiency and quality. This invention employs model quantization and inference acceleration technology, dynamically adjusting inference accuracy based on complexity. Simple songs are generated quickly, while complex songs maintain sound quality. Extremely complex songs can be generated in just 12 minutes with a signal-to-noise ratio of 41dB. It achieves an optimal balance between speed and sound quality at all levels of complexity, meeting the needs of efficient creation in various scenarios.

[0088] Figure 4 This chart verifies the effectiveness of the lyric emotional adaptation and refinement mechanism of this invention. Traditional generation methods fail to establish a deep connection between emotions and lyrics, relying solely on keyword matching. This results in an adaptation rate of less than 60% for complex emotions such as sadness and elation, leading to a disconnect between lyric content and emotional expression. This invention accurately calculates the emotional adaptation rate of lyrics, automatically optimizing vocabulary and sentence structure when it falls below a threshold. The adaptation rate for each emotional type exceeds 90%, with elation reaching 95%. The choice of vocabulary, sentence structure, and emotional expression in the lyrics are highly consistent. Exclamatory sentences, soothing sentences, and other sentence structures are precisely matched to corresponding emotions, making the lyrics more emotionally powerful and expressive.

[0089] Figure 5 This chart demonstrates the advantages of the multimodal feature fusion technology of this invention. Traditional generation methods have imperfect feature parsing and encoding mechanisms, failing to effectively map multiple types of features to a unified space. This results in a quality score that fails to exceed 80 points even when multi-dimensional features are input, and the synergistic effect between features is not fully realized. This invention, through a feature parsing and parameter encoding module, performs semantic modeling, labeling, and normalization on multiple types of features such as text, audio, emotion, and scene, generating high-dimensional feature vectors in a unified format. This fully explores the correlation information between features. When inputting full-dimensional features of text, audio, emotion, and scene, the quality score reaches 96 points, achieving a qualitative improvement in the consistency of song style, emotional coherence, and scene adaptability.

[0090] Figure 6 This chart demonstrates the high efficiency of the user interaction feedback iteration mechanism of this invention. Traditional generation methods do not establish a quantitative model of user preferences, requiring the re-analysis of all parameters in each iteration. This leads to a continuous increase in iteration time with the number of iterations, reaching 35 minutes for a single iteration after 5 iterations, resulting in high costs for users due to repetitive operations. This invention transforms user feedback into quantitative parameters, updates the semantic modeling results, and stores user preferences to build a personalized user profile. Subsequent iterations do not require repeated parsing of basic parameters, focusing only on optimizing adjustment items. One iteration takes only 5 minutes, and 5 iterations take only 9 minutes, significantly improving iteration efficiency, reducing user waiting time, and making personalized adjustments more convenient and efficient.

[0091] The above are merely preferred embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A multi-parameter-driven AIGC personalized song generation method, characterized in that, Includes the following steps: The user needs collection process supports three interaction methods: text, input voice description, and reference audio upload. It collects personalized needs parameters specified by the user and synchronously receives emotional reference information from the reference audio file or natural language description uploaded by the user. In the emotion feature extraction step, the speech representation model adopts a pre-trained audio emotion recognition architecture to analyze audio files and generate emotion tags. The large language model expands and generates an adjective set based on sentiment tags; based on the sentiment tags and the generated adjective set, it constructs a sentiment-theme related semantic network in conjunction with the lyrics theme to clarify the core emotional tone and expression direction of the song. The feature parsing and parameter encoding steps transform the user's input natural language description and structured parameters into multimodal feature vectors. Semantic modeling and style labeling are performed on sentiment tags, style types, rhythm parameters, and vocal range. The vectors are mapped to a high-dimensional feature space through an embedding layer, and feature normalization is used to generate model input data in a unified format. The deep generative model-driven process, based on an architecture combining Transformer and diffusion models, first uses an encoder to perform hierarchical feature fusion and semantic understanding of multimodal feature vectors, and then uses a decoder to automatically generate lyrics, simultaneously completing melody creation and harmony arrangement, generating single-track melody audio and harmony track audio and lyrics text that meet the requirements of emotional tone and style. The generation parameter dynamic adjustment step adopts a multi-task joint training strategy and a controllable generation mechanism to adjust the melody complexity, rhythm density and timbre characteristics in real time according to user needs. Combined with the intensity of emotional tags, the melody fluctuation amplitude and harmony richness are dynamically optimized so that the generated content is highly consistent with the user's emotional needs. The multi-track audio synthesis process involves high-fidelity vocal synthesis using a neural vocoder, followed by mixing the generated vocal track with the accompaniment track, harmony track, and special effects audio. An adaptive mixing algorithm is then used to adjust the volume balance, frequency distribution, and spatial sound effects of each track, resulting in a high-quality original song audio file.

2. The multi-parameter driven AIGC personalized song generation method according to claim 1, characterized in that, It also includes a dynamic adjustment step for melody complexity. In the deep generative model-driven process, it combines the intensity of sentiment tags with the user's demand for melody complexity, and uses a formula to adjust the melody complexity accordingly. Calculate the optimal melody complexity coefficient, where This is the melody complexity coefficient, ranging from 0.1 to 1.

0. This is the emotional complexity coefficient; the stronger the emotion, the closer the value is to 1.

0. This represents the emotional complexity weight, with a value ranging from 0.6 to 0.

8. This is the base value for rhythm complexity, set according to the rhythm speed. The value is 0.2-0.4, which is the weight for rhythmic complexity. The number of melodic notes, interval span, and proportion of ornaments are adjusted according to the calculation results. When the emotion is strong, the interval span and the number of ornaments are increased to enhance the expressiveness. When the rhythm is slow, the melody structure is simplified to enhance the smoothness.

3. The multi-parameter driven AIGC personalized song generation method according to claim 1, characterized in that, It also includes steps for adapting and refining lyrics to emotional content. After the large language model generates the initial draft of the lyrics, it combines emotional tags and adjective sets to perform semantic optimization of the lyrics, using formulas. Calculate the emotional fit of the lyrics, among which The emotional fit of the lyrics is rated from 0 to 1. The semantic similarity between the lyrics text and the emotion tags. This represents the semantic similarity weight, with a value between 0.7 and 0.

8. This represents the fit coefficient between lyric structure and emotional expression; the higher the proportion of exclamatory sentences, the closer the value is to 1. The sentence structure is assigned a weight of 0.2-0.

3. When the fit is lower than the preset threshold, words with higher emotional relevance are automatically replaced, sentence structure is adjusted, and rhythm is optimized to strengthen the consistency between the lyrics and the core emotional tone.

4. The multi-parameter driven AIGC personalized song generation method according to claim 1, characterized in that, It also includes a step for precise harmonic style adaptation. During the harmonic arrangement process, based on the user-specified music style and emotional tags, a pre-built style-harmonic mapping database is called to match the corresponding style's harmonic paradigm chord type and modulation rules. At the same time, the chord connection method is dynamically adjusted according to the melody direction and emotional changes. Ascending harmonics are used when the emotion progresses, smooth harmonic connections are used when the emotion is relaxed, and transitional chords are added when the emotion changes, so that the harmony and melody emotion form an organic whole.

5. The multi-parameter driven AIGC personalized song generation method according to claim 1, characterized in that, It also includes steps for customizing singing style and timbre, collecting user-specified singing style reference features, including articulation clarity, breath control intensity, vibrato frequency, transitions, and tail note processing. Combined with emotional tags, it adjusts timbre features: a bright and clear timbre corresponds to a cheerful emotion, a deep and resonant timbre corresponds to a sad emotion, and a high-pitched and powerful timbre corresponds to an excited emotion. It simulates the vocalization of different singing techniques through a speech synthesis model, adding details such as breath sounds, glissando, and vibrato. It also supports users uploading their own voice samples, extracting core timbre features through a timbre extraction algorithm and embedding them into the generation process to achieve a personalized singing effect with the user's individual timbre characteristics.

6. The multi-parameter driven AIGC personalized song generation method according to claim 1, characterized in that, It also includes a dynamic tempo optimization step, which receives the user-defined basic tempo parameters and adjusts the tempo density and beat intensity based on emotional tags and music style. By analyzing the rhythm patterns of classic songs of the same style, a rhythm template library is constructed. Upbeat emotions correspond to higher tempo density and emphasize strong beats, while soothing emotions have lower tempo density and weaken accents to highlight the smoothness of the melody. Exhilarating emotions use syncopated rhythms to enhance the sense of rhythm. At the same time, a tempo gradient effect is set at the transitions between song sections. The tempo density and speed intensity are gradually increased from the verse to the chorus, and the tempo pattern is appropriately adjusted from the chorus to the bridge to enhance the song's layering and expressiveness.

7. The multi-parameter driven AIGC personalized song generation method according to claim 1, characterized in that, It also includes multi-scene adaptation and adjustment steps, and presets parameter configuration templates for multiple application scenarios. Each template corresponds to a specific emotional tendency, style characteristics, rhythm speed, audio mixing standard and duration limit. Users can directly select a scene template, and the system will automatically match the corresponding generation parameters. At the same time, it supports users to fine-tune their personalized needs based on the template.

8. The multi-parameter driven AIGC personalized song generation method according to claim 1, characterized in that, It also includes steps to optimize the balance between generation speed and sound quality. It uses model quantization and inference acceleration technology to improve generation efficiency while maintaining sound quality. The balance is achieved by dynamically adjusting the model inference accuracy. Low-precision inference is used to speed up the generation of simple songs, while high-precision inference is used to ensure sound quality for complex songs. A multi-level sound quality evaluation standard is established to monitor in real time from five dimensions: frequency response, signal-to-noise ratio, vocal clarity, harmony, and accompaniment blending. If the sound quality is found to be substandard during the generation process, the model parameters are automatically adjusted to re-optimize the generation results of the corresponding track.

9. The multi-parameter driven AIGC personalized song generation method according to claim 1, characterized in that, It also includes user interaction feedback iteration steps. After generating the initial song audio, it provides interactive modification entry points for lyrics modification, melody adjustment, rhythm speed change, timbre change, and harmony optimization, and receives user feedback and adjustment instructions. User feedback is transformed into quantitative parameters, the semantic modeling results of the feature parsing and parameter encoding modules are updated, and the data is re-input into the deep generation model for secondary generation. At the same time, user modification preferences and aesthetic inclinations are recorded to establish personalized user profiles, store user preference parameters and generation history data, and automatically adapt to user preferences in subsequent generation processes to reduce repeated modification operations and gradually improve the fit between the generation results and user needs.

10. The multi-parameter driven AIGC personalized song generation method according to claim 1, characterized in that, It also includes exception handling and quality assurance steps, which check the completeness and rationality of parameters during the user input stage, automatically push default configuration suggestions when core parameters are missing, and prompt users to adjust conflicting items when parameters conflict. During the feature parsing stage, the polysemy of the fuzzy semantic description is resolved, and the optimal interpretation is determined through context association analysis. During the generation process, the model's running status is monitored, and if generation is stuck or results are abnormal, the backup model architecture is automatically switched to continue generation. After generation, a copyright risk detection is performed. The method combines melody features, hash comparison and lyric semantic similarity analysis to check the similarity between the lyrics and melody and existing works. When the similarity exceeds the threshold, the core melody direction, lyric semantic expression and harmony arrangement are automatically adjusted.

Citation Information

Patent Citations

  • Music video generation method and system based on AIGC

    CN118645123A

  • AIGC content generation method and system based on multi-modal fusion

    CN120578796A