Method for generating children's songs based on singing synthesis large model
By collecting, organizing and training children's song resources, generating children's song style models, and using large models to generate children's songs, the problem of insufficient children's song generation capabilities in existing technologies has been solved, and personalized and high-quality children's song generation has been achieved, which is suitable for a variety of commercial applications.
Patent Information
- Application Number
- CN202510984258.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-17
- Publication Date
- 2025-09-26
AI Technical Summary
The existing singing synthesis model has a weak ability in generating children's songs, which is difficult to meet the needs of children, and the number and effect of the generated children's songs are not good.
By collecting, organizing and screening children's song resources suitable for children, performing sound source separation and text alignment, training children's song style models, and using large models to generate children's songs, it provides default and user timbre modes and supports personalized customization.
It can generate children's songs similar to existing songs and create songs with unique timbres to meet the personalized needs of children. It is suitable for a variety of commercial application scenarios.
Smart Images

Figure CN120708568A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of singing synthesis, and in particular to a method for generating children's songs based on a singing synthesis large model. Background Art
[0002] Existing singing synthesis models are generated based on training of a large number of songs on the Internet, but children's songs account for a very small proportion of the data in the massive amount of songs on the Internet.
[0003] Therefore, existing technology can generate various types of music such as pop music, classical music, etc., and to a certain extent can reach the level of real creators, but the ability to generate children's songs is relatively weak. This is true for both open source models and closed source models, resulting in the number and effect of children's songs generated currently being unable to meet children's needs for children's songs.
[0004] Therefore, it is necessary to provide a method for generating children's songs based on a large singing synthesis model, so as to generate children's songs similar to existing songs and to create songs with a unique timbre of the generator. Summary of the Invention
[0005] The purpose of the present invention is to provide a method for generating children's songs based on a large singing synthesis model, so as to generate children's songs similar to existing songs and to create songs with a unique timbre of the generator.
[0006] In order to solve the problems existing in the prior art, the present invention provides a method for generating children's songs based on a large singing synthesis model, comprising the following steps:
[0007] Collect data, systematically organize and high-quality screen various children's song resources suitable for children, obtain vocal tracks and background music tracks, and achieve text and audio alignment;
[0008] Train a nursery rhyme style model for subsequent batch labeling;
[0009] The vocal track, background music track, audio-aligned text data, and batch labeled data are input into the pre-trained large model for training to obtain a large model for generating nursery rhymes.
[0010] Optionally, in the method for generating children's songs based on the large singing synthesis model, the method for organizing and screening children's song resources is as follows:
[0011] We conduct preliminary data cleaning and pre-screening on the existing children's song audio and lyrics texts on the market, eliminating audio samples that do not conform to the style of the children's songs, have inappropriate content, or have poor sound effects;
[0012] Based on the age groups of children that different nursery rhymes are suitable for, the nursery rhymes are classified in detail according to the age group mapping system;
[0013] Perform sound source separation on the collected children's song audio data to separate the human voice track and background music track;
[0014] Apply the automatic speech recognition model to transcribe the audio and generate lyrics text that is highly consistent with the audio, so as to improve the labeling system of the training data and achieve alignment between text and audio.
[0015] Optionally, in the method for generating children's songs based on the large singing synthesis model, the children's song style model is trained as follows:
[0016] Assign a global style label to each nursery rhyme through manual annotation;
[0017] During the training phase, the style model uses labeled samples to conduct deep learning on different style features to obtain a training nursery rhyme style model;
[0018] Train the nursery rhyme style model for subsequent batch labeling.
[0019] Optionally, in the method of generating children's songs based on the large singing synthesis model, the global style labels include: light and lively type, warm and lyrical type, sports and game type, fantasy and imagination type, natural exploration type, narrative story type, folk tradition type, educational and cognitive type, and interesting and humorous type.
[0020] Optionally, in the method for generating children's songs based on a large model of singing synthesis, the large model for generating children's songs has two modes, namely a default timbre mode and a user timbre mode.
[0021] Optionally, in the method for generating children's songs based on the singing synthesis model,
[0022] In the default timbre mode, the large model for generating nursery rhymes automatically generates song tracks with random default timbres;
[0023] In the user timbre mode, the large model that generates nursery rhymes uses the default timbre to generate complete nursery rhymes. Through the trained singing voice conversion model that supports zero-shot, the default timbre is converted into the user-specified target timbre with high fidelity.
[0024] Optionally, in the method for generating children's songs based on the large singing synthesis model, in the default timbre mode, a pure music gap is left in the song for inserting adult guiding voice through TTS technology.
[0025] Compared with the prior art, the present invention has the following advantages:
[0026] (1) The present invention can generate children's songs that are similar to existing songs, and can also create songs with a unique timbre of the generator.
[0027] (2) The present invention is designed with a default timbre mode and a user timbre mode. The large model can provide flexible product configuration for different commercial application scenarios (such as children's birthday wishes, parent-child interaction, personalized children's album customization, etc.), meeting the market characteristics of high willingness to pay and strong personalized needs in the field of children's songs. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] Figure 1 A flow chart of a method provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0029] The following is a more detailed description of the specific embodiments of the present invention with reference to schematic diagrams. The advantages and features of the present invention will become more apparent from the following description. It should be noted that the drawings are greatly simplified and not to exact scale, and are only used for the purpose of conveniently and clearly illustrating the embodiments of the present invention.
[0030] In the description of the present application, it should be understood that the terms "center", "longitudinal", "lateral", "length", "width", "thickness", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside", "clockwise", "counterclockwise" and the like to indicate orientations or positional relationships based on the orientations or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be understood as a limitation on the present application.
[0031] Hereinafter, if the method described herein includes a series of steps, the order in which the steps are presented herein is not necessarily the only order in which the steps may be performed, and some of the steps described may be omitted and / or some other steps not described herein may be added to the method.
[0032] Existing technology can generate various types of music such as pop music and classical music, and to a certain extent can reach the level of real creators, but the ability to generate children's songs is relatively weak. This is true for both open source models and closed source models, resulting in the number and effect of children's songs generated currently being unable to meet children's needs for children's songs.
[0033] In order to solve the problems existing in the prior art, the present invention provides a method for generating children's songs based on a large singing synthesis model, such as Figure 1 As shown, the following steps are included:
[0034] S1: Collect data, systematically organize and high-quality screen various children's song resources suitable for children, obtain vocal tracks and background music tracks, and achieve text-audio alignment;
[0035] Specifically, the methods for organizing and screening children's song resources are as follows:
[0036] S11: The data collection phase of this invention focuses on systematically organizing and high-quality screening of various children's song resources suitable for children. First, preliminary data cleaning and pre-screening are performed on the existing children's song audio and lyrics texts on the market, with a focus on eliminating audio samples that clearly do not conform to the children's song style, have inappropriate content, or are of low quality, to ensure the purity and effectiveness of subsequent training data.
[0037] S12: During the data collection process, we categorize children's songs according to the age groups they are suitable for. For example, we differentiate between simple, repetitive songs suitable for infants and toddlers aged 0-3 and rhythmic, interactive songs suitable for children aged 4-6. We then establish an age group mapping system to facilitate subsequent accurate recommendations.
[0038] S13: Perform sound source separation on the collected children's song audio data, effectively separating the vocal track from the background music track. This process helps the model subsequently accurately learn the independent attributes of the singing voice and background music.
[0039] S14: To improve the quality of text-audio alignment, we applied a highly accurate automatic speech recognition (ASR) model to transcribe the audio and generate lyrics text that is highly consistent with the audio. This further improved the labeling system of the training data, laying a solid foundation for subsequent style modeling and timbre conversion.
[0040] S21: Train a nursery rhyme style model for subsequent batch labeling;
[0041] Specifically, the method of practicing the nursery rhyme style model is as follows:
[0042] To efficiently identify and batch-label children's song styles, we designed and trained a Transformer-based children's song style classification model. Through manual annotation, we assigned a global style label to each song, covering eight common styles: lively and cheerful, warm and lyrical, sports and game-like, fantasy and imaginative, nature exploration-based, narrative and storytelling, folk and traditional, educational and cognitive, and humorous.
[0043] During the training phase, the model leverages labeled samples to deeply learn the characteristics of different styles, enhancing its ability to perceive complex information such as timbre, rhythm, and lyric semantics. After systematic training, the model achieved over 90% accuracy in style classification on the test set, meeting the practical application requirements of subsequent large-scale batch data automatic labeling.
[0044] This model can greatly reduce the cost of manual labeling, effectively improve the consistency and accuracy of style labels, and provide high-quality label data support for personalized children's song generation.
[0045] S22: Input the vocal track, background music track, audio-aligned text data, and batch-labeled data into a pre-trained large model for training to obtain a large model for generating children's songs. The present invention uses children's song data to train a timbre conversion model.
[0046] S23: The large model for generating children's songs has two modes: default timbre mode and user timbre mode. In default timbre mode, the large model automatically generates song tracks with random default timbres. In user timbre mode, the large model generates complete songs using the default timbre. Using a trained, zero-shot singing voice conversion model, the default timbre is converted to the user-specified target timbre with high fidelity, supporting personalized customization and improving user stickiness and paid conversion.
[0047] Preferably, in the default timbre mode, a pure music gap is left in the song for inserting adult guidance voice through TTS technology to provide a more humanized interactive experience.
[0048] S31: The present invention allows users to independently select the target children's song style during the inference process, covering a variety of style options, including light and lively, warm and lyrical, sports and game-based, fantasy and imaginative, nature exploration, narrative storytelling, folk and traditional, educational and cognitive, and fun and humorous. Leveraging large-scale text generation models (such as DeepSeek-R1), the system can generate personalized children's song lyrics that meet user needs based on the song theme, creative target (such as a specific child's name or scene), and target style set by the user, ensuring that the generated content is creative, interesting, and child-friendly.
[0049] S32: Default timbre children's song generation: Generate children's songs guided by user timbre using default sound + silent segment.
[0050] S33: User voice children's song generation: Use the default voice + utilize the singing voice conversion model to generate a children's song that is unique to the user's voice.
[0051] In summary, the present invention has the following advantages compared with the prior art:
[0052] (1) The present invention can generate children's songs that are similar to existing songs, and can also create songs with a unique timbre of the generator.
[0053] (2) The present invention is designed with a default timbre mode and a user timbre mode. The large model can provide flexible product configuration for different commercial application scenarios (such as children's birthday wishes, parent-child interaction, personalized children's album customization, etc.), meeting the market characteristics of high willingness to pay and strong personalized needs in the field of children's songs.
[0054] The above description is merely a preferred embodiment of the present invention and does not limit the present invention in any way. Any person skilled in the art who, without departing from the scope of the present invention, makes any equivalent substitution, modification, or other changes to the technical solution and technical content disclosed in the present invention shall be deemed to be within the scope of the present invention and still fall within the scope of protection of the present invention.
Claims
1. A method for generating children's songs based on a large singing synthesis model, characterized in that: The following steps are involved: Collect data, systematically organize and high-quality screen various children's song resources suitable for children, obtain vocal tracks and background music tracks, and achieve text and audio alignment; Train a nursery rhyme style model for subsequent batch labeling; The vocal track, background music track, audio-aligned text data, and batch labeled data are input into the pre-trained large model for training to obtain a large model for generating nursery rhymes.
2. The method for generating children's songs based on a large singing synthesis model according to claim 1, wherein: The methods for organizing and filtering children's song resources are as follows: We conduct preliminary data cleaning and pre-screening on the existing children's song audio and lyrics texts on the market, eliminating audio samples that do not conform to the style of the children's songs, have inappropriate content, or have poor sound effects; Based on the age groups of children that different nursery rhymes are suitable for, the nursery rhymes are classified in detail according to the age group mapping system; Perform sound source separation on the collected children's song audio data to separate the human voice track and background music track; Apply the automatic speech recognition model to transcribe the audio and generate lyrics text that is highly consistent with the audio, so as to improve the labeling system of the training data and achieve alignment between text and audio.
3. The method for generating children's songs based on a large singing synthesis model as claimed in claim 1, wherein: The way to train the nursery rhyme style model is as follows: Assign a global style label to each nursery rhyme through manual annotation; During the training phase, the style model uses labeled samples to conduct deep learning on different style features to obtain a training nursery rhyme style model; Train the nursery rhyme style model for subsequent batch labeling.
4. The method for generating children's songs based on a large singing synthesis model as claimed in claim 3, wherein: Global style tags include: light and lively type, warm and lyrical type, sports and games type, fantasy and imagination type, nature exploration type, narrative story type, folk traditional type, educational and cognitive type, and fun and humorous type.
5. The method for generating children's songs based on a large singing synthesis model as claimed in claim 1, wherein: The large model for generating children's songs has two modes: default timbre mode and user timbre mode.
6. The method for generating children's songs based on a large singing synthesis model as claimed in claim 5, characterized in that: In the default timbre mode, the large model for generating nursery rhymes automatically generates song tracks with random default timbres; In the user timbre mode, the large model that generates nursery rhymes uses the default timbre to generate complete nursery rhymes. Through the trained singing voice conversion model that supports zero-shot, the default timbre is converted into the user-specified target timbre with high fidelity.
7. The method for generating children's songs based on a large singing synthesis model as claimed in claim 5, characterized in that: In the default tone mode, there are pure music gaps in the song for inserting adult guidance voice through TTS technology.