Artificial intelligence-based composition method and device, intelligent terminal and storage medium

CN122531339APending Publication Date: 2026-08-07SHENZHEN COOCAA NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHENZHEN COOCAA NETWORK TECH CO LTD
Filing Date
2026-06-01
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

[0003]现有技术的AI作曲方案多为“歌词-旋律”串行生成,存在两大缺点:1)旋律生成时未实时反向校验歌词的韵律与重音,导致词曲咬合度低;2)采用固定风格嵌入,无法根据用户历史偏好及当前上下文情绪动态调整,个性化不足

Benefits of technology

[0016]本发明的有益效果:本发明提供了一种基于人工智能的作曲方法、装置、智能终端及存储介质,本发明通过引入“词曲咬合校验器”与“情绪-偏好双因子自适应嵌入”,在电视端实现高个性化、词曲同步优化的作曲,解决词曲错位及风格僵化问题。并且本发明通过专门新增的词曲咬合闭环校验模块,同时搭配情绪+用户偏好双动态适配风格算法,专门适配电视使用场景,做到词曲同步贴合校准、曲风贴合喜好和场景情绪,彻底解决词曲对不齐、作曲风格死板无新意的行业难题;并且本发明还具有如下优点:

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122531339A_ABST
    Figure CN122531339A_ABST
Patent Text Reader

Abstract

The application discloses a composition method and device based on artificial intelligence and a terminal, relates to the technical field of artificial intelligence, and comprises the following steps: obtaining a song generation requirement; extracting a song generation key feature, and calling corresponding emotion and user preference double-dimension linkage factors to perform adaptive style embedding fusion to obtain a fused song generation double-factor vector; generating a song lyric with a stress mark through a lyric generator; taking the stress mark as a hard constraint, adopting a syllable-music note one-to-one mapping network, and performing real-time reverse verification of rhyming and stress to compose a song, and when a word-music matching degree calculated by a word-music occlusion verifier is greater than a preset threshold, outputting the matched word-music; and inputting the matched word-music into a voice synthesis model with a self-defined personal voice to perform voice synthesis and play the music. The application solves the problems of word-music mispositioning and style rigidity, can quickly generate corresponding lyrics and music scores according to a song theme or emotional content input by a user, and can play music.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and in particular to a composition method, apparatus, smart terminal, and storage medium based on artificial intelligence. Background Technology

[0002] With the development of technology and the continuous improvement of people's living standards, various smart terminals and AI technologies are constantly developing, and some users have begun to use AI to compose music.

[0003] Existing AI composition solutions mostly generate lyrics-melody sequentially, which has two major drawbacks: 1) The rhythm and stress of the lyrics are not verified in real time during melody generation, resulting in low consistency between lyrics and music; 2) Fixed style embedding is used, which cannot be dynamically adjusted according to the user's historical preferences and current context and emotions, resulting in insufficient personalization.

[0004] Therefore, existing technologies still need to be improved and developed. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to provide a music composition method, device, smart terminal and storage medium based on artificial intelligence, which addresses the shortcomings of the prior art. The present invention solves the problems of misalignment of lyrics and music and rigidity of style. The present invention can quickly generate corresponding lyrics and musical scores based on the song theme or emotional content input by the user, and play the music.

[0006] The technical solution adopted by this invention to solve the problem is as follows: An artificial intelligence-based composition method, comprising: Obtain song generation requirements; The obtained song generation requirements are analyzed, key features of song generation are extracted, and adaptive style embedding and fusion are performed by calling the corresponding emotion and user preference dual-dimensional linkage factors to obtain the fused song generation dual-factor vector. Based on the fused song, a two-factor vector is generated, and lyrics with accent marks are generated through a lyrics generator. Based on the generated lyrics with accent marks, the system uses a preset melody generator with accent marks as hard constraints, employs a syllable-note one-to-one mapping network to perform real-time reverse verification of rhyme and accent, composes music, and outputs the matched lyrics and music when the matching degree of the lyrics and music is greater than a preset threshold, calculated by a preset lyrics and music matching verifier. The matched lyrics and music are input into a speech synthesis model with a custom personal voice, and then the speech is synthesized and played.

[0007] The aforementioned AI-based composition method, wherein the step of inputting the matched lyrics and melody into a speech synthesis model with a custom personal voice timbre for speech synthesis and playback includes: The original neural network model for speech synthesis was constructed and subjected to sensitivity progressive pruning and 4-bit integer quantization to obtain a speech synthesis model with a customizable personal voice.

[0008] The aforementioned AI-based composition method, wherein the step of inputting the matched lyrics and melody into a speech synthesis model with a custom personal voice timbre for speech synthesis and playback includes: A raw speech synthesis model with 20 Transformer layers and FP32 precision is pre-built; The original speech synthesis model is subjected to sensitivity progressive pruning, gradually removing unimportant weights to obtain the pruned speech synthesis model. The pruned speech synthesis model was layer-trimmed, compressed from 20 layers to 6 layers, and redundant layers were removed to obtain the trimmed speech synthesis model. The trimmed speech synthesis model is quantized using 4-bit integer quantization, and the weights are quantized from 32 bits to 4 bits to obtain the quantized speech synthesis model. The quantized speech synthesis model is fine-tuned and / or distilled to restore accuracy, resulting in a speech synthesis model with a customizable personal voice. The matched lyrics and music are input into a speech synthesis model with a custom personal voice, and then the speech is synthesized and played.

[0009] The aforementioned AI-based music composition method, wherein the step of obtaining song generation requirements includes: Get the song generation requirements from users via voice or text input.

[0010] The aforementioned AI-based composition method, wherein the steps of composing music based on generated lyrics with accent marks, using a preset melody generator with accent marks as hard constraints, employing a syllable-note one-to-one mapping network, and real-time reverse verification of rhyme and accent, include: If the verification fails, the composition process will be restarted.

[0011] The aforementioned AI-based music composition method, wherein the step of obtaining song generation requirements precedes the following: A pre-set lyric-melody matching checker is used to optimize the melody in real time with lyric accents as a hard constraint, preventing lyric-melody misalignment.

[0012] The aforementioned AI-based music composition method, wherein the step of obtaining song generation requirements precedes the following: A pre-set emotion-preference dual-factor module is used to dynamically adjust the style according to the user profile of the smart terminal. The emotion and user preference dual factors are adaptively embedded into the extracted key features of song generation to obtain the fused song generation dual-factor vector. A lightweight composition engine is pre-configured for the terminal, which enables real-time composition on the terminal through sensitivity pruning, 4-bit integer quantization, and sparse operator collaboration.

[0013] An artificial intelligence-based music composition device, wherein the device comprises: The acquisition module is used to acquire song generation requirements; The emotion preference dual-factor module is used to analyze the acquired song generation requirements, extract key features for song generation, and call the corresponding emotion and user preference dual-dimensional linkage factors for adaptive style embedding and fusion to obtain the fused song generation dual-factor vector. The lyrics generation module is used to generate two-factor vectors based on the fused song and generate lyrics with accent marks through the lyrics generator; The lightweight composition module is used to compose music based on generated lyrics with accent marks. It uses a preset melody generator with accent marks as hard constraints, a syllable-note one-to-one mapping network, and real-time reverse verification of rhyme and accent. When the matching degree of lyrics and music is greater than a preset threshold, the module outputs the matched lyrics and music. The speech synthesis output module is used to input the matched lyrics and music into a speech synthesis model with a custom personal voice for speech synthesis and playback.

[0014] A smart terminal includes a memory and one or more programs, wherein one or more programs are stored in the memory and configured to be executed by one or more processors, the one or more programs comprising the method for performing any one of the methods.

[0015] A computer-readable storage medium, wherein, when instructions in the storage medium are executed by a processor of an electronic device, the electronic device is enabled to perform any of the methods described above.

[0016] The beneficial effects of this invention are as follows: This invention provides a composition method, device, smart terminal, and storage medium based on artificial intelligence. By introducing a "lyric-music matching checker" and "emotion-preference dual-factor adaptive embedding," this invention achieves highly personalized, synchronized, and optimized composition on television, solving the problems of lyric-music misalignment and rigid style. Furthermore, this invention, through a specially added closed-loop lyric-music matching checker module, combined with an emotion + user preference dual dynamic adaptation style algorithm, is specifically adapted to television usage scenarios, achieving synchronized calibration of lyrics and music, and matching of musical style with preferences and scene emotions, completely solving the industry problems of misaligned lyrics and music and rigid, unoriginal composition styles. This invention also has the following advantages: 1) Improved the matching accuracy of lyrics and music; 2) Improved user satisfaction; 3) Accelerated the efficiency of generating a song offline on TV. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 This is a flowchart illustrating the AI-based composition method provided in Embodiment 1 of the present invention.

[0019] Figure 2 This is a flowchart illustrating the AI-based composition method provided in Embodiment 2 of the present invention.

[0020] Figure 3 This is a schematic diagram of the lightweight strategy process for the speech synthesis model of the AI-based composition method provided in this embodiment of the invention.

[0021] Figure 4 A schematic diagram of an embodiment of an artificial intelligence-based music composition device provided by the present invention.

[0022] Figure 5 This is a block diagram illustrating the internal structure of a smart terminal provided in an embodiment of the present invention. Detailed Implementation

[0023] To make the objectives, technical solutions, and advantages of this invention clearer and more explicit, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.

[0024] It should be noted that if the embodiments of the present invention involve directional indicators (such as up, down, left, right, front, back, etc.), the directional indicators are only used to explain the relative positional relationship and movement of the components in a certain specific posture (as shown in the figure). If the specific posture changes, the directional indicators will also change accordingly.

[0025] The current mainstream AI intelligent composition application solutions on the market generally adopt a serial linkage generation architecture of pre-lyric analysis and post-melody one-way deduction. When actually deployed on a large scale to adapt to diverse audio-visual adaptation scenarios such as multimedia audio-visual support and exclusive background music for large-screen TVs, the inherent structural technical shortcomings are gradually becoming apparent. The overall adaptability and practicality, as well as the quality of the composed product, are difficult to meet the requirements of the scenarios.

[0026] Firstly, in the entire process of frame-by-frame iterative generation of full-domain melody, the traditional algorithm architecture does not have a synchronous reverse closed-loop verification and control unit. It cannot specifically target the inherent rhythm and rhyme of the original lyrics, the rhythm of sentence segmentation, and the core semantic stress multi-dimensional core benchmark parameters to carry out real-time linkage verification and correction. This easily leads to high-frequency problems such as melody beat misalignment, word and sound mismatch, and musical phrase mismatch. Ultimately, this significantly reduces the accuracy of the integration of lyrics and music, resulting in a prominent sense of incongruity when adapted for simultaneous audio-visual broadcast on TV.

[0027] Secondly, traditional AI composition style empowerment adopts a fixed embedding mode with pre-set static parameters. It can only call the pre-set fixed style algorithm template in the background to carry out a single composition rendering. It cannot collect and accumulate user's full-cycle composition interaction history preference profile data, nor can it capture and match the current creative scene and the dynamic emotional tone of the whole context in real time. It lacks the ability to link and adapt to multi-dimensional variables, resulting in serious homogenization of the finished compositions, a prominent phenomenon of "one person, one face," and weak personalized scene adaptation empowerment efficiency. This addresses the core pain points of the industry that are common in the existing traditional serial AI composition architecture, namely poor matching of lyrics and music, rigid and inflexible style output, and lack of personalized adaptation capabilities.

[0028] To address the aforementioned issues, this application proposes an AI-based composition method. This application features a completely optimized and reconstructed AI composition underlying architecture, exclusively equipped with a lightweight dynamic closed-loop verifier for lyric-music alignment. It also includes a synchronously developed adaptive style embedding and control mechanism based on a dual-dimensional linkage factor of emotion and user preference, precisely adapting to the exclusive human-computer interaction operating conditions on television devices. This efficiently implements real-time optimization of lyric-music synchronization across the entire domain, empowering a personalized intelligent composition process tailored to each user. It fundamentally overcomes the industry's practical bottlenecks in traditional composition techniques, such as misalignment and disconnection of lyric-music alignment, rigid and inflexible style output, and insufficient adaptability to large-screen terminals.

[0029] like Figure 1 As shown, an artificial intelligence-based composition method according to Embodiment 1 of the present invention includes the following steps: S100, Obtain song generation requirements; In this embodiment of the invention, a smart terminal such as a smart TV can be used as an example. In this step, the user's input song generation request is obtained, for example, the user inputs "write a light and nostalgic song for graduation season". In specific implementation, the user's voice input song generation request can be obtained through the smart terminal's microphone. Of course, it can also be obtained directly through text input.

[0030] In this embodiment, the song generation requirement refers to all the requirements, specifications, styles, and content conditions proposed by the user who wants AI to automatically create / generate a complete song. It is a complete task instruction given to the AI ​​model to create a song. For example, generating a traditional Chinese style inspirational song with a gentle female vocal, 60 seconds in length, including a verse and chorus, arranged with traditional Chinese musical instruments, and lyrics about pursuing dreams and growing up.

[0031] The song generation requirements in this embodiment, i.e., the song creation requirements given by the user, include: musical style, mood, lyrical theme, tempo, vocal quality, whether it is Chinese style / pop, or whether it is sad / inspirational, etc.

[0032] S200. Analyze the obtained song generation requirements, extract key features for song generation, and call the corresponding emotion and user preference dual-dimensional linkage factors to perform adaptive style embedding and fusion to obtain the fused song generation dual-factor vector. In this embodiment, parsing the song generation requirements means converting the user's colloquial and fragmented song creation requests into structured data that a machine can understand. Extracting key features for song generation involves refining the core elements that determine the song's style from the user's requirements: such as mood (e.g., sad / upbeat), musical style (e.g., traditional Chinese / pop), tonality, BPM, narrative theme, and user preference tags.

[0033] In this embodiment, the emotion dimension factor is a model vector representing the emotional attributes of a song, such as emotional numerical features like healing, excitement, sadness, gentleness, and uplifting. The user preference dimension factor is a model vector representing a user's long-term listening and songwriting habits, such as a user's preference for traditional Chinese style, slow tempo, female vocals, and lyrical arrangements. The dual-dimensional linkage factor in this embodiment refers to a set of model parameters that bind and mutually constrain the emotion factor and the user preference factor; they do not function independently but are mutually matched and linked.

[0034] The adaptive style embedding and fusion in this embodiment refers to the automatic and dynamic adjustment of weights by the AI ​​model based on the current song characteristics, which integrates the emotional style and user preference style into a unified style representation, rather than a rigid splicing.

[0035] The song generation two-factor vector obtained in this step of the embodiment refers to a set of digital feature vectors output after fusion, which serve as the underlying input features for subsequent AI music composition, lyric writing, and melody generation.

[0036] In the specific implementation of this step, firstly, the song generation requirements submitted by the user are obtained; then, the song generation requirements are analyzed and understood, and key features for song generation are selected and extracted; then, pre-trained emotion factors and user preference factors are retrieved to establish a linkage between the two dimensions; then, through an adaptive style embedding fusion algorithm, the emotion style and user preference style are intelligently fused together; finally, a fused two-factor vector is generated, which serves as the core input basis for AI to generate melody, lyrics, arrangement, and vocal style.

[0037] For example, if the user's request is "to generate a gentle and healing ancient-style lyrical song, with a slow tempo and piano + guzheng arrangement," then in this embodiment, the spoken request is converted into a machine structure such as: style = ancient style, emotion = healing and gentle, tempo = slow, instrument = piano + guzheng. Then, key features are extracted, including: ancient style, healing, slow tempo, traditional Chinese instruments, and lyrical narrative. A two-dimensional linkage factor is then invoked, where the emotion factor includes: healing, gentle, and soothing vectors; and the user preference factor includes: preference for ancient style, slow tempo, traditional Chinese instruments, and lyrical style vectors. This two-factor linkage forces a match between the user's preference for ancient style and the healing emotion, preventing deviations to rock or electronic music.

[0038] Then, adaptive style embedding and fusion are performed, such as automatic model weighting: increasing the weight of traditional Chinese style and healing music, and decreasing the weight of fast-paced and electronic music, so as to smoothly blend the two styles. The resulting two-factor vector is a series of feature vectors. This makes it easy for subsequent steps to use the preset AI model to directly generate the corresponding melody, lyrics, and accompaniment. The finished product is both healing and in line with the user's preference for traditional Chinese style.

[0039] Thus, this invention avoids rigidly applying templates or forcing a fixed style onto the user. Instead, it adaptively matches user needs and personal preferences, resulting in greater customization. Furthermore, this invention considers both the emotions the song intends to express and the user's daily preferences, ensuring it's not merely pleasant to listen to but emotionally resonant. It also employs a two-dimensional factor linkage and fusion approach to avoid conflicts between melody, emotion, and user preferences, maintaining a unified and cohesive style. Moreover, this invention is scalable and reusable: features, factors, and vectors can be embedded into the model, becoming increasingly adept at understanding user preferences with repeated use.

[0040] S300: Generate a two-factor vector based on the fused song, and generate lyrics with accent marks through a lyrics generator; The fused song generation dual-factor vector in this embodiment is the standardized AI feature vector output from the previous steps. It carries two major pieces of information: the emotional features the song wants to express and the user's style preference features. It is the underlying style instruction for the AI-generated song in this embodiment.

[0041] The lyrics generator in this embodiment refers to a sub-model / functional module specifically designed for AI lyric writing, which can automatically create lyrics text that conforms to the rhythm, sentence structure, and length of a song based on style, mood, and theme.

[0042] The accent marks described in this embodiment refer to marking the positions on the lyrics that need to be emphasized, highlighted, paused, or adjusted in volume during reading and / or singing. These marks can be made using symbols, subscripts, or special fields; they are used for subsequent AI melody matching, automatic recitation, TTS singing synthesis, and rhythm alignment.

[0043] In the specific implementation of step S200, the combined emotion and user preference dual-factor vector from the previous step is used as input conditions to define the overall style, emotion, and tone of the lyrics generator; and a dedicated lyrics generator model is called to automatically create lyrics that fit the song structure, following the style, emotion, and theme constraints in the vector; in this embodiment of the invention, while generating lyrics, vocal accent marks are automatically added to the positions of words and sentences, and the structured lyrics with accent marks are output, rather than plain text.

[0044] For example, if the two-factor vector features are ancient style, gentle and healing, slow rhythm, and lyrical, then in this embodiment of the invention, the above two-factor vector is input into a preset lyrics generator; the lyrics generator automatically generates lyrics according to the style: "The evening breeze gently caresses the distant mountains, the moonlight falls upon the human world." In this embodiment of the invention, accent marks are automatically added simultaneously, for example, bolding is used to represent accents: "The evening breeze gently caresses the distant mountains, the moonlight falls upon the human world." In this specific embodiment, the meaning of the marks is that when composing and singing, "evening breeze, distant mountains, moonlight, human world" need to be emphasized, while the other words are sung softly, which is suitable for the rhythmic feel of ancient style and lyrical music. In this way, in this embodiment of the invention, when matching music with a large model later, the melody pitch and rhythm strength can be matched according to these accent positions to generate a song that is more in line with the effect of human vocal performance.

[0045] Thus, in this embodiment of the invention, relying on two-factor vectors for lyric creation, the lyrics' mood and musical style perfectly match the user's needs, ensuring no disconnect between the style and the requirements. Furthermore, this invention can generate labeled lyrics in one step, eliminating the need for manual accent marking later, producing standard lyrics directly usable for arrangement and vocal synthesis. Moreover, the accent markings in this invention directly provide a basis for subsequent melody generation, beat alignment, and AI singer performance, resulting in more natural rhythm and more nuanced vocal delivery. Furthermore, the lyrics of this invention are no longer plain text but structured data with rhythmic features, facilitating subsequent model processing and algorithm integration.

[0046] S400: Based on the generated lyrics with accent marks, the preset melody generator uses accent marks as hard constraints, adopts a syllable-note one-to-one mapping network, verifies rhyme and accent in real time, composes music, and outputs the matched lyrics and music when the matching degree of the lyrics and music is greater than the preset threshold by the preset lyrics and music matching verifier. The preset melody generator in this embodiment is a pre-trained and fixedly deployed AI composition sub-model, specifically designed to generate melodies that match lyrics. It has built-in composition rule templates for different styles, rhythms, and vocal ranges, and is specifically adapted for generating vocal melodies, unlike pure instrumental arrangement models.

[0047] In this embodiment, the accent marking refers to using the accent markings in the lyrics generated in the previous steps as a hard constraint. This means that the accent positions marked on the lyrics are unchangeable and must be followed as a mandatory rule. The arrangement of accents cannot be arbitrarily changed when the melody is generated, and the pitch, intensity, and rhythm of the notes are rigidly constrained.

[0048] The syllable-note one-to-one mapping network in this embodiment adopts a dedicated neural network structure. Its main function is to achieve precise binding between text and notes. For example, it strictly corresponds each Chinese syllable in the lyrics to an independent note, eliminating the disorder of multiple pronunciations for one character or multiple pronunciations for one space, and ensuring that the arrangement of lyrics and music is neat.

[0049] The real-time reverse verification in this embodiment refers to a dynamic detection mechanism that performs reverse verification while generating the melody, which is different from checking after generation. In this invention, during the melody generation process, the rhyme fluency and accent matching of the lyrics are checked in real time, which can correct unreasonable melodies in a timely manner.

[0050] The lyric-melody matching verifier used in this embodiment is a dedicated algorithm detection module used to quantitatively calculate the degree of fit between lyrics and melody. It generates a standardized matching score through multiple dimensions such as pitch fluctuation, tempo, accent position, and lip shape matching.

[0051] In this embodiment, the preset threshold is a pre-set pass / fail score standard, which is a critical value for determining whether the lyrics and music are compatible. For example, it is set to a match score greater than 0.92. Only when the match score is higher than this value is the lyrics and music considered to be compatible; if it is lower than the threshold, it is considered unacceptable and needs to be re-optimized and generated.

[0052] In the specific implementation of step S400, lyrics with accent marks are first used as the basic input, and the position of the accent marks in the lyrics is set as a hard constraint, preventing the melody from arbitrarily changing the accent rhythm. Then, a preset AI melody generator is called, relying on a one-to-one syllable-note mapping network to ensure that each Chinese character accurately corresponds to a single note, thus standardizing the arrangement of lyrics and music. During the real-time melody generation process, the rhyme of the lyrics is verified in reverse to ensure that the accent rhythm matches, and melody bugs are corrected in real time to avoid inconsistencies between lyrics and music. After the melody generation is completed, the lyrics and music matching checker is started to calculate the lyrics and music matching score from multiple dimensions. When the calculated matching score is greater than the pre-set qualified threshold (e.g., greater than 0.92), the lyrics and music matching is deemed to be up to standard, and the finally matched lyrics + melody combination is output; if it is not up to standard, the process returns to optimize the composition again.

[0053] For example, when the input lyrics (bold for stress marks) are: The evening breeze gently caresses the distant mountains, and the moonlight falls into the world; the preset matching threshold is 92 points. In the embodiments of the present invention, the hard constraint is executed, and the melody generator is forced to stipulate that: "evening breeze, distant mountains, moonlight, world" are stress words and must be placed on strong beats and paired with high notes, and the remaining ordinary Chinese characters are placed on weak beats and paired with gentle low notes, and the inversion of strong and weak rhythms is prohibited. Then, a syllable-note one-to-one mapping network is adopted, and they are strictly corresponding one by one as: evening (1 note), breeze (1 note), gently (1 note), caresses (1 note), without extra prolonged sounds or random empty beats, and the sentence pattern is neat.

[0054] In this embodiment, real-time reverse verification will be performed, that is, when generating the melody, it will be detected in real time whether: "mountains, world" rhyme smoothly and whether the volume of the stress notes is prominent; if the generated melody causes the rhyme to be rigid and the stress to be dull, the pitch of the notes will be fine-tuned on the spot without having to redo the whole song.

[0055] Then, bite verification and threshold determination will also be performed. Specifically, the verifier is used to detect the pitch fluctuation, beat matching, and stress fit, and finally the score is 96 points; 96 points > the preset threshold of 92 points, it is determined to be qualified, and this set of song and lyrics combination is output; if the score is 80 points, which is lower than the threshold, the melody will be automatically returned for modification.

[0056] As can be seen from the above, in the embodiments of the present invention, with stress as the hard constraint, it can ensure that the stressed words in the lyrics are paired with high notes and strong beat notes, and the weakly pronounced words are paired with soft notes and weak beat notes, which conforms to the human singing habit and will not have the awkward listening feeling of "should be stressed but not, should be soft but not". And the arrangement of the songs and lyrics in the present invention is regular. Relying on the syllable-note one-to-one mapping, it avoids problems such as multiple notes for one character and chaotic prolonged sounds. The rhythm of the lyrics is clean and tidy, suitable for AI voice synthesis and post-editing. Moreover, the present invention can also achieve real-time optimization and error reduction. By adopting real-time reverse verification, it is not necessary to wait until the whole melody is generated before modification, avoiding problems such as broken rhymes and misplaced stresses in advance, reducing the rework cost, and improving the generation efficiency. Furthermore, the present invention uses the verifier + preset threshold to determine the song and lyrics compatibility with data instead of fuzzy subjective judgment, ensuring the stable and unified bite quality of each generated song. And the present invention has strong coherence. Because it承接前文双因子向量、带重音歌词,全程链路连贯,牢牢贴合用户情绪、风格偏好,不出现曲风断裂、词曲割裂的情况。

[0057] S500. Input the matched songs and lyrics into the voice synthesis model with a custom personal voice, and perform voice synthesis and playback.

[0058] It should be noted that there is an unclear expression "承接前文双因子向量、带重音歌词,全程链路连贯,牢牢贴合用户情绪、风格偏好,不出现曲风断裂、词曲割裂的情况。" in the original text. You may need to check and correct it according to the actual situation. If you have any other questions, please feel free to let me know.In this embodiment, the matched lyrics and melody are the final lyrics + melody file that has passed the matching verification in step S400. Its rhythm, accents, and melody are perfectly matched and can be directly used for vocal performance. The custom personal timbre in this embodiment refers to a user-customized private vocal timbre, distinct from the general standard AI timbre, and is a personal, exclusive voice. The speech synthesis model in this embodiment refers to an AI singing TTS model, an algorithm model capable of automatically generating human voices from input lyrics and melody, simulating real human singing. The speech synthesis playback in this embodiment refers to the process of converting lyrics and melody data into human voice audio and completing the playback output.

[0059] In this step S500, the lyrics and melody that passed the verification and matching in the previous step are used as input and imported into an AI voice synthesis model equipped with a personalized voice. The model imitates the user's customized voice, generating a realistic singing voice according to the predetermined melody, rhythm, and accents, and finally completes the audio synthesis and playback output. For example, using the classical Chinese style lyrics and music output from the previous step: "The evening breeze gently caresses the distant mountains, the moonlight falls upon the human world," this embodiment of the step will input the qualified lyrics and music into a pre-recorded and trained user-specific gentle female voice model; the model replicates the user's voice, sings the complete song according to the original melody and accents, and finally directly plays the finished classical Chinese style vocal song.

[0060] Thus, this invention does not use generic mechanical timbres, but rather personalized timbres, resulting in highly unique and easily recognizable songs. Furthermore, this invention synthesizes vocals based on mature lyric and musical data, eliminating off-key singing and gaps in pitch, perfectly matching the melody and rhythm. Moreover, this invention can be directly synthesized and played in real time, adaptable to televisions, mobile devices, and other terminal devices.

[0061] In a further embodiment, the AI-based composition method includes step S500, which involves: applying sensitivity progressive pruning and 4-bit integer quantization to the constructed original neural network model for speech synthesis to obtain a speech synthesis model with a custom personal timbre.

[0062] In this embodiment of the invention, the original neural network model for speech synthesis is compressed from 20 layers to 6 layers by using sensitivity progressive pruning and 4-bit integer quantization (INT4 quantization), resulting in a size reduction of less than 80MB; thus achieving a lightweight music composition engine.

[0063] Furthermore, such as Figure 3 As shown, step S500 specifically includes: S501, pre-construct a raw speech synthesis model including 20 Transformer layers and FP32 precision; In this embodiment, the 20-layer Transformer refers to a large, stacked model with 20 layers of transformers. This deep network has strong feature extraction capabilities, making it suitable for high-precision human voice modeling. The FP32 precision refers to 32-bit single-precision floating-point storage, storing the original model at its highest precision without loss of weight information.

[0064] Specifically, this step involves building an initial basic model using a 20-layer Transformer network structure, with all weights stored in FP32 high-precision format. This initial model has a large number of parameters and strong feature capabilities, fully preserving all subtle features of the human voice, such as timbre, breath, and rhythm, and serves as the parent model for subsequent compression and optimization.

[0065] S502. Perform sensitivity progressive pruning on the constructed original speech synthesis model, gradually removing unimportant weights to obtain the pruned speech synthesis model. The sensitivity progressive pruning in this step gradually removes unimportant parameter weights that have minimal impact on the vocal effect based on the magnitude of the sensitivity weights. This is a refined slimming algorithm.

[0066] Specifically, this step performs sensitivity analysis on the constructed original large model to determine the degree of influence of each weight on the human voice synthesis effect; adopts a progressive approach to slowly remove redundant weights with low sensitivity and minimal contribution, retains the core key timbre weights, completes the sparsity of model parameters, and obtains the pruned speech synthesis model.

[0067] This step only removes invalid weights, with almost no loss in sound quality; it also generates a sparse model that is compatible with the TV NPU sparse operator, resulting in faster inference speed; furthermore, this step can reduce the computational load of the model and reduce inference latency.

[0068] S503. The pruned speech synthesis model is layer-trimmed from 20 layers to 6 layers, and redundant layers are removed to obtain the pruned speech synthesis model. Layer pruning in this step refers to directly removing redundant network layers from the model, reducing the computational load of model inference, and lowering structural complexity.

[0069] In this step, after weight pruning, the network structure of the pruned speech synthesis model is simplified by directly removing redundant and feature-repeating network layers, compressing the original 20-layer Transformer to 6 layers, greatly simplifying the model structure and obtaining a lightweight model.

[0070] In this way, the number of layers in the pruned speech synthesis model obtained through this step is greatly reduced, the reasoning logic is simpler, and the hardware computing power consumption is extremely low; it can also significantly reduce latency, and the present invention eliminates redundant layers, which can avoid the repeated calculation of features.

[0071] S504. The trimmed speech synthesis model is quantized using 4-bit integer quantization, and the weights are quantized from 32 bits to 4 bits to obtain the quantized speech synthesis model. In this embodiment, INT4 quantization (4-bit integer quantization) refers to compressing the model weights from 32-bit floating-point to 4-bit integer, which greatly reduces the model size and memory usage.

[0072] Specifically, this step involves performing low-bit-width quantization on the compressed model, forcibly compressing the original 32-bit floating-point weights into 4-bit integer storage, further minimizing the model size and generating a low-bit-quantized model. This reduces the model size to 1 / 8 of its original size, resulting in extremely low memory usage. Furthermore, it is compatible with TVs and low-end edge NPUs (Neural Network Processors), offering low hardware requirements and strong compatibility. Quantization significantly improves computational speed, making it suitable for real-time song generation and playback.

[0073] S505. Fine-tune and / or distill to restore the accuracy of the quantized speech synthesis model to obtain a speech synthesis model with a custom personal voice. The fine-tuning and distillation in this step refers to using the high-precision original large model as the teacher model to perform precision repair on the compressed small model, making up for the sound quality loss caused by pruning and quantization.

[0074] In the specific implementation of this step, because the model will produce slight sound quality distortion and timbre shift after pruning, layer cutting and quantization, this step uses a high-precision original large model as the teacher model. Through model distillation and parameter fine-tuning, the accuracy loss caused by compression is repaired, and finally a personal timbre speech synthesis model that retains the user's unique timbre, has low latency and is lightweight is generated.

[0075] Thus, this invention can solve the problems of deteriorated sound quality and timbre distortion in compressed models; it also balances lightweight design with high fidelity, preserving the distinctiveness of individual timbre; furthermore, this invention can complete model customization, ultimately producing a unique personal timbre singing model.

[0076] S506. Input the matched lyrics and music into the speech synthesis model with a custom personal voice, and perform speech synthesis and playback.

[0077] Specifically, in this step, the lyrics and melody that have been successfully matched and integrated are input into the final optimized lightweight personal voice model. The model generates personalized vocal audio according to a predetermined melody, accents, and rhythm, and then plays it back. This invention achieves fast edge-side inference speed, low latency, and smooth, stutter-free playback. Furthermore, this invention uses a personal, exclusive voice, resulting in a high degree of song personalization. Additionally, the model is small in size and can be directly deployed on home terminal devices such as televisions.

[0078] As can be seen from the above, this invention starts from a high-precision original model, and then achieves multi-level lightweight compression through pruning, layer trimming, and quantization. Then, it repairs compression loss through fine-tuning distillation, and finally obtains a personalized voice synthesis model that is adapted to the edge-side NPU, has low latency, and high fidelity. This enables the synthesis and playback of human voices in AI-customized songs, taking into account the three core advantages of model inference speed, hardware adaptability, and personal voice fidelity.

[0079] In a further embodiment of the present invention, the AI-based music composition method includes step S100, which includes: obtaining the song generation requirements of the user through voice input or text input.

[0080] In this embodiment of the invention, users can use the microphone of a smart terminal to input voice or use a keyboard or mobile phone to input text to generate songs, which is very convenient for users to operate and use.

[0081] In a further embodiment of the present invention, the composition method based on artificial intelligence further includes step S400: if the verification fails, the composition is re-executed.

[0082] In this embodiment of the invention, the melody generator uses accent marks as hard constraints and employs a one-to-one syllable-note mapping network to perform real-time reverse verification of rhyme and accent. If the verification fails, it backtracks and regenerates. The real-time reverse verification involves simultaneously detecting the fluency of the lyrics' rhyme and the reasonableness of the accent matching during the melody generation process. The backtracking and regeneration refers to reverting to the previous generation step when the verification fails, readjusting the melody parameters, and composing again, without retaining inferior melodies.

[0083] Specifically, in this step, when the melody generator is composing, it uses lyric accents as a mandatory constraint; it generates the melody using a network structure that maps syllables to notes, while simultaneously performing real-time reverse checks on rhyme and accent matching. If the checks detect issues such as awkward rhymes, misaligned accents, or disharmony between lyrics and melody, it returns to the previous node to regenerate the melody until it passes the checks.

[0084] In this way, the present invention can automatically backtrack and rewrite unqualified melodies without manual intervention, improve the pass rate of finished products, and has a strong self-correction ability; and achieve stable and controllable generation quality, avoiding the production of inferior songs with broken rhymes and disordered rhythms.

[0085] In a further embodiment of the present invention, the artificial intelligence-based composition method includes the following steps before step S100: A pre-set lyric-melody matching checker is used to optimize the melody in real time with lyric accents as a hard constraint, preventing lyric-melody misalignment and solving the problem of lyric-melody misalignment.

[0086] A pre-set emotion-preference dual-factor module is used to dynamically adjust the style according to the user profile of the smart terminal. The emotion and user preference dual factors are adaptively embedded into the extracted key features of song generation to obtain the fused song generation dual-factor vector. In this way, the present invention can realize the adaptive embedding of emotion-preference dual factors and dynamically adjust the style according to the user profile of the TV terminal, breaking through the limitations of static embedding.

[0087] A lightweight composition engine is pre-configured for the terminal, which enables real-time composition on the terminal through sensitivity pruning, 4-bit integer quantization, and sparse operator collaboration.

[0088] The present invention will be further described in detail below through another specific application embodiment: like Figure 2 As shown, the overall system block diagram implemented in this specific application embodiment includes a TV SoC (System-on-a-Chip). Lyrics and music matching checker Emotion-Preference Two-Factor Module Lightweight music composition engine The playback module; the technical problems to be solved in this specific embodiment are: 1) solving the problem of low word-music consistency and abrupt listening experience; 2) solving the problem of static style embedding that cannot change in real time with user emotions / preferences; 3) solving the problem of limited computing power on the TV and the need for lightweight inference.

[0089] This specific application embodiment provides an artificial intelligence-based composition method, such as... Figure 2 As shown, it includes the following steps: Step S11: Obtain the user's input song generation requirement as "Write a light and nostalgic song for graduation season"; Based on the input song generation requirement, use the emotion-preference module to call the TV user profile (historical viewing, voice emotion) to output a "light and nostalgic" dual-factor vector; Step S12: The lyrics generator first generates lyrics with accent marks; In this step of the embodiment, the output is a two-factor vector of "lighthearted + nostalgic". The lyrics generator first generates lyrics with accent marks, for example, the generated lyrics with accent marks are "Youth never fades, summer never says goodbye". Among them, the bolded "youth, fleeting years, summer, goodbye" are accented.

[0090] Step S13: The melody generator uses accent marks as hard constraints and adopts a one-to-one mapping network of "syllable-note" to back-check rhyme and accent in real time. If the check fails, it backtracks and regenerates. In this embodiment, the melody generator strictly follows the lyric stress marks during composition, with stress being an unchangeable hard requirement; a one-to-one mapping of syllables to notes is used to generate the melody, with one word corresponding to one note; during the generation process, the smoothness of the rhyme and the matching degree of the stress are checked in real time. If there is any awkward rhyme, misplaced stress, or incoherence between the lyrics and the melody, the composition is immediately restarted until it passes the check.

[0091] In this way, by strictly adhering to the position of the accented notes, it conforms to the human singing habits of varying volume, resulting in a smooth listening experience; it also ensures that the lyrics and music are neat and orderly, with each word paired with a note, creating a well-structured composition suitable for subsequent AI-generated individual voice synthesis; furthermore, this invention can achieve automatic error correction without manual intervention, automatically reverting to regeneration if an unqualified melody is found, thus improving the pass rate of the final product; through this invention, stable and controllable quality can be achieved, effectively preventing poor-quality generated effects such as broken rhymes, chaotic rhythms, and disconnected lyrics and music.

[0092] Step S14: The lyric-music matching checker calculates the "lyric-music matching degree" and outputs it only if it is greater than 0.92; This invention uses a lyric-music matching checker to calculate the "lyric-music matching degree" to be >0.92 before outputting, which can improve the lyric-music matching degree.

[0093] Step S15: The result is played back using a personal voice TTS (text-to-speech model).

[0094] In this embodiment of the invention, a lightweight strategy is employed, specifically using sensitivity progressive pruning and INT4 quantization to compress a 20-layer Transformer (speech synthesis model) to 6 layers, such as... Figure 3 As shown, this results in a compressed speech synthesis model size of less than 80MB. Furthermore, this invention introduces a sparse operator from a television NPU (Neural Processing Unit), achieving an inference latency of less than 300ms, thus improving song generation efficiency.

[0095] The core objective of this invention is to significantly reduce model size and computational overhead while maintaining model performance, enabling it to run on resource-constrained devices. By pruning the model to make it sparse (with many zeros), the NPU's sparse operator specifically skips these zeros, saving both memory and computational power. Ultimately, this allows the model to run extremely fast on TV chips, producing results within 300 milliseconds.

[0096] Of course, in other embodiments of the present invention, the hard constraint of accent can be changed to soft constraint loss, bypassing the validator structure; and the two-factor embedding can be split into a single emotion factor to reduce the degree of personalization; or the lightweight part can be generated purely in the cloud, with the TV only serving as the playback terminal, which can reduce the terminal pressure.

[0097] Exemplary device like Figure 4As shown, an embodiment of the present invention provides an artificial intelligence-based music composition device, comprising: Module 310 is used to obtain song generation requirements; The emotion preference dual-factor module 320 is used to analyze the acquired song generation requirements, extract key features of song generation, and call the corresponding emotion and user preference dual-dimensional linkage factors for adaptive style embedding and fusion to obtain the fused song generation dual-factor vector. The lyrics generation module 330 is used to generate a two-factor vector based on the fused song and generate lyrics with accent marks through the lyrics generator; The lightweight composition module 340 is used to compose music based on generated lyrics with accent marks. It uses a preset melody generator with accent marks as hard constraints, adopts a syllable-note one-to-one mapping network, and performs real-time reverse verification of rhyme and accent. When the matching degree of lyrics and music is greater than a preset threshold, it outputs the matched lyrics and music. The speech synthesis output module 350 is used to input the matched lyrics and music into a speech synthesis model with a custom personal voice for speech synthesis and playback, as described above.

[0098] Based on the above embodiments, the present invention also provides a smart terminal, the principle block diagram of which can be as follows: Figure 5 As shown, the smart terminal can be a smart large-screen TV or a smart computer. The smart terminal includes a processor, memory, network interface, display screen, and database connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface of the smart terminal is used to communicate with external terminals via a network connection. When the computer program is executed by the processor, it implements an artificial intelligence-based composition method. The database of the smart terminal stores the artificial intelligence-based composition program.

[0099] Those skilled in the art will understand that Figure 5 The block diagram shown is merely a partial structural diagram related to the present invention and does not constitute a limitation on the smart terminal to which the present invention is applied. A specific smart terminal may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0100] In one embodiment, a smart terminal is provided, including a memory and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by one or more processors. The one or more programs include instructions for performing the following operations: Obtain song generation requirements; The obtained song generation requirements are analyzed, key features of song generation are extracted, and adaptive style embedding and fusion are performed by calling the corresponding emotion and user preference dual-dimensional linkage factors to obtain the fused song generation dual-factor vector. Based on the fused song, a two-factor vector is generated, and lyrics with accent marks are generated through a lyrics generator. Based on the generated lyrics with accent marks, the system uses a preset melody generator with accent marks as hard constraints, employs a syllable-note one-to-one mapping network to perform real-time reverse verification of rhyme and accent, composes music, and outputs the matched lyrics and music when the matching degree of the lyrics and music is greater than a preset threshold, calculated by a preset lyrics and music matching verifier. The matched lyrics and music are input into a speech synthesis model with a custom personal voice, and then the speech is synthesized and played back, as described above.

[0101] The step of inputting the matched lyrics and music into a speech synthesis model with a custom personal voice for speech synthesis and playback includes: The original neural network model for speech synthesis was constructed and subjected to sensitivity progressive pruning and 4-bit integer quantization to obtain a speech synthesis model with a customizable personal voice.

[0102] The step of inputting the matched lyrics and music into a speech synthesis model with a custom personal voice for speech synthesis and playback includes: A raw speech synthesis model with 20 Transformer layers and FP32 precision is pre-built; The original speech synthesis model is subjected to sensitivity progressive pruning, gradually removing unimportant weights to obtain the pruned speech synthesis model. The pruned speech synthesis model was layer-trimmed, compressed from 20 layers to 6 layers, and redundant layers were removed to obtain the trimmed speech synthesis model. The trimmed speech synthesis model is quantized using 4-bit integer quantization, and the weights are quantized from 32 bits to 4 bits to obtain the quantized speech synthesis model. The quantized speech synthesis model is fine-tuned and / or distilled to restore accuracy, resulting in a speech synthesis model with a customizable personal voice. The matched lyrics and music are input into a speech synthesis model with a custom personal voice, and then the speech is synthesized and played.

[0103] The steps for obtaining song generation requirements include: Get the song generation requirements from users via voice or text input.

[0104] The aforementioned AI-based composition method, wherein the steps of composing music based on generated lyrics with accent marks, using a preset melody generator with accent marks as hard constraints, employing a syllable-note one-to-one mapping network, and real-time reverse verification of rhyme and accent, include: If the verification fails, the composition process will be restarted.

[0105] The step of obtaining song generation requirements includes the following prior steps: A pre-set lyric-melody matching checker is used to optimize the melody in real time with lyric accents as a hard constraint, preventing lyric-melody misalignment.

[0106] The step of obtaining song generation requirements includes the following prior steps: A pre-set emotion-preference dual-factor module is used to dynamically adjust the style according to the user profile of the smart terminal. The emotion and user preference dual factors are adaptively embedded into the extracted key features of song generation to obtain the fused song generation dual-factor vector. A lightweight music composition engine is pre-configured for the terminal, which enables real-time music composition on the terminal through sensitivity pruning, 4-bit integer quantization, and sparse operator collaboration, as described above.

[0107] In other embodiments, this application proposes a computer-readable storage medium that, when the instructions in the storage medium are executed by the processor of an electronic device, enables the electronic device to perform the above-described method.

[0108] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments described above. Any references to memory, storage, databases, or other media used in the embodiments provided by this invention can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0109] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.

Claims

1. An artificial intelligence-based composition method, characterized by, include: Obtain song generation requirements; The obtained song generation requirements are analyzed, key features of song generation are extracted, and adaptive style embedding and fusion are performed by calling the corresponding emotion and user preference dual-dimensional linkage factors to obtain the fused song generation dual-factor vector. Based on the fused song, a two-factor vector is generated, and lyrics with accent marks are generated through a lyrics generator. Based on the generated lyrics with accent marks, the system uses a preset melody generator with accent marks as hard constraints, employs a syllable-note one-to-one mapping network to perform real-time reverse verification of rhyme and accent, composes music, and outputs the matched lyrics and music when the matching degree of the lyrics and music is greater than a preset threshold, calculated by a preset lyrics and music matching verifier. The matched lyrics and music are input into a speech synthesis model with a custom personal voice, and then the speech is synthesized and played.

2. The composition method based on artificial intelligence according to claim 1, characterized in that, The steps of inputting the matched lyrics and music into a speech synthesis model with a custom personal voice for speech synthesis and playback include: The original neural network model for speech synthesis was constructed and subjected to sensitivity progressive pruning and 4-bit integer quantization to obtain a speech synthesis model with a customizable personal voice.

3. The composition method based on artificial intelligence according to claim 1, characterized in that, The steps of inputting the matched lyrics and music into a speech synthesis model with a custom personal voice for speech synthesis and playback include: A raw speech synthesis model with 20 Transformer layers and FP32 precision is pre-built; The original speech synthesis model is subjected to sensitivity progressive pruning, gradually removing unimportant weights to obtain the pruned speech synthesis model. The pruned speech synthesis model was layer-trimmed, compressed from 20 layers to 6 layers, and redundant layers were removed to obtain the trimmed speech synthesis model. The trimmed speech synthesis model is quantized using 4-bit integer quantization, and the weights are quantized from 32 bits to 4 bits to obtain the quantized speech synthesis model. The quantized speech synthesis model is fine-tuned and / or distilled to restore accuracy, resulting in a speech synthesis model with a customizable personal voice. The matched lyrics and music are input into a speech synthesis model with a custom personal voice, and then the speech is synthesized and played.

4. The composition method based on artificial intelligence according to claim 1, characterized in that, The steps for obtaining song generation requirements include: Get the song generation requirements from users via voice or text input.

5. The composition method based on artificial intelligence according to claim 1, characterized in that, The steps for composing music based on generated lyrics with accent marks, using a preset melody generator with accent marks as hard constraints, and employing a syllable-note one-to-one mapping network to verify rhyme and accent in real time, include: If the verification fails, the composition process will be restarted.

6. The composition method based on artificial intelligence according to claim 1, characterized in that, Prior to the step of obtaining song generation requirements, the following steps are included: A pre-set lyric-melody matching checker is used to optimize the melody in real time with lyric accents as a hard constraint, preventing lyric-melody misalignment.

7. The composition method based on artificial intelligence according to claim 1, characterized in that, Prior to the step of obtaining song generation requirements, the following steps are included: A pre-set emotion-preference dual-factor module is used to dynamically adjust the style according to the user profile of the smart terminal. The emotion and user preference dual factors are adaptively embedded into the extracted key features of song generation to obtain the fused song generation dual-factor vector. A lightweight composition engine is pre-configured for the terminal, which enables real-time composition on the terminal through sensitivity pruning, 4-bit integer quantization, and sparse operator collaboration.

8. A music composition device based on artificial intelligence, characterized in that, The device includes: The acquisition module is used to acquire song generation requirements; The emotion preference dual-factor module is used to analyze the acquired song generation requirements, extract key features for song generation, and call the corresponding emotion and user preference dual-dimensional linkage factors for adaptive style embedding and fusion to obtain the fused song generation dual-factor vector. The lyrics generation module is used to generate two-factor vectors based on the fused song and generate lyrics with accent marks through the lyrics generator; The lightweight composition module is used to compose music based on generated lyrics with accent marks. It uses a preset melody generator with accent marks as hard constraints, a syllable-note one-to-one mapping network, and real-time reverse verification of rhyme and accent. When the matching degree of lyrics and music is greater than a preset threshold, the module outputs the matched lyrics and music. The speech synthesis output module is used to input the matched lyrics and music into a speech synthesis model with a custom personal voice for speech synthesis and playback.

9. A smart terminal, characterized in that, It includes a memory and one or more programs, wherein one or more programs are stored in the memory and configured to be executed by one or more processors, wherein the one or more programs include methods for performing any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, When the instructions in the storage medium are executed by the processor of the electronic device, the electronic device is able to perform the method as described in any one of claims 1-7.