Lyrics generation method and apparatus, and computer-readable storage medium
By analyzing example lyrics to generate narrative text and utilizing natural language processing models, the problem of lack of coherence between paragraphs in AIGC lyrics was solved. This enabled the lyrics to be integrated with online novels and the rapid customization of background music, thereby improving the quality and dissemination effect of multimedia content.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- DOUYIN VISION CO LTD
- Filing Date
- 2024-10-29
- Publication Date
- 2026-05-07
AI Technical Summary
Existing lyrics generated based on AIGC technology lack coherence between paragraphs and are inconsistent in expression, making it difficult to meet the timely background music requirements of multimedia content such as online novels.
By analyzing example lyrics to obtain example features, generating narrative text and extracting target features, and using a natural language processing model to generate target lyrics, the logical coherence and subject matter matching of the lyrics are ensured by combining common features of online novels and external information. Category tags are also added to the generated lyrics for future use.
It improved the consistency of lyrics and their fit with online novels, shortened the music production cycle, enhanced the appeal of online novels and the effect of multimedia dissemination, and enabled the rapid generation and personalized customization of lyrics and music.
Smart Images

Figure CN2024128055_07052026_PF_FP_ABST
Abstract
Description
Lyrics generation method, apparatus and computer-readable storage medium Technical Field
[0001] This disclosure relates to the field of computer technology, and in particular to a method, apparatus and computer-readable storage medium for generating lyrics. Background Technology
[0002] AI-Generated Content (AIGC) refers to the automatic generation of various types of content, such as text, images, audio, and video, using artificial intelligence technology. The rapid development of this field is attributed to the significant advancements in artificial intelligence technology in recent years, particularly breakthroughs in natural language processing, computer vision, and speech synthesis.
[0003] The development of AIGC is a result of both technological advancements and market demand. As artificial intelligence technology matures, AIGC will be widely applied in more fields. In the future, AIGC will continue to drive the automation and intelligence of content generation, bringing more innovation and change.
[0004] Summary of the Invention
[0005] According to some embodiments of this disclosure, a lyrics generation method is provided, comprising: analyzing example lyrics of one or more example songs to obtain example features; generating narrative text based on a target theme and the example features; analyzing the narrative text to obtain target features, wherein the target features are different from the example features; and generating target lyrics based on the target features using a natural language processing model.
[0006] According to some embodiments of this disclosure, a lyrics generation apparatus is provided, comprising: a lyrics analysis module configured to analyze example lyrics of one or more example songs to obtain example features; a text generation module configured to generate narrative text based on a target theme and the example features; a text analysis module configured to analyze the narrative text to obtain target features, wherein the target features are different from the example features; and a lyrics generation module configured to generate target lyrics based on the target features using a natural language processing model.
[0007] According to some embodiments of the present disclosure, a lyrics generation apparatus is provided, including: a memory; and a processor coupled to the memory, the processor being configured to perform methods according to any of the embodiments described in the present disclosure based on instructions stored in the memory.
[0008] According to some embodiments of the present disclosure, a computer-readable storage medium is provided having computer program instructions stored thereon, which, when executed by a processor, implement the lyrics generation method according to any of the embodiments described in the present disclosure.
[0009] According to some embodiments of this disclosure, a computer program product is provided, including computer program instructions that, when executed by a processor, implement the lyrics generation method described in any of the embodiments of this disclosure.
[0010] Other features, aspects, and advantages of this disclosure will become clear from the following detailed description of exemplary embodiments with reference to the accompanying drawings. Attached Figure Description
[0011] Embodiments of this disclosure are described below with reference to the accompanying drawings. It should be understood that the drawings described below are merely illustrative of some embodiments of this disclosure and are not intended to limit the scope of this disclosure. In the drawings:
[0012] Figure 1 shows a schematic flowchart of a lyrics generation method according to some embodiments of the present disclosure;
[0013] Figure 2 illustrates a flowchart of generating narrative text according to some embodiments of the present disclosure;
[0014] Figure 3 illustrates a schematic diagram of lyrics generated using AIGC-based generation technology according to some embodiments of the present disclosure;
[0015] Figure 4 shows a block diagram of a lyrics generation apparatus according to some embodiments of the present disclosure;
[0016] Figure 5 shows a block diagram of a lyrics generation apparatus according to some embodiments of the present disclosure;
[0017] Figure 6 shows a block diagram of an electronic device according to other embodiments of the present disclosure.
[0018] It should be understood that, for ease of description, the dimensions of the various parts shown in the accompanying drawings are not necessarily drawn to actual scale. The same or similar reference numerals are used in the various drawings to denote the same or similar parts. Therefore, once an item is defined in one drawing, it may not be discussed further in subsequent drawings. Detailed Implementation
[0019] The technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. It should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein.
[0020] It should be understood that the various steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect. Unless otherwise specifically stated, the relative arrangement, numerical expressions, and values of components and steps set forth in these embodiments should be interpreted as merely exemplary and do not limit the scope of this disclosure.
[0021] As used in this disclosure, the term "comprising" and its variations are open-ended terms that include at least the following elements / features but do not exclude other elements / features, i.e., "including but not limited to". The term "based on" means "at least partially based on".
[0022] It should be noted that the concepts of "first," "second," etc., used in this disclosure are used only to distinguish different devices, modules, or units, and are not intended to define the order of functions performed by these devices, modules, or units or their interdependencies. Unless otherwise specified, the concepts of "first," "second," etc., are not intended to imply that the objects described herein must be in a given temporal, spatial, rank, or any other given order.
[0023] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0024] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0025] The embodiments of this disclosure are described in detail below with reference to the accompanying drawings; however, this disclosure is not limited to these specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. Furthermore, in one or more embodiments, specific features, structures, or characteristics can be combined in any suitable manner that will be apparent to those skilled in the art from this disclosure.
[0026] Among related technologies, lyrics generated based on AIGC technology lack coherence between paragraphs, and the expression may not be consistent.
[0027] Based on this, the present disclosure provides a lyrics generation method and apparatus, and a computer-readable storage medium, which can improve the consistency of the generated lyrics.
[0028] Figure 1 shows a schematic flowchart of a lyrics generation method according to some embodiments of the present disclosure.
[0029] As shown in Figure 1, the lyrics generation method includes: Step S1, analyzing the example lyrics of one or more example songs to obtain example features; Step S2, generating narrative text based on the target theme and example features; Step S3, analyzing the narrative text to obtain target features, wherein the target features and example features are different; Step S4, using a natural language processing model to generate target lyrics based on the target features.
[0030] The example songs and lyrics are all completed works. For example, the example songs and lyrics are those that have already been licensed for training. The text of the example lyrics can be directly parsed from the example song file. Alternatively, the lyrics can be recognized using methods such as speech recognition.
[0031] Example features are features that the model learns from example lyrics. Example features are, for example, in text form.
[0032] Narrative texts are a type of text used for storytelling, conveying a story by narrating one or more events, experiences, or plots.
[0033] Target features, for example, are features extracted from narrative text, and these target features can be used by the model to generate target lyrics.
[0034] The lyrics generation method in this embodiment can be executed on the client side or partially on the server side.
[0035] One or more steps S1-S4 described above can be implemented based on a natural language processing model. A natural language processing model is, for example, a generative model.
[0036] The lyrics generation method disclosed herein first learns the example features of example lyrics, generates a story based on these features, and then extracts target features from the story to generate the target lyrics. Compared to directly generating lyrics, the model generates target lyrics based on a story, resulting in stronger logical coherence in the content of the target lyrics, reducing fragmented and irrelevant sentences, and making the entire song more fluent and complete. Furthermore, target lyrics generated based on a story can convey a clearer core idea, a deeper connotation, and a more meaningful message, ensuring that the lyrics are not empty of content.
[0037] Figure 2 illustrates a flowchart of generating narrative text according to some embodiments of this disclosure.
[0038] As shown in Figure 2, step S2 includes: step S21, obtaining multiple original texts of the target topic; step S22, determining the common features of the multiple original texts; and step S23, generating a narrative text based on the common features and example features of the multiple original texts.
[0039] The target genres include fantasy, science fiction, urban, romance, adventure, youth, and humor. The original texts are, for example, popular online novels of a particular genre. The model can learn the common features of these popular online novels and, combining the example features with these common features, create lyrics for online novels of this genre.
[0040] In other words, given a specific theme, the model generates lyrics by learning from online novels of that theme.
[0041] The trained model analyzes the creative patterns in example lyrics, such as how metaphors and symbols are used to enhance the depth of the lyrics, how rhyme and rhythm are used to create a tense or lyrical atmosphere, and how the lyrics tell a compelling story. These creative techniques form the foundation of a song's lyrics, and these foundations remain constant regardless of how the lyrics and story change.
[0042] When generating lyrics, the model learns both the creation patterns of example lyrics and the characteristics of online novels. It can generate customized lyrics based on a certain theme, and because new elements are added, it can be different from the example lyrics.
[0043] Applying lyrics to online novels of a certain genre includes, for example, generating music that matches the target lyrics; and generating audiobooks based on the target lyrics, music, and the original text of the target genre.
[0044] Audio reading materials include audiobooks, radio dramas, etc., and the original text of the target subject is, for example, the text to be used to generate an audiobook.
[0045] In other words, after generating target lyrics based on a certain type of online novel, the target lyrics can be used to generate customized background music (BGM) for a specific online novel of that type, thereby generating an audiobook of the online novel based on the background music, thus increasing the attractiveness of the online novel.
[0046] In summary, the lyrics generation method according to some embodiments of this disclosure, which uses a natural language processing model to provide songs for online novels, innovates the way online novels are combined with music, improves the production efficiency of audiobooks, can quickly generate lyrics and music that match the target online novel, shortens the production cycle, reduces production costs, and at the same time improves the quality and attractiveness of audiobooks, thus promoting the automation and intelligent development of audiobook production.
[0047] In some embodiments, the lyrics generation method further includes: determining a category label for the target lyrics; and generating an audio reading based on text that matches the category label of the target lyrics.
[0048] For example, the model could initially be made more flexible by not restricting the categories of online novels to which the generated lyrics would apply. The applicable category could then be determined after the model generates the target lyrics.
[0049] Sometimes, a previously unknown work can suddenly gain immense popularity and attract a large readership due to a captivating plot twist or unique creative idea. Similarly, an author might unexpectedly announce a new book release when inspiration strikes, surprising readers. However, in the face of such sudden popularity or book releases, the production of accompanying songs is often difficult to keep pace with, which to some extent limits the dissemination and influence of online novels in the multimedia field.
[0050] According to some embodiments of this disclosure, after generating target lyrics, matching category tags are first added to them and kept for later use. These category tags can be set based on multiple dimensions such as the theme, style, and emotional tone of the online novel. By tagging the lyrics with these tags, they can be included in a large lyric library, ready to provide background music support for corresponding online novels at any time.
[0051] When a web novel suddenly gains popularity or an author releases a new book, lyrics matching the novel's theme and style can be quickly selected from a lyric library, and background music can be created based on these lyrics. Since the lyrics are pre-generated and tagged, the music production cycle can be significantly shortened, ensuring that matching background music is released promptly during the novel's peak popularity or the crucial period of a new book release, thereby further enhancing the novel's appeal and dissemination.
[0052] Furthermore, this method of generating lyrics first, tagging them, and keeping them for later use allows for more personalized background music services for online novels. Based on the novel's specific plot, character personalities, and reader feedback, the prepared lyrics can be fine-tuned or rewritten to ensure a perfect match between the music and the novel's content. In this way, each online novel can have a tailor-made soundtrack, further enhancing its uniqueness and appeal.
[0053] In summary, by adding matching category tags to the generated target lyrics and keeping them for later use, songs and background music can be generated for online novels more promptly. This effectively addresses the challenges of fluctuating popularity of online novels and the release of new books, enhancing the multimedia dissemination effect of online novels. It not only improves the efficiency and quality of background music production but also provides new ideas for the development of online novels in the multimedia field.
[0054] Example features include at least one of the following: character, character relationship, plot, time, place, environment, and narrative style; target features include extended features of at least some of the example features.
[0055] For example, a natural language processing model could be a conversational generation model. Prompt words can be input into the model, allowing it to extract elements such as characters, character relationships, plot, time, location, environment, and narrative style from example lyrics. The target theme and example features are then input into the model together, enabling it to create a story. Finally, target features are extracted from the story.
[0056] On the one hand, the example lyrics may not include complete information such as characters, character relationships, plot, time, place, environment, and narrative style. The model can supplement these elements during the process of generating the story.
[0057] On the other hand, because the target theme is included, the model may replace or modify the example features when generating the story, resulting in some of the extracted target features being extended or derived features of the example features, thereby avoiding the generated target lyrics being too similar to the original lyrics.
[0058] In summary, when generating target lyrics, the model can generate lyrics that are both consistent with the characteristics of the target theme and highly innovative by supplementing missing story elements in example lyrics, replacing or modifying example features in combination with the target theme, and balancing innovation with the preservation of original lyric features.
[0059] In some embodiments, the lyrics generation method further includes: determining the target theme based on the current season, weather, festival, or important event.
[0060] Current season, weather, holidays, and important events can serve as external information, which is also an important element of lyrics and a significant factor in their popularity. Generally, lyrics about returning home are popular in winter, while lyrics about sprouting and growth, full of hope, are popular in spring. By acquiring external information and analyzing the most suitable direction for lyric creation at that time, the probability of the target lyrics becoming popular can be increased.
[0061] In some embodiments, analyzing example lyrics of one or more example songs to obtain example features includes: selecting one or more example songs from multiple candidate songs based on user feedback; clustering one or more example songs to obtain one or more clusters; analyzing example songs belonging to the same cluster among one or more example songs to obtain common features of each cluster of one or more clusters; and determining the common features of the clusters corresponding to the target category as example features.
[0062] For example, feedback information reflects the popularity of a song. Example songs in the same cluster belong to the same category, such as love, family, friendship, etc.
[0063] Lyrics from candidate songs can be clustered using a text clustering model. From multiple candidate songs, popular songs are selected and clustered. The model can learn the creative characteristics of songs within the same category, i.e., example features, and thus generate customized lyrics for that category.
[0064] Feedback information may include at least one of play count and comment count. Based on user feedback on multiple candidate songs, one or more example songs are selected from the multiple candidate songs, including: sorting the multiple candidate songs according to the feedback information; and determining the one or more candidate songs that are ranked first as one or more example songs.
[0065] For example, you can select N songs that rank highly in terms of play count and / or comment count over a period of time as example songs, where N is a positive integer. Generating target lyrics based on these popular example songs can more accurately capture current trends, increase the play count and comment count of the generated target lyrics, and make the target lyrics easier to share and spread.
[0066] In some embodiments, the lyrics generation method further includes: determining a target category based on the number of example songs in each category.
[0067] For example, after clustering the top N songs by play count and / or comment count, the number of songs in each category is counted. The category with the most songs is the most popular category at present, and this is determined as the target category. Target lyrics generated based on this popular category are more in line with the current needs of users and are easier to share and spread.
[0068] In some embodiments, analyzing narrative text to obtain target features includes: generating a summary of the narrative text; analyzing the theme of the narrative text based on the summary; and determining the target features based on the theme and the narrative text.
[0069] For example, when narrative text is input into a natural language processing model, the model automatically generates a concise summary that retains key information and core plot points. Then, it extracts key information, identifies core themes, and analyzes emotions and motivations from the summary for use in generating target lyrics.
[0070] In some embodiments, determining target features based on a topic and narrative text includes: identifying named entities in the narrative text; selecting key entities based on the relevance of the named entities to the topic; and generating target features based on the key entities.
[0071] For example, narrative text is input into the model, which identifies and labels named entities in the text, such as names of people, places, and times. The model calculates the relevance and importance of each named entity to the topic. For instance, the model assesses the relevance of entities to the topic and determines their importance level in the text by analyzing their frequency of occurrence, contextual information, and association with other entities.
[0072] After identifying named entities closely related to the theme, the model begins extracting features from the text associated with these key entities. Taking love as an example, if the named entities related to the theme are Person A and Person B, the model will focus on extracting various features related to these two individuals. These features include, but are not limited to, their backgrounds (e.g., professions); personality traits (e.g., bravery, gentleness, independence); personal experiences (e.g., upbringing, significant events); and the development of their relationship (e.g., meeting, getting to know each other, falling in love). Furthermore, the model can also focus on analyzing and extracting the dialogue between Person A and Person B to extract information related to their emotions, motivations, and changes in their relationship.
[0073] In some embodiments, the lyrics generation method further includes: generating a text description of the corresponding musical accompaniment based on the target lyrics.
[0074] For example, the target lyrics can be input into a natural language processing model, which can then generate a text description of the corresponding musical accompaniment for use in subsequent composition. By generating a text description of the musical accompaniment, the efficiency of song generation can be improved, and the generated accompaniment can be made to better match the target lyrics.
[0075] In some embodiments, generating a text description of the corresponding musical accompaniment based on the target lyrics includes: dividing the target lyrics into multiple paragraphs according to the plot of the narrative text, and generating a text description of at least one of the instrument selection, harmony design, and rhythm design corresponding to each of the multiple paragraphs.
[0076] For example, the target lyrics can be divided into multiple paragraphs, each paragraph can be matched with the corresponding plot in the narrative text, and appropriate instruments, harmony design, rhythm design, etc. can be selected according to the corresponding plot, so that the generated accompaniment changes with the development of the plot and becomes richer in meaning.
[0077] In some embodiments, a natural language processing model is used to generate target lyrics based on target features, including: generating a first prompt text based on target features, target sentence structure, and rhetorical devices; and generating target lyrics based on the first prompt text using a natural language processing model.
[0078] For example, the prompt text includes statements that restrict the sentence structure and rhetoric of the lyrics. By taking the prompt text as input to the model, the model can generate lyrics content according to common sentence structure and rhetorical features of lyrics, such as lyrics containing some reduplicated words.
[0079] In some embodiments, generating target lyrics using a natural language processing model based on target features includes: determining rhymes based on target features; generating second prompt text based on target features and rhymes; and generating target lyrics using a natural language processing model based on the second prompt text.
[0080] For example, the second prompt text could include instructions to use certain vowel sounds at the end of sentences, enabling the model to rhyme in the target lyrics generated. Alternatively, instead of specifying a specific rhyme, the second prompt text could vaguely include a prompt such as "The generated lyrics need to rhyme," allowing the model to choose an appropriate rhyming method based on actual needs.
[0081] Figure 3 illustrates a schematic diagram of generating lyrics based on AIGC generation technology according to some embodiments of the present disclosure.
[0082] As shown in Figure 3, firstly, the lyrics of popular songs authorized by the original authors for AI training are obtained. The lyrics are analyzed to determine example features. Based on these features, and combined with the current season, trending events, etc., a new story is generated. Based on the newly generated story, the target lyrics are generated according to the rules and logic of the lyrics.
[0083] The following example illustrates the process of generating lyrics.
[0084] First, based on the number of songs of each genre and their play counts in the popular song list, select the currently popular song genres. Then, select popular lyrics for these popular genres. For the selected lyrics that have been authorized for training, analyze the stories behind the popular lyrics and summarize patterns. For example, the prompts input to the model might be as follows.
[0085] "You are a lyrics analysis expert. Please summarize the elements and characteristics of the entered lyrics and analyze and understand the story behind them. Specific elements include: theme, time, place, characters, events, environment, character relationships, event conflicts, and ending. Themes include: family, friendship, love, workplace, inspiration, etc."
[0086] Original lyrics:
[0087] People I Want to Contact
[0088] Does someone hold an umbrella for you when it rains?
[0089] Remember to add layers when it gets cold!
[0090] You were the most beautiful thing I ever missed in my life.
[0091] I dare not disturb them now.
[0092] I've edited the message repeatedly, but I don't dare send it..."
[0093] Based on the above prompts, the model outputs the following results.
[0094] Theme: Longing and helplessness in love.
[0095] Time: Winter.
[0096] Location: Somewhere.
[0097] Person: The singer himself.
[0098] Event: The singer misses someone she once cared about, but dares not easily contact them.
[0099] Environment: It was a cold winter day, with a cold wind blowing and snow falling outside.
[0100] Relationships: The singer may have had a relationship with someone she once cared about, but they are now separated, and the singer is either in unrequited love or unable to maintain normal contact.
[0101] The conflict stems from the singer's intense longing for the other person and their reluctance to disturb them.
[0102] The ending is left ambiguous, leaving the singer trapped in this painful state of longing, unable to escape.
[0103] After obtaining the above features, example stories can be generated based on these features, and example features can be extracted from the example stories. Alternatively, the steps of generating example stories and example features can be skipped, and the above features can be directly used as example features to generate narrative text.
[0104] The following describes the generation of example stories and example features. For example, the features mentioned above and the prompts below are input into the model.
[0105] "Based on the above characteristics, please summarize the input lyrics and provide the story behind the lyrics."
[0106] The model output is as follows.
[0107] "In a small town in the north, there lived a young man named A who was deeply in love with B. A and B met by chance and fell in love..."
[0108] Then, the model extracts example features based on the above story.
[0109] Besides the inherent characteristics of popular lyrics, current external information and trending elements are also crucial factors in determining their popularity. Generally, lyrics about returning home are popular in winter, while lyrics about sprouting and growth, full of hope, are popular in spring. External information includes, for example, the following:
[0110] Seasons, such as: spring, summer, autumn, winter
[0111] Weather, for example: rain, after rain, cloudy, blizzard, clear.
[0112] Time, such as: daytime, nighttime, early morning, midnight
[0113] Holidays, such as: Children's Day, Father's Day, Teacher's Day
[0114] Hot topics, such as: trending news: A mother finds her child who has been lost for many years and they are reunited.
[0115] The aforementioned external information can help determine the theme of the lyrics. Based on the patterns and characteristics of the stories behind popular lyrics obtained from model analysis, combined with the current season, hot events, etc., the AIGC generation model is used to generate new song stories, and the resulting narrative text is as follows.
[0116] "In that quaint little town, C lived. C harbored a special person in her heart, namely D, with whom she had a deep emotional entanglement."
[0117] Then, the target features are generated based on the narrative text using the model, as follows.
[0118] Theme: The struggle and helplessness of love and longing in the face of distance and time.
[0119] Time: From spring to winter.
[0120] Characters: C and D
[0121] Event: C and D met and fell in love in a small town, spending a wonderful time together.
[0122] Environment: The town has a rustic charm.
[0123] Relationship between characters: C and D were once lovers.
[0124] Conflict of Events: The physical distance between C and D prevents them from interacting as before, which is an external conflict. C's internal struggle—wanting to contact D yet fearing to disturb him—is an internal conflict.
[0125] The ending: C waits in longing, the future uncertain, leaving room for the reader's imagination.
[0126] Based on the generated target features, and using the model, lyrics are generated according to the characteristics of rhyme and reduplicated words in the lyrics. The lyrics generated by the model are as follows.
[0127] "Longing Blows in the Wind"
[0128] The street corner where we met
[0129] The beautiful time began with a heartbeat.
[0130] "Walking and laughing together..."
[0131] It can also match the generated lyrics to produce a description of the appropriate accompaniment. For example, based on the target lyrics, it can label the appropriate accompaniment category, such as upbeat, soothing, tense, or uplifting accompaniment. The input prompt text for the model is as follows.
[0132] "You are a music expert who is good at designing accompaniment for lyrics. Please design a unique accompaniment style based on the context and emotions of the lyrics of the song 'Longing Drifts in the Wind'."
[0133] The model generates the following textual description of the accompaniment.
[0134] "The opening is introduced with the gentle sound of wind... In the climax, the playing of all instruments intensifies, especially the strings and drums, creating a strong emotional impact, expressing the protagonist's intense longing and yearning for the departed, as well as his inner anguish and inability to let go of this relationship. Finally, the music gradually fades away amidst the sound effects of wind and falling snowflakes, leaving endless longing floating in the air."
[0135] Figure 4 shows a block diagram of a lyrics generation apparatus according to some embodiments of the present disclosure.
[0136] As shown in Figure 4, the lyrics generation device 4 includes: a lyrics analysis module 41, configured to analyze the lyrics of one or more example songs to obtain example features; a text generation module 42, configured to generate narrative text based on the target theme and example features; a text analysis module 43, configured to analyze the narrative text to obtain target features, wherein the target features and example features are different; and a lyrics generation module 44, configured to generate target lyrics based on the target features using a natural language processing model.
[0137] The lyrics analysis module 41 of the lyrics generation device 4 can be used to perform step S1 of Figure 1, the text generation module 42 can be used to perform step S2 of Figure 1, the text analysis module 43 can be used to perform step S3 of Figure 1, and the lyrics generation module 44 can be used to perform step S4 of Figure 1.
[0138] In some embodiments, the lyrics generation device 4 further includes: a speech reading generation module, configured to generate music matching the target lyrics; and to generate a speech reading based on the target lyrics, the music, and the original text of the target subject.
[0139] In some embodiments, the audio reading generation module is configured to determine the category tags of the target lyrics; and generate audio readings based on text that matches the category tags of the target lyrics.
[0140] In some embodiments, the lyrics generation device 4 further includes a theme determination module, configured to determine a target theme based on the current season, weather, festival, or important event.
[0141] In some embodiments, the lyrics generation apparatus 4 further includes a category determination module configured to determine a target category based on the number of example songs for each category.
[0142] In some embodiments, the lyrics generation device 4 further includes a text description generation module, configured to generate a text description of the corresponding musical accompaniment based on the target lyrics.
[0143] In some embodiments, the text description generation module is further configured to: divide the target lyrics into multiple paragraphs according to the plot of the narrative text; and generate a text description of at least one of the instrument selection, harmony design, and rhythm design corresponding to each of the multiple paragraphs.
[0144] Figure 5 shows a block diagram of a lyrics generation apparatus according to some embodiments of the present disclosure.
[0145] Memory 51 is used to store one or more computer-readable instructions. Memory 51 may include any combination of various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory, including but not limited to random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), read-only memory (ROM), and flash memory. Memory 51 may, for example, store operating systems, application programs, boot loaders, databases, and other programs, as well as various application programs and various data.
[0146] The processor 52 is configured to execute computer-readable instructions to implement the song selection method of any of the foregoing embodiments or the method of any of the foregoing embodiments. Specific implementations of each step of the method can be found in the above embodiments, and repeated details will not be elaborated here.
[0147] Processor 52 can be configured to execute the steps shown in Figures 1 and 2. Processor 52 can be various processing devices, such as a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The central processing unit (CPU) can be based on x86 or ARM architectures, etc.
[0148] The processor 52 and the memory 51 can communicate with each other directly or indirectly. For example, the processor 52 and the memory 51 can communicate via a network. The network can include a wireless network, a wired network, and / or any combination of wireless and wired networks. The processor 52 and the memory 51 can also communicate with each other via a system bus, which is not limited in this disclosure.
[0149] It should be noted that the components of the lyrics generation device 5 shown in Figure 5 are merely exemplary and not limiting. Depending on the specific application requirements, the lyrics generation device 5 may also have other components. The processor 52 can control other components in the lyrics generation device 5 to perform the desired functions.
[0150] The lyrics generation device 5 can be implemented by software, firmware and / or hardware, and can be integrated into a device with the relevant application installed.
[0151] This disclosure also provides a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, implement a lyrics generation method according to some embodiments of this disclosure.
[0152] This disclosure also provides a computer program product, including computer program instructions that, when executed by a processor, implement a lyrics generation method according to some embodiments of this disclosure.
[0153] Figure 6 shows a block diagram of an electronic device according to other embodiments of the present disclosure.
[0154] The electronic device 6 shown in Figure 6 can be a computer system with a dedicated hardware structure, capable of performing corresponding functions when relevant applications are installed.
[0155] Electronic devices include, but are not limited to, mobile terminals such as smartphones, laptops, personal digital assistants (PDAs), tablet computers (PCs), PMPs (portable multimedia players), in-vehicle terminals (such as in-vehicle navigation terminals), wearable devices, and fixed terminals such as digital televisions and desktop computers.
[0156] As shown in Figure 6, the Central Processing Unit (CPU) 61 performs various processes based on programs stored in the Read-Only Memory (ROM) 62 or programs loaded from the storage portion 68 into the Random Access Memory (RAM) 63. The RAM 63 stores data required as needed when the CPU 61 performs various processes. The CPU is merely exemplary and can also be other types of processors, such as the various processors described above. The ROM 62, RAM 63, and storage portion 68 can be various forms of computer-readable storage media. It should be noted that although the ROM 62, RAM 63, and storage portion 68 are shown separately in Figure 6, one or more of them can be combined or located in the same or different memories or storage modules.
[0157] CPU 61, ROM 62 and RAM 63 are interconnected via bus 64. Input / output interface 65 is also connected to bus 64.
[0158] The following components are connected to the input / output interface 65: input section 66, such as a touchscreen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; output section 67, including displays such as cathode ray tube (CRT), liquid crystal display (LCD), speakers, vibrators, etc.; storage section 68, including hard disks, magnetic tapes, etc.; and communication section 69, including network interface cards such as LAN cards, modems, etc. The communication section 69 allows communication processing to be performed via a network such as the Internet. It is readily understood that although parts of the electronic device 6 shown in Figure 6 communicate via bus 64, they can also communicate via a network or other means, wherein the network can include wireless networks, wired networks, and / or any combination of wireless and wired networks.
[0159] As needed, drive 610 is also connected to input / output interface 65. Removable media 611, such as disks, optical disks, magneto-optical disks, semiconductor memories, etc., are installed on drive 610 as needed, so that computer programs read from them can be installed into storage section 68 as needed.
[0160] When the above series of processes are implemented through software, the program constituting the software can be installed from a network such as the Internet or a storage medium such as removable medium 611.
[0161] According to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, some embodiments of this disclosure include a computer program product that, when run on a computer, causes the computer to perform the methods described in any of the foregoing embodiments. The computer program product includes computer instructions carried on a computer-readable medium, containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer instructions can be downloaded and installed from a network via communication section 69, or installed from storage section 68, or installed from ROM 62. When the computer program is executed by CPU 61, the methods of embodiments of this disclosure are performed.
[0162] It should be noted that, in the context of this disclosure, a computer-readable medium can be a tangible medium that may contain or store programs for use by or in conjunction with an instruction execution system, apparatus, or device.
[0163] A computer-readable medium may be a computer-readable storage medium, a computer-readable signal medium, or any combination thereof.
[0164] Computer-readable storage media include, but are not limited to, systems, apparatuses, or devices that are electrical, magnetic, optical, electromagnetic, infrared, or semiconductor, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to, electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. Computer instructions are stored on the computer-readable storage medium that, when executed by a processor, implement the methods described in any of the foregoing embodiments.
[0165] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media may also be any computer-readable medium other than computer-readable storage media, capable of sending, propagating, or transmitting programs for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0166] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0167] In some embodiments, a computer program is also provided, comprising: instructions that, when executed by a processor, cause the processor to perform the methods described in any of the foregoing embodiments. For example, the instructions may be embodied in computer program code.
[0168] In embodiments of this disclosure, computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof. These programming languages include, but are not limited to, object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network (including a local area network (LAN) or a wide area network (WAN)), or it can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0169] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0170] The functions described above can be performed, at least in part, by one or more hardware logic components. For example, without limitation, exemplary hardware logic components that can be used include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0171] While specific embodiments of this disclosure have been described in detail by way of example, those skilled in the art should understand that the examples are for illustrative purposes only and not intended to limit the scope of this disclosure. Those skilled in the art should understand that modifications can be made to the above embodiments without departing from the scope and spirit of this disclosure. The scope of this disclosure is defined by the appended claims.
Claims
1. A method for generating lyrics, comprising: Analyze the example lyrics of one or more example songs to obtain example features; Based on the target subject matter and the example features, generate narrative text; The narrative text is analyzed to obtain target features, wherein the target features are different from the example features; Using a natural language processing model, target lyrics are generated based on the target features.
2. The lyrics generation method according to claim 1, wherein, The step of generating narrative text based on the target subject matter and the example features includes: Obtain multiple original texts on the target topic; Identify the common features of the multiple original texts; Narrative text is generated based on the common features of the multiple original texts and the example features.
3. The lyrics generation method according to claim 1 or 2 further includes: Generate music that matches the target lyrics; An audiobook is generated based on the target lyrics, the music, and the original text of the target subject matter.
4. The lyrics generation method according to claim 1 further includes: Determine the category tags for the target lyrics; An audiobook is generated based on the text that matches the category tags of the target lyrics.
5. The lyrics generation method according to any one of claims 1-4, further comprising: Determine the target subject matter based on the current season, weather, holidays, and important events.
6. The lyrics generation method according to any one of claims 1-5, wherein, The analysis of the narrative text yields target features, including: Generate a summary of the narrative text; Based on the abstract, analyze the theme of the narrative text; Based on the stated theme and the stated narrative text, target features are determined.
7. The lyrics generation method according to any one of claims 1-6, wherein, The step of determining target features based on the theme and the narrative text includes: Identify named entities in the narrative text; Select key entities based on the relevance of the named entities to the topic; Based on the key entities, target features are generated.
8. The lyrics generation method according to any one of claims 1-7, wherein, The analysis of example lyrics from one or more example songs yields example features, including: Based on user feedback on multiple candidate songs, select one or more example songs from the multiple candidate songs; Cluster the one or more example songs to obtain one or more clusters; Analyze the example songs belonging to the same cluster among the one or more example songs to obtain the common features of each cluster of the one or more clusters; The common features of the clusters corresponding to the target category are determined as the example features.
9. The lyrics generation method according to claim 8, wherein, The feedback information includes at least one of play count and comment count. The step of selecting one or more example songs from the multiple candidate songs based on user feedback includes: The candidate songs are sorted according to the feedback information; The first or more candidate songs in the ranking are identified as the first or more example songs.
10. The lyrics generation method according to claim 8 or 9, further comprising: The target category is determined based on the number of example songs in each category.
11. The lyrics generation method according to any one of claims 1-10, further comprising: Based on the target lyrics, generate a text description of the corresponding musical accompaniment.
12. The lyrics generation method according to claim 11, wherein, The step of generating a text description of the corresponding musical accompaniment based on the target lyrics includes: Based on the plot of the narrative text, the target lyrics are divided into multiple paragraphs; Generate a text description of at least one of the instrument selection, harmony design, and rhythm design corresponding to each of the plurality of paragraphs.
13. The lyrics generation method according to any one of claims 1-12, wherein, The process of generating target lyrics using a natural language processing model based on the target features includes: Based on the target features, target sentence structure, and rhetorical devices, generate the first prompt text; Using the natural language processing model, the target lyrics are generated based on the first prompt text.
14. The lyrics generation method according to any one of claims 1-13, wherein, The process of generating target lyrics using a natural language processing model based on the target features includes: Based on the target features, determine the rhyme scheme; Based on the target features and rhymes, a second prompt text is generated; Using the natural language processing model, the target lyrics are generated based on the second prompt text.
15. The lyrics generation method according to any one of claims 1-14, wherein: The example features include at least one of the following: characters, character relationships, plot, time, place, setting, and narrative style; The target features include at least some of the extended features of the example features.
16. A lyrics generation device, comprising: The lyrics analysis module is configured to analyze the lyrics of one or more sample songs to obtain sample features; The text generation module is configured to generate narrative text based on the target subject matter and the example features; A text analysis module is configured to analyze the narrative text to obtain target features, wherein the target features are different from the example features; The lyrics generation module is configured to use a natural language processing model to generate target lyrics based on the target features.
17. A lyrics generation device, comprising: Memory; as well as A processor coupled to the memory, the processor being configured to execute the lyrics generation method according to any one of claims 1 to 15 based on instructions stored in the memory.
18. A computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, implement the lyrics generation method according to any one of claims 1 to 15.
19. A computer program product comprising computer program instructions that, when executed by a processor, implement the lyrics generation method according to any one of claims 1 to 15.