Multilingual Rhythmic Lyric Generation Method, System, Device and Storage Medium

The best rhyming word pairs are screened through speech generation and recognition technology, and combined with the autoregressive text generation algorithm, the shortcomings of the existing lyric generation model in terms of rhythm and multilingual rhyme are solved, and the multilingual lyric generation of higher quality and diversity are achieved.

CN114220408BActive Publication Date: 2025-07-25UNIV OF SCI & TECH OF CHINA
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111521191.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-13
Publication Date
2025-07-25
Estimated Expiration
2041-12-13

AI Technical Summary

Technical Problem

The existing lyric generation model has a big gap compared with human creation in terms of generative rhythm, especially in multilingual mixed rap lyrics, rhyme words are not ideal for generation, and the lexicon is poor in the diversity of the lexicon.

Method used

Generate multiple speech signals through speech generation and recognition technology, filter out the best rhyming word pairs, combine the autoregressive text generation algorithm to generate multilingual rhyming words, and introduce phoneme information during the generation process to improve rhythm and semantic coherence.

Benefits of technology

It improves the rhyme quality and semantic coherence of the generated words, supports the generation of multilingual lyrics, and makes the rhyme words more diverse.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114220408B_ABST
    Figure CN114220408B_ABST
Patent Text Reader

Abstract

The present invention discloses a method, system, device and storage medium for generating multilingual rhyming lyrics. A speech generation model is used to capture the pronunciation of words, and then rhyming word pairs are generated, greatly improving the rhyming quality of the generated words. At the same time, an autoencoder model with the rhyme as the starting input is adopted during the generation process, which can generate lyrics with more coherent semantics. Moreover, it supports the generation of multilingual lyrics, making the generated rhyming words more diverse.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of natural language processing, and in particular, to a method, system, device and storage medium for generating multi-language rhyming lyrics. Background Art

[0002] The speech recognition direction is currently relatively mature, and relatively accurate recognition results can be obtained in most languages. Text generation is also a popular field in recent natural language processing. With the support of large-scale corpus pre-trained language models, autoregressive language models such as GPT have been able to perform well in text generation tasks.

[0003] Through multi-language text generation, on the one hand, generation tasks can be carried out for people in different languages, and on the other hand, the mixture of multiple languages can bring more creativity and provide more inspiration for lyric creation.

[0004] However, the current lyric generation model lacks the soul of a song - rhythm. At present, text generation models have been able to generate some relatively fluent lyrics, but there is still a large gap in the rhythm of the lyrics compared with human creation. Especially in the face of some scenarios that emphasize the rhythm of lyrics, such as the currently popular hip-hop lyrics (rap lyrics) globally, the model's ability is slightly insufficient.

[0005] The current methods for generating rap lyrics are as Figure 1 shown.

[0006] During the process of generating lyrics, the song theme is first determined. According to the theme, the first sentence is selected from the predefined lyrics as the first sentence of the generated lyrics, and it is sent into the trained text generation model for lyric generation, and finally the whole lyrics are generated. The rap lyric modeling process is the main innovation point of its solution. The model adopts a neural network structure based on LSTM. The corpus comes from the NetEase rap lyric corpus. By using the initials and finals of the modern Chinese pinyin system, the pinyin of the last 1-5 characters of each lyric sentence is extracted, and the Jieba word segmentation tool is used to segment each lyric sentence to extract the key words of the lyrics. The word vectors of the lyrics are obtained through word2vec, and the obtained word vectors and the pinyin information of the last few words of each sentence are used as the training set to train the model.

[0007] In addition, an auxiliary lyric writing system is also implemented in this solution. By searching for knowledge-based words, candidate words in the thesaurus composed of corpus-based words and related rhyming words are used to replace the target word, so as to selectively provide some candidate operations.

[0008] The above scheme process is relatively simple and intuitive. The number of parameters of the entire generation model is small, the generation speed is fast, and it is easy to use online. At the same time, the generation process is divided into two modules, which improves the flexibility and reusability of the scheme. The lyrics generation module can be used to generate lyrics of any style. The auxiliary lyrics module can be used to replace non-rhyming words in the generation model to improve the generation quality of lyrics, and can be used to assist in modifying words to improve the diversity of generated sentences.

[0009] Although the above solution is simple and usable, it has the following defects:

[0010] 1) The ability to capture rhymes using pure text modeling is limited, so the generation effect of rhyming word pairs in the generative model is not very ideal, and an auxiliary lyric writing module is needed to make certain corrections.

[0011] 2) The solution is only for Chinese texts, but multi-language mixed rap is popular in the rap world, and some English words are often mixed in Chinese, which can also form a good rhyme and a higher style.

[0012] 3) Although using a vocabulary to submit words has high stability, its diversity is relatively poor due to the limitation of the vocabulary size. Summary of the invention

[0013] The purpose of the present invention is to provide a method, system, device and storage medium for generating rhythmic lyrics in multiple languages, which can generate multilingual lyrics with better rhythm.

[0014] The objective of the present invention is achieved through the following technical solutions:

[0015] A method for generating rhythmic lyrics in multiple languages, comprising:

[0016] Extracting a number of words to be rhymed from the previous lyrics that need to be rhymed; for each word to be rhymed, generating a speech signal of the word to be rhymed by using a speech generation technology, generating multiple new speech signals of the word to be rhymed by using a rhyme pair generation technology, and then generating multiple different candidate words corresponding to the multiple new speech signals of the word to be rhymed by using a speech recognition technology, forming a rhyme team with each candidate word and the word to be rhymed, and selecting the candidate word corresponding to the best rhyme team as the multilingual rhyme word for generating a sentence;

[0017] The lyrics text is generated by adopting an autoregressive text generation algorithm based on all the screened multilingual rhyming words, the previous lyrics and the language information of the identified previous lyrics.

[0018] A multilingual rhythmic lyrics generation system is implemented based on the above method, and the system includes:

[0019] A multilingual rhyming word generation module is used to extract a number of words to be rhymed from the sentences that need to rhyme in the previous lyrics; for each word to be rhymed, a voice signal of the word to be rhymed is generated through voice generation technology, and multiple new voice signals of the word to be rhymed are generated through rhyming pair generation technology, and then multiple different candidate words corresponding to the multiple new voice signals of the word to be rhymed are generated through voice recognition technology. Each candidate word and the word to be rhymed form a rhyming pair, and the candidate word corresponding to the best rhyming pair is selected as the multilingual rhyming word used to generate the sentence.

[0020] A multilingual rhyming lyrics text generation module is used to generate lyrics text by using an autoregressive text generation algorithm according to all the selected multilingual rhyming words, the previous lyrics, and the recognized language information of the previous lyrics.

[0021] A processing device includes: one or more processors; a memory for storing one or more programs;

[0022] Wherein, when the one or more programs are executed by the one or more processors, the one or more processors implement the foregoing method.

[0023] A readable storage medium stores a computer program, and is characterized in that when the computer program is executed by a processor, the foregoing method is implemented.

[0024] It can be seen from the technical solutions provided by the present invention above that a voice generation model is used to capture the pronunciation of words, and then rhyming word pairs are generated, which greatly improves the rhyming quality of the generated words. At the same time, an autoencoder model with the rhyme as the starting input (that is, the autoregressive text generation model introduced later) is used in the generation process, which can generate lyrics with more coherent semantics; moreover, it supports the generation of multilingual lyrics, making the generated rhyming words more diverse. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0026] Figure 1 It is a flowchart of the current rap lyrics generation modeling method provided for the background technology of the present invention;

[0027] Figure 2 It is a flowchart of a multilingual rhyming lyrics generation method provided for the embodiments of the present invention;

[0028] Figure 3A detailed flowchart of a multi - language rhyming lyrics generation method provided by an embodiment of the present invention;

[0029] Figure 4 A schematic diagram of the multi - language TTS model structure provided by an embodiment of the present invention;

[0030] Figure 5 A schematic diagram of a multi - language rhyming pair generation model with a Transformer structure provided by an embodiment of the present invention;

[0031] Figure 6 A schematic diagram of the structure of the LAS speech recognition model provided by an embodiment of the present invention;

[0032] Figure 7 A schematic diagram of the vector part design of the bert - structured language model for pronunciation information provided by an embodiment of the present invention;

[0033] Figure 8 A schematic diagram of the generation model provided by an embodiment of the present invention;

[0034] Figure 9 A schematic diagram of a multi - language rhyming lyrics generation system provided by an embodiment of the present invention;

[0035] Figure 10 A schematic diagram of a processing device provided by an embodiment of the present invention. Detailed implementation manners

[0036] Next, in combination with the accompanying drawings in the embodiments of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts belong to the protection scope of the present invention.

[0037] First, the following explanations are given for the terms that may be used in this article:

[0038] Descriptions with semantic meanings such as "including", "comprising", "containing", "having" or other similar ones should be interpreted as non - exclusive inclusion. For example: including a certain technical feature element (such as raw materials, components, ingredients, carriers, dosage forms, materials, dimensions, parts, components, mechanisms, devices, steps, processes, methods, reaction conditions, processing conditions, parameters, algorithms, signals, data, products or articles, etc.) should be interpreted as not only including the clearly listed certain technical feature element, but also including other well - known technical feature elements in the art that are not clearly listed.

[0039] The following is a detailed description of a multilingual rhythmic lyrics generation scheme provided by the present invention. The contents not described in detail in the embodiments of the present invention belong to the prior art known to professionals in the field. If no specific conditions are specified in the embodiments of the present invention, the conventional conditions in the field or the conditions recommended by the manufacturer shall be followed.

[0040] Embodiment 1

[0041] like Figure 1 As shown, a method for generating rhythmic lyrics in multiple languages mainly includes the following steps:

[0042] Step 1, extracting a number of words to be rhymed from the sentences that need to be rhymed in the previous lyrics; for each word to be rhymed, generating a speech signal of the word to be rhymed by using speech generation technology, and generating multiple new speech signals of the word to be rhymed by using rhyme pair generation technology, and then generating multiple different candidate words corresponding to the multiple new speech signals of the word to be rhymed by using speech recognition technology, forming a rhyme team with each candidate word and the word to be rhymed, and selecting the candidate word corresponding to the best rhyme team as the multilingual rhyming word for generating sentences.

[0043] Step 2: Generate lyrics text using an autoregressive text generation algorithm based on all the selected multilingual rhyming words, previous lyrics, and the identified language information of the previous lyrics.

[0044] For ease of understanding, the following Figure 3 The detailed flow chart shown introduces the preferred implementation methods of the above two steps in detail.

[0045] 1. Multilingual rhyming word generation.

[0046] This stage corresponds to the above step 1. It mainly selects rhyming words through rhyme rules, then the speech generation model generates word pronunciation, and the speech signal is sent to the speech generation module to generate rhyming word pronunciation. Finally, the multilingual speech recognition model generates the corresponding rhyming words and gives multiple different candidate words. All candidate words are judged by the rhyming word discrimination model to find the most suitable rhyming word.

[0047] In addition, considering the significance of the first word of a sentence, the first line of lyrics is input by the user. The present invention takes the creation of rap lyrics as an example to show the entire multilingual rhythmic lyrics generation process, and of course, it is also applicable to other types of song creation.

[0048] The preferred implementation of each part of this stage is as follows:

[0049] 1. Extract rhyming sentences.

[0050] In the embodiment of the present invention, the sentence that needs to rhyme is randomly selected from the previous lyrics, and the distance between the selection probability and the position of the current sentence to be generated satisfies the following relationship:

[0051]

[0052] Among them, dis(i) and dis(j) respectively represent the distances between sentence i and sentence j in the previous lyrics and the position of the current sentence to be generated, generally using the number of intervening sentences; α is a hyperparameter, and the larger the value, the more inclined to the sentences with a closer distance.

[0053] Exemplarily, if the position of the current sentence to be generated is the sixth sentence in the song, then a sentence is randomly selected from the previous five sentences for rhyming to increase the diversity of generation.

[0054] 2. Extraction of replacement words.

[0055] In the embodiments of the present invention, the extracted sentences that need to rhyme are segmented (which can be implemented using the Jieba segmentation tool), and the sentences that need to rhyme are divided into individual words separated by spaces, and then the words are processed. This processing method has a better effect than processing individual characters. During the creation of rap songs, different replacement words are extracted from the segmented sentences that need to rhyme according to different rhyming techniques as the words to be rhymed; such as double rhymes, triple rhymes in Chinese, and half rhymes, consonant rhymes in English, etc.

[0056] In the embodiments of the present invention, for the end words of sentences, a rhyming generation strategy is used. If the length of the end word after segmentation is 1, it is combined with the previous word to form a new word and then rhymed together.

[0057] The number of words to be rhymed extracted in this part can be determined according to the actual content of the sentences that need to rhyme, or the user can also set it according to experience or actual requirements. The present invention does not limit the number of words to be rhymed extracted.

[0058] 3. Speech generation model.

[0059] In the embodiments of the present invention, an end-to-end speech generation model is adopted to capture the pronunciation of the input text of words to be rhymed, obtain the corresponding vocalization, and form a speech signal.

[0060] Exemplarily, a multi-lingual speech generation model based on Tacotron can be adopted to directly synthesize the corresponding pronunciation through the multi-lingual text input in UTF-8 format. The speech generation model can be trained using a publicly available dataset or can also be trained using the company's internal private multi-lingual corpora such as English and Chinese.

[0061] As Figure 4 shown, the structure of the multi-lingual TTS model (i.e., the multi-lingual speech generation model based on Tacotron mentioned above) is shown.

[0062] 4. Rhyming pair generation model.

[0063] In the embodiment of the present invention, the rhyme pair generation model is used to encode and decode the input speech signal to obtain a new speech signal (ie, pronunciation) of the word to be rhymed; a is an integer greater than 1; illustratively, a=5 can be set.

[0064] Considering the short length of rhyme pairs, the rhyme pair generation model can be adopted as follows Figure 5 The Transformer structured multilingual rhyme pair generation model shown in the figure directly generates the corresponding multilingual rhyme pair pronunciations. Its input data uses the speech signal generated by the speech generation model, and the output directly generates multiple new speech signals of the words to be rhymed. This part can be understood as a speech-speech generation process. By directly modeling the speech signal, the rhythm information can be better captured. The multiple new speech signals can be of the same language or different languages. Since the generated speech signal is a speech signal, the language can be ignored at this time.

[0065] The training data can be collected from rhyming pairs in multiple languages through various rap lyrics, and then sent to the model for training to acquire the ability to capture the corresponding rhyming rules.

[0066] 5. Speech recognition model.

[0067] In an embodiment of the present invention, an end-to-end speech recognition model can be used to restore each new speech signal of the rhyming word, and each new speech signal of the rhyming word generates b words closest to it, generating a×b candidate words in total. The candidate words here can be of the same language or different languages, where a represents the total number of pronunciations of the rhyming words, and b is an integer greater than 1; illustratively, b=2 can be set.

[0068] In the embodiment of the present invention, the word closest to each new voice signal of the word to be rhymed can be judged accordingly according to the pronunciation rules. If the voice signal is "厉害", two corresponding Chinese words "厉害, 立害" etc. can be generated.

[0069] The speech recognition model is trained using labeled multilingual speech recognition data. Figure 6 As shown, the structure of the LAS speech recognition model is demonstrated.

[0070] 6. Rhyming word detector.

[0071] In an embodiment of the present invention, each candidate word is respectively combined with a word to be rhymed to form a rhyme pair, and a rhyme word discriminator is used to score all rhyme pairs, and the rhyme pair with the highest score is selected as the best rhyme team, and the candidate words in the best rhyme team are used as multilingual rhyme words for generating sentences; wherein, a number of rhyme words are extracted from songs in different languages, and some non-rhyme words and sensitive words are used as training data, and a rhyme word discriminator is trained using a neural network model.

[0072] Sections 3 to 6 above introduce a method for generating corresponding multilingual rhyming words for a single word to be rhymed. All words to be rhymed are processed by the above method to obtain multilingual rhyming words corresponding to all words to be rhymed.

[0073] The final multilingual rhyming words are output through the rhyming word discriminator. The core part (i.e. the aforementioned rhyming pair generation model) can effectively capture the pronunciation rules between rhyming words in different languages through the speech-to-speech generation method, and even realize the mixed rhyme generation of rap lyrics in different languages (such as Chinese-English rhyme).

[0074] In the embodiment of the present invention, multilingual rhyming words are generated at this stage. Multilingual rhyming words are relative to the aforementioned extracted rhyming words, that is, the rhyming words screened out by the rhyming word discriminator and the aforementioned extracted rhyming words can be in different languages; for example, the extracted words to be rhymed are in Chinese, but English rhyming words can still be used to form rhyming pairs with them. Taking the word to be rhymed as "揩油" as an example, the rhyming words screened out can be "Hello", and a rhyming pair is formed by "揩油-Hello"; in the later stage of the lyrics text generation process, the main language is identified as Chinese, and it will still be generated in Chinese, but other languages (for example, English as an example here) can be mixed in the sentence. Therefore, the present invention realizes the generation of rhythmic lyrics in multiple languages.

[0075] Those skilled in the art can understand that the sensitive words in the training data mainly refer to some illegal words or advertising words.

[0076] 2. Generation of rhythmic lyrics text in multiple languages.

[0077] In the embodiment of the present invention, an autoregressive text generation model is constructed based on an autoregressive text generation algorithm, a reverse text generation strategy from back to front is adopted, lyrics are generated reversely starting from the rhyming words at the end of the sentence, and the phoneme information of the words is introduced in the generation process.

[0078] This stage corresponds to the above step 2, which mainly includes two parts: the first part is to use the bert pre-training model to generate the previous lyrics and the corresponding word vector representations of all multilingual rhyming words obtained in the previous stage; the second part is to identify the language information of the previous lyrics (through Figure 3 The main language of the lyrics is identified and processed, and the two types of word vector representations output by the bert pre-training model are combined, and the lyrics text is generated by the autoregressive text generation model. The preferred implementation of each part is as follows:

[0079] 1. Generate the word vector representations corresponding to the lyrics of the previous text and the multilingual rhyming words obtained in the previous stage. To further improve the rhythm of the lyrics, the phoneme information of each character is utilized during the generation process, and the obtained word vector representations carry phoneme information.

[0080] In the embodiments of the present invention, during the pre-training process, a plurality of phoneme symbols (e.g., 66) are introduced based on the mbert pre-training model to describe the pronunciations of all words, and the embedding (word vector) part is improved. The segmentid of the original NSP part is used to distinguish the text part and the pronunciation part, and the position vector is used to align the pronunciations corresponding to words. The training is fine-tuned using a small amount of multilingual data based on mbert to obtain the vector representations of phoneme symbols. During the training process, the word vectors and pronunciation vectors at the same position vector are masked simultaneously.

[0081] As Figure 7 shown, during the usage process, the final word vector representation is obtained by weighted averaging the word vectors in the lyrics of the previous text and the multilingual rhyming words and the vectors of the corresponding phonemes, and it is used as the word vector representation for the encoding and decoding ends of the generation model, thereby introducing the phoneme information of the text. As shown in the following formula:

[0082] E final [you] = αE word [you] + β(E speech [n] + E speech [I])

[0083] where α and β are adjustable hyperparameters. E final refers to the embedding that fuses the text and phoneme information of the word, which will be used as the input of the generation model, and E word refers to the embedding corresponding to the lyrics text of the previous text or the multilingual rhyming word text, and E speech refers to the embedding corresponding to the corresponding phoneme.

[0084] It should be noted that the text "you" and the related phonemes "n" and "I" substituted in the above formula are only examples and do not constitute limitations. In actual applications, the corresponding text and phoneme information are substituted according to the actual text content.

[0085] 2. As Figure 8 shown, the autoregressive text generation model adopts a reverse text generation strategy from back to front, starting from the multilingual rhyming word at the end of the sentence and generating the lyrics in reverse from right to left. The autoregressive text generation model can adopt an encoder-decoder structure based on the mbart multilingual pre-training model, and input the processed language information, as well as the word vector representations with phoneme information of the lyrics of the previous text and all multilingual rhyming words into the model to generate the final lyrics text.

[0086] In the embodiments of the present invention, the method for the mbart multilingual pre-training model to process language information is described in the introduction of the existing mbart multilingual pre-training model, so it will not be elaborated here. Processing the language information can be to complete the language information by actively adding language information symbols in the input and output. For example, in the completed language information, Chinese is zh_CN, and English is en_XX, etc.

[0087] The language information input in the autoregressive text generation model refers to the language of the lyrics text to be generated, that is, the main language of the generated lyrics text needs to be consistent with the input language information, but it can include text in other languages. Exemplary: If the current position of the sentence to be generated is the sixth line of lyrics in the song, then identify the first five lines of lyrics to determine the language information. If the identified language information is Chinese, then the language of the generated lyrics text is Chinese. Since the multilingual rhyming words obtained at the current stage can be in multiple languages, assuming it is in English, then the generated lyrics text can include Chinese and English. For example, the generated lyrics text is "You are saying hi".

[0088] It should be noted that the text content (such as "Hello", etc.) that appears in the attached drawings corresponding to the above models and the specific text content of the lyrics that appear in the text are all examples and do not constitute limitations. In applications, relevant text content can be input into the relevant models according to actual situations.

[0089] In addition, since the models and their training methods are not described in detail, it means that they can all be implemented conventionally, so it will not be elaborated here.

[0090] The above solutions in the embodiments of the present invention mainly obtain the following beneficial effects:

[0091] 1) The speech generation model is used to capture the pronunciation of words and then generate rhyming word pairs, which greatly improves the rhyming quality of the generated words. At the same time, an autoencoder model with the rhyme as the starting input is used in the generation process, which can generate more semantically coherent lyrics.

[0092] 2) Based on the mbart pre-training model, a multilingual autoregressive model with large parameters is used to generate lyrics. At the same time, the phoneme information of words is introduced in the generation process to make the quality of the generated lyrics higher.

[0093] 3) It supports the generation of multilingual lyrics, and at the same time uses multiple languages to train the rhyming word matching model, making the generated rhyming words more diverse.

[0094] Embodiment 2

[0095] The present invention also provides a multilingual rhyming lyrics generation system, which is mainly implemented based on the method provided in the foregoing embodiments, asFigure 9 As shown, the system mainly includes:

[0096] A multilingual rhyming word generation module, which is used to extract several words to be rhymed from the sentences that need to rhyme in the previous lyrics; for each word to be rhymed, a voice signal of the word to be rhymed is generated through voice generation technology, and multiple new voice signals of the word to be rhymed are generated through rhyming pair generation technology, and then multiple different candidate words corresponding to the multiple new voice signals of the word to be rhymed are generated through voice recognition technology. Each candidate word and the word to be rhymed form a rhyming team, and the candidate word corresponding to the best rhyming team is selected as the multilingual rhyming word used to generate sentences;

[0097] A multilingual rhymed lyrics text generation module, which is used to generate lyrics text by using an autoregressive text generation algorithm according to all the selected multilingual rhyming words, the previous lyrics, and the recognized language information of the previous lyrics.

[0098] Those skilled in the art can clearly understand that, for the convenience and simplicity of description, only the above division of each functional module is used as an example. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the system is divided into different functional modules to complete all or part of the functions described above.

[0099] The relevant technical details involved in each module of the above system have been introduced in detail in the previous method embodiments, so they will not be repeated here.

[0100] Embodiment 3

[0101] The present invention also provides a processing device, as Figure 10 shown, which mainly includes: one or more processors; a memory for storing one or more programs; wherein, when the one or more programs are executed by the one or more processors, the one or more processors implement the method provided in the foregoing embodiment.

[0102] Furthermore, the processing device further includes at least one input device and at least one output device; in the processing device, the processor, the memory, the input device, and the output device are connected through a bus.

[0103] In the embodiments of the present invention, the specific types of the memory, the input device, and the output device are not limited; for example:

[0104] The input device can be a touch screen, an image acquisition device, a physical button, or a mouse, etc.;

[0105] The output device can be a display terminal;

[0106] The memory can be a Random Access Memory (RAM) or a non-volatile memory, such as a disk memory.

[0107] Embodiment 4

[0108] The present invention also provides a readable storage medium storing a computer program, which implements the method provided in the foregoing embodiments when executed by a processor.

[0109] In the embodiments of the present invention, as a computer-readable storage medium, the readable storage medium can be disposed in the foregoing processing device, for example, as the memory in the processing device. In addition, the readable storage medium can also be various media that can store program codes, such as a USB flash drive, a mobile hard disk, a Read-Only Memory (ROM), a magnetic disk, or an optical disc.

[0110] The above are only the preferred specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A method for generating multilingual rhyming lyrics, characterized in that, Including: Extracting several words to be rhymed from the sentences that need to rhyme in the previous lyrics; For each word to be rhymed, generating a speech signal of the word to be rhymed through speech generation technology, generating multiple new speech signals of the word to be rhymed through rhyme pair generation technology, and then generating multiple different candidate words corresponding to the multiple new speech signals of the word to be rhymed through speech recognition technology. Combining each candidate word with the word to be rhymed to form a rhyming pair, and screening out the candidate word corresponding to the best rhyming pair as the multi-lingual rhyming word used to generate sentences; According to all the screened multi-lingual rhyming words, the previous lyrics, and the identified language information of the previous lyrics, using an autoregressive text generation algorithm to generate the lyrics text, including: using the mbert pre-trained model to generate the word vector representations with phoneme information corresponding to the previous lyrics and all multi-lingual rhyming words; wherein, the word vectors of the words in the previous lyrics and the multi-lingual rhyming words and the vectors of the corresponding phonemes are weighted and averaged to obtain the word vector representations with phoneme information; processing the identified language information of the previous lyrics, and combining the two types of word vector representations with phoneme information output by the mbert pre-trained model, through the autoregressive text generation algorithm, adopting a reverse text generation strategy from back to front, starting from the multi-lingual rhyming word at the end of the sentence to generate the lyrics in reverse, and obtaining the corresponding lyrics text.

2. The method for generating multilingual rhythmic lyrics according to claim 1, wherein, The sentences that need to rhyme are randomly selected from the previous lyrics; The selection probability and the distance between the position where the current sentence to be generated is located satisfy the following relationship: wherein, dis(i) and dis(j) respectively represent the distances between sentence i and sentence j in the previous lyrics and the position where the current sentence to be generated is located; α is a hyperparameter.

3. A method for generating multilingual rhyming lyrics according to claim 1, characterized in that, Extracting the words to be rhymed and generating the speech signal of the word to be rhymed through speech generation technology includes: Extracting different replacement words as the words to be rhymed from the segmented sentences that need to rhyme according to different rhyming methods; for the end word of the sentence, using a rhyme generation strategy. If the length of the segmented end word is 1, then form a new word together with the previous word and then rhyme together; Adopting an end-to-end speech generation model to capture the pronunciation of the input text of the word to be rhymed, obtaining the corresponding pronunciation, and forming a speech signal; the speech generation model is trained using multi-lingual corpora.

4. A method for generating multilingual rhyming lyrics according to claim 1, characterized in that The generating multiple new speech signals of the word to be rhymed through rhyme pair generation technology and then generating multiple different candidate words corresponding to the multiple new speech signals of the word to be rhymed through speech recognition technology includes: Using a rhyme pair generation model to encode and decode the input speech signal to obtain corresponding a pronunciations of rhyming words; a is an integer greater than 1; the rhyme pair generation model is trained through pre-collected multi-lingual rhyme pairs, and the trained rhyme pair generation model has the ability to capture corresponding rhyming rules.

5. A method for generating rhyming lyrics in multiple languages according to claim 1, characterized in that, The generating multiple different candidate words corresponding to the multiple new speech signals of the word to be rhymed through speech recognition technology includes: An end-to-end speech recognition model is used to restore each new speech signal of the rhyming word. The b words closest to each new speech signal of the rhyming word are selected, and a total of a×b candidate words are generated, where a is an integer greater than 1, representing the total number of pronunciations of the rhyming word; b is an integer greater than 1; the speech recognition model is trained using labeled multi-lingual speech recognition data.

6. A method for generating rhyming lyrics in multiple languages according to claim 1, characterized in that, The method of forming a rhyming team with each candidate word and the rhyming word, and screening out the candidate word corresponding to the best rhyming team as the multi-lingual rhyming word for generating a sentence includes: Forming a rhyming pair with each candidate word and the rhyming word respectively, using a rhyming word discriminator to score all rhyming pairs, and the rhyming pair with the highest score is the best rhyming team, and the candidate word in the best rhyming team is used as the multi-lingual rhyming word for generating a sentence; Among them, several rhyming words are extracted from songs in different languages, and at the same time, some non-rhyming words and sensitive words are used as training data, and a rhyming word discriminator is trained using a neural network model.

7. A multilingual rhythmic lyric generation system, characterized in that, Implemented based on the method described in any one of claims 1 to 6, the system includes: A multi-lingual rhyming word generation module, which is used to extract several rhyming words from the sentences that need to rhyme in the previous lyrics; for each rhyming word, generate a speech signal of the rhyming word through speech generation technology, and generate multiple new speech signals of the rhyming word through rhyming pair generation technology, and then generate multiple different candidate words corresponding to the multiple new speech signals of the rhyming word through speech recognition technology, form a rhyming team with each candidate word and the rhyming word, and screen out the candidate word corresponding to the best rhyming team as the multi-lingual rhyming word for generating a sentence; A multi-lingual rhyming lyrics text generation module, which is used to generate lyrics text according to all the selected multi-lingual rhyming words, the previous lyrics, and the identified language information of the previous lyrics by using an autoregressive text generation algorithm.

8. A processing device, characterized in that, Including: One or more processors; A memory for storing one or more programs; Among them, when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any one of claims 1 to 6.

9. A readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, the method described in any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Information management system

    US20040133559A1