A method, apparatus, electronic device, and storage medium for song imitation
By segmenting the source song into source lyrics and source melody, generating imitation lyrics using visual words and generating imitation melody through autoregression, the problem of low efficiency in song imitation in existing technologies is solved, and similar imitation songs can be generated quickly.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- NETEASE (HANGZHOU) NETWORK CO LTD
- Filing Date
- 2023-03-29
- Publication Date
- 2026-05-26
AI Technical Summary
Current methods of song imitation mainly rely on musicians manually imitating the composition and lyrics, resulting in low imitation efficiency. Furthermore, existing automated technologies struggle to generate imitation songs that are similar to the original songs.
The source song is divided into source lyrics and source melody. Imitation lyrics are generated by visualizing words, and imitation melody is generated using an autoregressive model. Finally, the two parts are merged to obtain the target song.
It enables the intelligent generation of imitation songs similar to the source songs, greatly shortening the creation time for musicians and improving the efficiency of song imitation.
Smart Images

Figure CN116453489B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to a song imitation method, apparatus, electronic device, and storage medium. Background Technology
[0002] A song is an art form that combines lyrics and melody. It is an art form because it is both a language of emotion and an art of sound. A good song can provide people with a comprehensive auditory experience, expressing inner emotions and resonating with listeners.
[0003] Song imitation refers to creating a song with similar lyrics and melody by imitating the lyrics and melody of an existing song. Current song imitation methods primarily rely on musicians imitating the composition and lyrics, with relatively little automated song imitation technology. Therefore, this method, which depends on musicians imitating the composition and lyrics, requires a significant amount of creative time for musicians to produce a similar song, resulting in low efficiency. Summary of the Invention
[0004] In view of this, this application provides a song imitation method, apparatus, electronic device, and storage medium. First, a source song is acquired and segmented into source lyrics and a source melody. Then, imitation lyrics are generated based on visual words in the source lyrics, where visual words represent the images expressed in the source lyrics. Based on the source melody, an imitation melody is generated through autoregression. Finally, the imitation lyrics and the imitation melody are merged to obtain the target song. Thus, this application, by segmenting the source song into source lyrics and a source melody, generating imitation lyrics based on visual words in the source lyrics, generating an imitation melody through autoregression based on the source melody, and then merging the imitation lyrics and the imitation melody to obtain the target song, can quickly obtain imitation songs similar to the source song, achieving intelligent generation of imitation songs, greatly shortening the creator's creation time, and improving the efficiency of song imitation.
[0005] The first aspect of this application provides a song imitation method, which includes:
[0006] Obtain the source song and split it into source lyrics and source melody.
[0007] Generate paraphrased lyrics based on visual words in the source lyrics. Visual words are used to represent the images expressed in the source lyrics.
[0008] Based on the source melody, an autoregressive imitation melody is generated from the source melody.
[0009] By merging the imitated lyrics and the imitated melody, the target song is obtained.
[0010] A second aspect of this application provides a song imitation device, comprising:
[0011] The acquisition unit is used to acquire the source song.
[0012] The processing unit is used to segment the source song into source lyrics and source melody.
[0013] The generation unit is used to generate imitation lyrics based on the visual words in the source lyrics. The visual words are used to represent the images expressed in the source lyrics.
[0014] The generation unit is also used to autoregressively generate a copy of the source melody based on the source melody.
[0015] The processing unit is also used to merge the imitated lyrics and the imitated melody to obtain the target song.
[0016] A third aspect of this application provides an electronic device, including: a memory and a processor, wherein the memory and the processor are coupled.
[0017] The memory is used to store one or more computer instructions.
[0018] The processor is used to execute one or more computer instructions to implement the song imitation method described in the first aspect above.
[0019] A fourth aspect of this application also provides a computer-readable storage medium storing one or more computer instructions, characterized in that the instructions are executed by a processor to implement a song imitation method as described in any of the above technical solutions.
[0020] Compared with the prior art, the embodiments of this application have the following advantages:
[0021] The song imitation method provided in this application first obtains a source song and divides it into source lyrics and a source melody. Then, it generates imitation lyrics based on visual words in the source lyrics, where visual words represent the images expressed in the source lyrics. Finally, it generates an imitation melody based on the source melody using autoregression. The imitation lyrics and melody are then merged to obtain the target song. Thus, this application, by dividing the source song into source lyrics and a source melody, generating imitation lyrics based on visual words in the source lyrics, generating an imitation melody based on the source melody using autoregression, and then merging the imitation lyrics and melody to obtain the target song, can quickly obtain imitation songs similar to the source song. This achieves intelligent generation of imitation songs, greatly shortening the creator's creation time and improving the efficiency of song imitation. Attached Figure Description
[0022] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0023] Figure 1 A flowchart illustrating a song imitation method provided in an embodiment of this application;
[0024] Figure 2 This application provides a schematic diagram of a process for generating source lyrics based on visual words in a parody song.
[0025] Figure 3 This is a schematic flowchart of another lyric imitation method provided in an embodiment of this application;
[0026] Figure 4 A flowchart illustrating the process of determining visual words in source lyrics, provided in an embodiment of this application;
[0027] Figure 5 A schematic diagram of a pre-trained language model provided in an embodiment of this application;
[0028] Figure 6 A schematic diagram illustrating the initial melody imitation model training process provided in this application embodiment;
[0029] Figure 7 An audio pre-trained model architecture diagram based on a convolutional encoder and a transformer encoder is provided for embodiments of this application;
[0030] Figure 8 This is a schematic diagram of a song imitation method apparatus provided in an embodiment of this application;
[0031] Figure 9 This is a schematic diagram of an electronic device provided in an embodiment of this application. Detailed Implementation
[0032] This application provides a method, apparatus, electronic device, and storage medium for song imitation. First, a source song is acquired and segmented into source lyrics and a source melody. Then, imitation lyrics are generated based on visual imagery words in the source lyrics, where visual imagery words represent the images expressed in the source lyrics. Based on the source melody, an imitation melody is generated through autoregression. Finally, the imitation lyrics and the imitation melody are merged to obtain the target song. Thus, this application, by segmenting the source song into source lyrics and a source melody, generating imitation lyrics based on visual imagery words, generating an imitation melody through autoregression based on the source melody, and then merging the imitation lyrics and the imitation melody to obtain the target song, can quickly obtain imitation songs similar to the source song, achieving intelligent generation of imitation songs from source songs, greatly shortening the creation time for musicians, and improving the efficiency of song imitation.
[0033] To enable those skilled in the art to better understand the technical solutions of this application, the application will be clearly and completely described below with reference to the accompanying drawings of the embodiments. However, this application can be implemented in many other ways different from those described above. Therefore, based on the embodiments provided in this application, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this application.
[0034] It should be noted that the terms "first," "second," "third," etc., in the claims, specification, and drawings of this application are used to distinguish similar objects and are not used to describe a specific order or sequence. Such data are interchangeable where appropriate so that the embodiments of this application described herein can be implemented in a sequence other than that illustrated or described herein. Furthermore, the terms "comprising," "having," and their variations are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to these processes, methods, products, or apparatuses.
[0035] Before detailing the implementation methods of this application, the relevant technologies will be further introduced first.
[0036] Song imitation refers to creating a song with similar lyrics and melody by imitating the lyrics and melody of an existing song. For example, in some applications, imitating popular songs can generate similar songs, thus improving the efficiency of popular song production and possessing significant economic value.
[0037] Existing methods for creating lyrics by imitating a source song primarily rely on musicians imitating the source song's melody and lyrics to produce a similar imitation. Additionally, some existing technologies utilize automated techniques to automatically generate lyrics, mainly using lyric generation models. Specifically, users select any number of keywords and input them into the lyric generation model, which outputs lyrics containing those keywords. However, when the relevance between the selected keywords is low or the number of keywords is small, the lyrics generated by the model often result in disjointed lyrics, with each line based on a different keyword. This leads to severely fragmented context and a disorganized, incoherent song that fails to convey the complete meaning of the song. Furthermore, lyrics generated from keywords lack a source song for reference, resulting in primarily original lyrics rather than imitations of the source song. Automation techniques for imitating the melody of a source song are relatively limited. Therefore, to achieve the imitation of a source song, the current method mainly relies on musicians to imitate the composition and lyrics. This requires musicians to spend a lot of creative time to obtain an imitation song, resulting in low efficiency.
[0038] To address the aforementioned problems, this application provides a song imitation method, apparatus, electronic device, and storage medium. First, a source song is acquired and segmented into source lyrics and a source melody. Then, imitation lyrics are generated based on visual imagery words in the source lyrics, where visual imagery words represent the images depicted in the source lyrics. Based on the source melody, an imitation melody is generated through autoregression. Finally, the imitation lyrics and the imitation melody are merged to obtain the target song. Thus, this application, by segmenting the source song into source lyrics and a source melody, generating imitation lyrics based on visual imagery words, generating an imitation melody through autoregression based on the source melody, and then merging the imitation lyrics and the imitation melody to obtain the target song, can quickly obtain imitation songs similar to the source song, achieving intelligent generation of imitation songs from source songs, greatly shortening the creation time for musicians, and improving the efficiency of song imitation.
[0039] The methods, apparatus, terminals, and computer-readable storage media described in this application will be further described in detail below with reference to specific embodiments and accompanying drawings.
[0040] Figure 1This is a flowchart illustrating a song imitation method provided in an embodiment of this application. It should be noted that the steps shown in this flowchart can be executed in a computer system such as a set of computer-executable instructions. Furthermore, in some cases, the released steps can be executed in a logical order different from that shown in the flowchart.
[0041] The execution entity of this method can be a terminal device or a server. The terminal device can be a desktop computer, laptop computer, tablet computer, mobile phone, or other electronic device; this application does not specifically limit this. The server is used to provide background services for the client of the application in the terminal device. For example, the server can be the background server of the aforementioned application. The server can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms; this application does not specifically limit this.
[0042] like Figure 1 As shown, the method for imitating this song includes the following steps:
[0043] 101. Obtain the source song and divide it into source lyrics and source melody.
[0044] A source song is a reference song used to generate a song imitation. Source songs can be obtained through various channels such as the internet and music software. A source song consists of source lyrics and a source melody. Source lyrics refer to the text portion of the source song, while the source melody refers to the audio portion, which may include vocals and accompaniment. After obtaining the source song, it is divided into source lyrics and source melody. The source lyrics are used to generate the lyrics, and the source melody is used to generate the melody.
[0045] In one implementation, the source song is segmented into source lyrics and source melody. Specifically, the source lyrics can be obtained by crawling the lyrics. Crawling refers to simulating a client sending network requests, receiving request responses, and automatically collecting data from the Internet or other media according to certain rules.
[0046] The original melody of a song can be obtained using Spleeter. Spleeter is an open-source audio track splitting software from the music streaming company Deezer. Spleeter supports common audio formats such as mp3, wav, and ogg. For example, inputting a source song into Spleeter will yield its corresponding original melody. Spleeter can also separate the vocals and instrumental accompaniment from the original melody. In one implementation, the vocals and instrumental accompaniment of the original melody can be separated, and then the vocals and accompaniment can be separately transcribed and merged to obtain a transcribed version of the original melody. In another implementation, the vocals and instrumental accompaniment of the original melody are not separated; the original melody is directly transcribed to obtain a transcribed version of the original melody.
[0047] 102. Generate paraphrased lyrics based on the visual words in the source lyrics. The visual words are used to represent the images expressed in the source lyrics.
[0048] A complete song's lyrics often evoke one or more vivid images in the listener's mind, creating a strong sense of imagery. A song lacking imagery feels fragmented and disjointed. Therefore, the imagery conveyed by the lyrics is crucial for a complete song. The imagery conveyed by lyrics can be represented by a set of descriptive words. These descriptive words are the vocabulary used to describe the images expressed in the lyrics. The descriptive words in the source lyrics are used to represent the images expressed in the source lyrics.
[0049] For example, when we see the words "broken bridge," "snow," "oil-paper umbrella," and "girl," we might conjure up an image similar to the following: a beautiful girl holding an oil-paper umbrella, walking on a broken bridge in the snow. These words, which can form a complete picture, are called visual words.
[0050] This technology generates lyrics that mimic the imagery of the original lyrics, creating a similar visual experience. However, current lyric generation technologies rely on users selecting any number of keywords and inputting them into the model. The model then outputs lyrics containing those keywords. These keywords are key words representing the lyrics, not words describing the imagery they evoke. Therefore, keywords alone cannot conjure a complete picture in the user's mind. Furthermore, when the keywords selected by the user are poorly correlated or few in number, the generated lyrics often result in fragmented contexts, with each line based on a different keyword. This leads to disorganized and chaotic lyrics that fail to capture the intended imagery of the song.
[0051] In this application, since visual words can represent the imagery expressed by lyrics, the generated imitation lyrics of the source song are similar to the imagery expressed by the source lyrics. This solves the problem in the prior art where the method of generating lyrics by keywords results in a fragmentation of the lyrics' context and an inability to express the imagery of the lyrics.
[0052] 103. Based on the source melody, generate an autoregressive imitation melody of the source melody.
[0053] Autoregression is the process of using the currently predicted characters to predict the next information. In the audio prediction stage, given a segment of audio, the system predicts the next segment of audio information based on the current audio.
[0054] In one implementation, the source melody can be copied using a Transformer model, which is an autoregressive model. Specifically, to copy the source melody, firstly, a target sub-melody can be selected from the source melody; the target sub-melody refers to a subset of the sub-melody from the source melody.
[0055] For example, in one implementation, the source melody can be arbitrarily divided into multiple sub-melody sequences of varying lengths. For instance, a source melody can be divided into four sub-melody sequences: X1, X2, X3, and X4. X1, X2, X3, and X4 can be audio sequences of equal or different lengths. Generally, when selecting a target sub-melody, the melody at the beginning of the source melody can be chosen as the target sub-melody to generate the imitation sub-melody. For example, sub-melody X1 can be selected as the imitation sub-melody after the target sub-melody prediction.
[0056] In another implementation, a melody from the middle of the source melody or any other part can be selected as the target sub-melody to generate the imitation sub-melody. For example, sub-melodies X2, X3, and X4 can also be selected as target sub-melodies to generate the imitation sub-melody.
[0057] After selecting the target sub-melody, the target sub-melody is input into the target melody imitation model, which generates multiple imitation sub-melodies through autoregression. Then, the target sub-melody and multiple imitation sub-melodies are combined to form the imitation melody of the source song, and the imitation melody of the source song is output through the target melody imitation model.
[0058] Understandably, the target melody imitation model can output one imitation melody of the source song, or it can output multiple imitation melody of the source song.
[0059] For example, the target sub-melody is input into the target melody imitation model, and multiple imitation sub-melodies are generated through autoregression. Specifically, sub-melody X1 is selected as the target sub-melody, and then the target sub-melody X1 is input into the target melody imitation model. The imitation sub-melody X2' can be obtained through the target melody imitation model. X2' represents the imitation sub-melody of sub-melody X2. There can be one or more imitation sub-melodies of X2.
[0060] In one implementation, when there is only one imitation sub-melody for X2, X3' is generated based on X1 and X2', where X3' represents the imitation sub-melody of sub-melody X3. Then, X4' is generated based on X1, X2', and X3', where X4' represents the imitation sub-melody of sub-melody X4. In this way, three imitation sub-melody X2', X3', and X4' are generated in total through this autoregressive method. Then, X1, X2', X3', and X4' are combined to obtain the imitation melody of the source song. Finally, the target melody imitation model outputs an imitation melody of the source song.
[0061] In another implementation, when multiple imitation sub-melody X1 are input into the target melody imitation model to generate imitation sub-melody X2, for example, when the number of imitation sub-melody X2 is three, the corresponding imitation sub-melody X2 are X2', X2"', and X2"'. Assuming that the number of imitation sub-melody of the source song to be output is two, the two imitation sub-melody with the highest similarity to X2', X2"', and X2"' can be selected to generate subsequent imitation sub-melody.
[0062] For example, the similarity scores of the imitated sub-melodies X2', X2″, and X2″' with X2 are 0.9, 0.8, and 0.7, respectively. The two imitated sub-melodies with the highest similarity scores, X2' and X2″, are selected to generate subsequent imitated sub-melodies. X3' and X3″ are generated from X1 and X2', and X3″' and X3″' are generated from X1 and X2″. Based on the similarity scores of X3', X3″, X3″', and X3″' with X3, the two imitated sub-melodies with the highest similarity scores are selected to generate subsequent imitated sub-melodies, and so on. Finally, the target melody imitation model outputs two imitated melodies of the source song.
[0063] 104. Combine the imitated lyrics and imitated melody to obtain the target song.
[0064] This step involves merging the imitated lyrics and melody after obtaining them to obtain the target song. The target song is the imitated version of the source song.
[0065] In one implementation, for example, when the number of imitated lyrics is one and the number of imitated melodies is one, one imitated lyric and one imitated melody are merged to obtain the target song. When the number of imitated lyrics is multiple and the number of imitated melodies is multiple, any one of the multiple imitated lyrics and any one of the multiple imitated melodies is merged to obtain the target song.
[0066] In the technical solution provided in the above-described embodiments of this application, firstly, a source song is obtained and segmented into source lyrics and a source melody. Then, imitation lyrics are generated based on visual words in the source lyrics, where visual words represent the images expressed by the source lyrics. Based on the source melody, an imitation melody is generated through autoregression. Finally, the imitation lyrics and the imitation melody are merged to obtain the target song. Thus, this application, by segmenting the source song into source lyrics and a source melody, generating imitation lyrics based on visual words in the source lyrics, generating an imitation melody through autoregression based on the source melody, and then merging the imitation lyrics and the imitation melody to obtain the target song, can quickly obtain imitation songs similar to the source song, achieving intelligent generation of imitation songs from source songs, greatly shortening the creator's creation time and improving the efficiency of song imitation.
[0067] The following is combined Figure 2 and Figure 3 The method of generating imitation lyrics based on the image words in the embodiments of this application will be further explained and described.
[0068] Figure 2 This is a schematic diagram of a process for generating source lyrics based on visual words, provided in an embodiment of this application. Figure 3 This is a schematic flowchart of another lyric imitation method provided in the embodiments of this application.
[0069] 201. Obtain N visual words from the source lyrics.
[0070] In one implementation, N visual words of the source lyrics are obtained. The source lyrics can be divided into multiple lyric segments. A lyric segment refers to a portion of the lyrics in the source lyrics. There are several ways to divide the source lyrics into multiple lyric segments. It can be divided according to the paragraphs of the source lyrics. For example, if the source lyrics have three paragraphs, the source lyrics can be divided into three lyric segments. It can also be divided according to the main lyrics and the chorus of the source lyrics, dividing the source lyrics into two lyric segments: a main lyric segment and a chorus segment. Alternatively, the source lyrics can be divided into multiple lyric segments of equal or different lengths in other ways. The specific method is not limited.
[0071] After dividing the source lyrics into multiple lyric segments, the multiple lyric segments are input into a pre-trained language model. The pre-trained language model converts each lyric segment into a segment vector corresponding to each lyric segment. The segment vector refers to the vector corresponding to the lyric segment.
[0072] Furthermore, each lyric segment is segmented into words, that is, each line of lyrics in each lyric segment is divided into multiple word groups through word segmentation.
[0073] Understandably, since each Chinese character in a text is written continuously, we need to use specific methods to obtain each phrase, a method called word segmentation.
[0074] For example, in one implementation, word segmentation of each lyric segment can be performed using the jieba word segmentation tool. The jieba tool is primarily used for Chinese word segmentation. Its principle is to utilize a Chinese dictionary to determine the probability of association between Chinese characters. Characters with higher probabilities are grouped into word groups, forming the word segmentation result.
[0075] Jieba, a word segmentation tool, offers three segmentation modes. The first is the precise mode, which accurately segments a text into a number of Chinese words. These words are then combined to accurately reconstruct the original text, eliminating redundant words. The second is the full mode, which scans all possible words in a text. A text can be segmented into different words depending on the perspective. In full mode, the recombined information from the segmented words may contain redundancy, no longer resembling the original text. The third mode is the search engine mode, which further segments long words based on the precise mode to improve recall, making it suitable for search engine segmentation.
[0076] After obtaining the word segments corresponding to each lyric segment, the word segments are input into a pre-trained language model. The pre-trained language model then generates word vectors for each segmented lyric segment. The semantic similarity between the word vectors of each lyric segment and the segment vectors of the lyric segment is calculated. Based on the semantic similarity between the word vectors of each lyric segment and the segment vectors of the lyric segment, N visual words representing the source lyrics are determined.
[0077] In one implementation, the pre-trained language model processes the text corpus at the character level, obtaining segment vectors and word vectors for each lyric segment. For example, if a lyric segment contains 14 characters, the character vectors of each character are summed and averaged, and the resulting average vector is used as the segment vector for that lyric segment. Similarly, for instance, if the lyric segment is segmented into six words, the character vectors of each word are summed and averaged, and the resulting average vector is used as the word vector for each word. Then, the semantic similarity between the word vectors and the segment vectors is calculated.
[0078] To calculate the semantic similarity between the word vectors of a lyric segment and the segment vectors of the lyric segment, we can use cosine similarity. A higher semantic similarity indicates that the word better represents the entire sentence; that is, a larger cosine similarity indicates higher semantic similarity. Cosine similarity, also known as cosine similarity, measures the similarity between two vectors by measuring the cosine of the angle between them. Two vectors with identical directions have a cosine similarity of 1, while two vectors facing each other have a similarity of -1.
[0079] For example, when the lyrics segment is "I rush towards you, you are the stars and the sea," it is segmented into words, which can be divided into "I, towards you, rush towards, you, are, the stars and the sea." The lyrics segment is then input into a pre-trained language model, which obtains the segment vector. For instance, the segment vector is a 768-dimensional vector [1, 2, 1.3, 1.5…], and the segment vector corresponding to the word "rush towards" is a 768-dimensional vector [0.9, 0.8, 1.2, 1.4…]. After obtaining the segment vector of "I rush towards you, you are the stars and the sea" and the word vector corresponding to "rush towards," the semantic similarity between the two vectors is calculated using cosine similarity. The higher the semantic similarity between the two vectors, the more representative the segment is of the entire lyrics segment.
[0080] The N visual words of the source lyrics are determined based on the semantic similarity between the word vectors of each lyric segment and the segment vectors of the lyric segment. Specifically, the words corresponding to the top S word vectors with the highest semantic similarity to the segment vectors of the lyric segment are determined as the visual words of the lyric segment, where S is a positive integer greater than or equal to 1.
[0081] Based on the semantic similarity of the visual words in each lyric segment, the visual words in each lyric segment are sorted, and the N visual words with the highest semantic similarity are obtained. These N visual words are then determined as the N visual words of the source lyrics.
[0082] In one implementation, for example, the source lyrics are divided into two lyric segments: main lyrics and chorus lyrics. The N visual words of the source lyrics are determined based on the word vectors of the main lyrics segment and the chorus lyrics segment and the segment vectors of the main lyrics and the chorus lyrics.
[0083] Specifically, in combination Figure 4 The following content explains the method for determining the source lyrics of the visual text. Figure 4 This is a flowchart illustrating a method for determining the visual words of source lyrics provided in an embodiment of this application.
[0084] First, the source lyrics are divided into two segments: main lyrics and chorus lyrics. These segments are then input into a pre-trained language model, which generates segment vectors for each segment. Next, the jieba word segmentation tool is used to segment the main lyrics and chorus lyrics into individual words. These individual words are then input into the pre-trained language model, which generates word vectors for each word segment. Then, the semantic similarity between each word segment vector of the main lyrics segment and the segment vector of the main lyrics segment is calculated. The word segments corresponding to the word vectors of the top S main lyrics segments with the highest semantic similarity to the segment vectors of the main lyrics segment are selected as the visual words of the main lyrics segment. The semantic similarity between each word segment vector of the chorus lyrics segment and the segment vector of the chorus lyrics segment is calculated. The word segments corresponding to the word vectors of the top S chorus lyrics segments with the highest semantic similarity to the segment vectors of the chorus lyrics segment are selected as the visual words of the chorus lyrics segment. S is a positive integer greater than or equal to 1.
[0085] After obtaining the visual words for the main lyrics segment and the visual words for the chorus lyrics segment, the visual words for the main lyrics and the visual words for the chorus lyrics are sorted according to semantic similarity. The N visual words with the highest semantic similarity are obtained and determined as the N visual words of the source lyrics, where N is a positive integer greater than or equal to 1.
[0086] In one implementation, no more than five visual words can be selected as visual words for both the main lyrics and the chorus. For example, the selected visual words for the main lyrics are: visual word 1 (semantic similarity 0.9), visual word 2 (semantic similarity 0.7), visual word 3 (semantic similarity 0.6), and visual word 4 (semantic similarity 0.8); the selected visual words for the chorus are: visual word 5 (semantic similarity 0.45), visual word 6 (semantic similarity 0.75), visual word 7 (semantic similarity 0.65), and visual word 8 (semantic similarity 0.3). Visual words 1, 2, 3, 4, 5, 6, 7, and 8 are ranked according to their semantic similarity. In one implementation, the top five visual words with the highest semantic similarity can be selected as the visual words for the source lyrics, i.e., visual words 1, 4, 6, 2, and 7 are selected as the visual words for the source lyrics. When selecting the number of visual words for each lyric segment in the source lyrics, the selection can be based on the length of the source lyrics; no limit is imposed here.
[0087] 202. Input N visual words into the lyrics generation model, and generate reference lyrics containing the visual words through the lyrics generation model.
[0088] After determining the N visual words for the source lyrics, these N visual words are input into the lyrics generation model. The model can then output reference lyrics, which are lyrics containing the visual words. The model can generate one or multiple reference lyrics.
[0089] 203. Determine the lyrics information of the source lyrics.
[0090] Determining the lyric information of the source lyrics includes determining the word structure, style, and mood of the source lyrics.
[0091] The word structure of source lyrics refers to the format of the source lyrics. For example, it may include the number of paragraphs, the number of sentences in each paragraph, and the number of words in each sentence.
[0092] The style of the source lyrics refers to the musical genre of the source lyrics, for example, including rock, pop, classical, folk, rap, etc.
[0093] The emotion in the source lyrics refers to the emotion that the source song wants to express, such as happiness, joy, sadness, grief, and longing.
[0094] 204. The first corpus and the second corpus are defined as parallel corpora. The first corpus includes N visual words, lyric information of the source lyrics, and reference lyrics. The second corpus includes the source lyrics.
[0095] Parallel corpora refer to texts that are aligned sentence by sentence with another text to form parallel texts. In this application, the first and second corpora are aligned to form parallel corpora, and standard word pairs can be formed through corpus alignment.
[0096] In one implementation, N visual words from the source lyrics, the lyric information of the source lyrics, and reference lyrics are determined as the first corpus, the source lyrics are determined as the second corpus, and the first and second corpora are determined as parallel corpora, that is, the N visual words from the source lyrics, the lyric information of the source lyrics, and reference lyrics are aligned with the source lyrics. For example, constructing (N visual words from the source lyrics)... Lyrics information of the source lyrics Data pairs in the format of (reference lyrics, source lyrics).
[0097] 205. Input the parallel corpus into the initial lyrics imitation model and determine the initial lyrics imitation model as the target lyrics imitation model.
[0098] Parallel corpora are input into the initial lyrics imitation model for training. In one implementation, the maximum sequence length of the parallel corpora is set to 512. If the sequence length of the parallel corpora exceeds 512, the corpora exceeding 512 are pruned, and the corpora shorter than 512 are padded to form an N*512 matrix for training. It is understood that the maximum sequence length of the parallel corpora can also be set to 1024 or other values; the specific value is not limited. This processing of the parallel corpora facilitates rapid calculation of the initial lyrics imitation and more effectively trains the model.
[0099] The initial lyrics imitation model generates predicted lyrics based on parallel corpora, calculates the first loss function between the predicted lyrics and the source lyrics, adjusts the parameters of the initial lyrics imitation model according to the first loss function until the first loss function is lower than a first preset value, and determines the initial lyrics imitation model as the target lyrics imitation model.
[0100] In one implementation, it is necessary to achieve the imitation of source lyrics. Therefore, the initial lyrics imitation model can be trained by controlling the first loss function at a first preset value. The first preset value should not be set too small to avoid the output imitation lyrics being too similar to the source lyrics or the output imitation lyrics being the source lyrics due to an excessively small loss function. For example, the first preset value of the first loss function can be controlled at 25%. When the first loss function reaches the first preset value of 25%, the initial lyrics imitation model is determined to be the target lyrics imitation model. The first preset value of the first loss function can also be 20%, 30%, or other values, and there is no specific limitation.
[0101] 206. Input the source lyrics into the target lyrics imitation model, and generate the imitation lyrics of the source lyrics through the target lyrics imitation model.
[0102] The source lyrics are input into the target lyrics imitation model, which then generates imitation lyrics of the source lyrics. In essence, the target lyrics imitation model can generate one or multiple imitation lyrics of the source lyrics. The imitation lyrics generated by the target imitation model are similar to the imagery and lyrical information created by the source lyrics.
[0103] The following content will provide a brief introduction to the model training process involved in this application.
[0104] The execution entities of the following model training method and the above song imitation method can be different terminal devices or servers. For example, the execution entity of the song imitation method is laptop 1, and the execution entity of the model training method is laptop 2.
[0105] The following will introduce Figure 5 the pre-trained language model, Figure 5 which is a schematic diagram of the pre-trained language model in the embodiments of this application.
[0106] During the process of training the pre-trained language model, a large amount of text corpus needs to be collected. The corpus is usually a collection of resources with a certain quantity and scale, and the text corpus refers to a collection of language words. To collect the text corpus, as much text corpus as possible can be collected from the comment data of various social media websites and e-commerce websites, People's Daily, Chinese Wikipedia, Baidu Encyclopedia, etc.
[0107] Preprocess the collected text corpus, that is, clean the collected text corpus and filter out the dirty data in the text corpus. For example, the web links, tag information, website address names, lines with too few words, duplicate data, etc. existing in the text corpus. These dirty data can neither provide any useful information nor will they reduce the training effect of the pre-trained language model.
[0108] After filtering and cleaning the dirty data from the text corpus, process the text corpus into a format that the pre-trained model can parse. Since the pre-trained language model processes the text corpus at the character level internally, it is necessary to process the text corpus into a character-level format that the pre-trained language model can understand. Exemplarily, change the sentence "I come running towards you" to "I come running towards you" in the way of separating with spaces.
[0109] Input the processed text corpus into the pre-trained language model to train the pre-trained language model. Specifically, use the self-generative training method of GPT (Generative Pre-trained Transformer, a large-scale pre-trained model), that is, predict and generate the next word based on the existing words, and train the pre-trained language model through the Transformer transducer model structure. In the embodiments of this application, the pre-trained language model is a Transformer network structure including 12 layers, the dimension of the initialized word vector is 768, there are 12 attention layers, and the dictionary size is 21128.
[0110] The following will Figure 6 and Figure 7 explain the training process of the initial tune imitation model, Figure 6 which is a schematic diagram of the training process of the initial tune imitation model in the embodiments of this application. Figure 7 This is a diagram of the audio pre-trained model architecture based on convolutional encoders and transformer encoders.
[0111] Collecting audio corpora refers to a collection of audio files of a certain quantity and scale. For example, audio corpora can be collected through various channels such as the Internet, music software, social media websites, e-books, and podcasts.
[0112] Preprocessing the acquired audio corpus is necessary because the audio corpus collected through various channels may contain dirty data, such as external noises like horns, background noise, human voices, and car horns. Preprocessing removes this dirty data. The preprocessed audio corpus becomes cleaner and more conducive to model analysis and training. The preprocessed audio corpus is then input into the initial tone imitation model. This model randomly masks the audio corpus and generates predicted audio parts corresponding to the masked audio parts based on the unmasked parts. A second loss function is calculated between the predicted and masked audio parts. The parameters of the initial tone imitation model are adjusted based on this second loss function until it falls below a second preset value, thus identifying the initial tone imitation model as the target tone imitation model.
[0113] In one implementation, to achieve the imitation of a source melody, the initial melody imitation model can be trained by controlling the second loss function at a second preset value. The second preset value should not be set too small to avoid the output imitation melody being too similar to the source melody, or the output imitation melody being the source melody. For example, controlling the second loss function to 25%, when the second loss function reaches the second preset value of 25%, the initial melody imitation model is determined to be the target melody imitation model. The second preset value of the second loss function can also be 20%, 30%, or other values; there is no specific limitation.
[0114] Audio pre-trained models based on CNN convolutional encoders and Transformer encoders can be divided into CNN convolutional encoder modules and Transformer encoder modules. The CNN convolutional encoder has seven layers, each containing a temporal convolutional layer, a normalization layer, and a GELU (Gaussian error linear unit) activation function layer. The CNN convolutional encoder is used to recognize audio corpora, transforming them into vectors.
[0115] The Transformer encoder uses gated relative position encoding, introducing relative position into the attention network's computation for better modeling of local information. During training, the input audio file is randomly transformed, for example, by mixing two audio files or adding background noise. Then, approximately 50% of the audio signal is randomly masked, and the label corresponding to the masked position is predicted at the output. Formally, given an input audio file X, its label Y is first extracted. Then, X is noise-added and masked to generate X^. The Transformer needs to predict the label Y at the masked position using the input X^. Then, a mask prediction task is used for prediction, that is, a portion of the audio data is randomly masked through the mask prediction task. Finally, the masked audio is input to the Transformer encoder. The Transformer encoder predicts the masked portion of the audio from the unmasked audio, thus modeling the tune sequence and outputting a tune similar to the source FM corpus. Finally, it calculates the loss function between the predicted audio portion and the masked audio corpus portion, adjusts the parameters of the audio pre-training model based on the loss function, and continues training the audio pre-training model until the loss function is lower than a preset value. At this point, the audio pre-training model is considered to be a well-trained audio pre-training model.
[0116] In the technical solution provided in the above-described embodiments of this application, firstly, a source song is obtained and segmented into source lyrics and a source melody. Then, imitation lyrics are generated based on visual words in the source lyrics, where visual words represent the images expressed by the source lyrics. Based on the source melody, an imitation melody is generated through autoregression. Finally, the imitation lyrics and the imitation melody are merged to obtain the target song. Thus, this application, by segmenting the source song into source lyrics and a source melody, generating imitation lyrics based on visual words in the source lyrics, generating an imitation melody through autoregression based on the source melody, and then merging the imitation lyrics and the imitation melody to obtain the target song, can quickly obtain imitation songs similar to the source song, achieving intelligent generation of imitation songs from source songs, greatly shortening the creator's creation time and improving the efficiency of song imitation.
[0117] Figure 8 A schematic diagram of a song imitation method apparatus provided in this application embodiment is shown below. Figure 8 The embodiments described below are described in detail. The embodiments described below are used to explain the technical solutions of this application and are not intended to limit actual use.
[0118] The device includes:
[0119] Acquisition unit 801 is used to acquire the source song.
[0120] Processing unit 802 is used to divide the source song into source lyrics and source melody.
[0121] The generation unit 803 is used to generate imitation lyrics of the source lyrics based on the image words in the source lyrics. The image words are used to represent the image expressed by the source lyrics.
[0122] The generation unit 803 is also used to autoregressively generate a copy of the source melody based on the source melody.
[0123] The processing unit 802 is also used to merge the imitated lyrics and the imitated melody to obtain the target song.
[0124] In an optional implementation, the generation unit 802 is further configured to generate imitation lyrics of the source lyrics based on the visual words in the source lyrics.
[0125] The acquisition unit 801 is also used to acquire the target lyrics imitation model based on the words in the picture.
[0126] The generation unit 803 is also used to input the source lyrics into the target lyrics imitation model and generate the imitation lyrics of the source lyrics through the target lyrics imitation model.
[0127] In one optional implementation,
[0128] The processing unit 802 is also used to input parallel corpora into the initial lyrics imitation model.
[0129] The generation unit 803 is also used by the initial lyrics imitation model to generate predicted lyrics based on parallel corpora.
[0130] The processing unit 802 is also used to calculate a first loss function between the predicted lyrics and the source lyrics.
[0131] The processing unit 802 is further configured to adjust the parameters of the initial lyrics imitation model according to the first loss function until the first loss function is lower than a first preset value, and determine the initial lyrics imitation model as the target lyrics imitation model.
[0132] In an optional implementation, the processing unit 802 is further configured to input parallel corpora into the initial lyrics imitation model, and the generation unit 803 is further configured to, before the initial lyrics imitation model generates predicted lyrics based on the parallel corpora...
[0133] The acquisition unit 801 is also used to acquire N visual words of the source lyrics, where N is a positive integer greater than or equal to 1.
[0134] The processing unit 802 is also used to input N visual words into the lyrics generation model.
[0135] The generation unit 803 is also used to generate reference lyrics containing visual words through the lyrics generation model.
[0136] The processing unit 802 is also used to determine the lyrics information of the source lyrics.
[0137] The processing unit 802 is further configured to determine the first corpus and the second corpus as parallel corpus, wherein the first corpus includes the N picture words, the lyrics information of the source lyrics and the reference lyrics, and the second corpus includes the source lyrics.
[0138] In an optional implementation, the acquisition unit 801 is also used to acquire N visual words of the source lyrics.
[0139] The processing unit 802 is also used to divide the source lyrics into multiple lyric segments.
[0140] The processing unit 802 is also used to input lyrics segments into a pre-trained language model.
[0141] The acquisition unit 801 is also used to obtain the segment vector of the lyrics segment based on the pre-trained language model.
[0142] The processing unit 802 is also used to segment the lyrics into words.
[0143] The processing unit 802 is also used to input each word of the segmented lyrics into the pre-trained language model.
[0144] The acquisition unit 801 is also used to acquire word vectors of each lyric segment based on the pre-trained language model.
[0145] The processing unit 802 is also used to calculate the semantic similarity between each word vector of the lyrics segment and the segment vector of the lyrics segment.
[0146] The processing unit 802 is also used to determine N visual words of the source lyrics based on the semantic similarity between the word vectors of each lyric segment and the segment vectors of the lyric segment.
[0147] In an optional implementation, the processing unit 802 is further configured to determine N visual words of the source lyrics based on the semantic similarity between the word vectors of each lyric segment and the segment vectors of the lyric segment.
[0148] The processing unit 802 is also used to determine the top S word vectors with the highest semantic similarity to the segment vectors of the lyrics segment as the visual words of the lyrics segment, where S is a positive integer greater than or equal to 1.
[0149] The processing unit 802 is also used to sort the visual words of each lyrics segment according to the semantic similarity of the visual words of each lyrics segment, and obtain the N visual words with the highest semantic similarity.
[0150] The processing unit 802 is also used to determine the N scene words as the N scene words of the source lyrics.
[0151] In an optional implementation, the generation unit 803 is further configured to autoregressively generate a copy of the source melody based on the source melody.
[0152] The processing unit 802 is also used to select a target sub-melody from the source melody.
[0153] The generation unit 803 is also used to input the target sub-melody into the target melody imitation model and generate multiple imitation sub-melody through autoregression.
[0154] Processing unit 802 is also used to process the target sub-melody 、 Multiple imitation sub-melodies are combined to form an imitation melody of the source song, and the imitation melody of the source song is output through the target melody imitation model.
[0155] In one optional implementation,
[0156] Unit 801 is also used to acquire audio corpora.
[0157] The processing unit 802 is also used to input the audio corpus into the initial melody imitation model, the initial melody imitation model randomly masks the audio corpus, and generates the predicted audio part corresponding to the masked audio corpus part based on the unmasked audio corpus part.
[0158] The processing unit 802 is also used to calculate a second loss function between the predicted audio portion and the masked audio corpus portion.
[0159] The processing unit 802 is further configured to adjust the parameters of the initial melody imitation model according to the second loss function until the second loss function is lower than the second preset value, and determine the initial melody imitation model as the target melody imitation model.
[0160] In the technical solution provided in the above-described embodiments of this application, firstly, a source song is obtained and segmented into source lyrics and a source melody. Then, imitation lyrics are generated based on visual words in the source lyrics, where visual words represent the images expressed by the source lyrics. Based on the source melody, an imitation melody is generated through autoregression. Finally, the imitation lyrics and the imitation melody are merged to obtain the target song. Thus, this application, by segmenting the source song into source lyrics and a source melody, generating imitation lyrics based on visual words in the source lyrics, generating an imitation melody through autoregression based on the source melody, and then merging the imitation lyrics and the imitation melody to obtain the target song, can quickly obtain imitation songs similar to the source song, achieving intelligent generation of imitation songs from source songs, greatly shortening the creator's creation time and improving the efficiency of song imitation.
[0161] It should be noted that the information interaction and execution process between the various modules / units in the device are different from those in this application. Figures 1 to 7 The various method embodiments are based on the same concept, and the details can be found in the descriptions of the method embodiments shown above in this application, which will not be repeated here.
[0162] The following describes an electronic device provided by an embodiment of this application. Please refer to [link / reference]. Figure 9 , Figure 9 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. The electronic device 900 can specifically be a virtual reality (VR) device, a mobile phone, a tablet, a laptop computer, a smart wearable device, a monitoring data processing device, or a radar data processing device, etc., and is not limited thereto. The electronic device 900 may be equipped with... Figure 4 The song imitation device described in the corresponding embodiment is used to implement Figures 1 to 7 The functions correspond to those in the embodiments. Specifically, the electronic device 900 includes: a receiver 901, a transmitter 902, a processor 903, and a memory 904 (wherein the number of processors 903 in the execution device 900 may be one or more). Figure 5 (Taking a processor as an example), processor 903 may include application processor 9031 and communication processor 9032. In some embodiments of this application, receiver 901, transmitter 902, processor 903 and memory 904 may be connected via a bus or other means.
[0163] Memory 904 may include read-only memory and random access memory, and provides instructions and data to processor 903. A portion of memory 904 may also include non-volatile random access memory (NVRAM). Memory 904 stores processor and operation instructions, executable modules, or data structures, or subsets thereof, or extended sets thereof, wherein the operation instructions may include various operation instructions for implementing various operations.
[0164] Processor 903 controls the operation of the execution device. In specific applications, the various components of the execution device are coupled together through a bus system, which may include not only the data bus, but also power buses, control buses, and status signal buses. However, for clarity, all buses are referred to as the bus system in the diagram.
[0165] The methods disclosed in the embodiments of this application can be applied to or implemented by the processor 903. The processor 903 can be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by the integrated logic circuitry in the hardware of the processor 903 or by instructions in software form. The processor 903 can be a general-purpose processor, a digital signal processor (DSP), a microprocessor, or a microcontroller, and may further include an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The processor 903 can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly manifested as being executed by a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software module can reside in a mature storage medium in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or registers. This storage medium is located in memory 904, and processor 903 reads the information from memory 904 and, in conjunction with its hardware, completes the steps of the above method.
[0166] Receiver 901 can be used to receive input digital or character information, and to generate signal inputs related to the settings and function control of the execution device. Transmitter 902 can be used to output digital or character information through the first interface; transmitter 902 can also be used to send instructions to the disk group through the first interface to modify the data in the disk group; transmitter 902 may also include a display device such as a display screen.
[0167] In this embodiment of the application application, the application processor 9031 in the processor 903 is used to execute... Figures 1 to 3 The song imitation method in the corresponding embodiment. It should be noted that the specific manner in which the application processor 9031 executes each step is different from that in this application. Figures 1 to 7 The various method embodiments are based on the same concept, and the technical effects they bring are the same as those in this application. Figures 1 to 7 The corresponding method embodiments are the same, and for details, please refer to the description in the method embodiments shown above in this application, which will not be repeated here.
[0168] This application provides a computer-readable storage medium, which includes computer instructions. When executed by a processor, the computer instructions are used to implement the technical solution of any song imitation method in this application.
[0169] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0170] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0171] Computer-readable media, as defined herein, includes both permanent and non-permanent, removable and non-removable media, which can store information using any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include non-transitory computer-readable media, such as modulated data signals and carrier waves.
[0172] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0173] Although this application discloses preferred embodiments as described above, it is not intended to limit this application. Any person skilled in the art can make possible changes and modifications without departing from the spirit and scope of this application. Therefore, the scope of protection of this application should be determined by the scope defined in the claims of this application.
Claims
1. A method for song imitation, characterized in that, The method includes: Obtain the source song and divide it into source lyrics and source melody; Based on the visual words in the source lyrics, a similar lyric is generated; the visual words are used to characterize the imagery expressed by the source lyrics. Based on the source melody, an autoregressive rendition of the source melody is generated; The imitated lyrics and the imitated melody are combined to obtain the target song; The melody imitation model is obtained through training, and the training method for the melody imitation model includes: Obtain audio corpus; The audio corpus is input into the initial melody imitation model, which randomly masks the audio corpus and generates the predicted audio part corresponding to the masked audio corpus part based on the unmasked audio corpus part. Calculate a second loss function between the predicted audio portion and the masked audio corpus portion; The parameters of the initial melody imitation model are adjusted according to the second loss function until the second loss function is lower than the second preset value, and the initial melody imitation model is determined to be the target melody imitation model.
2. The method according to claim 1, characterized in that, The step of generating paraphrased lyrics based on visual words in the source lyrics includes: The target lyrics imitation model is obtained based on the visual words; The source lyrics are input into the target lyrics imitation model, and the target lyrics imitation model generates the imitation lyrics of the source lyrics.
3. The method according to claim 2, characterized in that, The target lyrics imitation model is obtained through training, and the training method of the target lyrics imitation model includes: Parallel corpora are input into an initial lyrics imitation model, which generates predicted lyrics based on the parallel corpora. Calculate the first loss function between the predicted lyrics and the source lyrics; The parameters of the initial lyrics imitation model are adjusted according to the first loss function until the first loss function is lower than a first preset value, and the initial lyrics imitation model is determined to be the target lyrics imitation model.
4. The method according to claim 3, characterized in that, The step of inputting parallel corpora into the initial lyrics imitation model, before the initial lyrics imitation model generates predicted lyrics based on the parallel corpora, further includes: Obtain N visual words from the source lyrics; where N is a positive integer greater than or equal to 1. The N visual words are input into the lyrics generation model, and the lyrics generation model generates reference lyrics containing the visual words. Determine the lyrics information of the source lyrics; The first corpus and the second corpus are defined as parallel corpora, wherein the first corpus includes the N image words, the lyrics information of the source lyrics, and the reference lyrics, and the second corpus includes the source lyrics.
5. The method according to claim 4, characterized in that, The acquisition of N visual words from the source lyrics includes: The source lyrics are divided into multiple lyric segments; The lyrics segments are input into a pre-trained language model, and the segmentation vectors of the lyrics segments are obtained based on the pre-trained language model. The lyrics are segmented into words; Each word segment of the lyrics after word segmentation is input into the pre-trained language model, and the word vectors of each lyric segment are obtained according to the pre-trained language model. Calculate the semantic similarity between each word vector of the lyrics segment and the segment vector of the lyrics segment; The N visual words of the source lyrics are determined based on the semantic similarity between the word vectors of each lyric segment and the segment vectors of the lyric segment.
6. The method according to claim 5, characterized in that, The step of determining the N visual words of the source lyrics based on the semantic similarity between the word vectors of each lyric segment and the segment vectors of the lyric segment includes: The word segments corresponding to the top S word vectors with the highest semantic similarity to the segment vectors of the lyrics segment are determined as the visual words of the lyrics segment; where S is a positive integer greater than or equal to 1. Based on the semantic similarity of the visual words in each of the lyrics segments, the visual words in each of the lyrics segments are sorted to obtain the N visual words with the highest semantic similarity. The N visual words are determined as the N visual words of the source lyrics.
7. The method according to claim 1, characterized in that, The step of generating a rewritten melody of the source melody through autoregression includes: Select the target sub-melody from the source melody; The target sub-melody is input into the target melody imitation model, and multiple imitation sub-melody are generated by autoregression; The target sub-melody and the multiple imitation sub-melody are combined to form the imitation melody of the source song, and the imitation melody of the source song is output through the target melody imitation model.
8. A song imitation device, characterized in that, include: The acquisition unit is used to acquire the source song; The processing unit is used to segment the source song into source lyrics and source melody; A generation unit is used to generate lyrics that mimic the source lyrics based on the imagery words in the source lyrics; the imagery words are used to characterize the imagery expressed by the source lyrics. The generation unit is also used to autoregressively generate a copy of the source melody based on the source melody; The processing unit is also used to merge the imitated lyrics with the imitated melody to obtain the target song; The melody imitation model is obtained through training, and the training method for the melody imitation model includes: Obtain audio corpus; The audio corpus is input into the initial melody imitation model, which randomly masks the audio corpus and generates the predicted audio part corresponding to the masked audio corpus part based on the unmasked audio corpus part. Calculate a second loss function between the predicted audio portion and the masked audio corpus portion; The parameters of the initial melody imitation model are adjusted according to the second loss function until the second loss function is lower than the second preset value, and the initial melody imitation model is determined to be the target melody imitation model.
9. An electronic device, characterized in that, The electronic device includes: a memory and a processor; the memory and the processor are coupled. The memory is used to store one or more computer instructions; The processor is used to execute one or more computer instructions to implement the song imitation method as described in any one of the technical solutions of claims 1-7.
10. A computer-readable storage medium storing one or more computer instructions thereon, characterized in that, The instruction is executed by the processor to implement the song imitation method as described in any one of the technical solutions of claims 1-7.