Song creation method, device, system, and storage medium

The method aligns lyric text with a singing melody and synthesizes it into a singing voice, allowing users to create songs efficiently without specialized music skills.

JP7760072B2Active Publication Date: 2025-10-24LEMON CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2024548539
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2022-05-07
Filing Date
2023-05-08
Publication Date
2025-10-24
Estimated Expiration
2043-05-08

AI Technical Summary

Technical Problem

Creating songs requires specialized skills, making it difficult for ordinary users to create songs efficiently.

Method used

A method for aligning target lyric text with a singing melody to determine correspondences between text units and notes, synthesizing the text into a singing voice, and integrating it with accompaniment audio to generate a target song, using a song generation device and system.

Benefits of technology

Enables users to create songs based on their own lyrics without professional music skills, improving efficiency and interest in song creation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007760072000005
    Figure 0007760072000005
  • Figure 0007760072000006
    Figure 0007760072000006
  • Figure 0007760072000007
    Figure 0007760072000007
Patent Text Reader

Abstract

The present disclosure relates to a method, device, system and storage medium for song creation, which obtains a target lyrics text input by a user, aligns the target lyrics text with the singing melody of an initial song, determines the correspondence between the text units in the target lyrics text and the notes in the singing melody, and synthesizes the target lyrics text based on the correspondence between the text units in the target lyrics text and the notes in the singing melody to obtain a singing voice that recites the target lyrics text with the singing melody, and further integrates the singing voice with the accompaniment audio of the initial song to generate a target song, which essentially allows the user to write lyrics by himself, and create a song based on the lyrics, singing melody and accompaniment audio created by the user, to form a completely new song. In this way, the user can create a song based on the lyrics he created, even if he does not have professional skills in creating music, and improves the efficiency and interest of the user in song creation.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] [CROSS-REFERENCE TO RELATED APPLICATIONS] This application is based on and claims priority from a Chinese application bearing application number 202210494217.7, filed on May 7, 2022, and entitled "Song generation method, device, system and storage medium," the entire disclosure of which is incorporated herein by reference.

[0002] The present disclosure relates to the field of artificial intelligence, and in particular to a song generation method, device, system, and storage medium. [Background technology]

[0003] Creating songs requires a certain level of specialized skills, and is extremely difficult for ordinary users who lack specialized skills. Therefore, there is an urgent need to find ways to help song creators create songs efficiently. Summary of the Invention

[0004] To solve the above technical problems, the present disclosure provides a song generation method, device, system, and storage medium.

[0005] According to a first aspect, the present disclosure provides a method for manufacturing a semiconductor device, comprising: obtaining target lyric text entered by a user; aligning the target lyric text with a singing melody of an initial song to determine a correspondence between text units in the target lyric text and notes in the singing melody, the singing melody being a singing melody for initial lyrics in the initial song; synthesizing the target lyrics text based on the correspondence between the text units in the target lyrics text and the notes in the singing melody to obtain a singing voice that sings the target lyrics text with the singing melody; and integrating the singing voice with the accompaniment audio of the initial song to generate a target song.

[0006] In some embodiments, prior to aligning the target lyric text with the vocal melody of the initial song as described above, the method further comprises: selecting an initial song from a plurality of preset songs in response to an initial song selection operation; and determining a corresponding singing melody and accompanying audio based on the initial song.

[0007] In some embodiments, aligning the target lyric text with the vocal melody of the initial song comprises: Dividing the singing melody into a plurality of melody paragraphs; Dividing the target lyric text into a plurality of lyric paragraphs, the number of the plurality of lyric paragraphs being the same as the number of the plurality of melody paragraphs; The method includes aligning the plurality of lyric paragraphs and the plurality of melody paragraphs one by one, and determining correspondences between text units in the lyric paragraphs and notes in the corresponding melody paragraphs.

[0008] In some embodiments, dividing the vocal melody into a plurality of melody paragraphs comprises: determining a paragraph division point for each predetermined measure in the singing melody; adjusting the number of paragraph division points based on the number of notes included in the melody paragraph corresponding to each paragraph division point; and adjusting the position of each paragraph division point based on a note attack distance, wherein the note attack distance includes a duration pitch of the note attack and / or a pitch pitch of the note attack.

[0009] In some embodiments, adjusting the number of paragraph division points based on the number of notes included in the melody paragraph corresponding to each paragraph division point includes: For any one paragraph division point, if the number of notes included in the melody paragraph corresponding to the one paragraph division point is smaller than a first threshold value, deleting the one paragraph division point; If the number of notes included in the melody paragraph corresponding to any one of the paragraph division points is greater than a second threshold value, the paragraph division point is increased by one.

[0010] In some embodiments, adjusting the position of each paragraph division point based on note attack distance includes: searching for a position where a note attack distance satisfies a predetermined condition within a predetermined range of beats around any one paragraph division point, and determining that position as the position of the any one paragraph division point; The specified condition is that the duration pitch of the note attack before and after any one of the paragraph division points is the largest, or that the duration pitch of the note attack before and after any one of the paragraph division points is equal and the pitch pitch is the largest.

[0011] In some embodiments, dividing the target lyric text into lyric paragraphs comprises: segmenting the target lyrics text into words and determining the part of speech corresponding to each word; and dividing the target lyric text into a plurality of lyric paragraphs based on the part of speech corresponding to each word, predefined linguistic rules, and the length of the sung melody. In some embodiments, aligning the plurality of lyric paragraphs with the plurality of melody paragraphs one by one comprises: For each of the melody paragraphs, obtaining a plurality of predetermined lyric alignment templates corresponding to the melody paragraph, each of the lyric alignment templates corresponding to a different number of lyric characters; selecting a target lyric alignment template from the plurality of lyric alignment templates, and the number of characters corresponding to the target lyric alignment template is the number of characters in a lyric paragraph corresponding to the melody paragraph; and aligning the melody paragraph with a lyric paragraph corresponding to the melody paragraph based on the target lyric alignment template.

[0012] In some embodiments, the plurality of lyric alignment templates corresponding to the melody paragraph include a first lyric alignment template, a second lyric alignment template and a third lyric alignment template; In the first lyric alignment template, each note in the melody paragraph corresponds to one text unit; In the second lyric alignment template, adjacent notes in the melody paragraph with the closest note attack distance are combined into one note pair, and one note pair corresponds to one text unit; In the third lyric alignment template, adjacent notes in the second lyric alignment template with the closest note attack distance are combined into one note pair, and one note pair corresponds to one text unit.

[0013] According to a second aspect, the present disclosure provides a method for manufacturing a semiconductor device, comprising: an acquisition unit for acquiring a target lyric text input by a user; an alignment unit for aligning the target lyric text with a singing melody of an initial song to determine a correspondence between text units in the target lyric text and notes in the singing melody, the singing melody being a singing melody corresponding to initial lyrics in the initial song; a synthesis unit for synthesizing the target lyrics text with a voice based on the correspondence between text units in the target lyrics text and notes in the singing melody to obtain a singing voice for singing the target lyrics text with the singing melody; A song generation device is further proposed, which includes a generation unit for integrating said singing voice and an accompaniment audio of said initial song to generate a target song.

[0014] According to a third aspect, the present disclosure further provides a system including at least one computing device and at least one storage device storing instructions that, when executed by the at least one computing device, cause the at least one computing device to perform steps of the song generation method described above.

[0015] According to a fourth aspect, the present disclosure further provides a computer-readable storage medium storing a program or instructions that, when executed by at least one computing device, cause the at least one computing device to perform the steps of the song generation method described above. [Brief explanation of the drawings]

[0016] The drawings herein are incorporated into the specification, constitute a part of this specification, illustrate embodiments consistent with the present disclosure, and together with the specification, serve to explain the principles of the present disclosure. In order to more clearly describe the embodiments of the present disclosure or the technical solutions in the prior art, the following briefly describes the drawings that need to be used in the description of the embodiments or the prior art. Obviously, those skilled in the art can derive other drawings based on these drawings without exerting creative efforts. [Figure 1] 1 is a flowchart of a song generation method according to an embodiment of the present disclosure. [Figure 2] 10 is a flowchart of another song generation method according to an embodiment of the present disclosure. [Figure 3] FIG. 1 is a schematic diagram of a musical staff of a melody according to an embodiment of the present disclosure. [Figure 4] FIG. 2 is a schematic staff diagram of one melody paragraph according to an embodiment of the present disclosure. [Figure 5] FIG. 1 is a schematic diagram of a display interface for realizing generating a song based on a target lyric text in a terminal according to an embodiment of the present disclosure; [Figure 6] 1 is a schematic diagram illustrating the configuration of a song generation device according to an embodiment of the present disclosure. [Figure 7]1 is an exemplary block diagram of a system including at least one computing device and at least one memory device that stores instructions, according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0017] In order to make the above-mentioned objects, features and advantages of the present disclosure more clearly understandable, the following further describes the aspects of the present disclosure. It should be noted that, unless contradictory, the embodiments and features in the embodiments of the present disclosure can be combined with each other.

[0018] In the following description, many specific details are set forth in order to fully understand the present disclosure; however, the present disclosure may be embodied in other forms different from those described herein. Obviously, the embodiments in the specification are only some of the embodiments of the present disclosure, but not all of the embodiments.

[0019] 1 is a flowchart of a song generation method according to an embodiment of the present disclosure. This embodiment is applicable to the creation of a song based on lyric text at a client side, where the song generation method may be executed by a song generation device, which may be implemented in software and / or hardware and located in an electronic device, such as a terminal, including, but not limited to, a smartphone, palmtop, tablet, wearable device with display, desktop, laptop, all-in-one device, smart home device, etc. Alternatively, this embodiment is applicable to the creation of a song based on lyric text at a service side, where the song generation method may be executed by a song generation device, which may be implemented in software and / or hardware and located in an electronic device, such as a server.

[0020] As shown in FIG. 1, the song generation method may specifically include steps S110 to S140.

[0021] S110: The target lyrics text input by the user is obtained.

[0022] There are multiple ways to implement this step, and the present application does not impose any limitations thereon. In some embodiments, in response to a trigger operation for generating a song, a song generation interface is displayed on the display of the terminal. The song generation interface includes a text input box. When a user needs to create a song, the user inputs target lyric text into the text input box. The song generation system achieves the purpose of obtaining the target lyric text by identifying the target lyric text input by the user.

[0023] S120: Align the target lyrics text with the singing melody of the initial song to determine a correspondence between text units in the target lyrics text and notes in the singing melody, where the singing melody is a singing melody for the initial lyrics in the initial song.

[0024] The singing melody refers to the main melody of the song that the user wants to create.

[0025] By performing alignment, correspondences between text units in the target lyric text and musical notes in the vocal melody can be determined. A text unit may be one or a combination of characters, words, sentences, phonetic units, etc.

[0026] In some embodiments, before this step, the song generation method further includes selecting an initial song from a preset plurality of songs in response to an initial song selection operation, and determining a corresponding singing melody and accompaniment audio based on the initial song. This setting essentially means that a song database containing a plurality of songs is pre-established, and a user is allowed to select a song of their choice from the song database as the initial song.

[0027] S130: Based on the correspondence between the text units in the target lyrics text and the notes in the singing melody, the target lyrics text is subjected to voice synthesis to obtain a singing voice that sings the target lyrics text with the singing melody.

[0028] Speech synthesis is singing voice synthesis (SVS), which synthesizes singing based on lyrics and song melody. Compared to text-to-speech (TTS), which makes a machine "speak," singing synthesis is more entertaining because it makes a machine sing. A singing synthesis model can be generated by pre-training. In this way, by inputting the timbre, custom text, and song melody into the singing synthesis model, the singing synthesis model can output audio containing pronunciation corresponding to the custom text. Conventional technology can be used to train the singing synthesis model, and no further explanation will be given.

[0029] S140: The singing voice and the accompaniment audio of the initial song are integrated to generate a target song.

[0030] The accompaniment audio refers to the chords of the song the user wants to create, and serves to complement the main melody.

[0031] In some embodiments, the vocal melody and the accompanying audio may actually correspond to each other, i.e., when the user selects a vocal melody, the accompanying audio is simultaneously selected due to the one-to-one correspondence between the two.

[0032] Alternatively, the singing melodies and the accompanying audio may not all correspond to each other, that is, the user must select the singing melody and then the accompanying audio.

[0033] The synthesis refers to the merging of vocals with accompaniment audio to form a new song. The synthesis method is prior art and will not be further described here.

[0034] The above technical solution essentially allows users to write lyrics themselves, and create songs based on the lyrics, vocal melody and accompanying audio they have created themselves, thereby forming completely new songs. In this way, users can create songs based on their own lyrics without having professional skills in music creation, thereby improving the efficiency and interest of users in song creation.

[0035] 2 is a flowchart of another song generation method according to an embodiment of the present disclosure, which is a specific example of FIG. 1. Referring to FIG. 2, the song generation method includes the following steps:

[0036] S210: The target lyrics text input by the user is obtained.

[0037] S220: Divide the vocal melody of the initial song into a plurality of melody paragraphs.

[0038] In some embodiments, the vocal melody may or may not actually be pre-processed.

[0039] If the singing melody has been preprocessed, the singing melody includes a plurality of paragraph division points, and the positions of the plurality of paragraph division points in the singing melody are constant. Based on the plurality of paragraph division points, the singing melody can be divided into a plurality of melody paragraphs.

[0040] If the singing melody is not pre-processed, the singing melody does not contain paragraph division points, in which case the singing melody needs to be processed to have paragraph division points at fixed positions, thereby achieving the purpose of dividing the singing melody into multiple melody paragraphs.

[0041] If the vocal melody is not pre-processed, there are multiple ways to implement this step, and this application does not impose any limitations on this. In some embodiments, the paragraph division points may be determined and further divided based only on the bar of the vocal melody, or the paragraph division points may be determined and further divided based only on the number of notes.

[0042] In some embodiments, a method for determining paragraph division points in a vocal melody is provided, which includes determining one paragraph division point for each predetermined number of bars, adjusting the number of paragraph division points based on the number of notes included in the melody paragraph corresponding to each paragraph division point, and adjusting the position of each paragraph division point based on the note attack distance, where the note attack distance includes the duration pitch of the note attack and / or the pitch pitch of the note attack. This setting essentially means that the paragraph division points are initially determined based on the number of bars, and then adjusted based on the number of notes and the note attack distance to obtain the final position of the paragraph division points. This setting can reduce the difficulty of subsequent alignment operations.

[0043] Adjusting the number of paragraph division points based on the number of notes contained in the melody paragraph corresponding to each paragraph division point includes deleting any one paragraph division point if the number of notes contained in the melody paragraph corresponding to that one paragraph division point is less than a first threshold, and adding one paragraph division point if the number of notes contained in the melody paragraph corresponding to that one paragraph division point is greater than a second threshold. By setting them in this way, the number of notes contained in each melody paragraph can be made uniform, which is advantageous for matching subsequent lyric characters with notes.

[0044] Adjusting the position of each paragraph division point based on the note attack distance includes searching for a position within a predetermined range of beats around any one paragraph division point where the note attack distance satisfies a predetermined condition, and setting that position as the position of any one paragraph division point, where the predetermined condition is that the duration pitch of the note attacks before and after any one paragraph division point is the largest, or that the duration pitch of the note attacks before and after any one paragraph division point are equal and the pitch pitch is the largest.

[0045] FIG. 3 is a schematic diagram of a staff of a melody according to an embodiment of the present disclosure. In some embodiments, FIG. 3 includes two lines of melody, each line having four bars. First, one paragraph division point is determined for every two bars. The first determined paragraph division point for the first line is at a1 in the figure, and the first determined paragraph division point for the second line is at b1 in the figure. Then, the first determined paragraph division point is corrected or adjusted based on the number of notes and the note attack distance. After the adjustment, the final determined paragraph division point for the first line is at a2 in the figure, and the final determined paragraph division point for the second line is at b2 in the figure. Finally, the first line of melody is divided into two melody paragraphs, 1 and 2, and the second line of melody is divided into two melody paragraphs, 3 and 4.

[0046] S230: Divide the target lyric text into a plurality of lyric paragraphs, and the number of the plurality of lyric paragraphs is the same as the number of the plurality of melody paragraphs.

[0047] There are multiple ways to implement this step, and this application does not impose any limitations thereon. In some embodiments, the target lyric text is segmented into words, the part of speech corresponding to each word is determined, and the target lyric text is divided into multiple lyric paragraphs based on the part of speech corresponding to each word, predefined linguistic rules, and the length of the singing melody.

[0048] Parts of speech refer to the classification of words based primarily on grammatical features (including syntactic function and morphological changes) while also taking into account lexical meaning, and include nouns, verbs, adjectives, pronouns, adverbs, and particles.

[0049] The linguistic rules are segmentation restriction rules established based on semantic completeness, such as indivisibility between adverbs and verbs, indivisibility between numerals and counter classifiers, indivisibility between pronouns and numerals, etc. Segmenting the target lyrics text into multiple lyrics paragraphs based on the parts of speech corresponding to each word and predefined linguistic rules means that after segmentation, indivisible phrases are located in the same lyrics paragraph.

[0050] Dividing the target lyrics text into a plurality of lyrics paragraphs based on the length of the singing melody means that after division, the number of characters in each lyrics paragraph is equal to or less than the number of notes in the melody paragraph corresponding to the lyrics paragraph.

[0051] In reality, there are often multiple linguistic rules. Considering that the number of notes in a melody paragraph is finite, we set up a dynamic optimization strategy, which aims to minimize the difference in length between each lyric phrase after division under as many linguistic rules as possible.

[0052] In some embodiments, the target lyric text entered by the user is (outside 1) TIFF0007760072000001.tif7127 (translation: "I learned to love freely only after meeting you") (length: 11 characters). The target lyrics text is segmented into words, (outside 2) TIFF0007760072000002.tif7127, the phrases "after," "talent," "academic," "love," "obtain," and "freedom" are obtained. The parts of speech for each of the obtained phrases are indicated, and the results are as follows: (Outside 3) TIFF0007760072000003.tif7127-pronoun, post-directive, talent-adverb, study-verb, love-verb, get-particle, free-adjective. Based on the dynamic optimization strategy, the target lyrics text is finally divided into the following parts:

[0053] Achievement (outside 4) TIFF0007760072000004.tif7127 (length: 4 characters), Talented and loved freely (length: 7 characters).

[0054] The linguistic rule used in this division process is the indivisibility between adverbs and verbs. After division, the difference in length between the two lyric phrases is three characters.

[0055] S240: A plurality of lyric paragraphs and a plurality of melody paragraphs are aligned one by one, and a correspondence between a text unit in a lyric paragraph and a note in the corresponding melody paragraph is determined.

[0056] In some embodiments, a specific implementation of this step includes:

[0057] First, a correspondence between the lyrics paragraph and the melody paragraph is established.

[0058] In some embodiments, if the target lyric text can be divided into eight lyric paragraphs, there will be a total of eight melody paragraphs, and a correspondence between the first lyric paragraph and the first melody paragraph, a correspondence between the second lyric paragraph and the second melody paragraph, ..., a correspondence between the eighth lyric paragraph and the eighth melody paragraph is established.

[0059] Then, the text units in the lyric paragraphs that have a corresponding relationship are associated with the notes in the melody paragraphs.

[0060] There are several specific implementation methods for achieving "associating text units in a lyric paragraph with corresponding notes in a melody paragraph," and this application does not impose any restrictions on these. In some embodiments, the nth character in a lyric text unit is aligned sequentially with the nth note in a melody paragraph.

[0061] Alternatively, a specific implementation method of this step includes: for each melody paragraph, obtaining a plurality of predetermined lyric alignment templates corresponding to the melody paragraph, each lyric alignment template corresponding to a different number of lyric characters; selecting a target lyric alignment template from the plurality of lyric alignment templates, the number of characters corresponding to the target lyric alignment template being the number of characters of the lyric paragraph corresponding to the melody paragraph; and aligning the melody paragraph with the lyric paragraph corresponding to the melody paragraph based on the target lyric alignment template.

[0062] In some embodiments, if a melody paragraph contains seven notes, three lyric alignment templates are set for the melody paragraph, as follows:

[0063] The first template is xxxxxxx and applies to lyric paragraphs containing seven characters.

[0064] The second template is xxx-xxx and applies to lyric paragraphs containing six characters.

[0065] The third template is xx--xxx, which applies to lyric paragraphs containing five characters.

[0066] An "x" represents a note with corresponding lyrics, and a "-" represents a note without corresponding lyrics.

[0067] Assume that "Caixuehui Aide Ziyou" in the above example corresponds to the melody paragraph. Since "Caixuehui Aide Ziyou" has a total of 7 characters, select the first template as the target lyric alignment template. Then, align the first note in the melody paragraph with "Cai", the second note with "Xue", ···, and the seventh note with "You".

[0068] Assume that for a certain melody paragraph, the lyric alignment template corresponding to the melody paragraph has not been determined in advance. In fact, based on the melody paragraph, the first lyric alignment template, the second lyric alignment template, and the third lyric alignment template can be generated. In the first lyric alignment template, each note in the melody paragraph corresponds to one text unit (for example, a character). Until the number of text units (for example, characters) that the finally obtained lyric alignment template can match equals the set threshold, repeat the following steps: Based on the previous lyric alignment template, integrate the adjacent note with the closest note attack distance in the melody paragraph as a note pair, regard the note pair as a new note, and associate it with only one character to obtain a new lyric alignment template. Here, the set threshold is a positive integer and is greater than or equal to 1 and less than or equal to the total number of notes in the melody paragraph.

[0069] In some embodiments, assume that a melody paragraph contains a total of seven notes and the set threshold is four. Each note in the melody paragraph is associated with one character to obtain a first lyric alignment template. The first lyric alignment template can match seven characters. Adjacent notes in the melody paragraph with the closest note attack distance are combined and associated with one character to obtain a second lyric alignment template that can match six characters. Adjacent notes in the second lyric alignment template with the closest note attack distance are combined and associated with one character to obtain a third lyric alignment template that can match five characters. The adjacent notes with the closest note attack distance in the third lyric alignment template are merged and then associated with one character to obtain a fourth lyric alignment template, which can match four characters and is equal to the set threshold, so no new lyric alignment templates are generated. Therefore, four lyric alignment templates are set for the melody paragraph.

[0070] In some embodiments, FIG. 4 is a schematic diagram of a staff of a melody paragraph according to an embodiment of the present disclosure. Referring to FIG. 4, each note in the melody paragraph is associated with one character to obtain a first lyric alignment template. Because the adjacent notes surrounded by rounded frame 1 have the closest attack distance, these two notes are combined and associated with one character to obtain a second lyric alignment template. At this time, the number of characters that can be matched is reduced by one compared to the first lyric alignment template. After combining the adjacent notes surrounded by rounded frame 1, the adjacent notes surrounded by rounded frame 2 have the closest attack distance, so these two notes are combined and associated with one character to obtain a third lyric alignment template. At this time, the number of characters that can be matched is reduced by one compared to the second lyric alignment template. After merging the adjacent notes surrounded by rounded frame 2, the adjacent note surrounded by rounded frame 3 has the closest attack distance. These two notes are merged and associated with one character to obtain the fourth lyric alignment template. At this time, the number of characters that can be matched is reduced by one compared to the third lyric alignment template. By repeating this process, multiple lyric alignment templates are obtained. In some embodiments, the top six merged note pairs are displayed in Figure 4.

[0071] In some embodiments, while performing the following procedure: "based on the previous lyric alignment template, merge adjacent notes in the melody paragraph with the closest note attack distance into one note pair, treat the note pair as one new note, and associate it with only one character to obtain a new lyric alignment template," if there are multiple pairs of adjacent notes in the melody paragraph, and the multiple pairs of adjacent notes satisfy the conditions of having the same note attack distance and the closest note attack distance, then merge the first beat, third beat, and first half beat of each note in a measure.

[0072] S250: Based on the correspondence between the text units in the target lyrics text and the notes in the singing melody, the target lyrics text is synthesized into a singing voice to sing the target lyrics text with the singing melody.

[0073] S260: The singing voice and the accompaniment audio of the initial song are integrated to generate a target song.

[0074] The above technical solution details the method for aligning the target lyric text and the singing melody, which can reduce the difficulty of aligning the target lyric text and the singing melody, which is conducive to realizing the purpose of generating a song based on the target lyric text, and improves the efficiency and interest of users in song creation.

[0075] 5 is a schematic diagram of a display interface for realizing song generation based on target lyric text in a terminal according to an embodiment of the present disclosure. Referring to FIG. 5, the display interface divides song generation into three steps. The first step is a custom option part, the second step is a text input part, and the third step is a result playback part.

[0076] In the first step, the user can configure the song setting item, the tone setting item, and the intelligent completion item, all of which are configured as pull-down lists.

[0077] When the user triggers the song setting item, a pull-down list of songs may be displayed for the user to select one, and each song may be assigned a unique melody and accompaniment.

[0078] When the user triggers the tone setting item, a plurality of tones, for example, a male tone, a female tone, a child tone, etc., may be displayed in a pull-down list format, and the user may select one tone.

[0079] When a user triggers an intelligent completion item, multiple intelligent completion methods may be displayed in a pull-down list format. Intelligent completion can be viewed as adjusting the custom text entered by the user to suit the selected melody. For example, if the custom text entered by the user can be divided into a maximum of five lyric phrases, but the target melody contains six melody paragraphs, intelligent completion can be performed on the custom text entered by the user, completing it with, for example, “la-la-la” to form a new lyric phrase, so that the final custom lyrics entered by the user contain six lyric phrases. Alternatively, after the custom text entered by the user is divided, if a lyric phrase contains five characters, but the lyric alignment template corresponding to the target melody has the fewest number of characters that can be matched: seven, for example, the lyric alignment template can be used to complete it with “la-la” so that the number of characters in the lyric phrase matches the number of characters that can be matched by the lyric alignment template. To help users quickly understand the meaning of the intelligent completion setting items, an explanation can be provided below the intelligent completion setting items. In some embodiments, the interpretation instruction is "If the number of characters in the entered lyrics is insufficient, an intelligent completion method will be used to complete the song."

[0080] In the second step, to help users quickly understand the lyric creation method, a message is added below the lyrics input box: "For best results, enter four Chinese texts, each containing 8 to 17 Chinese characters and separated by punctuation or breaks. If the entered text is too long, it will be intelligently restructured." The message can be changed as needed. The user can also arrange the audio format of the generated song. After completing the arrangement selection settings in both the first and second steps, the user can trigger the "Generate Song" control in the second step to display the lyrics of the generated song in the lyrics input box. At this time, the lyrics are divided into lyric phrases.

[0081] In a third step, the user can preview the generated song by clicking the play control on the music player. The user can download the generated song by clicking the download control.

[0082] Although the above-described method embodiments are described as a series of operations for ease of explanation, those skilled in the art will recognize that the present invention is not limited to the order of operations described, since some steps may be performed in other orders or simultaneously according to the present invention. Furthermore, those skilled in the art will recognize that the embodiments described in the specification are preferred embodiments, and the operations and modules involved are not necessarily essential to the present invention.

[0083] The technical solution according to the embodiment of the present disclosure obtains a target lyric text input by a user, aligns the target lyric text with the singing melody of an initial song, determines the correspondence between the text units in the target lyric text and the notes in the singing melody, and then performs voice synthesis on the target lyric text based on the correspondence between the text units in the target lyric text and the notes in the singing melody to obtain a singing voice that recites the target lyric text with the singing melody. The singing voice is then integrated with the accompaniment audio of the initial song to generate the target song. In essence, the technical solution allows the user to write lyrics themselves and create a song based on the lyrics, singing melody, and accompaniment audio created by the user, thereby forming a completely new song. In this way, a user can create a song based on lyrics they have written themselves without having specialized skills in music creation, thereby improving the efficiency and interest of users in song creation.

[0084] 6 is a schematic diagram of the configuration of a song generation device according to an embodiment of the present disclosure. The song generation device according to an embodiment of the present disclosure may be located on a client or a server. Referring to FIG. 6, the song generation device specifically includes: an acquisition unit 61 for acquiring a target lyrics text input by a user; an alignment unit 62 for aligning the target lyric text with the singing melody of the initial song to determine a correspondence between the text units in the target lyric text and the musical notes in the singing melody, the singing melody being a singing melody for the initial lyrics of the initial song; a synthesis unit 63 for synthesizing the target lyrics text with a voice based on the correspondence between the text units in the target lyrics text and the notes in the singing melody to obtain a singing voice for singing the target lyrics text with the singing melody; and a generating unit 64 for integrating the singing voice with the accompaniment audio of the initial song to generate the target song.

[0085] In some embodiments, the song generation device comprises: a selection unit for selecting an initial song from a plurality of preset songs in response to an initial song selection operation; The apparatus further includes a determining unit for determining a corresponding singing melody and accompaniment audio based on the initial song.

[0086] In some embodiments, alignment portion 62 includes: Dividing the singing melody into a plurality of melody paragraphs; Dividing the target lyric text into a plurality of lyric paragraphs, the number of the plurality of lyric paragraphs being the same as the number of the plurality of melody paragraphs; The alignment is used to align the plurality of lyric paragraphs and the plurality of melody paragraphs one by one, and determine the correspondence between the text units in the lyric paragraphs and the notes in the corresponding melody paragraphs.

[0087] In some embodiments, the alignment unit 62 dividing the singing melody into multiple melody paragraphs includes determining one paragraph division point for each predetermined measure in the singing melody, adjusting the number of paragraph division points based on the number of notes included in the melody paragraph corresponding to each paragraph division point, and adjusting the position of each paragraph division point based on a note attack distance, wherein the note attack distance includes the duration pitch of the note attack and / or the pitch pitch of the note attack.

[0088] In some embodiments, the alignment unit 62 adjusting the number of paragraph division points based on the number of notes contained in the melody paragraph corresponding to each paragraph division point includes, for any one paragraph division point, deleting the paragraph division point if the number of notes contained in the melody paragraph corresponding to the paragraph division point is smaller than a first threshold, and adding one paragraph division point if the number of notes contained in the melody paragraph corresponding to the paragraph division point is greater than a second threshold.

[0089] In some embodiments, the alignment unit 62 adjusting the position of each paragraph division point based on the note attack distance includes searching for a position within a predetermined range of beats around any one paragraph division point where the note attack distance satisfies a predetermined condition, and setting that position as the position of the paragraph division point, wherein the predetermined condition is that the duration pitch of the note attacks before and after the paragraph division point is the largest, or that the duration pitch of the note attacks before and after the paragraph division point are equal and the pitch pitch is the largest.

[0090] In some embodiments, the alignment unit 62 dividing the target lyric text into lyric paragraphs includes performing word segmentation on the target lyric text to determine the part of speech corresponding to each word, and dividing the target lyric text into the lyric paragraphs based on the part of speech corresponding to each word, predefined linguistic rules, and the length of the vocal melody.

[0091] In some embodiments, the alignment unit 62 aligning the plurality of lyric paragraphs with the plurality of melody paragraphs one by one includes: for each of the melody paragraphs, obtaining a plurality of predetermined lyric alignment templates corresponding to the melody paragraph, each of the lyric alignment templates corresponding to a different number of lyric characters; selecting a target lyric alignment template from the plurality of lyric alignment templates, the number of characters corresponding to the target lyric alignment template being the number of characters of the lyric paragraph corresponding to the melody paragraph; and aligning the melody paragraph with the lyric paragraph corresponding to the melody paragraph based on the target lyric alignment template.

[0092] In some embodiments, the plurality of lyric alignment templates corresponding to the melody paragraph include a first lyric alignment template, a second lyric alignment template and a third lyric alignment template; In the first lyric alignment template, each note in the melody paragraph corresponds to one text unit; In the second lyric alignment template, adjacent notes in the melody paragraph with the closest note attack distance are combined into one note pair, and one note pair corresponds to one text unit; In the third lyric alignment template, adjacent notes in the second lyric alignment template with the closest note attack distance are combined into one note pair, and one note pair corresponds to one text unit.

[0093] The song generation device according to the embodiments of the present disclosure can perform the steps performed by the client or server in the song generation method according to the method embodiments of the present disclosure, and the execution steps and beneficial effects thereof will not be further described here.

[0094] In some embodiments, the division of each means in the song generation device is only a division of logical functions, and when actually realized, there may be other division methods. For example, at least two means in the song generation device may be realized as one means, or each means in the song generation device may be divided into multiple sub-means. As can be understood, each means or sub-means may be realized by electronic hardware or a combination of computer software and electronic hardware. Whether these functions are implemented in a hardware manner or a software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art may realize the described functions using different methods for each specific application.

[0095] 7 is an exemplary block diagram of a system including at least one computing device and at least one storage device for storing instructions, according to an embodiment of the present disclosure. In some embodiments, the system may be used for big data processing, and the at least one computing device and the at least one storage device may be arranged in a distributed manner, making the system a distributed data processing cluster.

[0096] 7, the system includes at least one processing unit 51 and at least one memory unit 52 for storing instructions. As can be appreciated, the memory unit 52 in this embodiment may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memory.

[0097] In some embodiments, storage device 52 stores elements such as executable units or data structures, or a subset thereof, or an extended set thereof, such as an operating system and application programs.

[0098] The operating system includes various system programs, such as a framework layer, a core library layer, a driver layer, etc., and is used to realize various basic tasks and process hardware-based tasks. The application programs include various application programs, such as a media player, a browser, etc., and are used to realize various application tasks. A program for realizing the song generation method according to an embodiment of the present disclosure may be included in the application programs.

[0099] In the embodiments of the present disclosure, the at least one computing device 51 may call programs or instructions, particularly programs or instructions stored in an application program, stored in the at least one storage device 52. The at least one computing device 51 is used to perform steps of each embodiment of the song generation method according to the embodiments of the present disclosure.

[0100] The song generation method according to the embodiments of the present disclosure may be applied to or realized by a computing device 51. The computing device 51 may be an integrated circuit chip having signal processing capabilities. In the implementation process, each step of the above method may be completed by an integrated logic circuit of hardware in the computing device 51 or by instructions in the form of software. The above computing device 51 may be a general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component, or the like. The general-purpose processor may be a microprocessor, or the processor may be any conventional processor, etc.

[0101] The steps of the song generation method according to the embodiments of the present disclosure may be directly implemented to be executed and completed by a hardware decoding processor, or may be implemented and completed by a combination of hardware and software units in the decoding processor. The software units may be located in a storage medium well-established in the art, such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, an electrically erasable programmable memory, a register, etc. The storage medium is located in the storage device 52, and the calculation device 51 reads information in the storage device 52 and combines the hardware to complete the steps of the method.

[0102] An embodiment of the present disclosure further proposes a computer-readable storage medium storing a program or instructions that, when executed by at least one computing device, causes the at least one computing device to perform steps of each embodiment of the song generation method. To avoid repetition, no further description will be given here. Here, the computing device may be computing device 51 shown in FIG. 6. In some embodiments, the computer-readable storage medium is a non-transitory computer-readable storage medium.

[0103] An embodiment of the present disclosure further proposes a computer program product, the computer program product including a computer program, the computer program being stored in a non-transitory computer-readable storage medium, and at least one processor of a computer reading and executing the computer program from the storage medium to cause the computer to perform steps of each embodiment of the song generation method, which will not be further described herein to avoid repetition.

[0104] An embodiment of the present disclosure further proposes a computer program comprising instructions that, when executed by a processor, cause the processor to perform a song generation method according to any embodiment of the present disclosure.

[0105] It should be noted that, in this specification, the terms "comprises," "including," or any other variation thereof, are intended to cover the non-exclusive "comprises," whereby a process, method, article, or apparatus that includes a set of elements not only includes those elements, but also includes other elements not expressly listed or that are inherent in such process, method, article, or apparatus. In the absence of further limitations, an element qualified by the phrase "comprises" does not exclude the presence of other identical elements in a process, method, article, or apparatus that includes that element.

[0106] As will be understood by those skilled in the art, while some embodiments described herein include certain features and not other features included in other embodiments, it is understood that combinations of features from different embodiments are within the scope of the present disclosure and are meant to form different embodiments.

[0107] As can be understood by those skilled in the art, the description of each embodiment is focused on its own, and for the parts not described in detail in one embodiment, please refer to the description of other embodiments.

[0108] Although the embodiments of the present disclosure have been described with reference to the drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present disclosure, and such modifications and variations fall within the scope defined by the appended claims.

Claims

1. A computer-implemented song generation method comprising: obtaining target lyric text entered by a user; aligning the target lyric text with a singing melody of an initial song to determine a correspondence between text units in the target lyric text and notes in the singing melody, the singing melody being a singing melody for initial lyrics in the initial song; synthesizing the target lyrics text based on the correspondence between the text units in the target lyrics text and the notes in the singing melody to obtain a singing voice that sings the target lyrics text with the singing melody; and integrating the singing voice with an audio accompaniment of the initial song to generate a target song; Aligning the target lyric text with the vocal melody of the initial song includes: Dividing the singing melody into a plurality of melody paragraphs; Dividing the target lyric text into a plurality of lyric paragraphs, the number of the plurality of lyric paragraphs being the same as the number of the plurality of melody paragraphs; aligning the plurality of lyric paragraphs and the plurality of melody paragraphs one by one, and determining a correspondence between text units in the lyric paragraphs and notes in the corresponding melody paragraphs; Dividing the singing melody into a plurality of melody paragraphs determining a paragraph division point for each predetermined measure in the singing melody; adjusting the number of paragraph division points based on the number of notes included in the melody paragraph corresponding to each paragraph division point; 10. A song generation method comprising: adjusting the position of each paragraph division point based on a note attack distance, wherein the note attack distance includes a duration pitch of the note attack and / or a pitch pitch of the note attack.

2. Prior to aligning the target lyric text with the vocal melody of the initial song, the method further comprises: selecting an initial song from a plurality of preset songs in response to an initial song selection operation; 10. The song generation method of claim 1, further comprising: determining a corresponding singing melody and accompanying audio based on the initial song.

3. adjusting the number of paragraph division points based on the number of notes included in the melody paragraph corresponding to each of the paragraph division points; For any one paragraph division point, if the number of notes included in the melody paragraph corresponding to the one paragraph division point is smaller than a first threshold value, deleting the one paragraph division point; 2. The song generation method according to claim 1, further comprising: if the number of notes contained in the melody paragraph corresponding to any one of the paragraph division points is greater than a second threshold, increasing the paragraph division point by one.

4. Adjusting the position of each paragraph break point based on note attack distance searching for a position where a note attack distance satisfies a predetermined condition within a predetermined range of beats around any one paragraph division point, and determining that position as the position of the any one paragraph division point; 2. The song generation method according to claim 1, wherein the predetermined condition is that the duration pitch of the note attacks before and after any one of the paragraph division points is the largest, or that the duration pitch of the note attacks before and after any one of the paragraph division points is equal and the pitch pitch is the largest.

5. Dividing the target lyric text into a plurality of lyric paragraphs includes: segmenting the target lyrics text into words and determining the part of speech corresponding to each word; 2. The song generation method of claim 1, further comprising: dividing the target lyric text into a plurality of lyric paragraphs based on parts of speech corresponding to each word, predefined linguistic rules, and the length of the sung melody.

6. The alignment of the plurality of lyric paragraphs and the plurality of melody paragraphs one by one includes: For each of the melody paragraphs, obtaining a plurality of predetermined lyric alignment templates corresponding to the melody paragraphs, each of the lyric alignment templates corresponding to a different number of lyric characters; selecting a target lyric alignment template from the plurality of lyric alignment templates, and the number of characters corresponding to the target lyric alignment template is the number of characters in a lyric paragraph corresponding to the melody paragraph; 6. The method of claim 5, further comprising aligning the melody paragraph with a lyric paragraph corresponding to the melody paragraph based on the target lyric alignment template.

7. the plurality of lyric alignment templates corresponding to the melody paragraph include a first lyric alignment template, a second lyric alignment template and a third lyric alignment template; In the first lyric alignment template, each note in the melody paragraph corresponds to one text unit; In the second lyric alignment template, adjacent notes in the melody paragraph with the closest note attack distance are combined into one note pair, and one note pair corresponds to one text unit; 7. The song generation method of claim 6, wherein in the third lyric alignment template, adjacent notes in the second lyric alignment template with the closest note attack distance are combined into one note pair, and one note pair corresponds to one text unit.

8. The alignment of the plurality of lyric paragraphs and the plurality of melody paragraphs one by one includes: Establishing a correspondence between the lyric paragraph and the melody paragraph; 6. The method of claim 5, further comprising: associating text units in a lyric paragraph with corresponding notes in a melody paragraph.

9. Segmenting the target lyric text into multiple lyric paragraphs based on the parts of speech corresponding to each word and predefined linguistic rules, and after the segmentation, locating indivisible phrases in the same lyric paragraph; and / or 6. The song generation method according to claim 5, wherein the target lyrics text is divided into a plurality of lyrics paragraphs based on the length of the singing melody, and after the division, the number of characters in each lyrics paragraph is set to be equal to or less than the number of notes in the melody paragraph corresponding to the lyrics paragraph.

10. an acquisition unit for acquiring a target lyric text input by a user; an alignment unit for aligning the target lyric text with a singing melody of an initial song to determine a correspondence between text units in the target lyric text and notes in the singing melody, the singing melody being a singing melody corresponding to initial lyrics in the initial song; a synthesis unit for synthesizing the target lyrics text with a voice based on the correspondence between text units in the target lyrics text and notes in the singing melody to obtain a singing voice for singing the target lyrics text with the singing melody; a generating unit for integrating the singing voice and an accompaniment audio of the initial song to generate a target song; Aligning the target lyric text with the vocal melody of the initial song includes: Dividing the singing melody into a plurality of melody paragraphs; Dividing the target lyric text into a plurality of lyric paragraphs, the number of the plurality of lyric paragraphs being the same as the number of the plurality of melody paragraphs; aligning the plurality of lyric paragraphs and the plurality of melody paragraphs one by one, and determining a correspondence between text units in the lyric paragraphs and notes in the corresponding melody paragraphs; Dividing the singing melody into a plurality of melody paragraphs determining a paragraph division point for each predetermined measure in the singing melody; adjusting the number of paragraph division points based on the number of notes included in the melody paragraph corresponding to each paragraph division point; and adjusting the position of each paragraph division point based on a note attack distance, the note attack distance including a duration pitch of the note attack and / or a pitch pitch of the note attack.

11. 10. A system comprising at least one computing device and at least one storage device storing instructions that, when executed by said at least one computing device, cause said at least one computing device to perform the steps of the song generation method of any one of claims 1 to 9.

12. A computer readable storage medium storing a program or instructions which, when executed by at least one computing device, causes the at least one computing device to perform the steps of the song generation method of any one of claims 1 to 9.

13. A computer program comprising instructions which, when executed by a processor, cause the processor to carry out the song generation method of any one of claims 1 to 9.

Citation Information

Patent Citations

  • System and method for automatically generating musical output

    CN111213200A

  • Audio synthesis method, computer program therefor, computer device, and computer system configured with the computer device

    JP2021516787A