AI song editing method and apparatus, electronic device, and storage medium
By displaying and responding to user modification commands to generate second lyrics and the original song for AI songs, the problem of inaccurate lyrics modification in AI songs has been solved, achieving flexible and accurate lyrics modification and melody stability, thus improving the user experience.
Patent Information
- Application Number
- CN202411795467.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-06
- Publication Date
- 2025-12-19
- Estimated Expiration
- 2044-12-06
AI Technical Summary
Current technologies cannot accurately modify the lyrics of AI-generated songs, and the interaction process is complex, affecting the user experience.
By displaying the lyrics of an AI song and responding to the user's modification instructions, a second text containing a specific timestamp is generated. A second song is then generated based on this text, where the melody of the target song segment is the same as that of the original song segment, while the melody of the non-target song segments remains unchanged.
It enables precise modification of AI-generated song lyrics, improving the flexibility and accuracy of lyric modification, simplifying the interaction process, and ensuring the stability of the song melody.
Smart Images

Figure CN119580673B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of Internet technology, and in particular to an AI song editing method, apparatus, electronic device, and storage medium. Background Technology
[0002] Currently, artificial intelligence (AI) technology is increasingly being applied to various industries. For example, by training a song generation model using music samples, the model can be made capable of "creating" music, thereby generating "original songs" that meet user requirements. These are songs generated based on AIGC technology, or simply AI songs.
[0003] After generating AI songs using AIGC technology, the lyrics in AI songs are often inappropriate or inaccurate because they are generated by AI models. In the current technology, the only way to try to obtain a song that meets the user's requirements is to add corresponding restrictions to the prompts and regenerate the AI song.
[0004] However, existing solutions for adjusting lyrics in AI songs have limitations, such as the inability to accurately modify the lyrics and complex interaction processes, which negatively impact the user experience. Summary of the Invention
[0005] This disclosure provides an AI song editing method, apparatus, electronic device, and storage medium to overcome the problems of inaccurate lyrics modification and complex interaction processes in AI songs.
[0006] In a first aspect, embodiments of this disclosure provide an AI song editing method, including:
[0007] After generating a first song and its corresponding first lyrics, the first lyrics of the first song are displayed, wherein the first song is audio data generated based on artificial intelligence technology; in response to a modification instruction for a first character in the first lyrics, second lyrics are generated, wherein the second lyrics contain a second character corresponding to the playback timestamp of the first character; in response to the generation of the second lyrics, a second song is generated based on the second character, wherein the second song includes a target song segment and a non-target song segment, wherein the target song segment is the song segment corresponding to the second character, and the non-target song segment is the remaining song segment in the second song excluding the target song segment, and the melody of the non-target song segment is the same as the melody of the second original song segment in the first song corresponding to the playback timestamp of the non-target song segment.
[0008] Secondly, embodiments of this disclosure provide an AI song editing device, comprising:
[0009] The display module is used to display the first lyrics of the first song after the first song and the corresponding first lyrics are generated, wherein the first song is audio data generated based on artificial intelligence technology;
[0010] The processing module is used to generate second lyrics in response to a modification instruction for the first character in the first lyrics, wherein the second lyrics contain a second character corresponding to the playback timestamp of the first character;
[0011] A generation module is used to generate a second song based on the second text in response to the generation of the second lyrics. The second song includes a target song segment and a non-target song segment. The target song segment is the song segment corresponding to the second text. The non-target song segment is the remaining song segment in the second song excluding the target song segment. The melody of the non-target song segment is the same as the melody of the second original song segment in the first song corresponding to the playback timestamp of the non-target song segment.
[0012] Thirdly, embodiments of this disclosure provide an electronic device, including: a processor and a memory;
[0013] The memory stores computer-executed instructions;
[0014] The processor executes computer execution instructions stored in the memory, causing the at least one processor to perform the AI song editing method as described in the first aspect and various possible designs of the first aspect.
[0015] Fourthly, embodiments of this disclosure provide a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the AI song editing method described in the first aspect and various possible designs of the first aspect.
[0016] Fifthly, embodiments of this disclosure provide a computer program product, including a computer program that, when executed by a processor, implements the AI song editing method as described in the first aspect and various possible designs of the first aspect.
[0017] The AI song editing method, apparatus, electronic device, and storage medium provided in this embodiment, after generating a first song and corresponding first lyrics, displays the first lyrics of the first song, wherein the first song is audio data generated based on artificial intelligence technology; in response to a modification instruction for a first character in the first lyrics, second lyrics are generated, the second lyrics containing a second character corresponding to the playback timestamp of the first character; in response to the generation of the second lyrics, a second song is generated based on the second character, wherein the second song includes a target song segment and a non-target song segment, the target song segment being the song segment corresponding to the second character, and the non-target song segment being the remaining song segments in the second song excluding the target song segment, the melody of the non-target song segment being the same as the melody of the second original song segment in the first song corresponding to the playback timestamp of the non-target song segment. By displaying the first lyrics of the first song, modifying the first lyrics to the second lyrics in response to a modification instruction, and then generating the target song segment in the second song based on the second character in the second lyrics that has changed relative to the first lyrics, while the non-target song segments in the second song remain unchanged, precise modification of the lyrics is achieved without altering the melody of the song, thereby improving the flexibility and accuracy of AI song lyric modification and increasing interaction efficiency. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in the embodiments of this disclosure or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 This is an application scenario diagram of the AI song editing method provided in the embodiments of this disclosure;
[0020] Figure 2 Flowchart of the AI song editing method provided in this embodiment of the disclosure Figure 1 ;
[0021] Figure 3 A schematic diagram of the interface of a target application provided in an embodiment of this disclosure;
[0022] Figure 4 for Figure 2 A flowchart illustrating the specific implementation of step S102 in the illustrated embodiment;
[0023] Figure 5 This is a schematic diagram illustrating a process for generating second lyrics provided in an embodiment of the present disclosure;
[0024] Figure 6 This is a schematic diagram illustrating a process for generating a second song, provided as an embodiment of the present disclosure.
[0025] Figure 7 Flowchart of the AI song editing method provided in this embodiment. Figure 2 ;
[0026] Figure 8 for Figure 2 A flowchart of one possible implementation of step S205 in the illustrated embodiment;
[0027] Figure 9 for Figure 2 A flowchart of another possible implementation of step S205 in the illustrated embodiment;
[0028] Figure 10 This is a structural block diagram of the AI song editing device provided in the embodiments of this disclosure;
[0029] Figure 11 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present disclosure;
[0030] Figure 12 This is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation
[0031] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this disclosure, and not all embodiments. Based on the embodiments of this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.
[0032] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this disclosure are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation portals are provided for users to choose to authorize or refuse.
[0033] The application scenarios of the embodiments of this disclosure are explained below:
[0034] Figure 1This diagram illustrates an application scenario of the AI song editing method provided in this embodiment. The AI song editing method can be applied to applications (APPs) with song generation and editing functions, such as song applications and short video applications. More specifically, it can be applied to scenarios involving modifying lyrics in AI songs. The executing entity in this embodiment can be a terminal device running the aforementioned application with song generation and editing functions, a server deploying the server-side component corresponding to the aforementioned application, or other electronic devices performing similar functions. When the executing entity is a terminal device, the terminal device executes the method provided in this embodiment by running the aforementioned application. When the executing entity is a server, the server-side component of the aforementioned application with song generation or video generation functions can run partially or entirely on the server, executing the method provided in this embodiment on the server side, while the terminal device runs the client-side component of the application. Communication between the server and the terminal device is based on server-client communication, enabling the terminal device to obtain the execution result of the method provided in this embodiment and display it as needed.
[0035] In some embodiments, the terminal device or server can implement the AI song editing method provided in this disclosure by running various computer-executable instructions or computer programs. For example, computer-executable instructions can be program-level commands, machine instructions, or software instructions. Computer programs can be native programs or software modules in an operating system; they can be local applications, i.e., programs that need to be installed in the operating system to run, or mini-programs embedded in any app, i.e., programs that run in a browser environment. In summary, the aforementioned computer-executable instructions can be any form of instruction, and the aforementioned computer programs can be any form of application, module, or plugin; the specific implementation can be configured as needed. Furthermore, in implementing the AI song editing method provided in this disclosure, the terminal device can execute the method by running computer-executable instructions or computer programs set locally, or by calling computer-executable instructions or computer programs set in an external server. In some embodiments, the server may be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud storage, cloud communication, cloud database, cloud computing, cloud functions, network services, middleware services, domain name services, security services, content delivery network (CDN), and big data and artificial intelligence platforms. Among these, cloud services may be interactive processing services that can be invoked by terminal devices.
[0036] refer to Figure 1 As shown in the diagram, taking a terminal device as an example, a target application with song generation capabilities runs on the terminal device. After the target application is launched, the user inputs prompts for song generation into the target application by operating the terminal device, such as "sad," "lyrical," and "clear river," as shown in the diagram. Then, based on these prompts, the target application calls a song generation model deployed in the cloud to generate a corresponding "original song," i.e., an AI song (shown as Music_AI_No.001 in the diagram), and sends it back to the terminal device. Upon receiving the AI song, the terminal device displays the lyrics (represented by "X" in the diagram) and plays the AI song based on user commands.
[0037] In existing technologies, after generating AI songs using AI technology, the lyrics are generated by the AI model, which users cannot precisely edit. This often leads to problems such as inappropriate lyrics and inaccurate meaning. To address this, current technologies can only regenerate the AI song by adding appropriate restrictions to prompts to avoid inappropriate lyrics. However, since the melody and lyrics of the AI song are generated simultaneously by the model, there is a coupling between them. Therefore, regenerating the AI song may result in changes to the melody, making it impossible to modify only the lyrics without altering the melody. This results in the inability to accurately modify the lyrics of AI songs and a complex interaction process, negatively impacting the user experience.
[0038] This disclosure provides an AI song editing method to solve the above-mentioned problems.
[0039] refer to Figure 2 , Figure 2 Flowchart of the AI song editing method provided in this embodiment of the disclosure Figure 1 The method described in this embodiment can be applied to terminal devices. This AI song editing method includes:
[0040] Step S101: After generating the first song and the corresponding first lyrics, display the first lyrics of the first song, where the first song is audio data generated based on artificial intelligence technology.
[0041] Step S102: In response to the modification instruction for the first character in the first lyrics, generate the second lyrics, which contain the second character corresponding to the playback timestamp of the first character.
[0042] For example, refer to Figure 1The illustrated application scenario shows a terminal device running a target application with song generation capabilities. This application uses AIGC technology to generate an AI song, i.e., the first song. The process of generating such an AI song using AIGC technology typically includes receiving a prompt input from the user, and then calling a pre-trained song generation model to generate music matching the prompt. These steps will not be elaborated upon here. After generating the first song, the target application displays the first song and its corresponding lyrics.
[0043] Furthermore, after displaying the first lyrics of the first song, when the user needs to modify the lyrics of the first song, a modification command will be input to the terminal device. This modification command targets the first character in the first lyrics, specifically, adding, deleting, or changing one or more of the content of the first lyrics, where the changed content in the first lyrics is the first character. Figure 3 This is a schematic diagram of the interface of a target application provided in an embodiment of the present disclosure, such as... Figure 3 As shown, exemplarily, after the target application generates the first song, the song name is Music_AI_No.001. When playing or navigating to this first song, the first lyrics of the song are simultaneously displayed in the target application's playback interface, such as "AAAAA" and "BBBBB" as shown in the figure. The modification instruction modifies the 2nd, 3rd, and 4th characters "AAA" in the text segment "AAAAA," changing them to "CCC," resulting in the modified text segment becoming "ACCCA." The original character "AAA" is the first character; the modified character "CCC" is the second character. Correspondingly, the lyrics text containing the modified text segment is the second lyric. In another possible implementation, the modification instruction modifies multiple first characters, and at least two of the first characters are not consecutive in the first lyric; therefore, the corresponding generated second character is also not consecutive in the second lyric.
[0044] One possible implementation is, such as Figure 4 As shown, the specific implementation of step S102 includes:
[0045] Step S1021: In response to the modification instruction for the first character in the target lyrics text segment of the first lyrics, obtain the corresponding text segment to be determined, and detect the number of characters in the text segment to be determined.
[0046] Step S1022: If the number of characters in the text segment to be determined is equal to the number of characters in the target lyrics text segment, then the text segment to be determined is determined as the modified lyrics text segment.
[0047] Step S1023: Replace the target lyrics text segment in the first lyrics with the modified lyrics text segment to generate the second lyrics.
[0048] For example, the user-inputted modification command is used to modify the first character within a target lyric text segment in the first lyrics. The first lyrics consist of multiple lyric text segments, each corresponding to a sentence, which can be divided, for example, by punctuation marks. Each lyric text segment consists of one or more characters, such as a single Chinese character. After modifying one or more first characters based on the modification command, a modified text segment, i.e., the text segment to be determined, is first generated. Then, this text segment to be determined is checked to ensure that the number of Chinese characters in the modified text segment to be determined is consistent with the number of characters in the target lyric text segment; that is, the modification command cannot change the number of characters in the target lyric text segment. If the check confirms that the number of characters in the text segment to be determined is equal to the number of characters in the target lyric text segment, the generated modified text segment to be determined is determined as the modified lyric text segment. Next, the target lyrics text segment in the first lyrics is replaced with the modified lyrics text segment to generate the second lyrics, thus realizing the generation of the second lyrics; in another case, if the number of characters in the text segment to be determined is not equal to the number of characters in the target lyrics text segment after detection, a prompt message can be displayed to prompt the user to re-enter the modification command.
[0049] Due to the unique nature of AI songs, changes in the amount of lyrics can alter the vocal components (i.e., the vocal performance), potentially leading to a mismatch between the vocals and the melody. For example, increasing the amount of lyrics, while maintaining the melody, necessitates a faster vocal pace, resulting in a disharmony between the melody and the vocals, thus affecting the quality of the final generated song. In this embodiment, to address this issue, when the user modifies the lyrics (after the terminal device responds to the modification command), the terminal device checks the number of characters in the text segment to be determined based on the modification command. Only if the number of characters matches the original target lyrics text segment will subsequent steps be executed. This ensures consistency in the number of characters in the lyrics text before and after modification, preventing a mismatch between the melody and the vocals and improving the quality of the final generated second song.
[0050] Further, optionally, the steps in this embodiment also include:
[0051] Step S1024: If the number of characters in the text segment to be determined is not equal to the number of characters in the target lyrics text segment, then call the preset generation model to process the text segment to be determined, generate an optimized text segment, the number of characters in the optimized text segment is equal to the number of characters in the target lyrics text segment, and has the same semantics as the text segment to be determined.
[0052] Step S1025: Replace the target lyrics text segment in the first lyrics with the optimized text segment to generate the second lyrics.
[0053] For example, in another scenario, if the number of characters in the text segment to be determined is not equal to the number of characters in the target lyrics text segment, in addition to rejecting the modification instruction and displaying a prompt, the characteristics of generating and modifying AI songs on the same platform (i.e., the processes of generating the first song and generating the second song based on the first song are both implemented through the target application) can be utilized to optimize the user-input modification instruction, thereby obtaining a text segment with the same number of characters as the target lyrics text segment and having the same semantics as the text segment to be determined, i.e., an optimized text segment. This process can be implemented through a preset generation model, that is, processing the text segment to be determined through a preset generation model to generate the optimized text segment. In this embodiment, the preset generation model can be a song generation model that generates the first song and the corresponding first lyrics. By using the prior knowledge within the song generation model to optimize the text segment to be determined, the generated optimized text segment can have better consistency with the original lyrics (first lyrics) and better musical attributes, such as better rhythm, thereby giving the second song generated from the modified second lyrics better musical quality.
[0054] Figure 5 This is a schematic diagram illustrating a process for generating second lyrics provided in an embodiment of the present disclosure, such as... Figure 5 As shown, firstly, based on the modification instructions, the target lyric text segment (Seg_1) containing the first modified character is determined from the first lyric (represented by "X" in the figure). The target lyric text segment is then modified into a text segment to be determined (Seg_2 in the figure) by modifying the first character. Next, the number of characters in the text segment to be determined is checked. If the number of characters in the text segment to be determined is equal to the number of characters in the target lyric text segment (Y path in the figure), the original target lyric text segment is replaced with the text segment to be determined, thereby generating the second lyric. If the number of characters in the text segment to be determined is not equal to the number of characters in the target lyric text segment (N path in the figure), the text segment to be determined is processed using the song generation model. For example, the second character is modified into an optimized text segment (Seg_3 in the figure), and the optimized text segment is used to replace the original target lyric text segment, thereby generating the second lyric.
[0055] In this embodiment, when the number of characters in the text segment to be determined is not equal to the number of characters in the target lyrics text segment, the song generation module optimizes the text segment to be determined to generate an optimized text segment that better matches the original first song. Based on the optimized text segment, the second lyrics are generated. On the one hand, this ensures that the number of characters in the lyrics before and after modification is consistent, thereby avoiding the problem of disharmony between the melody and the vocal performance. On the other hand, it can improve the consistency between the modified second lyrics and the original first lyrics, as well as the musicality of the second lyrics themselves, so that the second song generated from the modified second lyrics has better musical quality.
[0056] Step S103: In response to the generation of the second lyrics, generate a second song based on the second text. The second song includes a target song segment and a non-target song segment. The target song segment is the song segment corresponding to the second text. The non-target song segment is the remaining song segment in the second song excluding the target song segment. The melody of the non-target song segment is the same as the melody of the second original song segment in the first song corresponding to the playback timestamp of the non-target song segment.
[0057] For example, further, after the second lyrics are generated (i.e., after the lyrics of the first song are modified), the terminal device will regenerate the second lyrics based on the second character in the second lyrics. Specifically, the terminal device uses the second character to regenerate the corresponding original song segment (i.e., the first original song segment) in the first song to generate the corresponding target song segment. The other song segments in the first song (i.e., the second original song segment) are not adjusted and are directly used as non-target song segments. Then, the target song segment and the non-target song segments are merged and combined to generate the second song. In one possible implementation, the above steps can be implemented by a song generation model. For example, the second character, the playback timestamp corresponding to the second character, and the first song are input into the song generation model, and the song generation model can output a second song with the above characteristics. In this process, only the song segment corresponding to the second character involved in the modification instruction is regenerated. Therefore, at least the melody of other parts of the song will not be affected, thereby achieving accurate modification of the specified lyrics in the AI song.
[0058] Figure 6 This is a schematic diagram illustrating a process for generating a second song according to an embodiment of this disclosure. The following is in conjunction with... Figure 6 For further details on the above process, please refer to [link / reference]. Figure 6As shown, firstly, based on the user's input modification command, the first text_1 in the first lyrics is determined. This first text_1 can consist of one or more Chinese characters, and its content is, for example, "He looked at me but didn't speak." Then, the first text_1 is modified to the second text_2, and the content of the second text_2 is, for example, "He didn't look at me and didn't speak." The playback timestamps of the first text include t1 and t2, representing the start and end times of the first text_1, respectively. Next, based on the second text_2, a new song segment, the target song segment (shown as Seg_1 in the diagram), is generated using a song generation model. On the other hand, based on the playback timestamps t1 and t2, a non-target song segment (shown as Seg_2 in the diagram) is extracted from the first song. Finally, the target song segment and the non-target song segment are combined to generate the second song.
[0059] In this embodiment, after generating a first song and its corresponding first lyrics, the first lyrics of the first song are displayed. The first song is audio data generated based on artificial intelligence technology. In response to a modification instruction for a first character in the first lyrics, second lyrics are generated. The second lyrics contain a second character corresponding to the playback timestamp of the first character. In response to the generation of the second lyrics, a second song is generated based on the second character. The second song includes a target song segment and non-target song segments. The target song segment is the song segment corresponding to the second character, and the non-target song segments are the remaining song segments in the second song excluding the target song segment. The melody of the non-target song segments is the same as the melody of the second original song segment in the first song corresponding to the playback timestamp of the non-target song segment. By displaying the first lyrics of the first song, modifying the first lyrics to the second lyrics in response to a modification instruction, and then generating the target song segment in the second song based on the second character in the second lyrics that has changed relative to the first lyrics, while the non-target song segments in the second song remain unchanged, precise modification of the lyrics is achieved without altering the melody of the song. This improves the flexibility and accuracy of lyric modification in AI songs and enhances interaction efficiency.
[0060] refer to Figure 7 , Figure 7 Flowchart of the AI song editing method provided in this embodiment. Figure 2 This embodiment is in Figure 2 Based on the illustrated embodiment, steps S102-S103 are further refined, and the AI song editing method includes:
[0061] Step S201: Obtain the first song and the corresponding first lyrics. The first lyrics include multiple lyric text segments, and each lyric text segment corresponds to an original song segment.
[0062] Step S202: Display the lyrics text segment of the first lyric in separate lines, and the corresponding editing control for the lyrics text segment. The editing control is configured to trigger the lyrics text to an editable state.
[0063] Step S203: In response to the first instruction of the modification control corresponding to the target lyrics text segment, the target lyrics text segment is triggered to an editable state.
[0064] Step S204: In response to the second instruction for the target lyrics text segment in the editable state, modify the first character in the target lyrics text segment to the second character to generate the second lyrics.
[0065] For example, after the terminal device generates a first song and corresponding first lyrics through the target application, it displays each lyric text segment of the first lyrics in a line-by-line manner. Each lyric text segment can be a paragraph or a sentence within the first lyrics. The lyric text segments can be displayed corresponding to the playback progress of the first song. Specifically, for example, as the first song plays or jumps to a specific point in time (T), the lyric text segment containing the lyrics at that point (T) is also displayed synchronously. Specifically, this can be done by displaying only that lyric text segment, or by displaying multiple lyric text segments, but highlighting the lyric text segment corresponding to point (T). The specific implementation method is not limited.
[0066] Furthermore, when displaying each lyric text segment, an editing control is also shown on the line containing that lyric text segment. When the editing control is triggered, for example, when the user clicks the editing control (triggers the first instruction), the lyric text segment in that line becomes editable. The user can then edit the lyric text segment in that line (triggers the second instruction). When editing is complete, for example, when the cursor leaves the line, the lyric text segment reverts to a non-editable state. In this embodiment, by setting an editing control for each lyric text segment, the editable state of the lyric text segment is controlled, thereby avoiding accidental editing and modification of the lyric text segment and improving the operational efficiency during the lyric modification process.
[0067] Optionally, it also includes:
[0068] Step S204A: Before and after responding to the modification instruction, display the playback timestamp and / or playback duration corresponding to the lyrics text segment.
[0069] For example, the playback timestamp corresponding to the lyrics text segment, such as the playback time in the first song corresponding to the start position of the lyrics text segment, is used to represent the position of the vocal component (vocal performance) corresponding to the lyrics text segment in the first song; while the playback duration represents the duration of the vocal component corresponding to the lyrics text segment in the first song. After the terminal device responds to the modification command, if the length of the lyrics text segment changes, the position and duration of the vocal component corresponding to the lyrics text segment may also change accordingly. In this embodiment, before and after responding to the modification command, the playback timestamp and / or playback duration corresponding to the lyrics text segment are displayed to show the changes in the song caused by the modification command, so that the user can control the content and duration of the final generated song based on this information, improving the interaction efficiency when the song length changes due to changes in lyrics.
[0070] Step S205: In response to the generation of the second lyrics, generate a second song based on the second text. The second song includes a target song segment and a non-target song segment. The target song segment is the song segment corresponding to the second text. The non-target song segments are the remaining song segments in the second song excluding the target song segment. The melody of the non-target song segment is the same as the melody of the original song segment in the first song corresponding to the playback timestamp of the non-target song segment.
[0071] For example, after modifying the first text and generating the second text, the terminal device will use the newly generated second text to synchronously modify the song content in the first song, thereby ensuring consistency between the song content and the lyrics. Here, the song content includes at least musical and vocal components. The vocal components refer to the parts sung by humans in the song; this part is related to the lyrics, and if the lyrics change, the corresponding vocal components also need to be adjusted based on AIGC technology. The musical components refer to various background music used to express the melody of the song, harmonies not related to the lyrics, etc. In one implementation, if the lyrics change, the musical components do not change, that is, the melody of the song does not change, only the vocal components change. In this case, such as Figure 8 As shown, the specific implementation of step S205 includes:
[0072] Step S205A-1: Based on the playback timestamp of the first text, determine the first original song segment and the second original song segment in the first song. The first original song segment is the original song segment corresponding to the playback timestamp of the first text, and the second original song segment is the remaining original song segment in the first song excluding the first original song segment.
[0073] Step S205A-2: Obtain the pronunciation data of the second character, and replace the human voice component in the first original song segment based on the pronunciation data to generate the target song segment. The melody of the target song segment is the same as the melody of the first original song segment in the first song corresponding to the playback timestamp of the target song segment.
[0074] Step S205A-3: Generate a second song based on the target song segment and the second original song segment in the first song.
[0075] For example, in the steps of this embodiment, firstly, the terminal device determines the first original song segment and the second original song segment in the first song based on the playback timestamp of the first text indicated by the modification instruction. The first original song segment is the original song segment corresponding to the playback timestamp of the first text, and the second original song segment is the remaining original song segment in the first song excluding the first original song segment; that is, the first song is divided into two parts. Next, the pronunciation data of the modified second text is obtained. This pronunciation data is either pre-recorded audio data or audio data generated by inputting the second text into a pre-trained speech model. Then, the pronunciation data of the second text replaces the human voice component in the first original song segment, thereby generating the target song segment. This process can be implemented using a pre-trained speech model. The melody of the target song segment is the same as the melody of the corresponding part in the first song; that is, the musical components remain unchanged. Finally, the generated target song segment is concatenated with the second original song segment in the first song to generate the second song. Among them, steps S205A-1 and S205A-2 can be executed at once based on the same speech mode. That is, after inputting the first song, the second text and the corresponding playback timestamp into the song generation model, the song generation model will obtain the pronunciation data corresponding to the second text by calling external data. Then, based on the speech data, the human voice components in the first original song segment are replaced to generate the target song segment. The specific implementation process will not be described in detail.
[0076] In this embodiment, the vocal components at corresponding positions in the first song are replaced using the pronunciation data of the second character to generate the target song segment. Then, the target song segment and the second original song segment in the first song are concatenated to generate the second song. This implementation method is simple to execute; it only requires adjusting the vocal components of the modified song segment to generate the second song. Simultaneously, it ensures that the melody of the generated second song is completely consistent with the melody of the first song, achieving the goal of precisely modifying the lyrics without changing the melody.
[0077] In another possible implementation, the musical components of the modified song segment change; that is, both the melody and vocals change simultaneously. In this case, such as... Figure 9 As shown, the specific implementation of step S205 includes:
[0078] Step S205B-1: Based on the playback timestamp of the first text, determine the first original song segment and the second original song segment in the first song;
[0079] Step S205B-2: Call the song generation model to generate a target song segment based on the second text and the second original song segment; wherein the melody of the target song segment is different from the melody of the first original song segment in the first song corresponding to the playback timestamp of the target song segment;
[0080] Step S205B-3: Generate a second song based on the target song segment and the second original song segment.
[0081] For example, similarly, firstly, the terminal device determines the first and second original song segments in the first song based on the playback timestamp of the first text indicated by the modification instruction. Then, it calls the song generation model to re-"create" the second text, while combining the second original song segment as contextual information to generate an audio segment that matches the second original song segment, i.e., the target song segment. Because the song segment corresponding to the second text is regenerated in two dimensions—vocal and musical components—through the song generation model, and because the second text has changed in relation to the first text, the melody of the target song segment is usually different from the melody of the first original song segment in the first song corresponding to the playback timestamp of the target song segment. However, it makes the melody of the target song segment more compatible with the modified second text. At the same time, by combining the contextual information (the second original song segment), the consistency of the generated second song's melody is also improved, thus enhancing the song quality of the second song.
[0082] Furthermore, in one possible implementation, the first song and the second song belong to the same target song project. Specifically, one AI song corresponds to one song project; that is, the first song and the second song can be understood as different versions of the same AI song, both located under the same song project. Optionally, this embodiment also includes:
[0083] Step S206: Display the generation record page corresponding to the target song project. The generation record page is used to display at least one of the following: the song generation record corresponding to the target song project, which includes at least the generation record of the first song and the generation record of the second song; and the changes made to the second song relative to the first song.
[0084] For example, after modifying the first song and generating the second song, a generation record page corresponding to the target song item to which the first and second songs belong can be further displayed within the target application. This generation record page records the song generation records corresponding to the target song item, such as the generation record of the first song, the generation record of the second song, and the generation record of the third song generated after further modifications to the second song. The generation record may include relevant information such as the generation time and the user who made the modification. Furthermore, the generation record page also records the changes made to the song before the modification, such as the changes made to the second song relative to the first song in this embodiment. These changes include, for example, the first and second text displayed side-by-side, and changes in melody due to changes in lyrics. This allows users to obtain the modification history of the target song item through the generation record page corresponding to the target song item, and to restore the song for different versions, improving the efficiency of modifying songs.
[0085] Corresponding to the AI song editing method in the above embodiments, Figure 10 This is a structural block diagram of the AI song editing device provided in the embodiments of this disclosure. The method described in the above embodiments can be executed by this AI song editing device, which can be implemented by software and / or hardware, and can be integrated into an electronic device with certain data processing capabilities. The electronic device may include, but is not limited to, mobile terminals with big data processing capabilities, as well as fixed terminals with big data processing capabilities such as desktop computers and supercomputers.
[0086] For ease of explanation, only the parts relevant to embodiments of this disclosure are shown. (Refer to...) Figure 10 The AI song editing device 3 includes:
[0087] Display module 31 is used to display the first lyrics of the first song after the first song and the corresponding first lyrics are generated, wherein the first song is audio data generated based on artificial intelligence technology;
[0088] Processing module 32 is used to generate second lyrics in response to a modification instruction for the first character in the first lyrics, wherein the second lyrics contain a second character corresponding to the playback timestamp of the first character;
[0089] The generation module 33 is used to generate a second song based on the second text in response to the generation of the second lyrics. The second song includes a target song segment and a non-target song segment. The target song segment is the song segment corresponding to the second text, and the non-target song segment is the remaining song segment in the second song excluding the target song segment. The melody of the non-target song segment is the same as the melody of the second original song segment in the first song corresponding to the playback timestamp of the non-target song segment.
[0090] According to one or more embodiments of this disclosure, the first lyrics include multiple lyric text segments, each lyric text segment corresponding to an original song segment; when displaying the first lyrics of the first song, the display module 31 is specifically used to: display the lyric text segments of the first lyrics in separate lines, and the modification control corresponding to the lyric text segments, the modification control being configured to trigger the lyric text to an editable state; the modification instruction includes a first instruction and a second instruction, and the processing module 32 is specifically used to: in response to the first instruction for the modification control corresponding to the target lyric text segment, trigger the target lyric text segment to an editable state; in response to the second instruction for the target lyric text segment in the editable state, modify the first character in the target lyric text segment to the second character to generate the second lyrics.
[0091] According to one or more embodiments of this disclosure, the display module 31 is further configured to: display the playback timestamp and / or playback duration corresponding to the lyrics text segment before and after responding to the modification instruction.
[0092] According to one or more embodiments of this disclosure, the processing module 32 is specifically configured to: in response to a modification instruction for a first character in a target lyrics text segment in the first lyrics, obtain a corresponding text segment to be determined; detect the number of characters in the text segment to be determined, and if the number of characters in the text segment to be determined is equal to the number of characters in the target lyrics text segment, determine the text segment to be determined as the modified lyrics text segment; replace the target lyrics text segment in the first lyrics with the modified lyrics text segment, and generate the second lyrics.
[0093] According to one or more embodiments of this disclosure, the processing module 32 is further configured to: if the number of characters in the text segment to be determined is not equal to the number of characters in the target lyrics text segment, then call a preset generation model to process the text segment to be determined, generate an optimized text segment, wherein the number of characters in the optimized text segment is equal to the number of characters in the target lyrics text segment, and has the same semantics as the text segment to be determined; replace the target lyrics text segment in the first lyrics with the optimized text segment, and generate the second lyrics.
[0094] According to one or more embodiments of this disclosure, when generating a second song based on a second character, the generation module 33 is specifically configured to: determine a first original song segment and a second original song segment in the first song based on the playback timestamp of the first character, wherein the first original song segment is the original song segment corresponding to the playback timestamp of the first character, and the second original song segment is the remaining original song segment in the first song excluding the first original song segment; obtain the pronunciation data of the second character, and replace the human voice component in the first original song segment based on the pronunciation data to generate a target song segment, wherein the melody of the target song segment is the same as the melody of the first original song segment in the first song corresponding to the playback timestamp of the target song segment; and generate a second song based on the target song segment and the second original song segment in the first song.
[0095] According to one or more embodiments of this disclosure, when generating a second song based on the second text, the generation module 33 is specifically used to: determine the first original song segment and the second original song segment in the first song based on the playback timestamp of the first text; call the song generation model to generate a target song segment based on the second text and the second original song segment; wherein the melody of the target song segment is different from the melody of the first original song segment in the first song corresponding to the playback timestamp of the target song segment; and generate a second song based on the target song segment and the second original song segment.
[0096] According to one or more embodiments of this disclosure, the first song and the second song belong to the same target song project. The display module 31 is further configured to: display the generation record page corresponding to the target song project. The generation record page is configured to display at least one of the following: the song generation record corresponding to the target song project, the song generation record including at least the generation record of the first song and the generation record of the second song; and the changes of the second song relative to the first song.
[0097] The display module 31, processing module 32, and generation module 33 are connected sequentially. The AI song editing device 3 provided in this embodiment can execute the technical solution of the above method embodiment, and its implementation principle and technical effect are similar, so it will not be described again here.
[0098] Figure 11 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present disclosure, such as... Figure 11 As shown, the electronic device 4 includes:
[0099] Processor 41, and memory 42 communicatively connected to processor 41;
[0100] Memory 42 stores instructions executed by the computer;
[0101] The processor 41 executes computer execution instructions stored in the memory 42 to achieve, for example, Figures 2-9 The AI song editing method in the illustrated embodiment.
[0102] Optionally, the processor 41 and the memory 42 are connected via a bus 43.
[0103] For relevant instructions, please refer to the corresponding text. Figures 2-9 The relevant descriptions and effects of the steps in the corresponding embodiments are understood, and will not be elaborated on here.
[0104] This disclosure provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement this disclosure. Figures 2-9 The AI song editing method provided in any of the corresponding embodiments.
[0105] This disclosure provides a computer program product, including a computer program, which, when executed by a processor, implements this disclosure. Figures 2-9 The AI song editing method provided in any of the corresponding embodiments.
[0106] To implement the above embodiments, this disclosure also provides an electronic device.
[0107] refer to Figure 12 The diagram illustrates a structural schematic of an electronic device 900 suitable for implementing embodiments of the present disclosure. The electronic device 900 can be a terminal device or a server. The terminal device can include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, personal digital assistants (PDAs), tablet computers, portable media players (PMPs), and in-vehicle terminals (e.g., in-vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers. Figure 12 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.
[0108] like Figure 12 As shown, the electronic device 900 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 901, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 902 or a program loaded from a storage device 908 into a random access memory (RAM) 903. The RAM 903 also stores various programs and data required for the operation of the electronic device 900. The processing unit 901, ROM 902, and RAM 903 are interconnected via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.
[0109] Typically, the following devices can be connected to I / O interface 905: input devices 906 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 907 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 908 including, for example, magnetic tapes, hard disks, etc.; and communication devices 909. Communication device 909 allows electronic device 900 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 12An electronic device 900 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0110] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 909, or installed from a storage device 908, or installed from a ROM 902. When the computer program is executed by a processing device 901, it performs the functions defined in the methods of embodiments of this disclosure.
[0111] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0112] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0113] The aforementioned computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to perform the methods shown in the above embodiments.
[0114] Computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, and conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0115] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0116] The units or modules described in the embodiments of this disclosure can be implemented in software or hardware. The names of the units or modules do not necessarily limit the specific unit itself.
[0117] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0118] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0119] In a first aspect, according to one or more embodiments of this disclosure, an AI song editing method is provided, comprising:
[0120] After generating a first song and its corresponding first lyrics, the first lyrics of the first song are displayed, wherein the first song is audio data generated based on artificial intelligence technology; in response to a modification instruction for a first character in the first lyrics, second lyrics are generated, wherein the second lyrics contain a second character corresponding to the playback timestamp of the first character; in response to the generation of the second lyrics, a second song is generated based on the second character, wherein the second song includes a target song segment and a non-target song segment, wherein the target song segment is the song segment corresponding to the second character, and the non-target song segment is the remaining song segment in the second song excluding the target song segment, and the melody of the non-target song segment is the same as the melody of the original song segment in the first song corresponding to the playback timestamp of the non-target song segment.
[0121] According to one or more embodiments of this disclosure, the first lyrics include a plurality of lyric text segments, each of which corresponds to an original song segment; displaying the first lyrics of the first song includes: displaying the lyric text segments of the first lyrics line by line, and an editing control corresponding to each lyric text segment, the editing control being configured to trigger the lyric text to an editable state; the modification instruction includes a first instruction and a second instruction; generating the second lyrics in response to a modification instruction for a first character in the first lyrics includes: in response to a first instruction for the editing control corresponding to a target lyric text segment, triggering the target lyric text segment to an editable state; in response to a second instruction for the target lyric text segment in the editable state, modifying the first character in the target lyric text segment to the second character to generate the second lyrics.
[0122] According to one or more embodiments of this disclosure, the method further includes: displaying the playback timestamp and / or playback duration corresponding to the lyrics text segment before and after responding to the modification instruction.
[0123] According to one or more embodiments of this disclosure, generating second lyrics in response to a modification instruction for a first character in the first lyrics includes: obtaining a corresponding text segment to be determined in response to a modification instruction for a first character in a target lyrics text segment of the first lyrics; detecting the number of characters in the text segment to be determined, and if the number of characters in the text segment to be determined is equal to the number of characters in the target lyrics text segment, then determining the text segment to be determined as the modified lyrics text segment; replacing the target lyrics text segment in the first lyrics with the modified lyrics text segment to generate the second lyrics.
[0124] According to one or more embodiments of this disclosure, the method further includes: if the number of characters in the text segment to be determined is not equal to the number of characters in the target lyrics text segment, then calling a preset generation model to process the text segment to be determined, generating an optimized text segment, wherein the number of characters in the optimized text segment is equal to the number of characters in the target lyrics text segment, and has the same semantics as the text segment to be determined; replacing the target lyrics text segment in the first lyrics with the optimized text segment, generating the second lyrics.
[0125] According to one or more embodiments of this disclosure, generating a second song based on the second text includes: determining a first original song segment and a second original song segment in the first song based on the playback timestamp of the first text, wherein the first original song segment is the original song segment corresponding to the playback timestamp of the first text, and the second original song segment is the remaining original song segment in the first song excluding the first original song segment; obtaining the pronunciation data of the second text, and replacing the vocal components in the first original song segment based on the pronunciation data to generate a target song segment, wherein the melody of the target song segment is the same as the melody of the first original song segment in the first song corresponding to the playback timestamp of the target song segment; and generating a second song based on the target song segment and the second original song segment in the first song.
[0126] According to one or more embodiments of this disclosure, generating a second song based on the second text includes: determining a first original song segment and a second original song segment in the first song based on the playback timestamp of the first text; calling a song generation model to generate a target song segment based on the second text and the second original song segment; wherein the melody of the target song segment is different from the melody of the first original song segment in the first song corresponding to the playback timestamp of the target song segment; and generating a second song based on the target song segment and the second original song segment.
[0127] According to one or more embodiments of this disclosure, the first song and the second song belong to the same target song project, and the method further includes: displaying a generation record page corresponding to the target song project, the generation record page being used to display at least one of the following: song generation records corresponding to the target song project, the song generation records including at least the generation record of the first song and the generation record of the second song; changes in the second song relative to the first song.
[0128] Secondly, according to one or more embodiments of this disclosure, an AI song editing device is provided, comprising:
[0129] The display module is used to display the first lyrics of the first song after the first song and the corresponding first lyrics are generated, wherein the first song is audio data generated based on artificial intelligence technology;
[0130] The processing module is used to generate second lyrics in response to a modification instruction for the first character in the first lyrics, wherein the second lyrics contain a second character corresponding to the playback timestamp of the first character;
[0131] A generation module is used to generate a second song based on the second text in response to the generation of the second lyrics. The second song includes a target song segment and a non-target song segment. The target song segment is the song segment corresponding to the second text. The non-target song segment is the remaining song segment in the second song excluding the target song segment. The melody of the non-target song segment is the same as the melody of the second original song segment in the first song corresponding to the playback timestamp of the non-target song segment.
[0132] According to one or more embodiments of this disclosure, the first lyrics include multiple lyric text segments, each of which corresponds to an original song segment; when displaying the first lyrics of the first song, the display module is specifically configured to: display the lyric text segments of the first lyrics in lines, and the modification control corresponding to the lyric text segments, wherein the modification control is configured to trigger the lyric text to an editable state; the modification instruction includes a first instruction and a second instruction; the processing module is specifically configured to: in response to the first instruction for the modification control corresponding to the target lyric text segment, trigger the target lyric text segment to an editable state; in response to the second instruction for the target lyric text segment in the editable state, modify the first character in the target lyric text segment to the second character to generate the second lyrics.
[0133] According to one or more embodiments of this disclosure, the display module is further configured to: display the playback timestamp and / or playback duration corresponding to the lyrics text segment before and after responding to the modification instruction.
[0134] According to one or more embodiments of this disclosure, the processing module is specifically configured to: in response to a modification instruction for a first character in a target lyrics text segment of the first lyrics, obtain a corresponding text segment to be determined; detect the number of characters in the text segment to be determined, and if the number of characters in the text segment to be determined is equal to the number of characters in the target lyrics text segment, determine the text segment to be determined as the modified lyrics text segment; replace the target lyrics text segment in the first lyrics with the modified lyrics text segment, and generate the second lyrics.
[0135] According to one or more embodiments of this disclosure, the processing module is further configured to: if the number of characters in the text segment to be determined is not equal to the number of characters in the target lyrics text segment, then call a preset generation model to process the text segment to be determined, generate an optimized text segment, wherein the number of characters in the optimized text segment is equal to the number of characters in the target lyrics text segment, and has the same semantics as the text segment to be determined; replace the target lyrics text segment in the first lyrics with the optimized text segment, and generate the second lyrics.
[0136] According to one or more embodiments of this disclosure, when the generation module generates a second song based on the second text, it is specifically configured to: determine a first original song segment and a second original song segment in the first song based on the playback timestamp of the first text, wherein the first original song segment is the original song segment corresponding to the playback timestamp of the first text, and the second original song segment is the remaining original song segment in the first song excluding the first original song segment; obtain the pronunciation data of the second text, and replace the vocal components in the first original song segment based on the pronunciation data to generate a target song segment, wherein the melody of the target song segment is the same as the melody of the first original song segment in the first song corresponding to the playback timestamp of the target song segment; and generate the second song based on the target song segment and the second original song segment in the first song.
[0137] According to one or more embodiments of this disclosure, when the generation module generates a second song based on the second text, it is specifically configured to: determine a first original song segment and a second original song segment in the first song based on the playback timestamp of the first text; call a song generation model to generate a target song segment based on the second text and the second original song segment; wherein the melody of the target song segment is different from the melody of the first original song segment in the first song corresponding to the playback timestamp of the target song segment; and generate the second song based on the target song segment and the second original song segment.
[0138] According to one or more embodiments of this disclosure, the first song and the second song belong to the same target song project. The display module is further configured to: display a generation record page corresponding to the target song project, wherein the generation record page is configured to display at least one of the following: the song generation record corresponding to the target song project, wherein the song generation record includes at least the generation record of the first song and the generation record of the second song; and the changes made by the second song relative to the first song.
[0139] Thirdly, according to one or more embodiments of the present disclosure, an electronic device is provided, comprising: at least one processor and a memory;
[0140] The memory stores computer-executed instructions;
[0141] The at least one processor executes computer execution instructions stored in the memory, causing the at least one processor to perform the AI song editing method as described in the first aspect and various possible designs of the first aspect.
[0142] Fourthly, according to one or more embodiments of the present disclosure, a computer-readable storage medium is provided, wherein computer-executable instructions are stored therein, and when a processor executes the computer-executable instructions, the AI song editing method described in the first aspect and various possible designs of the first aspect is implemented.
[0143] Fifthly, according to one or more embodiments of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the AI song editing method as described in the first aspect and various possible designs of the first aspect.
[0144] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.
[0145] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.
[0146] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.
Claims
1. An AI song editing method, characterized in that, include: After generating the first song and its corresponding first lyrics, the first lyrics of the first song are displayed, wherein the first song is audio data generated based on artificial intelligence technology; In response to a modification instruction for a first character in the first lyrics, second lyrics are generated, the second lyrics containing a second character corresponding to the playback timestamp of the first character; wherein, the playback timestamp includes the start time and end time corresponding to the first character; The step of generating second lyrics in response to a modification instruction for a first character in the first lyrics includes: In response to a modification instruction for the first character within the target lyrics text segment in the first lyrics, the corresponding text segment to be determined is obtained; The number of characters in the text segment to be determined is detected; if the number of characters in the text segment to be determined is not equal to the number of characters in the target lyrics text segment, a preset generation model is called to process the text segment to be determined, and a text segment with the same number of characters as the target lyrics text segment and the same semantics as the text segment to be determined is generated; the target lyrics text segment in the first lyrics is replaced with the text segment to generate the second lyrics; In response to the generation of the second lyrics, a second song is generated based on the second text. The second song includes a target song segment and a non-target song segment. The target song segment is the song segment corresponding to the second text. The non-target song segment is the remaining song segment in the second song excluding the target song segment. The melody of the non-target song segment is the same as the melody of the original song segment in the first song corresponding to the playback timestamp of the non-target song segment.
2. The method according to claim 1, characterized in that, The first lyrics include multiple lyric text segments, each of which corresponds to an original song segment; The display of the first lyrics of the first song includes: The lyrics text segment of the first lyrics is displayed in separate lines, along with the corresponding editing control for the lyrics text segment. The editing control is configured to trigger the lyrics text to an editable state. The modification instructions include a first instruction and a second instruction. The step of generating second lyrics in response to the modification instruction for the first character in the first lyrics includes: In response to a first instruction to the modification control corresponding to the target lyrics text segment, the target lyrics text segment is triggered to an editable state; In response to a second instruction for a target lyrics text segment in the editable state, the first character within the target lyrics text segment is modified to the second character to generate the second lyrics.
3. The method according to claim 2, characterized in that, The method further includes: Before and after responding to the modification instruction, display the playback timestamp and / or playback duration corresponding to the lyrics text segment.
4. The method according to claim 1, characterized in that, The step of generating second lyrics in response to a modification instruction for a first character in the first lyrics further includes: If the number of characters in the text segment to be determined is equal to the number of characters in the target lyrics text segment, then the text segment to be determined is determined as the modified lyrics text segment; The target lyrics text segment in the first lyrics is replaced with the modified lyrics text segment to generate the second lyrics.
5. The method according to claim 1, characterized in that, The process of generating a second song based on the second text includes: Based on the playback timestamp of the first text, the first original song segment and the second original song segment in the first song are determined. The first original song segment is the original song segment corresponding to the playback timestamp of the first text, and the second original song segment is the remaining original song segment in the first song excluding the first original song segment. Obtain the pronunciation data of the second character, and replace the human voice component in the first original song segment based on the pronunciation data to generate a target song segment, wherein the melody of the target song segment is the same as the melody of the first original song segment in the first song corresponding to the playback timestamp of the target song segment; The second song is generated based on the target song segment and the second original song segment in the first song.
6. The method according to claim 1, characterized in that, The process of generating a second song based on the second text includes: Based on the playback timestamp of the first text, the first original song segment and the second original song segment in the first song are determined. The song generation model is invoked to generate a target song segment based on the second text and the second original song segment; wherein the melody of the target song segment is different from the melody of the first original song segment in the first song corresponding to the playback timestamp of the target song segment; The second song is generated based on the target song segment and the second original song segment.
7. The method according to claim 1, characterized in that, The first song and the second song belong to the same target song project, and the method further includes: The generated record page corresponding to the target song project is displayed, and the generated record page is used to display at least one of the following: The song generation record corresponding to the target song project includes at least the generation record of the first song and the generation record of the second song; The changes made to the second song compared to the first song.
8. An AI song editing device, characterized in that, include: The display module is used to display the first lyrics of the first song after the first song and the corresponding first lyrics are generated, wherein the first song is audio data generated based on artificial intelligence technology; The processing module is configured to generate second lyrics in response to a modification instruction for a first character in the first lyrics, wherein the second lyrics contain a second character corresponding to the playback timestamp of the first character; wherein the playback timestamp includes the start time and end time corresponding to the first character; A generation module is used to generate a second song based on the second text in response to the generation of the second lyrics. The second song includes a target song segment and a non-target song segment. The target song segment is the song segment corresponding to the second text. The non-target song segment is the remaining song segment in the second song excluding the target song segment. The melody of the non-target song segment is the same as the melody of the second original song segment in the first song that corresponds to the playback timestamp of the non-target song segment. The processing module is specifically configured to respond to a modification instruction for a first character within a target lyric text segment in the first lyrics, obtain a corresponding text segment to be determined; detect the number of characters within the text segment to be determined; if the number of characters within the text segment to be determined is not equal to the number of characters in the target lyric text segment, then call a preset generation model to process the text segment to be determined, generate a text segment with the same number of characters as the target lyric text segment and having the same semantics as the text segment to be determined; replace the target lyric text segment in the first lyrics with the text segment to generate the second lyrics.
9. An electronic device, characterized in that, include: Processor and memory; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory, causing the processor to perform the AI song editing method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by the processor, implement the AI song editing method as described in any one of claims 1 to 7.
11. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the AI song editing method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Audio production method and device, equipment and storage medium
CN111899706A