Singing video generation method and device, electronic equipment and storage medium

By configuring video generation controls and material pages within the application, performance videos that match the music can be generated directly, solving the problems of low efficiency and low matching degree in intelligent music video generation, and achieving efficient and high-definition video generation.

CN119653203BActive Publication Date: 2025-10-24BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411775029.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-04
Publication Date
2025-10-24
Estimated Expiration
2044-12-04

AI Technical Summary

Technical Problem

In the process of intelligently generating music and corresponding videos in the existing technology, the interaction efficiency is low and the matching degree between the generated video and the music is low, resulting in poor communication effect.

Method used

Configure a video generation control within the target application. The publishing page will directly redirect to the video generation page, display the video footage, and generate a singing video based on the selected video footage, so that the lip movements of the target object in the video match the music content.

Benefits of technology

It improves the efficiency of generating performance videos, ensures the matching and clarity of video and music, avoids compression issues caused by cross-platform transmission, and enhances video quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119653203B_ABST
    Figure CN119653203B_ABST
Patent Text Reader

Abstract

The method and device for generating a singing video, the electronic device and the storage medium provided in the embodiments of the present disclosure can display a publishing page corresponding to target music after generating the target music through a target application program, and the publishing page is configured with a video generation control for generating a singing video matching the target music; in response to a first trigger operation on the video generation control, a first video generation page is displayed in the target application program, and the first video generation page is configured to display at least one video material; in response to a second trigger operation on the first video generation page, a target video material is determined, and a singing video is generated based on the target video material, and a lip movement of a target object in the singing video matches music content of the target music. The method and device for generating a singing video, the electronic device and the storage medium provided in the embodiments of the present disclosure can generate a "lip-synch" video matching target music, and the process does not need to be implemented through other third-party software, thereby improving the generation efficiency of the singing video.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] Embodiments of the present disclosure relate to the technical field of Internet, and particularly, to a singing video generation method and device, electronic equipment and storage medium. BACKGROUND

[0002] Currently, artificial intelligence generated content (AIGC) technology has been increasingly applied in various industries. For example, a music generation model can be trained by using music samples, so that the model has the ability to "create" music, thereby generating "original music" that meets user requirements, i.e., intelligent music generation.

[0003] After the intelligent music generation with melody and corresponding lyrics is generated by using the intelligent content generation technology, the intelligent music generation can be spread through platform publishing, sharing, etc. However, since there is no matching singing video, the spread effect of the intelligent music generation is poor. In related technologies, the intelligent music generation can be processed by using an additional third-party software to generate a corresponding video, and then returned to the original intelligent music generation platform for publishing or sharing.

[0004] However, in the process of generating a corresponding video for the intelligent music generation, there are problems of low interaction efficiency and low matching degree between the generated video and the intelligent music generation. SUMMARY

[0005] Embodiments of the present disclosure provide a singing video generation method and device, electronic equipment and storage medium to overcome the problems of low interaction efficiency and low matching degree between the generated video and the intelligent music generation in the process of generating a corresponding video for the intelligent music generation.

[0006] In a first aspect, embodiments of the present disclosure provide a singing video generation method, comprising:

[0007] After generating a target music by using a target application, a publishing page corresponding to the target music is displayed in the target application, wherein the publishing page is configured with a video generation control for generating a singing video matching the target music; in response to a first trigger operation on the video generation control, a first video generation page is displayed in the target application, and the first video generation page is configured to display at least one video material; in response to a second trigger operation on the first video generation page, a target video material is determined, and the singing video is generated based on the target video material, wherein the singing video and the target video material contain a same target object, and a lip movement action of the target object in the singing video matches music content of the target music.

[0008] In a second aspect, the embodiments of the present disclosure provide a singing video generation apparatus, comprising:

[0009] a first display module configured to display a publishing page after generating a target music, wherein the publishing page is configured with a video generation control for generating a singing video matching the target music;

[0010] a second display module configured to display a first video generation page in the target application in response to a first trigger operation on the video generation control, wherein the first video generation page is configured to display at least one video material;

[0011] a generation module configured to determine a target video material and generate the singing video based on the target video material in response to a second trigger operation on the first video generation page, wherein the singing video and the target video material contain a same target object, and a lip movement of the target object in the singing video matches a music content of the target music.

[0012] In a third aspect, the embodiments of the present disclosure provide an electronic device, comprising a processor and a memory.

[0013] The memory stores computer-executable instructions.

[0014] The processor executes the computer-executable instructions stored in the memory, so that the at least one processor executes the singing video generation method as described in the first aspect and various possible designs of the first aspect.

[0015] In a fourth aspect, the embodiments of the present disclosure provide a computer-readable storage medium, wherein the computer-readable storage medium stores computer-executable instructions, and when a processor executes the computer-executable instructions, the singing video generation method as described in the first aspect and various possible designs of the first aspect is implemented.

[0016] In a fifth aspect, the embodiments of the present disclosure provide a computer program product, comprising a computer program, and when a processor executes the computer program, the singing video generation method as described in the first aspect and various possible designs of the first aspect is implemented.

[0017] The singing video generation method, device, electronic equipment and storage medium provided by the embodiment display a publishing page corresponding to the target music in the target application after the target music is generated by the target application, wherein the publishing page is configured with a video generation control for generating a singing video matching the target music; in response to a first trigger operation on the video generation control, a first video generation page is displayed in the target application, and the first video generation page is configured to display at least one video material; in response to a second trigger operation on the first video generation page, a target video material is determined, and the singing video is generated based on the target video material, wherein the singing video and the target video material contain the same target object, and the lip movement action of the target object in the singing video matches the music content of the target music. By configuring the video generation control in the publishing page after the intelligent generated music is generated, the client can directly jump to the first video generation page for generating a singing video matching the intelligent generated music after the intelligent generated music is generated, and the responsive video material is displayed, and then the target music containing the target object consistent with the target video material and capable of performing the lip movement action matching the music content, that is, the "lip movement" video matching the target music, is generated based on the selected target video material. This process does not need to be realized by other third-party software, so that the generation efficiency of the singing video is improved, and since no cross-platform image transmission is involved, the audio and image compression process is avoided, so that the matching degree and video clarity of the generated singing video and the intelligent generated music are higher, and the video effect is better. BRIEF DESCRIPTION OF DRAWINGS

[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description are some embodiments of the present disclosure, and those skilled in the art can also obtain other drawings according to these drawings without creative labor.

[0019] Figure 1 An application scenario diagram of the singing video generation method provided by the embodiment of the present disclosure;

[0020] Figure 2 A flowchart of the singing video generation method provided by the embodiment of the present disclosure Figure 1 ;

[0021] Figure 3 A flowchart of the specific implementation of step S102 in the embodiment shown in Figure 2

[0022] Figure 4 ​A schematic diagram of a first video generation page provided by an embodiment of the present disclosure

[0023] Figure 5 A flowchart of a specific implementation of step S207 in the embodiment shown in

[0024] Figure 6 A flowchart of a singing video generation method provided by an embodiment of the present disclosure Figure 2 ;

[0025] Figure 7 A flowchart of a specific implementation of step S207 in the embodiment shown in Figure 6

[0026] Figure 8 A process schematic diagram of a singing video generation provided by an embodiment of the present disclosure

[0027] Figure 9 A structural block diagram of a singing video generation apparatus provided by an embodiment of the present disclosure

[0028] Figure 10 A structural schematic diagram of an electronic device provided by an embodiment of the present disclosure

[0029] Figure 11 A hardware structural schematic diagram of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION

[0030] To make the objects, technical solutions and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be described clearly and completely below with reference to the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are some but not all of the embodiments of the present disclosure. Based on the embodiments in the present disclosure, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of the present disclosure.

[0031] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present disclosure are all information and data authorized by the user or authorized by all parties, and the collection, use and processing of related data need to comply with relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation portal for user to choose authorization or refusal.

[0032] The application scenarios of the embodiments of the present disclosure are explained as follows:

[0033] Figure 1 ​An application scenario diagram of the singing video generation method provided by the embodiments of the present disclosure, the singing video generation method provided by the embodiments of the present disclosure can be applied to an application (APP, Application) having a music generation function and a video generation function, for example, a music application, a short video application and the like. More specifically, it can be applied to an application scenario of intelligently generating a singing video corresponding to the music and making an intelligent music video. The execution subject of the present embodiment can be a terminal device running the above-mentioned application having a music generation function or a video generation function, or a server of a server corresponding to the above-mentioned application, or other electronic devices having similar functions. When the execution subject is a terminal device, the terminal device executes the method provided by the embodiments of the present disclosure by running the above-mentioned application; when the execution subject is a server, the server of the application having a music generation function or a video generation function can be partially or entirely run on the server, and the method provided by the embodiments of the present disclosure is executed on the server side, while the client of the terminal device running the application, the communication between the server and the terminal device is based on the server-client, so that the terminal device can obtain the execution result of the method provided by the embodiments of the present disclosure, and display it as needed.

[0034] In some embodiments, the terminal device or the server can implement the singing video generation method provided by the embodiments of the present disclosure by running various computer executable instructions or computer programs. For example, the computer executable instructions can be program-level commands, machine instructions or software instructions. The computer program can be a native program or a software module in the operating system; it can be a local application, that is, a program that needs to be installed in the operating system to run, or it can be a small program embedded in any APP, that is, a program running based on a browser environment. In summary, the above-mentioned computer executable instructions can be any form of instructions, and the above-mentioned computer programs can be any form of application programs, modules or plug-ins, and the specific implementation form can be configured as needed. Further, the terminal device can execute the singing video generation method provided by the embodiments of the present disclosure by running the computer executable instructions or computer programs set locally, or by calling the computer executable instructions or computer programs set in the server outside. In some embodiments, the server can be a standalone physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing cloud services, cloud storage, cloud communication, cloud database, cloud computing, cloud function, network service, middleware service, domain name service, security service, content distribution network (CDN), and big data and artificial intelligence platform, etc. Basic cloud computing services, among which the cloud service can be an interactive processing service for calling by the terminal device.

[0035] Reference Figure 1 As shown in the figure, taking a terminal device as an example, a target application program with a music generation function is running in the terminal device. After the target application program is started, the user inputs keywords for generating music into the target application program by operating the terminal device, for example, as shown in the figure, including: “sadness”, “lyricism”, “clear river”, etc. Then, the target application program calls a music generation model deployed in the cloud based on the above keywords to generate a corresponding “original song”, that is, an intelligent generated music (shown as Music AI No. 001 in the figure), and sends it back to the terminal device for playing and displaying. Then, the user can further add a cover picture to the intelligent generated music and publish it to the Internet platform (music platform) corresponding to the target application program to realize the propagation of the intelligent generated music.

[0036] In the prior art, in the above application scenario, after the user generates an intelligent generated music through intelligent generation technology, the intelligent generated music can be propagated through platform publishing, sharing and other means, but since there is no matching singing video, the propagation effect of the intelligent generated music is poor. In related technologies, an additional third-party image generation software can be used to generate a singing video matching the intelligent generated music based on image intelligent generation technology, which can also be called a “lip-sync video”. Then, the user can publish or share the singing video on the original intelligent generated music generation platform, thereby improving the propagation effect of the created intelligent generated music. However, since the generated intelligent generated music needs to be generated and transmitted across platforms, on the one hand, it will cause additional user operations and waiting time, affecting the efficiency of video production, and on the other hand, during the transmission of the intelligent generated music, audio data needs to be transmitted to the third-party software for processing, which requires the intelligent generated music to be compressed during the process, resulting in the loss of music details, thereby causing the generated singing video to have a low matching degree with the intelligent generated music. That is, in the process of generating a corresponding video for the intelligent generated music, there is a low interaction efficiency and a low matching degree between the generated video and the intelligent generated music.

[0037] The embodiment of the present disclosure provides a singing video generation method to solve the above problems.

[0038] Reference Figure 2 , Figure 2 Flowchart of the singing video generation method provided by the embodiment of the present disclosure Figure 1 The method of the present embodiment can be applied in a terminal device. The singing video generation method comprises:

[0039] Step S101: After the target music is generated through the target application, a publishing page corresponding to the target music is displayed in the target application, wherein the publishing page is configured with a video generation control for generating a performance video matching the target music.

[0040] Step S102: In response to a first trigger operation on the video generation control, a first video generation page is displayed in the target application, where the first video generation page is configured to display at least one video material.

[0041] Step S103: In response to the second trigger operation for the first video generation page, the target video material is determined, and a singing video is generated based on the target video material, wherein the singing video and the target video material contain the same target object, and the lip movements of the target object in the singing video match the music content of the target music.

[0042] For example, refer to Figure 1 As shown in the schematic diagram of the application scenario, the terminal device generates target music based on intelligent generation technology by running a target application with a music generation function, that is, intelligently generated music. Among them, the process of generating such intelligently generated music based on intelligent generation technology usually includes receiving keywords input by the user, and calling a pre-trained music generation module in combination with the keyword to generate music that matches the keyword. The steps will not be elaborated this time. Afterwards, after the target music is generated, the target application jumps to the publishing page corresponding to the target music. In the publishing page, a video generation control for generating a singing video that matches the target music is configured. The user can generate a singing video of the target music in the current target application by generating the video control.

[0043] Specifically, after the terminal device receives the first trigger operation for the video generation control, it controls the target application to display a first video generation page within the target application. The first video generation page is configured to display at least one video material, which is a material file used to generate a singing video, such as a picture containing facial features.

[0044] Afterwards, the terminal device can determine a target video material in the first video generation page in response to a second trigger operation of the user for the first video generation page. In one possible implementation, the second trigger operation is a selection operation for a video material in the first video generation page, and the target video material is determined from the first video generation page in response to the selection operation. In another possible implementation, the terminal device automatically selects a default video material as the target video material, or determines a recommended video material as the target video material by calculation, and the second trigger operation is a confirmation operation for the default video material or the recommended video material, for example, an operation of clicking a "confirm" control in the first video generation page, so that the terminal device determines the final target video material in response to the second trigger operation.

[0045] Further, after determining the target video material, the terminal device generates a singing video based on the target video material, and completes the process of generating a singing video matching the target music. The target video material includes a target object, for example, a person, a small animal, or another object with facial features. The terminal device processes the target video material by calling a video generation model to generate a video with the same length as the target music, containing the same target object, and the target object having lip movements matching the lyrics and / or melody of the target music, i.e., a singing video corresponding to the target music. Specifically, the singing video and the target video material contain the same target object, and the lip movements of the target object in the singing video match the music content of the target music, i.e., the video content of the singing video includes lip movements of the target object (a pet dog) in the target video material (a photo of the pet dog) dynamically changing with the music melody and content to represent the target object (the pet dog) singing the current target song.

[0046] The video generation model can be deployed entirely locally on the terminal device, entirely in the cloud, or partially locally and partially in the cloud, and the specific implementation can be set as needed. The video generation model is a pre-trained model that has the ability to output a video from a piece of music (corresponding file) and a picture. The specific implementation and training process of the video generation model are not described here.

[0047] Further, in one possible implementation, the first video generation page is configured with a first material area and a second material area. The first material area is used to display a video generation template in the cloud or cached locally, and the second material area is used to display a local image stored in the terminal device.

[0048] Exemplarily, the video generation template displayed by the first material area is a file set or information set for generating a singing video, and at least contains a template picture, and can further contain video special effect parameters and the like. The video generation template can be a template uploaded by another user, and is used to generate a singing video with a specific style and with a specific target object (for example, a specific person or animal) to achieve the purpose and effect of "reproducing the same video". The template can be stored in the cloud and previewed locally on the terminal device. The local image displayed in the second material area is the image stored in the terminal device, and is displayed after obtaining the authorization of the user or in the second material area.

[0049] Further, for the content displayed in the second material area of the first video generation page, in a possible implementation manner, as shown in Figure 3 the specific implementation manner of step S102 includes:

[0050] Step S1021: After obtaining the authorization of the user, the local image stored in the terminal device is obtained.

[0051] Step S1022: From the local image stored in the terminal device, a candidate local image containing a target object is selected, wherein the target object is an object with facial features.

[0052] Step S1023: The candidate local image is displayed in the second material area of the first video generation page.

[0053] In this embodiment, before the terminal device displays the first video generation page in the target application, in view of the problem that the number of local images is large and not all of them are suitable for generating a singing video, after obtaining the authorization of the user, the local image stored in the terminal device is obtained, and the local image is selected. Specifically, the local image is feature detected, and the image containing facial features is determined as a candidate local image containing a target object. Then, the candidate local image is displayed in the second material area of the first video generation page, so as to complete the display process of the first video generation page. Since the lip movement action needs to be displayed in the generated singing video, the image for generating the singing video also needs to clearly show the face and the lip shape in a static state. In the step of this embodiment, the candidate local image is an image containing a target object with facial features, which is selected by the terminal device, and can avoid the problem that the lip area cannot be matched in the subsequent process, so as to improve the authenticity of the generated singing video.

[0054] On the other hand, the content displayed in the first material area of the first video generation page, i.e., the video generation template, can be personalized pushed by the cloud and displayed in the first material area of the first video generation page. This will not be described again.

[0055] Further, after displaying the first video generation page, the specific implementation of step S103 includes two kinds, one of which includes:

[0056] Step S103A: in response to the second trigger operation for the first material area, determining the target video generation template, and generating the singing video based on the target video generation template.

[0057] Another possible implementation of step S103 includes:

[0058] Step S103B: in response to the second trigger operation for the second material area, determining the target local image, and generating the singing video based on the target local image.

[0059] Figure 4 A schematic diagram of a first video generation page provided by the embodiment of the present disclosure is shown below. Figure 4 The above process will be described in detail as follows. Figure 4 As shown in the figure, exemplarily, first, after jumping from the publishing page to the first video generation page, in the upper half of the first video generation page, the first material area is used to display the video generation template of the cloud, and below each video generation template, the use method description information of the video generation template is displayed, such as the "shoot the same video" shown in the figure, to realize the use guidance of the video generation template; in the lower half of the first video generation page, the second material area is used to display the local image stored in the terminal device. Then, when the user clicks the target video generation template of the first material area, such as the video generation template M1 shown in the figure, the terminal device calls the video generation model deployed in the cloud to process the video generation template M1 and the target music, thereby generating a singing video based on the style of the video generation template M1; and when the user clicks the target local image of the second material area, such as the local image P1 shown in the figure, the terminal device uploads the local image P1 to the cloud server, and then calls the video generation model to process the local image P1 and the target music, thereby generating a singing video based on the local image P1.

[0060] Of course, in another possible implementation, the first video generation page can only include the first material area or the second material area described above, i.e., only display the video generation template or the local image. The specific implementation and the corresponding mutual mode are similar to the above-described content, which will not be described again.

[0061] Furthermore, in a possible implementation, after determining the target video material, the method further includes:

[0062] Step S103C: Obtain the image content features of the target video material; if the image content features do not meet the preset feature requirements, display a prompt message.

[0063] Exemplarily, after determining the target video material, for example, after determining the target local image in the above embodiment, in order to ensure that the target local image meets the requirements for generating a singing video (i.e., the target object is contained in the picture), in this embodiment, the terminal device determines whether the target local image selected by the user can be used as the source data for generating a singing video by detecting the image content characteristics of the target local image. If, according to the image content characteristics of the target local image, the target local image contains preset features, such as facial contour features and mouth features, the subsequent steps are continued; if the preset feature requirements are not met, a prompt message is displayed to prompt the user to make a replacement, thereby avoiding the problem of not being able to generate a singing video due to user misoperation.

[0064] Furthermore, in a possible implementation, as Figure 5 As shown, after determining the target video material, the specific implementation method of generating a singing video based on the target video material includes:

[0065] Step S1031: Acquire music type information corresponding to the target music. The music type information is generated by the target application in the process of generating the target music.

[0066] Step S1032: Generate a second prompt word based on the music type information, where the second prompt word represents the action features of the lip movement.

[0067] Step S1033: Input the target music, the target local image and the second prompt word into the video generation model to generate a singing video, wherein the lip movements of the target object in the singing video have the movement characteristics represented by the second prompt word.

[0068] Exemplarily, taking the case of a local image as the target video material as an example, after the terminal device determines the target local image, the terminal device first acquires the music type information corresponding to the target music generated in the previous step. The music type information is information for classifying the type and style of music, and more specifically, the music type information corresponding to the target music includes, for example, pop, rock, opera, etc. (according to song classification); or includes lyric, dynamic, sad, etc. (according to emotional style). The music type information can be a specific classification identifier in specific implementation, which is not specifically limited. Then, the terminal device converts the music category information into a corresponding second prompt word according to the preset configuration information. The second prompt word represents the action characteristics of the mouth movement, such as the frequency and amplitude of the mouth movement, so as to realize the mapping of the music category information to the mouth movement. Finally, the second prompt word representing the action characteristics of the mouth movement is input into the video generation model together with the target music and the target local image to generate a singing video. The second prompt word controls the singing video generation process, so that the mouth movement of the target object in the generated singing video has the action characteristics represented by the second prompt word, thereby improving the matching degree and authenticity between the singing video and the target music.

[0069] Meanwhile, it should be noted that the scheme provided in this embodiment is completed in the same software or platform (target application), which avoids the process of cross-platform transmission of the related information of the target music compared with other schemes in the prior art. Therefore, in this embodiment, since the music type information is generated by the target application in the process of generating the target music, the terminal device can efficiently and accurately implement step S1031 to obtain the music type information. Compared with the music type information obtained by analyzing and identifying the target music in the conventional scheme, the music type information has higher accuracy and execution efficiency.

[0070] In an embodiment of the present disclosure, after generating target music through a target application, a publishing page corresponding to the target music is displayed in the target application, wherein the publishing page is configured with a video generation control for generating a singing video matching the target music; in response to a first trigger operation on the video generation control, a first video generation page is displayed in the target application, and the first video generation page is configured to display at least one video material; in response to a second trigger operation on the first video generation page, the target video material is determined, and a singing video is generated based on the target video material, wherein the singing video and the target video material contain the same target object, and the lip movements of the target object in the singing video match the musical content of the target music. By configuring a video generation control in the publishing page after generating the intelligently generated music, the client can directly jump to the first video generation page for generating a singing video matching the intelligently generated music after generating the intelligently generated music, and display the corresponding video material. Then, based on the selected target video material, target music is generated that contains a target object that is consistent with the target video material and can perform lip movements that match the music content, that is, a "lip-sync" video that matches the target music. This process does not require the use of other third-party software, thereby improving the generation efficiency of the singing video. Moreover, since it does not involve cross-platform image transmission, it avoids the audio and image compression process, so that the generated singing video can have a higher degree of match with the intelligently generated music and a higher video clarity, resulting in a better video effect.

[0071] refer to Figure 6 , Figure 6 Schematic diagram of the process of the singing video generation method provided by the embodiment of the present disclosure Figure 2 In this embodiment Figure 2 Based on the embodiment shown, the scheme is further refined, and a potential step of generating target music is optionally added. The singing video generation method includes:

[0072] Step S201: obtaining a first prompt word input by a user, where the first prompt word is used to represent a musical feature of a target music.

[0073] Step S202: The first prompt word is processed by the target application to generate target music.

[0074] For example, this embodiment adds a specific implementation step for generating the target music. This involves obtaining a first prompt word input by the user and then invoking a music generation model using the target application to generate the target music. In this step, the first prompt word input by the user is cached locally on the terminal device and used in subsequent steps to select the default material, thereby improving the efficiency of generating the performance video.

[0075] Step S203: After the target music is generated by the target application, a publishing page corresponding to the target music is displayed in the target application, wherein the video generation control for generating a singing video matching the target music is configured in the publishing page.

[0076] Step S204: In response to a first trigger operation on the video generation control, a first video generation page is displayed in the target application, and the first video generation page is configured to display at least one video material.

[0077] Step S205: In the first video generation page, based on the first prompt word, a corresponding default video material is selected.

[0078] Exemplarily, in this embodiment, in combination with the schemes of steps S201-S202, after the first video generation page is displayed in the target application, one possible implementation is to select one as the target video material by responding to the trigger operation input by the user; and in this step, the terminal device automatically determines a video material, i.e., the default video material, in combination with the first prompt word obtained in the process of previously generating the target music. The first prompt word represents the music characteristics of the target music, therefore, the video material generated based on the first prompt word is also a video material matching the music characteristics of the target music, and selecting the default video material as the target video material and performing the subsequent step of generating a singing video can make the generated singing video more matched with the music characteristics (e.g., music melody, music lyrics) of the target music, and provide the authenticity of the singing video.

[0079] Step S206: In response to a second trigger operation on the first video generation page, the default video material is determined as the target video material.

[0080] Step S207: A singing video is generated based on the target video material and the target music, wherein the singing video and the target video material contain the same target object, and the lip movement of the target object in the singing video is matched with the music content of the target music.

[0081] Exemplarily, in combination with the previous steps, after the terminal device selects the corresponding default video material, the second trigger operation is, for example, a confirmation operation, i.e., an operation of confirming the default video material as the target video material, and then the terminal device generates the corresponding singing music by calling the video generation model based on the obtained target video material.

[0082] Further, in one possible implementation, as shown in Figure 7 the specific implementation of step S207 includes:

[0083] Step S2071: Obtain the target music and divide the target music into a vocal segment and a non-vocal segment;

[0084] Step S2072: Processing the target video material and the vocal segment using the video generation model to generate a first video segment corresponding to the vocal segment, wherein the lip movements of the target subject in the first video segment match the musical content of the target music;

[0085] Step S2073: generating a second video segment corresponding to the non-human voice segment based on the target video material, wherein the second video segment is a static video with no changes in the video picture;

[0086] Step S2074: Generate a singing video based on the combination of the first video clip and the second video clip based on the playback timestamps.

[0087] After the terminal device obtains the previously generated target music through the target application, it divides the target music into vocal segments and non-vocal segments, wherein the vocal segments are segments corresponding to the vocal part of the target music, in other words, segments in which the target object in the corresponding singing video needs to perform lip movements. On the contrary, the non-vocal segments are segments corresponding to the non-vocal part of the target music, in other words, segments in which the target object in the corresponding singing video does not need to perform lip movements. Vocal segments and non-vocal segments can be represented by paired timestamps. Furthermore, with regard to the specific implementation method of dividing vocal segments and non-vocal segments, one possible implementation method is to divide the vocal segments and non-vocal segments by parsing the audio features of the target music and performing vocal recognition; another possible implementation method is to divide the vocal segments and non-vocal segments by reading the music information corresponding to the target music (such as lyrics text and corresponding timestamps), which can be set according to needs. Next, the video generation model processes the target video material and the vocal clip to generate a first video clip corresponding to the vocal clip, which is the "lip sync video" for that vocal clip. For the non-vocal clip, the target video material is directly used to generate a static video, the second video clip. Finally, the first and second video clips are combined based on their playback timestamps to generate the performance video.

[0088] Figure 8 A schematic diagram of a process for generating a singing video provided by an embodiment of the present disclosure is shown below. Figure 8 For detailed description of the above steps, refer to Figure 8As shown, the target music is an audio with a length of 10 seconds; by analyzing the target music, the target music is divided into a vocal segment (shown as segment A in the figure) and a non-vocal segment (shown as segment B in the figure), for example, the vocal segment is the first 4 seconds of the target music, and the non-vocal segment is the last 6 seconds of the target music. Then, the terminal device uploads the vocal segment of the target music and the target video material to the video generation model in the cloud, generates a corresponding first video segment, i.e., video segment Video_1, through the video generation model. On the other hand, the terminal device generates a second video segment, i.e., video segment Video_2, by using the target video material to generate a video picture that does not change locally; or, the terminal device sends the target video material to the cloud server to generate a second video segment (this case is not shown in the figure) whose video picture does not change. Finally, the terminal device locally combines the video segment Video_1 and the video segment Video_2 to generate a singing video. Alternatively, the terminal device uploads the video segment Video_1 locally to the cloud server, and the cloud server combines the video segment Video_1 and the video segment Video_2 to generate a singing video and sends it back to the terminal device (this case is not shown in the figure).

[0089] In the step of the embodiment, by dividing the target music into a vocal segment and a non-vocal segment and processing them respectively to generate corresponding video segments, and by combining the video segments to generate a singing video, the model resource overhead for generating a static picture corresponding to the non-vocal segment in the singing video is saved, the network resource overhead of the server and the model pressure of the video generation model are reduced, and the generation efficiency of the singing video is improved.

[0090] Optionally, in the process of performing step S207, the embodiment further includes:

[0091] Step S208: displaying a second video generation page in the target application, the second video generation page being configured to display the current generation progress of the singing video.

[0092] Optionally, after step S207 is performed, the embodiment further includes:

[0093] Step S209: displaying a preview video of the singing video in the second video generation page after the singing video is generated.

[0094] Illustratively, since the singing video needs to call the video generation model for generation, considering the load of model resources, the video generation model usually needs to queue the requests for generating the singing video sent by each terminal device for processing when there are high concurrency tasks. Therefore, the user usually needs to wait for the generation of the singing video to be completed. In the embodiment, during the generation of the singing video, the second video generation page is displayed in the target application program, and the current generation progress of the singing video is displayed through the second video generation page, and after the singing video is generated, the preview video of the singing video is displayed in the second video generation page, so that the user can know the remaining time, thereby implementing other operations within the remaining time, and improving the operation efficiency of the target application program.

[0095] Optionally, the embodiment further includes:

[0096] Step S209A: in response to a third triggering operation on the second video generation page, switching the second video generation page to run in the background;

[0097] Step S209B: in response to a fourth triggering operation on the second video generation page, canceling the generation of the singing video, and returning to the first video generation page or the publishing page.

[0098] Illustratively, in combination with the steps of the previous embodiments, during the generation of the singing video, the terminal device can further switch the second video generation page to run in the background in response to a third triggering operation on the second video generation page, so as to execute other functions in the target application program, thereby improving the efficiency of operating the target application program; or, in response to a fourth triggering operation on the second video generation page, canceling the generation of the singing video, and returning to the first video generation page or the publishing page, thereby realizing the interruption of the generation process of the singing video, and the subsequent steps of modifying and regenerating the singing video, which can also realize the improvement of the efficiency of operating the target application program and generating the final singing video.

[0099] Optionally, the embodiment further includes:

[0100] Step S210: in response to a modification instruction for the target music, generating the changed target music, and returning to step S207.

[0101] Exemplarily, after the target music is generated, the terminal device can re-modify the target music through the target application program in response to a modification instruction input by the user, for example, modify the prompt word (for example, the first prompt word in the previous embodiment step) when the target music is generated, or modify the lyrics of the target music, and the like, to re-generate the target music in whole or in part, that is, to generate the changed target music. Then, by taking advantage of the feature that the target music and the generated singing video are in the same target application program, the terminal device can re-generate the singing video based on the changed target music and the target video material (that is, return to execute step S207), so as to quickly re-generate the singing video without repeating the above steps such as selecting the target video material, thereby achieving the effect of synchronously modifying the corresponding singing video after the intelligent generated music is modified, and improving the operation efficiency of adjusting the singing video.

[0102] In this embodiment, the implementation manner of step S203 is the same as that of step S101 in the embodiment shown in the above Figure 2 The implementation manner of step S101 in the embodiment shown in the above

[0103] Corresponding to the singing video generation method of the above embodiment, Figure 9 A structural block diagram of a singing video generation apparatus provided by the embodiment of the present disclosure is shown in FIG. 3. The method introduced in the above embodiment can be executed by the singing video generation apparatus. The apparatus can be implemented in the form of software and / or hardware. The apparatus can be integrated in an electronic device having a certain data processing function. The electronic device can include but is not limited to a mobile terminal having a large data processing capacity, and a desktop computer, a supercomputer, and the like having a large data processing capacity.

[0104] For ease of illustration, only parts related to the embodiment of the present disclosure are shown. For parts not related to the embodiment of the present disclosure, refer to the description of the above Figure 9 The singing video generation apparatus 3 includes:

[0105] The first display module 31 is configured to display a publishing page after the target music is generated, wherein the publishing page is configured with a video generation control for generating a singing video matching the target music.

[0106] The second display module 32 is configured to display a first video generation page in the target application program in response to a first trigger operation on the video generation control. The first video generation page is configured to display at least one video material.

[0107] The generation module 33 is configured to determine a target video material in response to a second trigger operation on the first video generation page, and generate a singing video based on the target video material. The singing video and the target video material contain a same target object, and the lip movement of the target object in the singing video matches the music content of the target music.

[0108] According to one or more embodiments of the present disclosure, the first video generation page is configured with a first material area and a second material area, wherein the first material area is configured to display a video generation template; the second material area is configured to display a local image stored in the terminal device; the generation module 33 is specifically configured to: in response to a second triggering operation on the first material area, determine a target video generation template, and generate a singing video based on the target video generation template; or in response to a second triggering operation on the second material area, determine a target local image, and generate a singing video based on the target local image.

[0109] According to one or more embodiments of the present disclosure, when the second display module 32 displays the first video generation page in the target application, it is specifically configured to: after obtaining the authorization of the user, obtain a local image stored in the terminal device; from the local image stored in the terminal device, filter a candidate local image containing a target object, wherein the target object is an object with facial features; and display the candidate local image in the second material area of the first video generation page.

[0110] According to one or more embodiments of the present disclosure, when the generation module 33 generates a singing video based on a target local image, it is specifically configured to: obtain music type information corresponding to the target music; generate a second prompt word according to the music type information, the second prompt word representing the action feature of the lip movement; input the target music, the target local image and the second prompt word into the video generation model to generate a singing video, wherein the lip movement of the target object in the singing video has the action feature represented by the second prompt word.

[0111] According to one or more embodiments of the present disclosure, before the generation module 33 generates the target music, it is further configured to: obtain a first prompt word input by the user, the first prompt word being used to represent the music feature of the target music; and generate the target music by processing the first prompt word through the target application; after the second display module 32 displays the first video generation page in the target application in response to a first triggering operation on the video generation control, it is further configured to: in the first video generation page, select a corresponding default video material based on the first prompt word.

[0112] According to one or more embodiments of the present disclosure, after the generation module 33 generates a singing video based on a target video material, it is further configured to: in response to a modification instruction for the target music, generate a changed target music; and regenerate a singing video based on the changed target music and the target video material.

[0113] According to one or more embodiments of the present disclosure, the generation module 33 is further configured to, after determining the target video material, display a second video generation page in the target application, the second video generation page being configured to display a current generation progress of the singing video; and display a preview video of the singing video in the second video generation page after the singing video is generated.

[0114] According to one or more embodiments of the present disclosure, the generation module 33 is further configured to perform at least one of the following: in response to a third triggering operation on the second video generation page, switching the second video generation page to run in the background; and in response to a fourth triggering operation on the second video generation page, canceling the generation of the singing video and returning to the first video generation page or the publishing page.

[0115] According to one or more embodiments of the present disclosure, the generation module 33 is further configured to, after determining the target video material, acquire an image content feature of the target video material; and display a prompt information if the image content feature does not meet a preset feature requirement.

[0116] According to one or more embodiments of the present disclosure, the generation module 33 is further configured to, when generating the singing video based on the target video material, acquire the target music and divide the target music into a vocal segment and a non-vocal segment; process the target video material and the vocal segment by using a video generation model to generate a first video segment corresponding to the vocal segment, a mouth movement of a target object in the first video segment matching music content of the target music; generate a second video segment corresponding to the non-vocal segment based on the target video material, the second video segment being a static video in which a video picture does not change; and combine the first video segment and the second video segment based on a playback timestamp to generate the singing video.

[0117] The first display module 31, the second display module 32, and the generation module 33 are connected in sequence. The singing video generation device 3 provided in this embodiment can execute the technical solutions of the above-mentioned method embodiments, and has similar implementation principles and technical effects. Details are not described herein again.

[0118] Figure 10 A structural schematic diagram of an electronic device provided in an embodiment of the present disclosure is shown in FIG. 4. Figure 10 As shown in FIG. 4, the electronic device 4 includes:

[0119] a processor 41 and a memory 42 connected with the processor 41 in communication;

[0120] The memory 42 stores computer execution instructions.

[0121] The processor 41 executes the computer execution instructions stored in the memory 42 to implement the singing video generation method in the embodiment shown in FIG. 3. Figures 2-8 The processor 41 executes the computer execution instructions stored in the memory 42 to implement the singing video generation method in the embodiment shown in FIG. 3.

[0122] wherein, optionally, the processor 41 and the memory 42 are connected by the bus 43.

[0123] The relevant description can be understood by referring to Figures 2-8 The relevant description and effects of the steps in the corresponding embodiments are not described in detail here.

[0124] The embodiment of the present disclosure provides a computer readable storage medium, the computer readable storage medium stores computer execution instructions, and the computer execution instructions are used for realizing the method for generating a singing video Figures 2-8 The singing video generation method provided by any one of the embodiments.

[0125] The embodiment of the present disclosure provides a computer program product, comprising a computer program, and the computer program is executed by a processor to realize the method for generating a singing video Figures 2-8 The singing video generation method provided by any one of the embodiments.

[0126] In order to realize the above-mentioned embodiments, the embodiment of the present disclosure further provides an electronic device.

[0127] Reference Figure 11 , which shows a structural schematic diagram of an electronic device 900 suitable for realizing the embodiments of the present disclosure. The electronic device 900 can be a terminal device or a server. The terminal device can include but is not limited to mobile terminals such as mobile phones, notebook computers, digital broadcast receivers, personal digital assistants (PDA), tablet computers, portable multimedia players (PMP), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), and the like, and fixed terminals such as digital TVs, desktop computers, and the like. Figure 11 The electronic device shown is only an example, and should not bring any limitation to the functions and use range of the embodiments of the present disclosure.

[0128] As Figure 11As shown, the electronic device 900 can include a processing device (e.g., a central processor, a graphics processor, etc.) 901 that can perform various suitable actions and processes in accordance with programs stored in a Read Only Memory (ROM) 902 or loaded from a storage device 908 into a Random Access Memory (RAM) 903. Various programs and data required by the electronic device 900 for operation are also stored in the RAM 903. The processing device 901, the ROM 902, and the RAM 903 are connected to each other by a bus 904. An Input / Output (I / O) interface 905 is also connected to the bus 904.

[0129] Generally, the following devices can be connected to the I / O interface 905: input devices 906 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; output devices 907 including, for example, a Liquid Crystal Display (LCD), a speaker, a vibrator, etc.; storage devices 908 including, for example, a magnetic tape, a hard disk, etc.; and communication devices 909. The communication devices 909 can allow the electronic device 900 to communicate wirelessly or wired with other devices to exchange data. Although Figure 11 The electronic device 900 is shown with various devices, but it should be understood that not all of the shown devices are required to be implemented or present. More or fewer devices can alternatively be implemented or present.

[0130] In particular, the processes described above with reference to the flowcharts can be implemented as a computer software program according to embodiments of the present disclosure. For example, embodiments of the present disclosure include a computer program product comprising a computer program carried on a computer readable medium, the computer program containing program code for performing the methods illustrated by the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network through the communication devices 909, or installed from the storage devices 908, or installed from the ROM 902. When the computer program is executed by the processing device 901, the above-mentioned functions defined in the methods of embodiments of the present disclosure are performed.

[0131] It should be noted that the computer-readable medium in the above disclosure can be a computer-readable signal medium or a computer-readable storage medium or any combination of the two. The computer-readable storage medium may, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or apparatus, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device or apparatus. In the present disclosure, the computer-readable signal medium can include a data signal carried in a baseband or as a part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take many forms, including but not limited to an electromagnetic signal, an optical signal or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, which can send, propagate or transmit a program for use by or in conjunction with an instruction execution system, device or apparatus. The program code contained in the computer-readable medium can be transmitted by any suitable medium, including but not limited to a wire, a cable, an RF (radio frequency) or the like, or any suitable combination of the above.

[0132] The computer-readable medium described above can be contained in the electronic device described above; or can exist separately and not be assembled into the electronic device.

[0133] The computer-readable medium described above carries one or more programs, which, when executed by the electronic device, cause the electronic device to perform the methods shown in the above embodiments.

[0134] Computer program code for carrying out operations of the present disclosure can be written in any combination of one or more programming languages, including an object oriented programming language such as Java, Smalltalk, C++ or the like and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).

[0135] The flow diagrams and the block diagrams in the drawings are illustrations of architectures, functionalities, and operations of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flow diagrams or block diagrams can represent a module, a procedure, or a part of code, which comprises one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently or in the reverse order, depending on the functionality involved. It is also noted that each block in the block diagrams and / or flow diagrams and combinations of blocks in the block diagrams and / or flow diagrams can be implemented by special-purpose hardware-based systems that perform the specified functions or operations, or combinations of special-purpose hardware and computer instructions.

[0136] The units or modules described in the embodiments of the present disclosure can be implemented by means of software, or by means of hardware. In some cases, the name of the unit or module does not constitute a limitation on the unit itself.

[0137] The functions described above in the specification of the present disclosure can be performed by one or more hardware logic components. For example, and without limitation, illustrative types of hardware logic components that can be used include Field-programmable Gate Arrays (FPGAs), Program-specific Integrated Circuits (ASICs), Program-specific Standard Products (ASSPs), System-on-a-chip systems (SOCs), Complex Programmable Logic Devices (CPLDs), etc.

[0138] In the context of this disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium will include one or more of: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0139] In a first aspect, according to one or more embodiments of the present disclosure, a singing video generation method is provided, comprising:

[0140] After generating the target music through the target application, a publishing page corresponding to the target music is displayed in the target application, wherein the target music is an intelligently generated music, and the publishing page is configured with a video generation control for generating a singing video matching the target music; in response to a first triggering operation on the video generation control, a first video generation page is displayed in the target application, and the first video generation page is configured to display at least one video material; in response to a second triggering operation on the first video generation page, a target video material is determined, and the singing video is generated based on the target video material, wherein the singing video and the target video material contain the same target object, and the lip movement of the target object in the singing video matches the music content of the target music.

[0141] According to one or more embodiments of the present disclosure, the first video generation page is configured with a first material area and a second material area, wherein the first material area is used to display a cloud video generation template, and the second material area is used to display a local image stored in a terminal device; in response to a second triggering operation on the first video generation page, a target video material is determined, and the singing video is generated based on the target video material, comprising: in response to a second triggering operation on the first material area, a target video generation template is determined, and the singing video is generated based on the target video generation template; or, in response to a second triggering operation on the second material area, a target local image is determined, and the singing video is generated based on the target local image.

[0142] According to one or more embodiments of the present disclosure, the displaying the first video generation page in the target application includes: obtaining a local image stored in the terminal device after obtaining authorization of a user; selecting a candidate local image containing a target object from the local image stored in the terminal device, wherein the target object is an object having a facial feature; and displaying the candidate local image in a second material area of the first video generation page.

[0143] According to one or more embodiments of the present disclosure, the generating the singing video based on the target local image includes: obtaining music type information corresponding to the target music; generating a second prompt word based on the music type information, the second prompt word representing a motion feature of a lip motion; and inputting the target music, the target local image, and the second prompt word into a video generation model to generate the singing video, wherein the lip motion of the target object in the singing video has the motion feature represented by the second prompt word.

[0144] According to one or more embodiments of the present disclosure, before the generating the target music, the method further includes: obtaining a first prompt word input by a user, the first prompt word being used to represent a music feature of the target music; and generating the target music by processing the first prompt word through the target application; and after the displaying the first video generation page in the target application in response to a first trigger operation on the video generation control, the method further includes: selecting a corresponding default video material based on the first prompt word in the first video generation page.

[0145] According to one or more embodiments of the present disclosure, after the generating the singing video based on the target video material, the method further includes: generating a changed target music in response to a modification instruction on the target music; and regenerating a singing video based on the changed target music and the target video material.

[0146] According to one or more embodiments of the present disclosure, after the determining the target video material, the method further includes: displaying a second video generation page in the target application, the second video generation page being configured to display a current generation progress of the singing video; and displaying a preview video of the singing video in the second video generation page after the singing video is generated.

[0147] According to one or more embodiments of the present disclosure, the method further includes at least one of: switching the second video generation page to run in the background in response to a third trigger operation on the second video generation page; and canceling the generation of the singing video and returning to the first video generation page or the publishing page in response to a fourth trigger operation on the second video generation page.

[0148] According to one or more embodiments of the present disclosure, after the target video material is determined, the method further includes: obtaining image content features of the target video material; and displaying prompt information if the image content features do not meet preset feature requirements.

[0149] According to one or more embodiments of the present disclosure, the generating the singing video based on the target video material includes: obtaining the target music and dividing the target music into a vocal segment and a non-vocal segment; processing the target video material and the vocal segment by using a video generation model to generate a first video segment corresponding to the vocal segment, wherein a lip movement of a target object in the first video segment matches music content of the target music; generating a second video segment corresponding to the non-vocal segment based on the target video material, wherein the second video segment is a static video with no change in video frames; and combining the first video segment and the second video segment based on their playback time stamps to generate the singing video.

[0150] In a second aspect, according to one or more embodiments of the present disclosure, a singing video generation apparatus is provided, and the apparatus includes:

[0151] A first display module is configured to display a publishing page after the target music is generated, wherein the publishing page is configured with a video generation control for generating a singing video matching the target music;

[0152] A second display module is configured to display a first video generation page in the target application in response to a first trigger operation on the video generation control, wherein the first video generation page is configured to display at least one video material.

[0153] A generation module is configured to determine a target video material in response to a second trigger operation on the first video generation page, and generate the singing video based on the target video material, wherein the singing video and the target video material contain a same target object, and a lip movement of the target object in the singing video matches music content of the target music.

[0154] According to one or more embodiments of the present disclosure, the first video generation page is configured with a first material area and a second material area, wherein the first material area is configured to display a video generation template; and the second material area is configured to display a local image stored in the terminal device; and the generation module is specifically configured to: in response to a second triggering operation on the first material area, determine a target video generation template, and generate the singing video based on the target video generation template; or in response to a second triggering operation on the second material area, determine a target local image, and generate the singing video based on the target local image.

[0155] According to one or more embodiments of the present disclosure, when the second display module displays the first video generation page in the target application, it is specifically configured to: after obtaining the authorization of the user, obtain a local image stored in the terminal device; from the local image stored in the terminal device, filter a candidate local image containing a target object, wherein the target object is an object with facial features; and display the candidate local image in a second material area of the first video generation page.

[0156] According to one or more embodiments of the present disclosure, when the generation module generates the singing video based on the target local image, it is specifically configured to: obtain music type information corresponding to the target music; generate a second prompt word according to the music type information, wherein the second prompt word represents the action feature of the lip movement; input the target music, the target local image, and the second prompt word into a video generation model to generate the singing video, wherein the lip movement of the target object in the singing video has the action feature represented by the second prompt word.

[0157] According to one or more embodiments of the present disclosure, before the generation module generates the target music, it is further configured to: obtain a first prompt word input by the user, wherein the first prompt word is used to represent the music feature of the target music; and generate the target music by processing the first prompt word through the target application; and after the second display module displays the first video generation page in the target application in response to the first triggering operation on the video generation control, it is further configured to: in the first video generation page, select a corresponding default video material based on the first prompt word.

[0158] According to one or more embodiments of the present disclosure, after the generation module generates the singing video based on the target video material, it is further configured to: in response to a modification instruction for the target music, generate a changed target music; and regenerate a singing video based on the changed target music and the target video material.

[0159] According to one or more embodiments of the present disclosure, the generation module is further configured to: display a second video generation page in the target application after determining the target video material, the second video generation page being configured to display a current generation progress of the singing video; and display a preview video of the singing video in the second video generation page after the singing video is generated.

[0160] According to one or more embodiments of the present disclosure, the generation module is further configured to at least one of: switch the second video generation page to run in the background in response to a third triggering operation on the second video generation page; and cancel the generation of the singing video and return to the first video generation page or the publishing page in response to a fourth triggering operation on the second video generation page.

[0161] According to one or more embodiments of the present disclosure, the generation module is further configured to: obtain image content features of the target video material; and display a prompt message if the image content features do not meet preset feature requirements.

[0162] According to one or more embodiments of the present disclosure, the generation module is further configured to: obtain the target music and divide the target music into a vocal segment and a non-vocal segment when generating the singing video based on the target video material; process the target video material and the vocal segment by using a video generation model to generate a first video segment corresponding to the vocal segment, a lip movement of a target object in the first video segment matching music content of the target music; generate a second video segment corresponding to the non-vocal segment based on the target video material, the second video segment being a static video with no changes in video frames; and combine the first video segment and the second video segment based on their playback timestamps to generate the singing video.

[0163] In a third aspect, according to one or more embodiments of the present disclosure, an electronic device is provided, including: at least one processor and a memory.

[0164] The memory stores computer-executable instructions.

[0165] The at least one processor executes the computer-executable instructions stored in the memory, so that the at least one processor performs the singing video generation method as described in the first aspect and various possible designs of the first aspect.

[0166] In a fourth aspect, according to one or more embodiments of the present disclosure, a computer readable storage medium is provided, and the computer readable storage medium has stored therein computer executing instructions which, when executed by a processor, implement the singing video generation method according to the first aspect and various possible designs of the first aspect.

[0167] In a fifth aspect, according to one or more embodiments of the present disclosure, a computer program product is provided, and the computer program product includes a computer program which, when executed by a processor, implements the singing video generation method according to the first aspect and various possible designs of the first aspect.

[0168] The above description is merely illustrative of the exemplary embodiments of the present disclosure and the principles of the technology involved. It is understood that the disclosed scope of the present disclosure is not limited to the technical solutions formed by the specific combinations of the above technical features, and also covers other technical solutions formed by the combinations of the above technical features or equivalent features thereof without departing from the above disclosed concept. For example, the technical solutions formed by the mutual replacement of the above features and the technical features disclosed in the present disclosure (but not limited to) having similar functions.

[0169] In addition, although each operation is depicted in a particular order, this should not be understood as requiring the operations to be performed in the particular order shown or in sequential order. In certain circumstances, multitasking and parallel processing can be advantageous. Likewise, the specific sequence of operations described above is merely illustrative of the embodiments of the present disclosure. Certain features described in the context of separate embodiments can also be combined in single embodiments. Conversely, various features described in the context of single embodiments can also be separated into separate embodiments. In addition, although specific implementation details are discussed in the above discussion, these should not be construed as limiting the scope of the present disclosure. Certain features described in the context of separate embodiments can also be combined in single embodiments. Conversely, various features described in the context of single embodiments can also be separated into separate embodiments.

[0170] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely illustrative of exemplary forms of implementing the claims.

Claims

1. A method for generating a singing video, characterized in that, The method comprises the following steps: After generating the target music through the target application, a publishing page corresponding to the target music is displayed in the target application, wherein the publishing page is configured with a video generation control for generating a singing video matching the target music; In response to a first trigger operation on the video generation control, a first video generation page is displayed in the target application, and the first video generation page is configured to display at least one video material; In response to a second trigger operation on the first video generation page, a target video material is determined, and the singing video is generated based on the target video material, wherein the singing video and the target video material contain the same target object, and the lip movement of the target object in the singing video matches the music content of the target music.

2. The method of claim 1, wherein, The first video generation page is configured with a first material area and a second material area, wherein the first material area is used to display a video generation template, and the second material area is used to display a local image stored in a terminal device; In response to a second trigger operation on the first video generation page, a target video material is determined, and the singing video is generated based on the target video material, comprising: In response to a second trigger operation on the first material area, a target video generation template is determined, and the singing video is generated based on the target video generation template; or in response to a second trigger operation on the second material area, a target local image is determined, and the singing video is generated based on the target local image.

3. The method of claim 2, wherein, The first video generation page is displayed in the target application, comprising: After obtaining the authorization of a user, a local image stored in the terminal device is obtained; From the local image stored in the terminal device, a candidate local image containing a target object is screened, wherein the target object is an object with facial features; In the second material area of the first video generation page, the candidate local image is displayed.

4. The method of claim 2, wherein, The singing video is generated based on the target local image, comprising: Music type information corresponding to the target music is obtained; According to the music type information, a second prompt word is generated, and the second prompt word represents the action feature of the lip movement; The target music, the target local image, and the second prompt word are input into a video generation model to generate the singing video, wherein the lip movement of the target object in the singing video has the action feature represented by the second prompt word.

5. The method of claim 1, wherein, The target music is intelligently generated music, and before generating the target music, the method further comprises the following steps: A first prompt word input by a user is obtained, and the first prompt word is used to represent the music feature of the target music; The target music is generated by processing the first prompt word through the target application; After displaying the first video generation page in the target application in response to the first trigger operation on the video generation control, the method further comprises the following steps: In the first video generation page, the corresponding default video material is selected based on the first prompt word.

6. The method of claim 5, wherein, After the singing video is generated based on the target video material, the method further includes: in response to a modification instruction for the target music, generating a changed target music; based on the changed target music and the target video material, regenerating a singing video.

7. The method of claim 1, wherein, After the target video material is determined, the method further includes: displaying a second video generation page in the target application, the second video generation page being configured to display a current generation progress of the singing video; after the singing video is generated, displaying a preview video of the singing video in the second video generation page.

8. The method of claim 7, wherein, The method further includes at least one of: in response to a third triggering operation for the second video generation page, switching the second video generation page to run in the background; in response to a fourth triggering operation for the second video generation page, canceling the generation of the singing video and returning to the first video generation page or the publishing page.

9. The method of claim 1, wherein, After the target video material is determined, the method further includes: obtaining image content features of the target video material; if the image content features do not meet preset feature requirements, displaying a prompt message.

10. The method of claim 1, wherein, The singing video is generated based on the target video material, including: obtaining the target music and dividing the target music into a vocal segment and a non-vocal segment; processing the target video material and the vocal segment using a video generation model to generate a first video segment corresponding to the vocal segment, the lip movement of a target object in the first video segment matching the music content of the target music; based on the target video material, generating a second video segment corresponding to the non-vocal segment, the second video segment being a static video with no changes in video frames; combining the playback time stamps of the first video segment and the second video segment to generate the singing video.

11. A singing video generation apparatus characterized by comprising: including: a first display module configured to display a publishing page after generating a target music through a target application, wherein the publishing page is configured with a video generation control for generating a singing video matching the target music; a second display module configured to display a first video generation page in the target application in response to a first triggering operation for the video generation control, the first video generation page being configured to display at least one video material; a generation module configured to determine a target video material in response to a second triggering operation for the first video generation page and generate a singing video based on the target video material, wherein the singing video and the target video material contain a same target object, and the lip movement of the target object in the singing video matches the music content of the target music.

12. An electronic device, comprising: including: a processor and a memory; the memory stores computer execution instructions; the processor executes the computer execution instructions stored in the memory, so that the processor executes the singing video generation method according to any one of claims 1 to 10.

13. A computer-readable storage medium, characterized in that, The computer readable storage medium stores computer execution instructions, and when the processor executes the computer execution instructions, the computer execution instructions implement the singing video generation method in any one of claims 1 to 10.

14. A computer program product comprising a computer program, characterized in that, The computer program is executed by the processor to implement the singing video generation method in any one of claims 1 to 10.

Citation Information

Patent Citations

  • Rassing audio and video synthesis method, system and equipment and readable storage medium

    CN116030785A

  • Digital human multimedia resource generation method and device, equipment and storage medium

    CN118860233A