Video generation method and apparatus, electronic device, and storage medium
By acquiring media content to generate storyboard data and using a pre-trained model to generate videos, the problem of low efficiency, poor quality, and high repetition in existing video generation technologies is solved, achieving automated, efficient, and diversified video generation.
Patent Information
- Application Number
- CN202411899323.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-20
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2044-12-20
AI Technical Summary
Existing technologies suffer from low efficiency, poor quality, and high video similarity during video generation, especially when generating complex video content, which requires complex descriptive terms and repeated debugging.
By acquiring the first media content, the first storyboard data is generated, including the second media content and the first storyboard text. Based on this data, the third media content is generated and added to the storyboard fragment. Finally, the video is generated using a pre-trained image processing model and a neural network model.
It achieves fully automated video generation, improving efficiency and quality, avoiding video duplication, and enhancing video diversity.
Smart Images

Figure CN119788936B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments of the present disclosure relate to the technical field of image processing, and particularly relate to a video generation method and device, electronic equipment and storage medium. BACKGROUND
[0002] Currently, with the development of artificial intelligence generated content (AIGC) technology, in the process of video content creation, users can generate various media materials by using tools provided in video editing software, thereby improving the production efficiency of video content.
[0003] In the prior art, in the process of producing video content, a common implementation is that a user generates a video meeting a picture requirement by inputting a picture description word into an image processing model.
[0004] However, since the picture description word requires input skills and experience, the prior art solution has difficulty in generating a long video with complex video content, that is, has problems of low video generation efficiency, poor video quality, and high video similarity. SUMMARY
[0005] Embodiments of the present disclosure provide a video generation method and device, electronic equipment and storage medium to overcome the problems of low video generation efficiency, poor video quality, and high video similarity.
[0006] In a first aspect, embodiments of the present disclosure provide a video generation method, comprising:
[0007] obtaining first media content; obtaining first shot script data according to the first media content, the first shot script data including a first shot script segment, the first shot script segment including second media content and first shot text, the first shot text being used to describe at least one picture element in the second media content and a shot feature of media content generated based on the first shot script segment, the second media content including at least one media content in the first media content; generating third media content according to the first shot script segment, and adding the third media content to the first shot script segment to obtain second shot script data; and generating a video based on the second shot script data.
[0008] In a second aspect, embodiments of the present disclosure provide a video generation device, comprising:
[0009] The acquisition module is configured to acquire first media content, and acquire first split script data according to the first media content, wherein the first split script data comprises a first split script segment, the first split script segment comprises second media content and first split text, the first split text is used to describe at least one picture element in the second media content and a shot feature of media content generated based on the first split script segment, and the second media content comprises at least one media content in the first media content.
[0010] The processing module is configured to generate third media content according to the first split script segment, and add the third media content to the first split script segment to obtain second split script data.
[0011] The generation module is configured to generate a video based on the second split script data.
[0012] In a third aspect, an electronic device is provided, which comprises a processor and a memory.
[0013] The memory stores computer-executable instructions.
[0014] The processor executes the computer-executable instructions stored in the memory, so that the at least one processor executes the video generation method according to the first aspect and various possible designs of the first aspect.
[0015] In a fourth aspect, a computer-readable storage medium is provided, which stores computer-executable instructions, and when a processor executes the computer-executable instructions, the video generation method according to the first aspect and various possible designs of the first aspect is implemented.
[0016] In a fifth aspect, a computer program product is provided, which comprises a computer program, and when a processor executes the computer program, the video generation method according to the first aspect and various possible designs of the first aspect is implemented.
[0017] The video generation method, apparatus, electronic device, and storage medium provided in this embodiment obtain first media content; based on the first media content, obtain first storyboard data, the first storyboard data including a first storyboard fragment, the first storyboard fragment including second media content and first storyboard text, the first storyboard text being used to describe at least one scene element in the second media content and the shot features of media content generated based on the first storyboard fragment, the second media content including at least one media content in the first media content; based on the first storyboard fragment, generate third media content, and add the third media content to the first storyboard fragment to obtain second storyboard data; based on the second storyboard data, generate a video. By directly generating first storyboard data based on the first media content provided by the user, and then using the image features of the second media content corresponding to each storyboard fragment in the first storyboard data and the shot features represented by the first storyboard text, corresponding third media content is generated and added to the corresponding storyboard fragment to generate second storyboard data. Finally, a video is generated based on the storyboard fragments in the generated second storyboard data, realizing a full-process video generation. Users do not need to manually control the media material generation process, layout, and editing process, which improves video generation efficiency and video quality. At the same time, since the first media content input by the user is different, the final video generated based on the first media content is also diverse, avoiding the problem of video duplication. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in the embodiments of this disclosure or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 This is an application scenario diagram of the video generation method provided in the embodiments of this disclosure;
[0020] Figure 2 Flowchart of the video generation method provided in the embodiments of this disclosure Figure 1 ;
[0021] Figure 3 for Figure 2 A flowchart illustrating the specific implementation of step S102 in the illustrated embodiment;
[0022] Figure 4 for Figure 3 A flowchart illustrating the specific implementation of step S1021 in the illustrated embodiment;
[0023] Figure 5 For Figure 3 a flowchart of the specific implementation of step S1022 in the embodiment shown in
[0024] Figure 6 a process schematic diagram of generating a video provided by an embodiment of the present disclosure;
[0025] Figure 7 a flowchart of a video generation method provided by an embodiment of the present disclosure Figure 1 ;
[0026] Figure 8 a schematic diagram of an interactive interface provided by an embodiment of the present disclosure;
[0027] Figure 9 For Figure 8 a flowchart of the specific implementation of step S205 in the embodiment shown in
[0028] Figure 10 a process schematic diagram of generating an optimized screenplay provided by an embodiment of the present disclosure;
[0029] Figure 11 a process schematic diagram of another embodiment of the present disclosure for generating an optimized screenplay;
[0030] Figure 12 a structural block diagram of a video generation apparatus provided by an embodiment of the present disclosure;
[0031] Figure 13 a structural schematic diagram of an electronic device provided by an embodiment of the present disclosure;
[0032] Figure 14 a hardware structural schematic diagram of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION
[0033] To make the purposes, technical solutions and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be described clearly and completely below with reference to the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are part of the embodiments of the present disclosure, rather than all the embodiments of the present disclosure. Based on the embodiments in the present disclosure, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the scope of protection of the present disclosure.
[0034] It should be noted that the user information (including but not limited to user equipment information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present disclosure are all information and data authorized by the user or authorized by all parties, and the collection, use and processing of related data need to comply with relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation portal for user to choose authorization or refusal.
[0035] The application scenarios of the embodiments of the present disclosure are explained as follows:
[0036] The video generation method provided by the embodiments of the present disclosure can be applied in an application (APP, Application) with video editing function, such as a video editing application, a short video application, etc. More specifically, it can be applied in an application scenario of generating personalized video content based on AIGC technology. The execution subject of the present embodiment can be a terminal device running the above-mentioned application with video editing function, or a server deploying the server side of the above-mentioned application, or other electronic devices with similar functions. When the execution subject is a terminal device, the terminal device executes the method provided by the present embodiment by running the above-mentioned application; when the execution subject is a server, the server side of the above-mentioned application with video editing function can be partially or entirely run on the server, and the method provided by the present embodiment is executed on the server side, while the client side of the terminal device running the application, the server and the terminal device communicate based on the server-client mode, so that the terminal device can obtain the execution result of the method provided by the present embodiment and display it as needed.
[0037] In some embodiments, the terminal device or the server can implement the video generation method provided by the embodiments of the present disclosure by running various computer-executable instructions or computer programs. For example, the computer-executable instructions can be program-level commands, machine instructions, or software instructions. The computer program can be a native program in the operating system or a software module; it can be a local application program, that is, a program that needs to be installed in the operating system to run, or it can be a small program embedded in any APP, that is, a program running based on a browser environment. In summary, the above computer-executable instructions can be any form of instructions, and the above computer programs can be any form of application programs, modules, or plug-ins, and the specific implementation form can be configured as needed. Further, the terminal device can execute the method by running the computer-executable instructions or computer programs set locally or by calling the computer-executable instructions or computer programs set in the server outside the terminal device in the process of implementing the video generation method provided by the embodiments of the present disclosure. In some embodiments, the server can be a standalone physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud services, cloud storage, cloud communication, cloud database, cloud computing, cloud function, network service, middleware service, domain name service, security service, content distribution network (Content Delivery Network, CDN), and big data and artificial intelligence platform, etc. Basic cloud computing services, wherein the cloud service can be an interactive processing service for calling by the terminal device.
[0038] Figure 1 An application scenario diagram of the video generation method provided by the embodiments of the present disclosure is shown in Figure 1 For example, the terminal device runs a target application program with a video editing function, and the user can input a description word to the target application program to describe the picture content of the video to be generated, such as "generate a video of a sunset". Then, the terminal device sends the above description word to the image generation model based on the AIGC technology deployed on the server side to generate a corresponding video, and returns the video to the terminal device for display, thereby completing the video generation process.
[0039] However, in the prior art, Figure 1In the illustrated video generation scenario, only a video segment with simple picture content and short duration can be generated, but for a video with more complex content, such as an advertisement, a promotional video, a video log (vlog), or a complex content video with multiple shots and camera movement logic, a very complex description word is required to construct a video outline and a shot script, and after repeated debugging, a video content meeting the expectation can be generated, thus resulting in that most ordinary users cannot directly generate the above complex content video through the description word, but still create the video content through the traditional manual editing and layout, thereby causing low video generation efficiency and poor video quality.
[0040] Embodiments of the present disclosure provide a video generation method to solve the above problems.
[0041] Reference Figure 2 , Figure 2 Flowchart of a video generation method provided by embodiments of the present disclosure Figure 1 The method of the present embodiment can be applied in a terminal device or a server, and the video generation method comprises:
[0042] Step S101: acquiring first media content;
[0043] Step S102: acquiring first shot script data according to the first media content, the first shot script data comprising a first shot script segment, the first shot script segment comprising second media content and a first shot text, the first shot text being used to describe at least one picture element in the second media content and a shot feature of media content generated based on the first shot script segment, and the second media content comprising at least one media content in the first media content.
[0044] Exemplarily, referring to the application scenario diagram as shown in Figure 1 In the present embodiment, a terminal device is taken as an execution subject to introduce the provided video generation method, and a target application program with an image editing function is running in the terminal device. Through an interactive interface of the target application program, a user can manually select media content, i.e., first media content, stored in the terminal device locally or in the cloud, so as to make the terminal device load the first media content. The media content, i.e., media data, is, for example, a video or a picture taken or downloaded by the user. Specifically, the first media content comprises one or more media contents, for example, the first media content is 4 pictures stored in the terminal device locally and selected by the user, and again for example, the first media content is 2 pictures and 1 video stored in the terminal device locally and selected by the user.
[0045] After obtaining the first media content, the terminal device processes the first media content to generate first shot script data. The shot script data can be a data structure in the target application for generating a final output video. For example, the first shot script data is a draft data in the target application. More specifically, the first shot script data is, for example, a draft video_1. The specific implementation process of step S102 is a process of creating a corresponding draft data based on the first media content selected by the user.
[0046] Further, the first shot script data includes one or more shot script segments. Each shot script segment corresponds to a shot script and a shot segment of the finally generated video. The first shot script segment is a shot script segment generated and created based on the first media content. The first shot script segment includes second media content and first shot text. The second media content includes at least one media content in the first media content. For example, the second media content is a picture or a video. The second media content is one or more media contents in the first media content, i.e., a subset of the first media content. More specifically, the first media content includes pictures p1, p2, and p3. The second media content is, for example, the picture p1. The first shot text is used to describe at least one picture element in the second media content and the shot characteristics of the media content generated based on the first shot script segment. Specifically, the first shot text is a description text, which describes at least one picture element in the second media content and the shot characteristics of the media content generated based on the first shot script segment. For example, the content of the first shot text is: “wide-angle panoramic shot, the background is a clear sky, and the lens slowly advances to a corner”. The first shot text describes the shot characteristics of a video shot. In the subsequent steps, a corresponding shot video segment will be generated based on the first shot text, so that the generated shot video segment has the above-mentioned shot characteristics.
[0047] Further, in addition to the first shot script segment, the first shot script data described above can also include other shot script segments, such as a second shot script segment and a third shot script segment, which are used to generate other shot video segments. Similar to the first shot script segment, the second shot script segment and the third shot script segment also have the characteristics of the first shot script segment, i.e., include corresponding second media content and first shot text. The second media content is one or more media contents in the first media content, which are different from the second media content corresponding to the first shot script segment. The creation process of the first shot script segment (i.e., the second media content and the first shot text) can be implemented by a pre-trained model.
[0048] Further, as Figure 3As shown in a possible implementation, the specific implementation of step S102 includes:
[0049] Step S1021: obtaining at least one second media content according to the content features of each media content in the first media content, the content feature of the second media content corresponding to a preset content scenario, the preset content scenario representing an image picture composed of picture elements with specific correlation.
[0050] For example, first, the terminal device analyzes each media content in the first media content through an image processing model to obtain the content feature corresponding to each media content, such as an image feature matrix or vector; then, based on the above content feature, the media content with a preset content scenario is determined as the second media content, where the preset content scenario represents an image picture composed of picture elements with specific correlation. The preset content scenario may be, for example, “coffee store exploration”, “sunset and dusk”, “humanistic scenic spot”, and the like. For example, the first media content includes pictures p1, p2 and p3. The image processing model is used to extract and classify the content features of the pictures p1, p2 and p3 to determine that the image picture composed of picture elements (such as a coffee table, a coffee table, and a billboard) in the picture p1 corresponds to the preset content scenario “coffee store exploration”, and thus the picture p1 is determined as the second media content (one of). Similarly, the image processing model is used to determine that the preset content scenario corresponding to the picture p2 is “humanistic scenic spot”, and thus the picture p2 is also determined as the second media content. The image processing model does not detect the preset content scenario corresponding to the picture p3, and thus the picture p3 is not determined as the second media content. The case where the media content is a video is similar, and thus is not described herein.
[0051] Further, in the above case, when the first media content only includes a unique media content, such as the picture p1, if the picture p1 corresponds to the preset content scenario, it is determined as the second media content; if the picture p1 does not correspond to any preset content scenario, a prompt information is returned and displayed to prompt the user to reselect, otherwise the final video cannot be generated. When the first media content includes more than one media content, and more than one media content corresponds to the preset content scenario, in a possible implementation, one of them can be determined as the second media content by further responding to the selection operation of the user. In another possible implementation, the second media content can also be determined by the system recommendation.
[0052] Specifically, in a possible implementation, as shown in Figure 4 the specific implementation of step S1021 includes:
[0053] Step S1021-1: According to the content features of each media content in the first media content, at least one preset content scene is identified.
[0054] Step S1021-2: At least one target content scene is determined from the at least one preset content scene, and the target content scene is a hot content scene in the video content platform.
[0055] Step S1021-3: The media content corresponding to the target content scene is determined as the second media content.
[0056] For example, first, at least one preset content scene is identified according to the content features of each media content in the first media content, and then the hot content scene in the video content platform is obtained, for example, the video of the "coffee store exploration" type is the "hot video" video in the current video content platform, and then the "coffee store exploration" scene (hot content scene) in the at least one preset content scene is determined as the target content scene. Then, the media content corresponding to the target content scene is determined as the second media content.
[0057] In the steps of this embodiment, by combining the popular information of the video content platform, a video content scene with higher propagation rate and larger viewing volume, i.e., a hot content scene, is identified from multiple media contents, and then the second media content is determined based on the hot content scene, and the final video is generated based on the second media content, thereby improving the video propagation property of the final generated video.
[0058] Step S1022: The second media content is processed by calling a script generation model to generate a first script text corresponding to the second media content.
[0059] Step S1023: The first script data is generated according to the second media content and the corresponding first script text.
[0060] For example, after the second media content is determined, the second media content is processed by calling a script generation model to generate a corresponding first script text, which is a description word (prompt word, prompt) text. The script generation model can be a pre-trained neural network model, which can generate a corresponding description word, i.e., a first script text, according to the input media content (picture, video). The ability of the script generation model is obtained by training, and the implementation principle is not described here.
[0061] In one possible implementation, as shown in Figure 5 The specific implementation of step S1022 includes:
[0062] Step S1022-1: According to the preset content scene corresponding to the second media content, a corresponding first prompt word is generated.
[0063] Step S1022-2: input the first prompt word and the second media content into the script generation model to generate the first shot text.
[0064] Exemplarily, before inputting the second media content into the script generation model, first, the scene name of the preset content scene corresponding to the second media content is obtained, which is determined in the previous step by manual or automatic means, for example, the target content scene determined based on the hot content scene. Then, the scene name of the preset content scene corresponding to the second media content is converted into the corresponding first prompt word, which is used to control the process of generating the first shot text by the script generation model, so that the first shot text generated by the script generation model matches the scene name, for example, the content of the first prompt word is "this is a video clip with the theme of coffee store exploration". Then, the first prompt word and the second media content are input into the script generation model, and the script generation model will generate the corresponding first shot text in combination with the image content features of the second media content and the content of the above-mentioned first prompt word. In this embodiment, the scene name of the preset content scene corresponding to the second media content is converted into an additional prompt word (first prompt word) to control the process of generating the first shot text by the script generation model, which makes full use of the prior knowledge obtained in the previous step, so that the content of the generated first shot text matches the content scene (for example, the hot content scene) corresponding to the second media content, and improves the accuracy and content richness of the generated first shot text.
[0065] Further, optionally, in this embodiment, after step S101, it further includes:
[0066] Step S100: obtaining a feature label of a target user, the feature label being used to represent an interest point of the target user.
[0067] Correspondingly, the specific implementation manner of step S102 includes:
[0068] obtaining first shot script data according to the first media content and the feature label of the target user.
[0069] Specifically, in one possible implementation manner, after the terminal device obtains the first media content selected by the user, in addition to the above-mentioned embodiment steps, the second media content is selected by the image content features of the first media content, and the second media data combined with the personal interest point is determined by combining the feature label of the target user, and then the first shot script data is generated based on the second media data combined with the personal interest point. That is, the specific implementation manner of step S1021 is: obtaining at least one second media content according to the content features of each media content in the first media content and the feature label of the target user.
[0070] In another possible implementation, after the terminal device determines the second media data, the second media content and the feature label are processed by calling the script generation model to generate the first breakdown text that combines the personal interest point and the second media content. That is, the specific implementation of step S1022 is that the second media content and the feature label of the target user are processed by calling the script generation model to generate the first breakdown text corresponding to the second media content.
[0071] The feature label can be obtained by the terminal device from the video content platform after being authorized by the target user. The specific implementation of the first breakdown script data based on the feature label can be configured as needed, which is not limited herein.
[0072] Step S103: generating third media content according to the first breakdown script segment, and adding the third media content to the first breakdown script segment to obtain second breakdown script data.
[0073] Step S104: generating a video based on the second breakdown script data.
[0074] Further, after the first breakdown script segment is generated, the second media content corresponding to the first breakdown script segment and the first breakdown text can be used to further call a video generation model constructed based on a neural network model to generate a corresponding breakdown video segment. The first breakdown text is a description word matching the video generation model. In one possible implementation, the first breakdown text can be input into the video generation model, and the video generation capability of the video generation model is used to generate a video matching the picture features and the shot features described in the first breakdown text, that is, the third media content. In another possible implementation, the first breakdown text and the second media content can be input into the video generation model, and the video generation capability of the video generation model is used to generate a video matching the picture features of the second media content, the picture features and the shot features described in the first breakdown text. In this case, the first breakdown text contains processing operation instructions for the second media content, for example, including: continuing to write the second media content (video), generating a video based on the second media content (picture), and the like, so as to generate the third media content.
[0075] Afterwards, the third media content is added to the first script segment, the third media content can be a picture or a video, the first script segment corresponds to a script segment in the draft for generating a video, after one or more times of the above steps, the first script segment contains more than one third media content, at the same time, after a script segment (the first script segment) in the first script data changes, an updated script data, that is, the second script data, is formed, finally, the updated first script data, that is, the second script data, is used to perform video fusion, combination, and set to generate a final video.
[0076] Figure 6 A process diagram for generating a video is provided for the embodiments of the present disclosure, and the following will be described in combination with Figure 6 The above embodiment steps are introduced in more detail, as shown in Figure 6 Based on the selection operation of the user, the terminal device acquires the first media content and creates the first script data, which includes the picture P1, the picture P2, the picture P3, and the picture P4. Afterwards, the first media content is filtered by acquiring and using the user tags and the hot content scene to determine that the second media content is the picture P1 and the picture P2. Afterwards, the picture P1 and the picture P2 are processed by the pre-trained model to generate corresponding first script texts, that is, the script text T1 and the script text T2. In the first script data, corresponding script segment data is created for the picture P1 and the picture P2, that is, the script segment A and the script segment B, wherein the script segment A or the script segment B is the first script segment. In the data structure corresponding to the script segment A, the picture P1 and the script text T1 are contained; in the data structure corresponding to the script segment B, the picture P2 and the script text T2 are contained. Afterwards, the picture P1 and the script text T1 corresponding to the script segment A and the picture P2 and the script text T2 corresponding to the script segment B are processed by the video generation model to obtain the video V1 corresponding to the script segment A and the video V2 corresponding to the script segment B. The video V1 and the video V2 are the third media content referred to in the above embodiment. Afterwards, the video V1 and the video V2 are added to the corresponding first script segment, that is, the video V1 is added to the script segment A and the video V2 is added to the script segment B to form two script video segments, and finally, the script segment A and the script segment B are fused to generate a final generated video.
[0077] The method comprises the following steps: acquiring first media content; acquiring first split script data according to the first media content, wherein the first split script data comprises a first split script segment, the first split script segment comprises second media content and first split text, the first split text is used for describing at least one picture element in the second media content and a shot feature of media content generated based on the first split script segment, and the second media content comprises at least one media content in the first media content; generating third media content according to the first split script segment, and adding the third media content to the first split script segment to obtain second split script data; and generating a video based on the second split script data. According to the method, the first split script data is generated directly based on the first media content provided by a user, the image feature of the second media content corresponding to each split script segment in the first split script data and the shot feature represented by the first split text are used to generate corresponding third media content, and the third media content is added to the corresponding split script segment to generate second split script data. Finally, the video is generated based on the split script segment in the generated second split script data. The method realizes the whole-process video generation, does not need the user to manually control the media material generation process, the layout and the editing process, improves the video generation efficiency and the video quality, and avoids the video duplication problem because the first media content input by the user is different.
[0078] Reference Figure 7 , Figure 7 Flowchart of a video generation method provided by the embodiments of the present disclosure Figure 1 The embodiments of the present disclosure provide a video generation method. Figure 2 The embodiments of the present disclosure provide a video generation method.
[0079] Step S201: acquiring first media content.
[0080] Step S202: acquiring first split script data according to the first media content, wherein the first split script data comprises at least a first split script segment and a second split script segment.
[0081] Step S203: determining a segment order of the first split script segment and the second split script segment according to second media content corresponding to the first split script segment and fourth media content corresponding to the second split script segment.
[0082] Exemplarily, in the embodiment, the first media content selected by the user includes at least two media contents, based on the first media content, the second media content is determined, and the fourth media content is also determined, wherein the fourth media content is the media content in the first media content used to generate the second split script segment. The determination scheme of the fourth media content is the same as the scheme of determining the second media content, which has been described in detail in the previous embodiments, and will not be repeated here. Correspondingly, the first split script data created based on the second media content contains at least two split script segments, that is, the first split script data at least includes the first split script segment and the second split script segment, that is, since the first split script data contains at least two split script segments, the video generated based on the first split script data is composed of at least two video split segments (corresponding to the split script segments). In the related technology, when there are multiple split scripts or multiple video split segments need to be generated, the playing order of the split video segments needs to be manually designed, so there is the problem of low efficiency and affecting the quality of the video. In the embodiment, after determining the second media content according to the first media content, the picture features of the second media content corresponding to the first split script segment and the picture features of the fourth media content corresponding to the second split script segment are used to determine the segment order of the first split script segment and the second split script segment by using the model. Specifically, for example, the picture features of the second media content and the fourth media content are converted into description texts, and then the inference ability of the model is used to determine the order of each description text based on the semantics of the description texts, and further determine the segment order. Thus, the playing order of the split video segments is adaptively matched, and the video generation efficiency and quality are improved.
[0083] Step S204: according to the segment order, display the first split script text, and for the script score corresponding to the first split script text, the script score is used to represent the content richness of the media content described by the first split script text.
[0084] Furthermore, after determining the first storyboard segment, the second storyboard segment, and other storyboard segments, the content of each storyboard segment is displayed according to the segment order. Specifically, for example, the first storyboard segment corresponds to the first video storyboard segment, i.e., it is located at the first position in the segment order. The terminal device displays the first storyboard segment by default, specifically, the first storyboard text and the script score corresponding to the first storyboard text. Optionally, the corresponding second media content can also be displayed. The script score is used to characterize the richness of the media content described by the first storyboard text. The higher the script score, the richer the content of the media content described by the first storyboard text, and vice versa. In one possible implementation, the script score can be determined by the number of words in the first storyboard text; the more words, the higher the script score. Of course, other pre-configured models can also be used to evaluate the first storyboard text to obtain the script score, which can be set as needed. The script score can prompt and guide users to optimize the content of the current first storyboard text to generate a higher quality video.
[0085] Figure 8 This is a schematic diagram of an interactive interface provided in an embodiment of the present disclosure, such as... Figure 8 As shown, the target application's interactive interface displays the second media content corresponding to the first storyboard fragment, the first storyboard text, and the rating corresponding to the first storyboard text. The second media content is image P1 as shown in the figure, and the first storyboard text reads, "First Scene Storyboard: Wide-angle panoramic shot, background is a clear sky, the camera slowly zooms in on a corner." Simultaneously, the interactive interface also displays the first storyboard text corresponding to the second storyboard fragment, which reads, "Second Scene Storyboard: Static mid-shot in natural light, slowly zooming in from the window to the protagonist, who is holding a cup of coffee and gazing out the window." This first storyboard text corresponding to the second storyboard fragment describes the fourth media content corresponding to the second storyboard fragment and will not be elaborated further. When the user clicks on the first storyboard text corresponding to the second storyboard fragment, the interactive interface displays the fourth media content corresponding to the second storyboard fragment. Meanwhile, for example, when a certain storyboard script segment (such as the currently selected first storyboard script segment) is selected, clicking the "Rating" button will display the script score of the first storyboard text corresponding to that storyboard script segment. For example, as shown in the figure, the script score corresponding to the first storyboard script segment is 50 points. If the highest score for a script is 100 points, it means that the content richness of the media content described by the current first storyboard text is low and can be further optimized.
[0086] Step S205: In response to the first user instruction, an optimized screenplay text corresponding to the first screenplay text is generated, the optimized screenplay text contains the first picture element, and the optimized screenplay text is used to represent the camera movement feature based on the first picture element.
[0087] For example, further, by responding to the first user instruction input by the user, the first screenplay text can be optimized to generate the optimized screenplay text. In one possible implementation, a trigger control is arranged in the interactive interface, and when the trigger control is triggered, the first user instruction is generated, so as to realize automatic optimization of the first screenplay text, that is, a "one-key optimization" function. The function can be realized by a pre-configured prompt word model, which is not described herein. In another possible implementation, the first user instruction input manually can be used to modify the content of the first screenplay text, so as to generate the optimized screenplay text. The optimized screenplay text obtained after optimization is used to represent the camera movement feature based on the first picture element. For example, an example of the content of the optimized screenplay text is: "wide-angle panoramic shot, the background is a clear sky, and the lens slowly advances to a corner. A handheld camera shot at a high speed, which quickly passes through the lake surface and enters the mountains in a dynamic motion". In the above example, "lake surface" and "mountains" are picture elements in the second media content, that is, the first picture element as described above. In the lens feature described in the optimized screenplay text, the specific first picture element in the second media content is combined to describe the camera movement feature, so as to realize more rich and dynamic description of the third media content to be generated, improve the richness of the content description, and further make the picture content of the generated third media content more rich and accurate.
[0088] Further, as shown in Figure 9 , for example, the specific implementation of step S205 includes:
[0089] Step S2051: In response to the first user instruction, the element identifier of the picture element in the second media content is displayed.
[0090] Step S2052: In response to the click operation on the element identifier, the first picture element is determined, and the optimized screenplay text is constructed based on the first picture element.
[0091] For example, in combination with Figure 8The schematic diagram of the interaction interface is shown. The terminal device displays the corresponding second media content while displaying the first split script text. In this embodiment, the terminal device displays the element identifiers of the picture elements in the second media content in response to the first user instruction, for example, displays the element identifiers (for example, abbreviations of the picture elements) of the picture elements such as "mountain range", "lake surface", and "pleasure boat" on the second media content (picture). Then, the first picture element is determined in response to the point selection operation (that is, the trigger operation) of the element identifier applied by the user, so as to construct the optimized split script text.
[0092] Further, in a possible implementation, the specific implementation of step S2052 includes:
[0093] Step S2052A-1: According to the content features of the second media content, a corresponding prompt word template is obtained and displayed. The prompt word template contains a blank position for writing an identifier indicating a target object, and the prompt word template is used to describe a target camera operation feature realized based on at least one target object.
[0094] Step S2052A-2: In response to the point selection operation of the element identifier of the first picture element, the element identifier of the first picture element is written into the blank position in the prompt word template, and the optimized split script text is generated.
[0095] For example, in the step of this embodiment, the terminal device determines a corresponding prompt word template according to the content features of the second media content. The prompt word template contains a part of fixed description text, which is determined by the content features of the second media content and is used to realize the picture description of the second media content. On the other hand, the prompt word template also contains a blank position for writing an identifier indicating a target object. The user can determine the first picture element through a trigger operation. Then, the terminal device writes the element identifier of the first picture element into the blank position in the prompt word template, so as to generate the optimized split script text.
[0096] Figure 10 A process schematic diagram for generating an optimized split script text provided by an embodiment of the present disclosure is shown in FIG. 25. Figure 10As shown, in the interactive interface of the target application, the second media content corresponding to the first split script fragment and the first split text are displayed, wherein the second media content is the picture P1 shown in the figure, and the positions of the image elements on the picture P1 also display the corresponding element identifiers, such as "mountain", "lake surface", etc. The first split text is generated based on the prompt word template, and the content is: "wide-angle panoramic shot, the background is a clear sky, the lens slowly advances a corner. A fast handheld camera shot, quickly passes through ___ in a dynamic motion, and enters into ___". In the above first split text, "___" is a blank, and other text content is provided by the prompt word template. Subsequently, in response to the triggering operation of the user for the element identifier, "mountain" and "lake surface" are filled into the above "___", thereby forming the optimized split text, that is, "wide-angle panoramic shot, the background is a clear sky, the lens slowly advances a corner. A fast handheld camera shot, quickly passes through the lake surface in a dynamic motion, and enters into the mountain".
[0097] In this embodiment, the prompt word template matched with the second media content is determined through the image content features of the second media content, and then the starting point of the lens operation is selected from the image elements existing in the second media content itself in combination with the selection operation of the user, so as to realize the design of the video lens operation. This process only needs the user to implement a simple triggering operation, and realizes the high degree of freedom of video content creation while greatly improving the interactive efficiency and video generation efficiency.
[0098] In another possible implementation, the specific implementation of step S2052 includes:
[0099] Step S2052B-1: In response to the second user instruction, selecting the target platform media content in the video content platform.
[0100] Step S2052B-2: Acquiring and displaying the split text template corresponding to the target platform media content, the split text template containing a first blank for writing an identifier indicating a target object, the split text template being used to describe the picture features of the platform media content and the target lens operation features realized based on the target object.
[0101] Step S2052B-3: In response to the triggering operation of the element identifier for the first picture element, writing the element identifier of the first picture element into the first blank in the split text template to generate an optimized split text.
[0102] Exemplarily, in another possible implementation, the user can also imitate the shot features of the target platform media content in the video content platform to generate a video by selecting the target platform media content in the video content platform, specifically, first, the terminal device selects the target platform media content in the video content platform, for example, a popular video in the video content platform, by responding to the second user instruction input by the user. Then, the corresponding shot script template of the target platform media content is obtained and displayed. The shot script template is similar to the first shot script, and is used to describe the picture elements and shot features (for example, camera movement features) of the corresponding media content. The shot script template contains fixed description text and a first blank position. The fixed description text describes the image features (for example, picture elements, positions between picture elements, visual effects, etc.) of the target platform media content, and the first blank position is used to write an identifier indicating a target object. The shot script template is used to describe the picture features of the platform media content and the target camera movement features realized based on the target object. The shot script template is pre-generated, for example, it can be generated in the process of generating the target platform media content, and the generation process is not described herein. Then, the element identifier of the first picture element is written into the first blank position in the shot script template by responding to the triggering operation of the element identifier of the first picture element, to generate an optimized shot script. The process is similar to the process of writing the element identifier of the first picture element in the blank position in the prompt word template in the embodiment shown in FIG. 20 to generate the optimized shot script, and is not described herein. Figure 10
[0103] In this embodiment, the optimized shot script is generated by obtaining the target platform media content in the video content platform and based on the shot script template of the target platform media content, which realizes the imitation of the camera movement features of the target platform media content. In this process, the user can also select different picture elements to flexibly change the camera movement features, further improving the efficiency of generating a video by the user and the diversity of video content.
[0104] Further, optionally, the shot script template further contains a second blank position, and the second blank position is used to write a visual effect identifier representing the visual effect of the picture element. In this embodiment, the steps further include:
[0105] Step S2052B-4: In response to the first triggering instruction for the second blank position, display at least two alternative visual effect identifiers corresponding to the second blank position.
[0106] Correspondingly, the specific implementation of step S2052B-3 includes:
[0107] In response to a trigger operation for the element identifier of the first picture element, the element identifier of the first picture element is written into the first blank in the screenplay text template, and in response to a trigger operation for a target visual effect identifier in the at least two alternative visual effect identifiers, the target visual effect identifier is written into the second blank in the screenplay text template, to generate the optimized screenplay text.
[0108] Exemplarily, in another possible implementation, the screenplay text template further includes a second blank, the second blank being used to write a visual effect identifier representing a visual effect of a picture element, in short, the second blank is used to write a relevant descriptor of the visual effect, for example, "layered", "glassy". After the terminal device acquires the screenplay text template, according to the screenplay text template, the terminal device can obtain at least two alternative visual effect identifiers corresponding to each second blank. In response to a first trigger instruction for the second blank, the terminal device can display the alternative visual effect identifiers corresponding to the second blank. Then, the user can select an alternative visual effect identifier as a target visual effect identifier according to needs, and write the target visual effect identifier into the second blank. Then, in combination with the fixed description text of the screenplay text template, the content written in the first blank, and the content written in the second blank, the optimized screenplay text is generated. Thus, the content richness of the optimized screenplay text is further improved while ensuring the interaction efficiency.
[0109] Figure 11 Another process schematic diagram for generating an optimized screenplay text provided by the embodiments of the present disclosure is as follows: Figure 11As shown, in the interactive interface of the target application, the second media content (P1 in the figure) corresponding to the first split script segment and the first split text are displayed, and the content of the first split text is, for example, "First split script: wide-angle panoramic shot, the background is a clear sky, the lens slowly pushes to a corner". Then, in response to a second user instruction, a platform media content page of a video content platform is displayed, and the target platform media content (P2 in the figure) is obtained based on the selection operation of the user. Then, based on the split text template of the target platform media content, the text content including the first and second positions is inserted into the first split text, for example, as shown in the figure, the inserted content in the first split text is "Sunlight is inclined from the ___ of the ___ and forms a natural light column like a visual focus. A fast handheld camera lens quickly passes through the ___ of the ___ in a dynamic motion and enters the ___". Then, the user determines the element identifier of the first picture element or the target visual effect identifier from the element identifier or the alternative visual effect identifier of the picture element displayed in the second media content by sequentially selecting (first trigger instruction) each position, and writes into the corresponding position, thereby generating an optimized split text, for example, as shown in the figure, the content of the optimized split text corresponding to the first split script segment is "First split script: wide-angle panoramic shot, the background is a clear sky, the lens slowly pushes to a corner. Sunlight is inclined from the dense cloud layer and forms a natural light column like a visual focus. A fast handheld camera lens quickly passes through the layer upon layer of mountains in a dynamic motion and enters the forest"
[0110] Step S206: Display the second split text corresponding to the second split script segment, and generate an optimized split text corresponding to the second split text in response to a second user instruction.
[0111] Exemplarily, the second split text for the second split script segment can be optimized in the same way based on the above steps, thereby obtaining the optimized split text corresponding to the second split text. For details, refer to the process of generating the optimized split text corresponding to the first split text, which will not be described here.
[0112] Step S207: Generate the third media content corresponding to the first split script segment and the third media content corresponding to the second split script segment according to the optimized split text corresponding to the first split script segment, the optimized split text corresponding to the second split script segment, and the segment order of the first split script segment and the second split script segment.
[0113] Step S208: Add the third media content to the corresponding first split script segment and second split script segment respectively to obtain second split script data, and generate a video based on the second split script data.
[0114] Further, after generating the optimized screenplay text corresponding to the first screenplay script segment and the optimized screenplay text corresponding to the second screenplay script segment, the third media content corresponding to the first screenplay script segment and the second screenplay script segment is generated in combination with the segment order of the first screenplay script segment and the second screenplay script segment. This process can be performed by inputting the optimized screenplay text corresponding to each screenplay script segment (the first screenplay script segment, the second screenplay script segment, etc.) and the segment order of each screenplay script segment into a video generation model. The video generation model takes the above information as context information, comprehensively reasons and fuses, and generates the third media content corresponding to each screenplay script segment. Since the context information such as the segment order is combined, the third media content corresponding to each screenplay script segment has content continuity, and the content of the finally generated video is more coherent and real, and has better video quality.
[0115] In this embodiment, the implementation manner of step S201 is the same as that of step S101 in the embodiment of the disclosure shown in the above Figure 2 The implementation manner of step S201 is the same as that of step S101 in the embodiment of the disclosure shown in the above
[0116] The video generation method corresponding to the above embodiment, Figure 12 A structural block diagram of a video generation apparatus provided by an embodiment of the disclosure is shown in FIG. 3. The method introduced in the above embodiment can be performed by the video generation apparatus. The apparatus can be implemented in a software and / or hardware manner, and can be integrated in an electronic device having a certain data processing function. The electronic device can include but is not limited to a mobile terminal having a large data processing capacity, and a desktop computer, a supercomputer, and other fixed terminals having a large data processing capacity.
[0117] For ease of illustration, only parts related to the embodiments of the disclosure are shown. For details, refer to the above Figure 12 The video generation apparatus 3 includes:
[0118] The acquisition module 31 is configured to acquire the first media content, and acquire the first screenplay script data according to the first media content. The first screenplay script data includes the first screenplay script segment. The first screenplay script segment includes the second media content and the first screenplay text. The first screenplay text is used to describe at least one picture element in the second media content and the shot feature of the media content generated based on the first screenplay script segment. The second media content includes at least one media content in the first media content.
[0119] The processing module 32 is configured to generate the third media content according to the first screenplay script segment, and add the third media content to the first screenplay script segment to obtain the second screenplay script data.
[0120] The generation module 33 is configured to generate the video based on the second screenplay script data.
[0121] According to one or more embodiments of the present disclosure, the obtaining module 31, in the process of obtaining the first script data according to the first media content, is specifically configured to: obtain at least one second media content according to the content features of each media content in the first media content, the content features of the second media content corresponding to a preset content scenario, the preset content scenario representing an image picture composed of picture elements with specific relevance; generate a first script text corresponding to the second media content by processing the second media content through a script generation model; and generate the first script data according to the second media content and the corresponding first script text.
[0122] According to one or more embodiments of the present disclosure, in the process of obtaining at least one second media content according to the content features of each media content in the first media content, the obtaining module 31 is specifically configured to: identify at least one preset content scenario according to the content features of each media content in the first media content; determine at least one target content scenario from the at least one preset content scenario, the target content scenario being a hot content scenario in the video content platform; and determine the media content corresponding to the target content scenario as the second media content.
[0123] According to one or more embodiments of the present disclosure, in the process of generating a first script text corresponding to the second media content by processing the second media content through a script generation model, the obtaining module 31 is specifically configured to: generate a first prompt word corresponding to the preset content scenario of the second media content; and input the first prompt word and the second media content into the script generation model to generate the first script text.
[0124] According to one or more embodiments of the present disclosure, the obtaining module 31 is further configured to: obtain a feature label of the target user, the feature label being used to represent the interest points of the target user; and in the process of obtaining the first script data according to the first media content, the obtaining module 31 is specifically configured to: obtain the first script data according to the first media content and the feature label of the target user.
[0125] According to one or more embodiments of the present disclosure, the first script data further includes at least a second script segment, and the processing module 32 is further configured to: determine a segment order of the first script segment and the second script segment according to the second media content corresponding to the first script segment and the fourth media content corresponding to the second script segment; and generate a third media content based on the segment order, the first script segment, the second script segment, and the segment order of the first script segment and the second script segment.
[0126] According to one or more embodiments of the present disclosure, the processing module 32 is further configured to: display the first breakdown text, and a script score corresponding to the first breakdown text, the script score being used to represent content richness of the media content described by the first breakdown text; and in response to a first user instruction, generate an optimized breakdown text corresponding to the first breakdown text, the optimized breakdown text containing the first picture element, and the optimized breakdown text being used to represent a camera movement feature based on the first picture element; and the processing module 32 is specifically configured to generate the third media content according to the first breakdown script segment, by generating the third media content according to the optimized breakdown text.
[0127] According to one or more embodiments of the present disclosure, the processing module 32 is specifically configured to, in response to the first user instruction, display an element identifier of a picture element in the second media content, and in response to a triggering operation on the element identifier, determine the first picture element and construct the optimized breakdown text based on the first picture element.
[0128] According to one or more embodiments of the present disclosure, the processing module 32 is specifically configured to, in response to the triggering operation on the element identifier, determine the first picture element and construct the optimized breakdown text based on the first picture element, by: obtaining and displaying a corresponding prompt word template according to a content feature of the second media content, the prompt word template containing a blank position for writing an identifier of a target object, and the prompt word template being used to describe a target camera movement feature realized based on at least one target object; and in response to the triggering operation on the element identifier of the first picture element, writing the element identifier of the first picture element into the blank position in the prompt word template to generate the optimized breakdown text.
[0129] According to one or more embodiments of the present disclosure, the processing module 32 is specifically configured to, in response to the triggering operation on the element identifier, determine the first picture element and construct the optimized breakdown text based on the first picture element, by: in response to a second user instruction, selecting target platform media content in a video content platform; obtaining and displaying a breakdown text template corresponding to the target platform media content, the breakdown text template containing a first blank position for writing an identifier of a target object, and the breakdown text template being used to describe a picture feature of the platform media content and a target camera movement feature realized based on the target object; and in response to the triggering operation on the element identifier of the first picture element, writing the element identifier of the first picture element into the first blank position in the breakdown text template to generate the optimized breakdown text.
[0130] According to one or more embodiments of this disclosure, the storyboard text template further includes a second empty space, which is used to write a visual effect identifier representing the visual effect of a scene element; the processing module 32 is further configured to: in response to a first trigger command for the second empty space, display at least two alternative visual effect identifiers corresponding to the second empty space; when the processing module 32 writes the element identifier of the first scene element into the first empty space in the storyboard text template in response to a trigger operation for the element identifier of the first scene element to generate optimized storyboard text, it is specifically configured to: in response to a trigger operation for the element identifier of the first scene element, write the element identifier of the first scene element into the first empty space in the storyboard text template, and in response to a trigger operation for the target visual effect identifier among the at least two alternative visual effect identifiers, write the target visual effect identifier into the second empty space in the storyboard text template to generate optimized storyboard text.
[0131] The acquisition module 31, processing module 32, and generation module 33 are connected sequentially. The video generation device 3 provided in this embodiment can execute the technical solution of the above method embodiment, and its implementation principle and technical effect are similar, so it will not be described again here.
[0132] Figure 13 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present disclosure, such as... Figure 13 As shown, the electronic device 4 includes:
[0133] Processor 41, and memory 42 communicatively connected to processor 41;
[0134] Memory 42 stores instructions executed by the computer;
[0135] The processor 41 executes computer execution instructions stored in the memory 42 to achieve, for example, Figures 2-11 The video generation method in the illustrated embodiment.
[0136] Optionally, the processor 41 and the memory 42 are connected via a bus 43.
[0137] For relevant instructions, please refer to the corresponding text. Figures 2-11 The relevant descriptions and effects of the steps in the corresponding embodiments are understood, and will not be elaborated on here.
[0138] This disclosure provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement this disclosure. Figures 2-11 The video generation method provided in any of the corresponding embodiments.
[0139] This disclosure provides a computer program product, including a computer program, which, when executed by a processor, implements this disclosure.Figures 2-11 The video generation method according to any one of the embodiments.
[0140] To implement the above-mentioned embodiments, the electronic device according to the embodiments of the present disclosure is also provided.
[0141] Reference Figure 14 , which shows a structural diagram of an electronic device 900 suitable for implementing the embodiments of the present disclosure. The electronic device 900 can be a terminal device or a server. The terminal device can include, but is not limited to, a mobile terminal such as a mobile phone, a notebook computer, a digital broadcast receiver, a Personal Digital Assistant (PDA), a Tablet Computer, a Portable Media Player (PMP), a vehicle terminal (e.g., a car navigation terminal), and the like, and a stationary terminal such as a digital TV, a desktop computer, and the like. Figure 14 The electronic device shown is only an example and should not impose any limitation on the functions and use range of the embodiments of the present disclosure.
[0142] As shown in Figure 14 , the electronic device 900 can include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 901 that can perform various appropriate actions and processes according to programs stored in a Read Only Memory (ROM) 902 or loaded into a Random Access Memory (RAM) 903 from a storage device 908. Various programs and data required for the operation of the electronic device 900 are also stored in the RAM 903. The processing device 901, the ROM 902, and the RAM 903 are connected to each other through a bus 904. An Input / Output (I / O) interface 905 is also connected to the bus 904.
[0143] Generally, the following devices can be connected to the I / O interface 905: an input device 906 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, and the like; an output device 907 including, for example, a Liquid Crystal Display (LCD), a speaker, a vibrator, and the like; a storage device 908 including, for example, a magnetic tape, a hard disk, and the like; and a communication device 909. The communication device 909 can allow the electronic device 900 to communicate with other devices wirelessly or via a wire to exchange data. Although Figure 14 The electronic device 900 with various devices is shown, but it should be understood that all the devices shown are not required to be implemented or possessed. More or fewer devices can be alternatively implemented or possessed.
[0144] In particular, according to embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as a computer software program. For example, embodiments of the present disclosure include a computer program product comprising a computer program carried on a computer readable medium, the computer program comprising program code for performing the methods illustrated by the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via the communication device 909, or installed from the storage device 908, or installed from the ROM 902. When the computer program is executed by the processing device 901, the above-mentioned functions defined in the methods of the embodiments of the present disclosure are performed.
[0145] It should be noted that the computer readable medium described above in the present disclosure can be a computer readable signal medium or a computer readable storage medium or any combination thereof. The computer readable storage medium, for example, can be, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or apparatus, or any suitable combination of the above. More specific examples of the computer readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer readable storage medium can be any tangible medium that contains or stores a program that can be used by or in connection with an instruction execution system, apparatus, or device. In the present disclosure, the computer readable signal medium can include a data signal carried in a baseband or as part of a carrier wave, in which the computer readable program code is carried. Such a propagated data signal can take any of a variety of forms, including but not limited to electro-magnetic, optical, or any suitable combination thereof. The computer readable signal medium can also be any computer readable medium that is not a computer readable storage medium and that can communicate, propagate or transport a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer readable medium can be transmitted by any suitable medium, including but not limited to wire, cable, RF (radio frequency), etc., or any suitable combination of the above.
[0146] The computer readable medium described above can be included in the electronic device described above; or can exist separately from the electronic device and be not assembled into the electronic device.
[0147] The computer readable medium described above carries one or more programs, which, when executed by the electronic device, cause the electronic device to perform the methods illustrated by the embodiments described above.
[0148] Computer program code for carrying out operations of the present disclosure can be written in any combination of one or more programming languages, including an object oriented programming language such as Java, Smalltalk, C++ or the like and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).
[0149] The flow diagrams and the block diagrams in the drawings are illustrations of architectures, functionalities, and operations of possible implementations of systems, methods, and computer program products according to various embodiments of present disclosure. In this regard, each block in the flow diagrams or block diagrams can represent a module, a procedure, or a part of code, which comprises one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently or the blocks may
[0150] The units or modules described in the embodiments of the present disclosure can be implemented by software, or by hardware. In some cases, the name of the unit or module does not constitute a limitation on the unit itself.
[0151] The functions described above in the detailed description of embodiments of the present disclosure can be performed by one or more hardware logic components. For example, and without limitation, illustrative types of hardware logic components that can be used include Field-programmable Gate Arrays (FPGAs), Program-specific Integrated Circuits (ASICs), Program-specific Standard Products (ASSPs), System-on-a-chip systems (SOCs), Complex Programmable Logic Devices (CPLDs), etc.
[0152] In the context of this disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include but is not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium will include one or more of: an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0153] In a first aspect, according to one or more embodiments of the present disclosure, a video generation method is provided, comprising:
[0154] obtaining first media content; obtaining first shot script data according to the first media content, the first shot script data including a first shot script segment, the first shot script segment including second media content and first shot text, the first shot text being used to describe at least one picture element in the second media content and a shot feature of media content generated based on the first shot script segment, the second media content including at least one media content in the first media content; generating third media content according to the first shot script segment, and adding the third media content to the first shot script segment to obtain second shot script data; and generating a video based on the second shot script data.
[0155] According to one or more embodiments of the present disclosure, the obtaining first shot script data according to the first media content comprises: obtaining at least one second media content according to content features of each media content in the first media content, the content features of the second media content corresponding to a preset content scenario, the preset content scenario representing an image picture composed of picture elements with specific correlation; generating first shot text corresponding to the second media content by calling a script generation model to process the second media content; and generating the first shot script data according to the second media content and the corresponding first shot text.
[0156] According to one or more embodiments of the present disclosure, the obtaining the at least one second media content according to the content features of the media content in the first media content comprises: identifying at least one preset content scene according to the content features of the media content in the first media content; determining at least one target content scene from the at least one preset content scene, the target content scene being a hot content scene in the video content platform; and determining the media content corresponding to the target content scene as the second media content.
[0157] According to one or more embodiments of the present disclosure, the generating the first breakdown text corresponding to the second media content by processing the second media content through the script generation model comprises: generating a corresponding first prompt word according to a preset content scene corresponding to the second media content; and inputting the first prompt word and the second media content into the script generation model to generate the first breakdown text.
[0158] According to one or more embodiments of the present disclosure, the method further comprises: obtaining a feature label of a target user, the feature label being used to represent an interest point of the target user; and the obtaining the first breakdown script data according to the first media content comprises: obtaining the first breakdown script data according to the first media content and the feature label of the target user.
[0159] According to one or more embodiments of the present disclosure, the first breakdown script data further comprises at least a second breakdown script segment, and the method further comprises: determining a segment order of the first breakdown script segment and a second breakdown script segment according to second media content corresponding to the first breakdown script segment and fourth media content corresponding to the second breakdown script segment; and the generating the third media content according to the first breakdown script segment comprises: generating the third media content according to the first breakdown script segment, the second breakdown script segment, and the segment order. According to one or more embodiments of the present disclosure, the method further comprises: displaying the first breakdown text, and a script score corresponding to the first breakdown text, the script score being used to represent a content richness of media content described by the first breakdown text; in response to a first user instruction, generating an optimized breakdown text corresponding to the first breakdown text, the optimized breakdown text comprising a first picture element, the optimized breakdown text being used to represent a panning feature based on the first picture element; and the generating the third media content according to the first breakdown script segment comprises: generating the third media content according to the optimized breakdown text.
[0160] According to one or more embodiments of the present disclosure, the generating the optimized screenplay text corresponding to the first screenplay text in response to the first user instruction comprises: displaying an element identifier of a picture element in the second media content in response to the first user instruction; determining the first picture element in response to a triggering operation on the element identifier, and constructing the optimized screenplay text based on the first picture element.
[0161] According to one or more embodiments of the present disclosure, the determining the first picture element in response to the triggering operation on the element identifier, and constructing the optimized screenplay text based on the first picture element comprises: acquiring and displaying a corresponding prompt word template according to a content feature of the second media content, the prompt word template containing a blank position for writing an identifier of a target object, and the prompt word template being used to describe a target camera operation feature realized based on at least one target object; and writing the element identifier of the first picture element into the blank position in the prompt word template in response to the triggering operation on the element identifier of the first picture element, to generate the optimized screenplay text.
[0162] According to one or more embodiments of the present disclosure, the determining the first picture element in response to the triggering operation on the element identifier, and constructing the optimized screenplay text based on the first picture element comprises: selecting target platform media content in a video content platform in response to a second user instruction; acquiring and displaying a screenplay text template corresponding to the target platform media content, the screenplay text template containing a first blank position for writing an identifier of a target object, the screenplay text template being used to describe a picture feature of the platform media content and a target camera operation feature realized based on the target object; and writing the element identifier of the first picture element into the first blank position in the screenplay text template in response to the triggering operation on the element identifier of the first picture element, to generate the optimized screenplay text.
[0163] According to one or more embodiments of the present disclosure, the split-screen text template further comprises a second position for writing a visual effect identifier representing a visual effect of a picture element; the method further comprises: in response to a first trigger instruction for the second position, displaying at least two alternative visual effect identifiers corresponding to the second position; and the writing, in response to the trigger operation on the element identifier of the first picture element, of the element identifier of the first picture element into the first position in the split-screen text template to generate the optimized split-screen text comprises: writing, in response to the trigger operation on the element identifier of the first picture element, of the element identifier of the first picture element into the first position in the split-screen text template, and writing, in response to a trigger operation on a target visual effect identifier of the at least two alternative visual effect identifiers, of the target visual effect identifier into the second position in the split-screen text template to generate the optimized split-screen text.
[0164] In a second aspect, according to one or more embodiments of the present disclosure, a video generation apparatus is provided, comprising:
[0165] an acquisition module configured to acquire first media content, and acquire first split-screen script data according to the first media content, the first split-screen script data comprising a first split-screen script segment, the first split-screen script segment comprising second media content and a first split-screen text, the first split-screen text being used to describe at least one picture element in the second media content and a shot feature of media content generated based on the first split-screen script segment, the second media content comprising at least one media content in the first media content;
[0166] a processing module configured to generate third media content according to the first split-screen script segment, and add the third media content to the first split-screen script segment to obtain second split-screen script data;
[0167] a generation module configured to generate a video based on the second split-screen script data.
[0168] According to one or more embodiments of the present disclosure, when acquiring the first split-screen script data according to the first media content, the acquisition module is specifically configured to: obtain at least one second media content according to content features of each media content in the first media content, the content feature of the second media content corresponding to a preset content scenario, the preset content scenario representing an image picture composed of picture elements having a specific correlation; generate first split-screen text corresponding to the second media content by invoking a script generation model to process the second media content; and generate the first split-screen script data according to the second media content and the corresponding first split-screen text.
[0169] According to one or more embodiments of the present disclosure, the obtaining module is specifically configured to: according to the content features of each media content in the first media content, identify at least one preset content scene; determine at least one target content scene from the at least one preset content scene, the target content scene being a hot content scene in the video content platform; and determine the media content corresponding to the target content scene as the second media content.
[0170] According to one or more embodiments of the present disclosure, the obtaining module is specifically configured to: according to the preset content scene corresponding to the second media content, generate a corresponding first prompt word; and input the first prompt word and the second media content into the script generation model to generate the first shot text.
[0171] According to one or more embodiments of the present disclosure, the obtaining module is further configured to: obtain a feature label of a target user, the feature label being used to represent an interest point of the target user; and the obtaining module is specifically configured to: according to the first media content and the feature label of the target user, obtain the first shot script data.
[0172] According to one or more embodiments of the present disclosure, the first shot script data further includes at least a second shot script segment, and the processing module is further configured to: according to the second media content corresponding to the first shot script segment and the fourth media content corresponding to the second shot script segment, determine a segment order of the first shot script segment and the second shot script segment; and the processing module is specifically configured to: according to the first shot script segment, the second shot script segment, and the segment order, generate the third media content.
[0173] According to one or more embodiments of the present disclosure, the processing module is further configured to: display the first shot text, and a script score corresponding to the first shot text, the script score being used to represent a content richness of the media content described by the first shot text; and in response to a first user instruction, generate an optimized shot text corresponding to the first shot text, the optimized shot text including a first picture element, the optimized shot text being used to represent a panning feature based on the first picture element; and the processing module is specifically configured to: according to the optimized shot text, generate the third media content.
[0174] According to one or more embodiments of the present disclosure, the processing module, in response to the first user instruction, generates the optimized screenplay text corresponding to the first screenplay text, specifically configured to: in response to the first user instruction, display an element identifier of a picture element in the second media content; in response to a trigger operation on the element identifier, determine the first picture element, and construct the optimized screenplay text based on the first picture element.
[0175] According to one or more embodiments of the present disclosure, the processing module, in response to the trigger operation on the element identifier, determines the first picture element, and constructs the optimized screenplay text based on the first picture element, specifically configured to: according to the content characteristics of the second media content, obtain and display a corresponding prompt word template, the prompt word template containing a blank, the blank being used to write an identifier indicating a target object, the prompt word template being used to describe a target camera operation feature realized based on at least one target object; in response to the trigger operation on the element identifier of the first picture element, write the element identifier of the first picture element into the blank in the prompt word template to generate the optimized screenplay text.
[0176] According to one or more embodiments of the present disclosure, the processing module, in response to the trigger operation on the element identifier, determines the first picture element, and constructs the optimized screenplay text based on the first picture element, specifically configured to: in response to a second user instruction, select target platform media content in a video content platform; obtain and display a screenplay text template corresponding to the target platform media content, the screenplay text template containing a first blank, the first blank being used to write an identifier indicating a target object, the screenplay text template being used to describe picture characteristics of the platform media content and target camera operation features realized based on the target object; in response to the trigger operation on the element identifier of the first picture element, write the element identifier of the first picture element into the first blank in the screenplay text template to generate the optimized screenplay text.
[0177] According to one or more embodiments of the present disclosure, the split-screen text template further comprises a second placeholder for writing a visual effect identifier representing a visual effect of a picture element; the processing module is further configured to: in response to a first trigger instruction for the second placeholder, display at least two alternative visual effect identifiers corresponding to the second placeholder; and when generating the optimized split-screen text in response to the trigger operation on the element identifier of the first picture element, the processing module is specifically configured to: write the element identifier of the first picture element into the first placeholder in the split-screen text template in response to the trigger operation on the element identifier of the first picture element, and write a target visual effect identifier of the at least two alternative visual effect identifiers into the second placeholder in the split-screen text template in response to a trigger operation on the target visual effect identifier, to generate the optimized split-screen text.
[0178] In a third aspect, according to one or more embodiments of the present disclosure, an electronic device is provided, comprising: at least one processor and a memory;
[0179] The memory stores computer-executable instructions;
[0180] The at least one processor executes the computer-executable instructions stored in the memory, so that the at least one processor performs the video generation method according to the first aspect and various possible designs of the first aspect.
[0181] In a fourth aspect, according to one or more embodiments of the present disclosure, a computer-readable storage medium is provided, and the computer-readable storage medium stores computer-executable instructions, when a processor executes the computer-executable instructions, the video generation method according to the first aspect and various possible designs of the first aspect is implemented.
[0182] In a fifth aspect, according to one or more embodiments of the present disclosure, a computer program product is provided, comprising a computer program, when a processor executes the computer program, the video generation method according to the first aspect and various possible designs of the first aspect is implemented.
[0183] The above description is merely preferred embodiments of the present disclosure and a description of the principles of the technology used. Those skilled in the art should understand that the disclosed scope of the present disclosure is not limited to the technical solutions formed by the specific combinations of the above technical features, and should also cover other technical solutions formed by any combinations of the above technical features or their equivalent features without departing from the disclosed concept. For example, the above features are replaced with the technical features disclosed in the present disclosure (but not limited to) having similar functions to form technical solutions.
[0184] Moreover, while operations are depicted in a particular order, this should not be understood as requiring such an order nor infringing on the scope of the disclosure. Certain of the operations described in the discussion are combinable into a single operation, and certain operations can be separated into several operations. In some embodiments, the operations described in the discussion can be performed in an order different than presented in the discussion. In some embodiments, the operations described in the discussion can be performed concurrently. Also, while several specific implementation details are discussed in the discussion, these should not be interpreted as limiting the scope of the disclosure. Rather, certain features described in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable sub-combination.
[0185] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are disclosed as example forms of implementing the claims.
Claims
1. A method of video generation, the method comprising: The method comprises: obtaining first media content; obtaining first split script data according to the first media content, the first split script data comprising a first split script segment, the first split script segment comprising second media content and first split script text, the first split script text being used for describing at least one picture element in the second media content and shot characteristics of media content generated based on the first split script segment, the second media content comprising at least one media content in the first media content; generating third media content according to the first split script segment and adding the third media content to the first split script segment to obtain second split script data; generating a video based on the second split script data.
2. The method of claim 1, wherein, The obtaining of the first split script data according to the first media content comprises: obtaining at least one second media content according to content characteristics of each media content in the first media content, the content characteristics of the second media content corresponding to a preset content scenario, the preset content scenario representing an image picture composed of picture elements with specific correlation; generating first split script text corresponding to the second media content by processing the second media content through a script generation model; generating the first split script data according to the second media content and the corresponding first split script text.
3. The method of claim 2, wherein, The obtaining of the at least one second media content according to the content characteristics of each media content in the first media content comprises: identifying at least one preset content scenario according to the content characteristics of each media content in the first media content; determining at least one target content scenario from the at least one preset content scenario, the target content scenario being a hot content scenario in a video content platform; determining media content corresponding to the target content scenario as the second media content.
4. The method of claim 2, wherein, The generating of the first split script text corresponding to the second media content by processing the second media content through the script generation model comprises: generating a corresponding first prompt word according to the preset content scenario corresponding to the second media content; inputting the first prompt word and the second media content into the script generation model to generate the first split script text.
5. The method of claim 1, wherein, The method further comprises: obtaining a feature label of a target user, the feature label being used for representing an interest point of the target user; The obtaining of the first split script data according to the first media content comprises: obtaining the first split script data according to the first media content and the feature label of the target user.
6. The method of claim 1, wherein, The first split script data further comprises at least a second split script segment, and the method further comprises: determining a segment order of the first split script segment and the second split script segment according to second media content corresponding to the first split script segment and fourth media content corresponding to the second split script segment; The generating of the third media content according to the first split script segment comprises: generating the third media content according to the first split script segment, the second split script segment, and the segment order.
7. The method of claim 1, wherein, The method further comprises: displaying the first screenplay text and a script score corresponding to the first screenplay text, the script score being used to represent content richness of the media content described by the first screenplay text; in response to a first user instruction, generating an optimized screenplay text corresponding to the first screenplay text, the optimized screenplay text containing a first picture element, the optimized screenplay text being used to represent a camera movement feature based on the first picture element; the generating the third media content according to the first screenplay script segment comprises: generating the third media content according to the optimized screenplay text.
8. The method of claim 7, wherein, the generating the optimized screenplay text corresponding to the first screenplay text in response to the first user instruction comprises: in response to a first user instruction, displaying an element identifier of a picture element in the second media content; in response to a triggering operation on the element identifier, determining the first picture element and constructing the optimized screenplay text based on the first picture element.
9. The method of claim 8, wherein, the determining the first picture element and constructing the optimized screenplay text based on the first picture element in response to the triggering operation on the element identifier comprises: according to a content feature of the second media content, obtaining and displaying a corresponding prompt word template, the prompt word template containing a blank position, the blank position being used to write an identifier indicating a target object, the prompt word template being used to describe a target camera movement feature realized based on at least one of the target objects; in response to a triggering operation on the element identifier of the first picture element, writing the element identifier of the first picture element into the blank position in the prompt word template to generate the optimized screenplay text.
10. The method of claim 8, wherein, the determining the first picture element and constructing the optimized screenplay text based on the first picture element in response to the triggering operation on the element identifier comprises: in response to a second user instruction, selecting target platform media content in a video content platform; obtaining and displaying a screenplay text template corresponding to the target platform media content, the screenplay text template containing a first blank position, the first blank position being used to write an identifier indicating a target object, the screenplay text template being used to describe a picture feature of the platform media content and a target camera movement feature realized based on the target object; in response to a triggering operation on the element identifier of the first picture element, writing the element identifier of the first picture element into the first blank position in the screenplay text template to generate the optimized screenplay text.
11. The method of claim 10, wherein, the screenplay text template further contains a second blank position, the second blank position being used to write a visual effect identifier representing a visual effect of a picture element; the method further comprises: in response to a first triggering instruction on the second blank position, displaying at least two alternative visual effect identifiers corresponding to the second blank position; the writing the element identifier of the first picture element into the first blank position in the screenplay text template to generate the optimized screenplay text in response to the triggering operation on the element identifier of the first picture element comprises: In response to a trigger operation for the element identifier of the first picture element, the element identifier of the first picture element is written into a first blank in the shot list template, and in response to a trigger operation for a target visual effect identifier in the at least two alternative visual effect identifiers, the target visual effect identifier is written into a second blank in the shot list template, to generate the optimized shot list.
12. A video generating apparatus characterized by comprising: Comprising: An acquisition module, configured to acquire first media content, and acquire first shot list script data according to the first media content, the first shot list script data comprising a first shot list script segment, the first shot list script segment comprising second media content and first shot list text, the first shot list text being used to describe at least one picture element in the second media content and shot characteristics of media content generated based on the first shot list script segment, the second media content comprising at least one media content in the first media content; A processing module, configured to generate third media content according to the first shot list script segment, and add the third media content to the first shot list script segment to obtain second shot list script data; A generation module, configured to generate a video based on the second shot list script data.
13. An electronic device, comprising: Comprising: A processor and a memory; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory, so that the processor executes the video generation method in any one of claims 1 to 11.
14. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, and when the processor executes the computer-executable instructions, the video generation method in any one of claims 1 to 11 is implemented.
15. A computer program product comprising a computer program, characterized in that, The computer program is executed by the processor to implement the video generation method in any one of claims 1 to 11.
Citation Information
Patent Citations
Video generation method, video generation device, electronic equipment and readable storage medium
CN118509618A
Video processing method, computing device, computer storage medium and computer program product
CN118972671A