Video generation method, display method, device, storage medium and program product

By working together between the server and the client, the image generation model is used to generate videos based on storyboard text clips and shooting parameters, the problems of low processing efficiency and long production cycle of traditional video generation process are solved, and efficient and high-quality video generation is achieved.

CN120017923APending Publication Date: 2025-05-16RAJAX NETWORK &TECHNOLOGY (SHANGHAI) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510053908.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-13
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The traditional video generation process is inefficient in processing and long production cycle, making it difficult to meet the market's demand for efficient video production.

Method used

A video generation method is proposed, through the collaborative work of the server and the client, the image generation model is used to generate images based on storyboard text clips and shooting parameters, and video is generated based on these images.

Benefits of technology

It improves the processing efficiency of video generation, shortens the video production cycle, ensures that the generated video is consistent with the storyboard text in semantic expression, and the visual effect is in line with expectations, improving the video quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120017923A_ABST
    Figure CN120017923A_ABST
Patent Text Reader

Abstract

The invention discloses a video generation method, a display method, a device, a storage medium and a program product, the video generation method comprises the following steps: receiving a first image generation instruction of a client, the first image generation instruction comprising a split text fragment obtained from a split text in advance and a first shooting parameter; performing image generation processing on the split text fragment and the first shooting parameter by using an image generation model to obtain at least one first image; and generating a corresponding video based on each first image. According to the method, the processing efficiency of the generated video can be improved, and the video production period can be shortened.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a video generation method, display method, device, storage medium and program product. Background Art

[0002] Video is an important way to convey information. It can combine visual and auditory elements to attract the audience's attention and convey brand or product information. It plays an important role in maintaining user attention and enhancing brand loyalty.

[0003] The traditional video generation process requires a series of work processes including market research, script writing, actual shooting, post-shooting, review and modification, publishing and distribution. The processing efficiency of video generation is low and the production cycle is long. Summary of the invention

[0004] In response to the problems of low processing efficiency and long production cycle of traditional video generation processes in the above-mentioned related technologies, the present application proposes a video generation method, display method, device, storage medium and program product to improve the processing efficiency of generated videos and shorten the video production cycle.

[0005] The first aspect of the present application proposes a video generation method, which is applied to a server, and the method includes: receiving a first image generation instruction from a client, wherein the first image generation instruction includes a storyboard text fragment and a first shooting parameter obtained in advance from a storyboard text; using an image generation model, performing image generation processing on the storyboard text fragment and the first shooting parameter to obtain at least one first image; and generating a corresponding video based on each first image.

[0006] The second aspect of the present application proposes a video generation method, which is applied to a client, and the method includes: receiving a storyboard text fragment and a first shooting parameter input by a user through a first information input interface; generating a first image generation instruction including the storyboard text fragment and the first shooting parameter; sending the first image generation instruction to a server, so that the server generates a corresponding video based on the storyboard text fragment and the first shooting parameter; and receiving the corresponding video sent by the server.

[0007] The third aspect of the present application proposes a video display method, which includes: displaying an operation interface of a predetermined application, wherein the operation interface includes an interactive entry element of a member page; in response to an operation instruction for the interactive entry element, opening the member page and playing a predetermined video, wherein the predetermined video is a video generated according to the method of the first aspect.

[0008] The fourth aspect of the present application proposes a video generation system, which includes a client and a server; the client is used to receive a storyboard text fragment and a first shooting parameter in a storyboard text, generate a first image generation instruction according to the storyboard text fragment and the first shooting parameter, and send the first image generation instruction to the server; the server is used to obtain the storyboard text fragment and the first shooting parameter from the first image generation instruction, perform image generation processing on the storyboard text fragment and the first shooting parameter using an image generation model, obtain at least one first image, and generate a corresponding video based on each first image.

[0009] The fifth aspect of the present application proposes an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the method described in the first aspect, the second aspect, or the third aspect above.

[0010] The sixth aspect of the present application proposes a computer-readable storage medium having a computer program stored thereon, wherein the program is executed by a processor to implement the method described in the first aspect, the second aspect or the third aspect above.

[0011] The embodiment of the seventh aspect of the present application provides a computer program product, including a computer program, and the computer program is executed by a processor to implement the method described in the first aspect, the second aspect or the third aspect above.

[0012] Based on the method of the first aspect above, the present application has at least the following beneficial effects or advantages:

[0013] According to the method of the embodiment of the present application, the server can obtain the storyboard text fragment used to generate the first image from the first image generation instruction from the client, use the image generation model to process the storyboard text fragment and the first shooting parameter, generate the first image, and then generate the corresponding video based on each first image. In this method, the first image is generated by using the storyboard text fragment obtained by screening from the storyboard text information, so that the first image and the storyboard text fragment are consistent in semantic expression, and the interference of irrelevant text information outside the text fragment on the generated image is avoided. Compared with the amount of information contained in the storyboard text, the amount of information contained in the text fragment is less, which is conducive to improving processing efficiency. In addition, the generation of the first image needs to be based not only on the text fragment, but also on the first shooting parameter, so that the first image can have a visual effect corresponding to the first shooting parameter, so that when the corresponding video is generated based on each first image, it is helpful to generate the video to achieve the expected visual appeal, thereby improving the video quality. Through the automated video generation processing flow in the example of the present application, it is conducive to improving the processing efficiency of the generated video and shortening the video production cycle.

[0014] The above description is only an overview of the technical solution of the present application. In order to more clearly understand the technical means of the present application, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the specific implementation methods of the present application are listed below. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0016] Figure 1 A schematic diagram of the architecture of a video generation system according to an embodiment of the present application;

[0017] Figure 2 is a flowchart of a video generation method according to an embodiment of the present application;

[0018] Figure 3 A schematic diagram showing the generation of storyboard text according to an exemplary embodiment of the present application;

[0019] Figure 4a exemplarily showing each first image corresponding to the first text segment;

[0020] Figure 4b exemplarily showing each first image corresponding to the second text segment;

[0021] Figure 4c Schematic diagrams showing respective second images of exemplary embodiments of the present application;

[0022] Figure 5 A schematic diagram showing an exemplary embodiment of the present application for performing image editing on a target storyboard image;

[0023] Figure 6a is a schematic diagram of a processing process of generating a first storyboard video according to a first storyboard image and corresponding motion parameters according to an exemplary embodiment of the present application;

[0024] Figure 6b is a schematic diagram of a processing process of generating a second storyboard video according to a second storyboard image and corresponding motion parameters according to an exemplary embodiment of the present application;

[0025] Figure 6c is a schematic diagram of a processing process of generating a third storyboard video according to a third storyboard image and corresponding motion parameters according to an exemplary embodiment of the present application;

[0026] Figure 7 A schematic diagram showing an exemplary video generation result of the present application is shown;

[0027] Figure 8A flowchart showing a video generation method of an exemplary embodiment of the present application;

[0028] Fig. 9 A flowchart showing a video generation method according to another embodiment of the present application;

[0029] Fig.10 A flowchart of a video display method provided in an embodiment of the present application;

[0030] Fig.11 It is a structural schematic diagram of a video generating device according to an embodiment of the present application;

[0031] Fig.12 It is a structural schematic diagram of a video generating device according to an embodiment of the present application;

[0032] Fig.13 It is a structural schematic diagram of a video display device according to an embodiment of the present application;

[0033] Fig.14 This is a schematic diagram of the hardware structure of an electronic device according to an exemplary embodiment of the present application;

[0034] Fig.15 The figure is a schematic diagram of the structure of a storage medium according to an exemplary embodiment of the present application. DETAILED DESCRIPTION

[0035] Exemplary embodiments will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The implementations described in the following exemplary embodiments do not represent all implementations consistent with the present application. Instead, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.

[0036] The terms used in this application are only for the purpose of describing specific embodiments and are not intended to limit this application. The singular forms of "a", "a", "an" and "the" used in this application and the appended claims are also intended to include plural forms unless the context clearly indicates other meanings. It should also be understood that the term "and / or" used in this article refers to and includes any or all possible combinations of one or more associated listed items.

[0037] It should be understood that, although the terms first, second, third, etc. may be used in the present application to describe various information, these information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determination", etc.

[0038] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0039] In the related art, in the process of organizing and promoting marketing activities, it is often necessary to produce various creative videos or advertising videos to demonstrate the effects and advantages of products. The traditional workflow of video production may include a series of processes such as market research, script writing, actual shooting, post-shooting, review and modification, publishing and distribution. The production cycle of video generation is long and inefficient.

[0040] Based on this, an embodiment of the present application provides a video generation method for improving the processing efficiency of video generation and shortening the video production cycle.

[0041] The technical solution of the present application and how the technical solution of the present application solves the above-mentioned technical problems are described in detail below with specific embodiments. The several specific embodiments listed can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described in detail below with reference to the accompanying drawings.

[0042] Figure 1 FIG. 1 is a schematic diagram of the architecture of the video generation system according to an embodiment of the present application. Figure 1 As shown, the architecture may include: a client 10 , a server 20 and a network 30 .

[0043] Among them, the client 10 can be various electronic devices with a display screen, including but not limited to user equipment (UE), mobile devices, user terminals, laptops, desktop computers, personal digital assistants (PDA), handheld devices, vehicle-mounted devices, wearable devices, etc.

[0044] The server 20 may include an independent physical server, a server cluster consisting of a plurality of servers, or a cloud server capable of cloud computing.

[0045] The network 30 is used to provide a medium for a communication link between the client 10 and the server 20. Specifically, the network 30 may include various connection types, such as wired or wireless communication links or optical fiber cables.

[0046] It should be understood that Figure 1 The number of devices is only for illustration. It can be adjusted flexibly according to the actual application needs. In addition, this architecture can also include some auxiliary devices, such as routers, etc. It can be flexibly configured according to the needs, and there is no restriction on this aspect.

[0047] Figure 2 A flowchart of a video generation method provided in an embodiment of the present application, such as Figure 2 As shown, the method specifically includes the following steps S210-S230.

[0048] S210, receiving a first image generation instruction from a client, where the first image generation instruction includes a storyboard text segment obtained in advance from the storyboard text and a first shooting parameter.

[0049] S220: Using an image generation model, perform image generation processing on the storyboard text segment and the first shooting parameter to obtain at least one first image.

[0050] S230: Generate a corresponding video based on each first image.

[0051] According to the method of the embodiment of the present application, the server can obtain the storyboard text fragment used to generate the first image from the first image generation instruction from the client, use the image generation model to process the storyboard text fragment and the first shooting parameter, generate the first image, and then generate the corresponding video based on each first image. In this method, the first image is generated by using the storyboard text fragment obtained by screening from the storyboard text information, so that the first image and the storyboard text fragment are consistent in semantic expression, and the interference of irrelevant text information outside the text fragment on the generated image is avoided. Compared with the amount of information contained in the storyboard text, the amount of information contained in the text fragment is less, which is conducive to improving processing efficiency. In addition, the generation of the first image needs to be based not only on the text fragment, but also on the first shooting parameter, so that the first image can have a visual effect corresponding to the first shooting parameter, so that when the corresponding video is generated based on each first image, it helps the generated video to achieve the expected visual appeal, thereby improving the video quality. Through the automated video generation processing flow in the example of the present application, it is conducive to improving the processing efficiency of the generated video and shortening the video production cycle.

[0052] In the description of the embodiments of the present application, the video may include various types of videos. For example, it may be an advertising video, a product introduction video, a brand promotion video, a tutorial video, etc. In order to simplify the description, the following multiple embodiments of this document take the advertising video as an example to describe the processing flow of the video generation method. However, this description cannot be interpreted as limiting the scope or implementation possibility of this solution. The generation method of other types of videos other than the advertising video is consistent with the processing flow of the advertising video generation method.

[0053] In step S210, the storyboard text can be understood as a visual script. Taking the advertising storyboard text as an example, the advertising storyboard text can be used to plan the shooting and production of the advertisement, and is an important link in the advertising production process. The storyboard text is text information that needs to be conveyed to the target user through the storyboard. For example, the storyboard text can describe at least one of the following information: scene description, role, action, behavior, dialogue, sound effect, visual effect, camera angle, motion trajectory and duration of each storyboard, etc. The source of the storyboard text can include at least one of the following: independent creation, model generation, customer submission and network download.

[0054] In some embodiments, before receiving the first input in step S210, the method includes: receiving a text generation instruction from the client, wherein the text generation instruction includes storyboard description information; and generating storyboard text based on the storyboard description information in response to the text generation instruction.

[0055] Exemplarily, the storyboard description information may include one or more semantic units, where a semantic unit refers to at least one of a word and a sentence. Exemplarily, the storyboard description information may be processed in any of the following ways to generate the storyboard text, such as a model-based generation method and a knowledge base query-based method.

[0056] For example, in a model-based generation method, the storyboard description information can be input into a pre-trained text generation model for text generation processing to obtain the storyboard text output by the text generation model.

[0057] In the embodiment of the present application, the text generation model is used to characterize: the correspondence between the input storyboard description information and the output storyboard text. The text generation model can be a language model obtained by pre-training the model based on a large number of text samples. The language model can learn the language patterns in the text samples in advance, so as to generate a coherent and grammatically correct text based on the input description information.

[0058] For another example, in a method based on knowledge base query, text keywords can be extracted from the storyboard description information, and storyboard case text corresponding to the text keywords can be searched in the knowledge base, and the searched storyboard case text can be used as the storyboard text. The knowledge base includes but is not limited to an online resource library or a pre-established storyboard script library, etc., which is not specifically limited in the embodiments of the present application.

[0059] In this embodiment, the storyboard text is generated based on the storyboard description information, and the storyboard text can be automatically generated, which reduces the time of manual writing and improves the processing efficiency. Figure 3 , describing the specific process of generating storyboard text.

[0060] Figure 3 A schematic diagram showing the generation of storyboard text according to an exemplary embodiment of the present application.

[0061] In some embodiments, the server receives a text generation instruction from the client and obtains the storyboard description information in the text generation instruction, such as "Help me write a 3-second cookie video advertisement storyboard". In response to the text generation instruction, the server can use the text generation model to process the description information to obtain the advertisement storyboard text.

[0062] In some embodiments, the server may display the generated storyboard text. Figure 3 , the text shown in the area below the text input box is an example of a storyboard text generated based on the above-mentioned storyboard description information.

[0063] In some scenarios, when a user accesses a server through a client through a remote desktop, the user can intuitively see the complete storyboard text through the server's display. The user can directly extract the required storyboard text fragments from the displayed storyboard text, thereby reducing additional searches.

[0064] Indicatively, Figure 3 The box in shows the storyboard text segment required by the user in the storyboard text. The storyboard text segment can be a word, a sentence or a paragraph, which is not specifically limited in the embodiment of the present application.

[0065] As an example, the number of storyboard texts generated in the embodiment of the present application is greater than or equal to 1, that is, at least one storyboard text can be generated based on the storyboard description information.

[0066] In some embodiments, the server may also receive a new text generation instruction from the client, and obtain new storyboard description information contained in the new text generation instruction, such as "rewrite a 3-second cookie video advertisement storyboard for me, and add sound effects." In response to the new text generation instruction, the server uses the text generation model to process the new storyboard description information and obtain a new advertisement storyboard text.

[0067] In some scenarios, after sending the text generation instruction to the server, the client may display a first information input interface, where the first information input interface is used to receive the storyboard text fragment and the first shooting parameter input by the user. The client generates a first image generation instruction based on the storyboard text fragment and the first shooting parameter, and sends the first image generation instruction to the server.

[0068] exist Figure 3 In the example, the boxes shown are only for schematically illustrating the storyboard text segments in the storyboard text. In actual scenes, it is not necessary to frame the text segments.

[0069] In some embodiments, the server may receive multiple text generation instructions from the client, and the storyboard description information in multiple text generation instructions may be the same or different. By responding to each text generation instruction, multiple storyboard texts may be obtained, providing more options for selecting storyboard text segments.

[0070] In step S210, the first shooting parameter is used to indicate the visual effect of at least one first image. As an example, the first shooting parameter includes but is not limited to at least one of the following: shutter speed, aperture, focal length, aspect ratio, resolution, exposure compensation, etc.

[0071] In step S220, the storyboard text segment and the first shooting parameter may be input into the image generation model, and the image generation model performs image generation processing on the input storyboard text segment and the first shooting parameter, and outputs at least one first image. The image generation model is used to characterize the correspondence between the input text segment and the shooting parameter and the output image. The image generation model may be a model pre-trained based on a large number of text samples and shooting parameter samples.

[0072] As an example, assume that the storyboard text segments include a first text segment, a second text segment, and a third text segment. Figure 4a exemplarily showing each first image corresponding to the first text segment; Figure 4b The first images corresponding to the second text segments are shown as examples.

[0073] exist Figure 4a , a plurality of first images generated according to the first text segment "a golden dragon" and the corresponding first shooting parameters are shown. Figure 4b , a plurality of first images generated according to the second text segment “cookies sticky drinks” and the corresponding first shooting parameters are shown.

[0074] In step S230, there are many ways to generate a corresponding video based on the image. For example, a video editing tool can be called to automatically add animation effects to the image, such as automatically scaling, rotating and translating the image, adding fade-in and fade-out animation effects to the image, etc., to obtain a video corresponding to the image. For another example, a video generation model can be used to perform video generation processing on the input image and motion parameters to obtain a storyboard video corresponding to the image.

[0075] In the embodiment of the present application, the video generation model is used to characterize the correspondence between each input image and the output video. The video generation model can be a model obtained by pre-training with a large number of image samples using a machine learning algorithm.

[0076] The text generation model, image generation model and video generation model in the above embodiments of the present application can all be multimodal models, that is, they can process text, images and videos at the same time. The embodiments of the present application do not limit the specific implementation form of the corresponding models.

[0077] In some embodiments, step S230 may specifically include: determining a set of candidate storyboard images based on each first image; receiving an image selection instruction from the client, the image selection instruction being used to indicate at least one target storyboard image in the set of candidate storyboard images; and generating a corresponding video based on each target storyboard image.

[0078] In this embodiment, the candidate storyboard image set includes all first images. According to the image selection instruction, the target storyboard image is screened out from each first image, and the corresponding video is generated according to the target storyboard image. In this embodiment, the storyboard image required for generating the video can be flexibly selected by selecting the storyboard image, which is conducive to improving the quality of the generated video, so that the subsequent generation of the video is more in line with the user's needs.

[0079] In some embodiments, before receiving the image selection instruction from the client, it also includes: receiving a second image generation instruction from the client, the second image generation instruction including detail description information and description information of the second shooting parameters, the detail description information being used to describe detail features of storyboard image constituent elements in natural language; using an image generation model, performing image generation processing on the detail description information and the description information of the second shooting parameters to obtain at least one second image; and using each second image increment to update the set of alternative storyboard images.

[0080] Exemplarily, the detail description information is used to indicate the detail features of the image elements contained in the image. The image elements are components that constitute the visual picture of the image, and may include but are not limited to objects, icons, text, lines, backgrounds, etc. in the image.

[0081] Figure 4cSchematic diagrams showing various second images of exemplary embodiments of the present application. In some embodiments, assuming that the information in the second image generation instruction is "Chinese New Year, red and festive, family sitting around the dining table, enjoying a hot New Year's Eve dinner, top view shooting, shutter speed 1 / 250 second, aspect ratio of the picture is 16:9", the second image generation instruction includes description information of the image details of the second image required by the user, and also includes description information of the second shooting parameters, which are used to indicate the visual effect that the user wants the second image to present.

[0082] In this embodiment, the image generation model is used to perform image generation processing on the above-mentioned detail description information and the description information of the second shooting parameter to obtain the following Figure 4c At least one second image is shown.

[0083] In the embodiment of the present application, when the set of candidate storyboard images is updated by using each second image increment, the set of candidate storyboard images includes each first image and each second image. The updated set of candidate storyboard images expands the selection range of the target storyboard images, provides richer image materials for subsequent video generation, is conducive to further improving the quality of the generated video, and is conducive to the subsequent generation of videos that better meet customer needs.

[0084] In some embodiments, before receiving the image selection instruction in the above embodiments, it also includes: receiving at least one third image submitted by a client from the client; and incrementally updating the set of candidate storyboard images using each third image.

[0085] For example, the third image is an image submitted by the client. When the client is an advertiser, the third image can be an image that the advertiser requires to be used in the advertisement generation process. For example, the opening and ending of an advertisement video are two very important parts of an advertisement, and need to contain at least one of the product key visual (KV) information, product logo, and daily banner. The images used in this part of the video can be submitted by the client, which helps to ensure that the final generated video meets the client's needs.

[0086] Exemplarily, when the set of alternative storyboard images is updated by incrementally using each third image, each third image is added to the set of alternative storyboard images on the basis of existing images, further expanding the selection range of the target storyboard images, providing richer image materials for subsequent video generation, which is conducive to further improving the video quality and facilitating the subsequent generation of videos that better meet customer needs.

[0087] In some embodiments, the above-mentioned step of generating a corresponding video according to each target storyboard image may specifically include: performing image editing on each target storyboard image to obtain each post-processed storyboard image; and generating a corresponding video according to each post-processed storyboard image.

[0088] Exemplarily, image editing refers to adjusting and modifying an image. In this embodiment, image editing can improve image quality, ensure the effect and quality of the generated video, reduce modifications and rework in post-production, and thus speed up the processing flow of video generation, save costs and improve processing efficiency.

[0089] In some embodiments, image editing of each target storyboard image may include: performing image editing on each target storyboard image according to image editing operations preset by the server; or, performing image editing on each target storyboard image according to an external image editing application called by the server; or, after performing image editing on each target storyboard image to obtain each first edited image according to image editing operations preset by the server, image editing on each first edited image again according to an external image editing application called by the server.

[0090] In this embodiment, the server-side preset image editing operation and the server-side calling external image editing application are two different server-side editing methods for the target storyboard image. The former can pre-develop the corresponding editing function according to the specific editing needs, which is highly flexible and does not require the image to be sent to an external application, and has high data security; the latter can use professional image processing services to provide more professional editing functions, which is easy to develop and easy to use. Users can choose one of the editing methods according to their needs, or combine the two to meet various image editing needs.

[0091] Exemplarily, the image editing operations preset by the server include but are not limited to at least one of the following: image expansion, exposure adjustment, and white balance. Among them, image expansion can be understood as: expanding the peripheral scene of the image by generating additional image pixels. For example, when the constituent elements of the image are shot with a longer focal length and the focal length is expanded to 1.5 times the current focal length, more scene space can be captured. By expanding the image, a layer of scene can be generated outside the scene in the image to maintain the naturalness and continuity of the image. Among them, exposure adjustment refers to changing the brightness level of the image to correct an image that is too dark or too bright. White balance refers to adjusting the colors in the image.

[0092] Exemplarily, the image editing performed by the server calling the external image editing application includes but is not limited to at least one of the following: color adjustment, image extraction, repairing defects in the image or removing unnecessary elements, etc. Image extraction refers to selecting a specific area from the image to remove the background outside the area.

[0093] In the embodiment of the present application, the target storyboard image can be edited by using the image editing operation preset by the server according to actual needs; the target storyboard image can also be edited by calling an external image editing application by the server; or the two can be combined to perform image editing. This is conducive to users choosing flexible editing methods according to their needs and improving image editing efficiency.

[0094] Figure 5 A schematic diagram showing image editing performed on a target storyboard image according to an exemplary embodiment of the present application.

[0095] In some scenarios, it is assumed that the server receives a second image generation instruction, and the detail description information and the description information of the second shooting parameters contained in the second image generation instruction are "golden scales on the body of a Chinese dragon, close-up, shimmering waves, blurred focus, shutter speed 1 / 250 second, and aspect ratio of the picture is 16:9". The image generation model is used to perform image generation processing on the detail description information and the description information of the second shooting parameters to obtain at least one second image; each second image is added to the candidate storyboard set. Select the target storyboard image from the backup storyboard set, that is, Figure 5 “Storyboard image 1” shown in .

[0096] like Figure 5 As shown, after performing image expansion processing (image editing operation preset by the server) and color adjustment (image editing performed by the server calling an external image editing application) on "storyboard image 1", a post-processed storyboard image is obtained, for example Figure 5 "Post-processed storyboard image 1" in.

[0097] In some scenarios, it is assumed that the server receives a second image generation instruction, and the detailed description information and the description information of the second shooting parameters contained in the second image generation instruction are "waves hitting the lens, leaving splashes, underwater camera, close-up, shooting at eye level, shutter speed 1 / 250 second, and aspect ratio of the picture is 16:9". The image generation model is used to perform image generation processing on the detailed description information and the description information of the second shooting parameters to obtain at least one second image; each second image is added to the candidate storyboard set. Select the target storyboard image from the backup storyboard set, such as Figure 5 "Storyboard Image 2" in the video.

[0098] like Figure 5 As shown, after color adjustment is performed on "storyboard image 2" (the server calls an external image editing application to perform image editing), a post-processed storyboard image is obtained, that is, Figure 5 “Post-processed storyboard image 2” shown in FIG.

[0099] In some scenarios, it is assumed that the server receives a second image generation instruction, and the detailed description information and the description information of the second shooting parameters contained in the second image generation instruction are "the moment when the cookie falls into the drink, top view shooting, light and shadow effects in the video, blurred focus, close-up shot, shutter speed 1 / 250 second, and aspect ratio of the picture is 16:9", and the image generation model is used to perform image generation processing on the detailed description information and the description information of the second shooting parameters to obtain at least one second image; each second image is added to the candidate storyboard set. Select the target storyboard image from the backup storyboard set, such as Figure 5 "Storyboard Image 3" in the video.

[0100] like Figure 5 As shown, after performing image extraction and color adjustment on "storyboard image 3" (both are image editing performed by the server calling an external image editing application), a post-processed storyboard image is obtained, that is, Figure 5 “Post-processed storyboard image 3” shown in FIG.

[0101] It should be understood that the image editing operations preset by the server and the image editing operations performed by the external image editing application called by the server are merely illustrative. In actual application scenarios, more types of editing operations may be included, and the embodiments of the present application do not make specific limitations.

[0102] In some embodiments, generating corresponding videos according to each post-processed storyboard image may specifically include: receiving a video generation instruction from a client, the video generation instruction is used to indicate at least one specified storyboard image among each post-processed storyboard image and motion parameters corresponding to each specified storyboard image, the motion parameters are used to indicate motion characteristics of elements contained in the corresponding specified storyboard image; performing video generation processing on each specified storyboard image and the corresponding motion parameters to obtain storyboard videos corresponding to each specified storyboard image; generating corresponding videos according to the storyboard videos corresponding to each specified storyboard image.

[0103] Exemplarily, the specified storyboard images and motion parameters can be processed for video generation in a variety of ways. For example, the server can process the specified storyboard images and motion parameters according to a pre-trained video generation model to obtain a corresponding storyboard video. An external video generation application can also be called to automatically generate a storyboard video according to the input storyboard images and motion parameters.

[0104] Exemplarily, the motion parameter is used to control the dynamic effect of the image-generated video. For example, the motion parameter includes but is not limited to at least one of the following parameter items: a lens motion parameter and a motion trajectory parameter.

[0105] Exemplarily, lens motion parameters refer to a series of parameters used to describe and control how the camera perspective moves in video production, including but not limited to horizontal parameters, vertical parameters, proximity parameters, and ambient parameters.

[0106] Among them, the horizontal motion parameter is used to describe the amount of movement of the object in the horizontal direction. The vertical motion parameter is used to describe the amount of movement of the object in the vertical direction. The proximity parameter is used to adjust the proximity of the object to increase dynamic effects, such as lens push-in, lens pull-out, etc. The environmental parameter is used to create and enhance elements related to the scene atmosphere in the video. For example, add environmental sound effects (the sound of rolling waves, crowd noise, etc.).

[0107] Exemplarily, the motion trajectory parameter is used to indicate the path of the image elements from the starting point to the end point in the generated video. In the video production process, the motion trajectory is used to define and control the movement of the image elements in the video, which is conducive to building a smooth visual effect.

[0108] exist Figure 6a-6c In the figure, the following motion parameters are schematically shown: horizontal motion parameter, vertical motion parameter, proximity parameter, and environmental parameter. It should be understood that motion parameters may also include more other types. For example, rotation parameters, speed parameters, etc. Rotation parameters are used to describe the rotation of an object around a certain axis or center point in a three-dimensional space or a two-dimensional plane. Speed ​​parameters are used to describe how fast an object in an image moves in a video. The specific settings can be customized according to actual needs, and the embodiments of the present application do not make specific limitations.

[0109] For example, as described in the above embodiments, there are multiple implementations of generating corresponding videos based on images. A video editing tool can be called to automatically add animation effects to a specified storyboard image, such as automatically scaling, rotating, and translating the specified storyboard image, adding fade-in and fade-out animation effects to the specified storyboard image, etc., to obtain a video corresponding to the image. For another example, a video generation model can be used to perform video generation processing on the input specified storyboard image and the corresponding motion parameters to obtain a storyboard video corresponding to the image.

[0110] In the embodiment of the present application, the server can perform video generation processing on the specified storyboard image and motion parameters in response to the video generation instruction of the client, and obtain the storyboard video corresponding to the specified storyboard image. In this method, the constituent elements of the storyboard image can be accurately dynamically controlled by the motion parameters. After receiving the video generation instruction, the server can automatically perform video generation processing on the specified storyboard image and motion parameters, thereby improving the video generation efficiency.

[0111] Figure 6a is a schematic diagram of a processing process of generating a first storyboard video according to a first storyboard image and corresponding motion parameters according to an exemplary embodiment of the present application; Figure 6b is a schematic diagram of a processing process of generating a second storyboard video according to a second storyboard image and corresponding motion parameters according to an exemplary embodiment of the present application; Figure 6c 1 is a schematic diagram of a processing process for generating a third storyboard video according to a third storyboard image and corresponding motion parameters of an exemplary embodiment of the present application, wherein the first storyboard image, the second storyboard image and the third storyboard image are only different references to different storyboard images.

[0112] In some embodiments, the server may receive a video generation instruction from the client, and obtain at least one designated storyboard image and motion parameters corresponding to each designated storyboard image from the video generation instruction.

[0113] like Figure 6a As shown, the server can provide a video production operation interface, which is used to display the specified storyboard image and motion parameters. The video production operation interface also provides interactive elements for editing the input motion parameters. Through the operation instructions of these interactive elements, the motion parameters are adjusted and the adjusted motion parameters are displayed. In response to the video generation instruction, the video generation model is used to perform video generation processing on the specified storyboard image and the currently displayed motion parameters to obtain the storyboard video corresponding to the specified storyboard image.

[0114] exist Figure 6b and 6c In the process, the server can receive a new video generation instruction from the client, obtain a new designated storyboard image and corresponding motion parameters from the new video generation instruction, perform video generation processing on the new designated storyboard image and the corresponding motion parameters, and obtain a storyboard video corresponding to the designated storyboard image.

[0115] exist Figure 6a-6c , the following motion parameters are schematically shown: horizontal motion parameter, vertical motion parameter, proximity parameter and environment parameter. In practical applications, the types of motion parameters can be customized according to actual conditions, and are not specifically limited here.

[0116] In some embodiments, generating corresponding videos according to the storyboard videos corresponding to each designated storyboard image includes: receiving a video processing instruction from a client, the video processing instruction being used to indicate at least one designated storyboard video among the storyboard videos; performing post-processing on each designated storyboard video to obtain each post-processed storyboard video; and generating corresponding videos according to each post-processed storyboard video.

[0117] Exemplarily, the server may respond to the received video processing instruction by first performing post-processing on each designated storyboard video, and then integrating the post-processed storyboard videos to obtain a corresponding video.

[0118] Exemplarily, the video processing instruction may include a storyboard sequence, which is used to indicate the display order of each storyboard video in the video production process. In response to the video processing instruction, each storyboard video may be integrated according to the storyboard sequence to obtain a corresponding complete video.

[0119] Exemplarily, the types of video processing include not only storyboard video integration, but also many other types, including but not limited to at least one of adding video special effects and video color correction. This embodiment of the present application does not make specific limitations.

[0120] Exemplarily, the server may perform post-processing on the storyboard video in a variety of ways. For example, the storyboard video may be post-processed according to a video post-processing operation preset by the server; the preset video post-processing operation may include but is not limited to: at least one of video compression, video forwarding, and resolution adjustment; for another example, the target storyboard image may be automatically edited according to an external video editing application called by the server.

[0121] In this embodiment, post-processing the storyboard video is beneficial to improving the visual effect of the generated video and improving the video quality.

[0122] Figure 7 A schematic diagram showing an exemplary video generation result of the present application is shown.

[0123] exist Figure 7 In the final generated video, there are five storyboard videos, such as “Storyboard 1, golden dragon goes out to sea”, “Storyboard 2, sea water floods the lens” and “Storyboard 3, biscuits appear”. These three storyboard videos are obtained by using the video generation model to process different target storyboard images and corresponding motion parameters in sequence.

[0124] exist Figure 7 As shown in "Storyboard 4, revealing the main visual information" and "Storyboard 5, folding into a daily banner", these two storyboard videos can be storyboard videos that can be added to the end of the corresponding video based on the main visual information and daily banner submitted by the customer during the post-processing of each storyboard video.

[0125] According to the video generation method of the embodiment of the present application, the server obtains the storyboard text fragment used to generate the first image from the first image generation instruction from the client, uses the image generation model to process the storyboard text fragment and the first shooting parameter, generates the first image, and then generates the corresponding video based on each first image. In this method, the first image is generated by using the text fragment obtained by screening from the storyboard text information, so that the first image and the storyboard text fragment are consistent in semantic expression, and the interference of irrelevant text information outside the text fragment on the generated image is avoided. Compared with the amount of information contained in the storyboard text, the amount of information contained in the text fragment is less, which is conducive to improving processing efficiency. In addition, the generation of the first image needs to be based not only on the text fragment, but also on the first shooting parameter, so that the first image can have a visual effect corresponding to the first shooting parameter, so that when the corresponding video is generated based on each first image, it is helpful to generate the video to achieve the expected visual appeal, thereby improving the video quality. Through the automated video generation processing flow in the example of the present application, it is conducive to improving the processing efficiency of the generated video and shortening the video production cycle.

[0126] Figure 8 A flow chart showing a video generation method of an exemplary embodiment of the present application is shown. The video generation method is applied to the server. Figure 8 As shown, the video generation method includes the following steps.

[0127] S801, write storyboard script.

[0128] In this step, the server receives a text generation instruction from the client, the text generation instruction including storyboard description information. The server generates a storyboard text based on the storyboard description information in response to the text generation instruction.

[0129] S802: Generate a static image.

[0130] In this step, the server receives a first image generation instruction from the client, the first image generation instruction including a storyboard text segment and a first shooting parameter; in response to the first image generation instruction, the server uses an image generation model to perform image generation processing on the storyboard text segment and the first shooting parameter to obtain at least one first image. Each first image is a static image.

[0131] S803, post-processing the static image.

[0132] In this step, the target storyboard image is edited on the server side to obtain a post-processed storyboard image.

[0133] S804, converting static image to video.

[0134] In this step, the server receives the video generation instruction from the client, the video generation instruction is used to indicate each storyboard image selected for post-processing and the motion parameters corresponding to each storyboard image, and performs video generation processing on each storyboard image and the motion parameters to obtain the storyboard video corresponding to each storyboard image.

[0135] S805, post-processing video.

[0136] In this step, the server receives the video processing instruction from the client, and in response to the video processing instruction, performs post-processing on each storyboard video to obtain each post-storyboard video; and generates a corresponding video according to each post-storyboard video.

[0137] The video generation method according to the example of the present application is beneficial to improving the processing efficiency of video generation and shortening the video production cycle.

[0138] Fig. 9 A flow chart of a video generation method according to another embodiment of the present application is shown. The method is applied to a client, and in some embodiments, the method includes the following steps.

[0139] S910, receiving a storyboard text segment and a first shooting parameter input by a user through a first information input interface.

[0140] S920, generating a first image generation instruction including a storyboard text segment and a first shooting parameter.

[0141] S930, sending a first image generation instruction to the server, so that the server generates a corresponding video based on the storyboard text segment and the first shooting parameter.

[0142] S940: Receive the corresponding video sent by the server.

[0143] According to the video generation method of the embodiment of the present application, the client can generate a first image generation instruction based on the storyboard text and shooting parameters received from the user, and then send the first image generation instruction to the server, the storyboard text fragment and the first shooting parameter are used to generate at least one first image on the server, and each generated first image can be used to generate a corresponding video. The automated image generation instruction is conducive to accurately transmitting the storyboard text fragment and the first shooting parameter input by the user to the server, so as to generate a corresponding video based on the storyboard text fragment and the first shooting parameter on the server. The storyboard text and shooting parameters input by the user are sent to the server for processing in the form of instructions, which is conducive to saving the use of local computing resources on the client and improving the processing efficiency of video generation.

[0144] In some embodiments, before step S910, the method further includes: receiving storyboard description information input by a user through a second information input interface; sending a text generation instruction to a server, wherein the text generation instruction includes the storyboard description information; the text generation instruction is used to instruct the server to generate a storyboard text according to the storyboard description information.

[0145] In this embodiment, the client can receive the storyboard description information input by the user, generate and send a text generation instruction to the server, so as to generate the storyboard text on the server, and provide optional text materials for the storyboard text fragments.

[0146] In some embodiments, after sending the first image generation instruction to the server in S930, the method further includes: determining a set of alternative storyboard images based on each first image; displaying a user interaction interface, the user interaction interface being used to display predetermined interaction elements, the predetermined interaction elements being used to trigger the selection of an image from the set of alternative storyboard images; determining a target storyboard image in response to a selection operation performed on an image in the set of alternative storyboard images; and sending an image selection instruction to the server, the image selection instruction being used to indicate the target storyboard image.

[0147] Exemplarily, the candidate storyboard image set includes all first images. The predetermined interactive element in the user interaction interface may be, for example, a selection control button, such as a single-selection control button, a multiple-selection control button, etc. for each candidate storyboard image.

[0148] In this embodiment, the client displays a user interaction interface for selecting an image from each first image, and uses the selected image as a target storyboard image, and sends an image selection instruction to the server, so that the server obtains the target storyboard image selected by the user after receiving the image selection instruction.

[0149] In some embodiments, before displaying the user interaction interface, the method also includes: generating a second image generation instruction based on the received detail description information and the second shooting parameters through a third information input interface, the detail description information being used to describe the detail features of the storyboard image constituent elements through natural language; sending the second image generation instruction to the server; and updating the set of alternative storyboard images according to each second image increment.

[0150] In this embodiment, the client can send a second image generation instruction to the server, so that the server can execute the step of generating the second image according to the detailed description information and the second shooting parameters, and can update the set of alternative storyboard images according to each second image increment, thereby helping to expand the selection range of the target storyboard image and provide richer image materials for subsequent video generation, which is conducive to further improving the quality of the generated video and generating a video that better meets customer needs.

[0151] In some embodiments, before displaying the user interaction interface, the method further includes: sending at least one third image submitted by the client in advance to the server; and incrementally updating the set of candidate storyboard images according to each third image.

[0152] In this embodiment, the client can send at least one third image submitted by the client to the server for incremental updating of the candidate storyboard image set. Adding each third image submitted by the client to the existing candidate storyboard image set further expands the selection range of the target storyboard image, provides the image material required by the client for subsequent video generation, and is conducive to the subsequent generation of videos that better meet the needs of the client.

[0153] In some embodiments, the corresponding video of the target storyboard image is a video generated based on each post-processed storyboard image obtained after image editing of the target storyboard image; after the step of sending an image selection instruction to the server, the method also includes: accessing each post-processed storyboard image of the server; obtaining at least one specified storyboard image in response to a selection instruction for at least one of the accessed storyboard images; receiving motion parameters for each specified storyboard image, the motion parameters being used to indicate motion characteristics of elements contained in each specified storyboard image; sending a video generation instruction to the server, the video generation instruction being used to indicate each specified storyboard image and the motion parameters.

[0154] Exemplarily, the client can access the post-processed storyboard images of the server in a variety of ways. For example, the server displays an image display interface to display the post-processed storyboard images. The user can perform remote desktop access to the server through the client to view the post-processed storyboard images displayed by the server. For another example, the server does not need to display the image display interface, but can save the subsequently processed storyboard images in a first file sharing directory. The client can access the first file sharing directory of the server through a browser based on the hypertext transfer protocol to view the post-processed storyboard images stored in the first file sharing directory.

[0155] In this embodiment, the client can access each storyboard image processed by the server, generate a video generation instruction according to at least one designated storyboard image and the motion parameters input by the user, and send the video generation instruction to the server. After receiving the video generation instruction, the server performs video generation processing on the designated storyboard image and the motion parameters to obtain the storyboard video corresponding to each designated storyboard image, thereby realizing automatic video generation.

[0156] In some embodiments, at least one specified storyboard image is used on the server side to generate a corresponding storyboard video; after sending a video generation instruction to the server side, the method further includes: accessing each storyboard video on the server side; in response to receiving a selection instruction for at least one specified storyboard video, generating a video processing instruction, the video processing instruction being used to indicate at least one specified storyboard video; and sending the video processing instruction to the server side.

[0157] Exemplarily, the client can access the storyboard video of the server in a variety of ways. For example: the server displays a video display interface, and the video display interface is used to display icons corresponding to multiple storyboard videos, such as image thumbnails. The user can perform remote desktop access to the server through the client to view the thumbnails corresponding to the multiple storyboard videos displayed by the server. The user clicks on any image thumbnail to display the corresponding storyboard video. For another example, the server does not need to display the video display interface, but can save the multiple storyboard videos in a second file sharing directory. The client can access the second file sharing directory of the server through a browser based on the Hypertext Transfer Protocol to view the multiple storyboard videos stored in the second file sharing directory.

[0158] In this embodiment, the client can access multiple storyboard videos of the server. The user accesses the server through the client and determines at least one designated storyboard video; then a video processing instruction can be generated according to each designated storyboard video, and the video processing instruction can be sent to the server. After receiving the video processing instruction, the server can perform post-processing on each designated storyboard video to obtain each post-storyboard video, and generate a corresponding video according to each post-storyboard video. Post-processing can convert the storyboard video into a higher quality video. Sending the storyboard video selected by the user to the server in the form of an instruction for post-processing is conducive to saving the use of local computing resources of the client and improving the processing efficiency of video generation.

[0159] According to the video display method of the embodiment of the present application, the client can generate a first image generation instruction based on the storyboard text and shooting parameters received from the user, and then send the first image generation instruction to the server, the storyboard text fragment and the first shooting parameter are used to generate at least one first image on the server, and each generated first image can be used to generate a corresponding video. The automated image generation instruction is conducive to accurately transmitting the storyboard text fragment and the first shooting parameter input by the user to the server, so as to generate the first image on the server and generate a corresponding video based on each first image. Sending the storyboard text and shooting parameters input by the user to the server for processing in the form of instructions is conducive to saving the use of local computing resources on the client, reducing the cost of using local computing resources on the client, and improving the processing efficiency of video generation.

[0160] Fig.10A flowchart of a video display method provided in an embodiment of the present application, such as Fig.10 As shown, the method specifically includes the following steps S1010-S1020.

[0161] S1010, displaying an operation interface of a predetermined application, wherein the operation interface includes an interactive entry element of a member page.

[0162] S1020, in response to an operation instruction for the interactive entry element, opening a member page and playing a predetermined video, where the predetermined video is a video generated according to the above-mentioned video generation method applied to the server.

[0163] In step S1010, the predetermined application may be any application running on the electronic device, for example, at least one of a shopping application, a media player, a social application, an educational application, etc.

[0164] In step S1020, the interactive entry element may be an interface element for triggering the opening of the membership page, such as a button, icon, link or other clickable interface element.

[0165] According to the video display method, on the displayed operation interface of the predetermined application, in response to the operation instruction for the interactive entry element, the member page can be opened and the predetermined video can be played. The predetermined video is a video generated by the video generation method of the aforementioned embodiment, so the video has a higher quality and the video production efficiency is higher, which helps to make the video display method based on the video generation method more efficient.

[0166] In the description of the above embodiments, regarding the processing steps in each method, the step numbers do not impose a unique restriction on the order of execution. In the absence of conflict, the steps can be executed sequentially, simultaneously, or one after the other.

[0167] Corresponding to the aforementioned embodiment of the video generating method applied to the server, the present application also provides an embodiment of the video generating device applied to the server. Fig.11 FIG. 1 is a schematic diagram of a structure of a video generating device according to an embodiment of the present application, wherein the video generating device is used to execute the video generating method applied to the server provided in any of the above embodiments, such as Fig.11 As shown, the video generating device comprises:

[0168] The receiving module 1110 is used to receive a first image generation instruction from a client, where the first image generation instruction includes a storyboard text segment obtained in advance from the storyboard text and a first shooting parameter.

[0169] The image generation module 1120 is used to use the image generation model to perform image generation processing on the storyboard text segment and the first shooting parameter to obtain at least one first image.

[0170] The video generation module 1130 is configured to generate a corresponding video based on each first image.

[0171] In some embodiments, the receiving module 1110 is also used to receive a text generation instruction from the client before receiving the first image generation instruction, and the text generation instruction includes storyboard description information; the video generation device also includes: a text generation module, which is used to generate storyboard text based on the storyboard description information in response to the text generation instruction.

[0172] In some embodiments, the video generation module 1130 is specifically used to: determine a set of candidate storyboard images based on each first image; receive an image selection instruction from the client, the image selection instruction is used to indicate at least one target storyboard image in the set of candidate storyboard images; and generate a corresponding video based on each target storyboard image.

[0173] In some embodiments, the receiving module 1110 is further used to receive a second image generation instruction from the client, the second image generation instruction includes detail description information and description information of the second shooting parameters, the detail description information is used to describe the detail features of the elements constituting the storyboard image through natural language; the image generation module 920 is further used to use the image generation model to perform image generation processing on the detail description information and the description information of the second shooting parameters to obtain at least one second image; the video generation device also includes: an update module, which is used to update the set of alternative storyboard images using each second image increment.

[0174] In some embodiments, the receiving module 1110 is further used to receive at least one third image submitted by a client from a client; the updating module is further used to incrementally update the set of candidate storyboard images using each third image.

[0175] In some embodiments, when the video generation module 1130 is used to generate a corresponding video according to each target storyboard image, it is specifically used to: perform image editing on each target storyboard image to obtain each post-processed storyboard image; and generate a corresponding video according to each post-processed storyboard image.

[0176] In some embodiments, when the video generation module 1130 is used to perform image editing on each target storyboard image, it is specifically used to: perform image editing on each target storyboard image according to image editing operations preset by the server; or, perform image editing on each target storyboard image according to an external image editing application called by the server; or, after performing image editing on each target storyboard image according to image editing operations preset by the server to obtain each first edited image, image editing is performed on each first edited image again according to the external image editing application called by the server.

[0177] In some embodiments, when the video generation module 1130 is used to generate corresponding videos based on each post-processed storyboard image, it is specifically used to: receive a video generation instruction from a client, the video generation instruction is used to indicate at least one specified storyboard image among each post-processed storyboard image and motion parameters corresponding to each specified storyboard image, the motion parameters are used to indicate the motion characteristics of the elements contained in the corresponding specified storyboard image; perform video generation processing on each specified storyboard image and the corresponding motion parameters to obtain the storyboard videos corresponding to each specified storyboard image; generate a corresponding video based on the storyboard videos corresponding to each specified storyboard image.

[0178] In some embodiments, when the video generation module 1130 is used to generate a corresponding video based on the storyboard videos corresponding to each specified storyboard image, it is specifically used to: receive a video processing instruction from the client, the video processing instruction is used to indicate at least one specified storyboard video among the storyboard videos; perform post-processing on the specified storyboard video to obtain a post-storyboard video; and generate a corresponding video based on the post-storyboard video.

[0179] The video generation device applied to the server provided in the embodiment of the present application and the video generation method applied to the server provided in the embodiment of the present application are based on the same inventive concept and have the same beneficial effects as the methods adopted, run or implemented therein.

[0180] Corresponding to the aforementioned embodiment of the video generating method applied to the client, the present application also provides an embodiment of the video generating device applied to the client. Fig.12 FIG. 1 is a schematic diagram showing a structure according to an embodiment of the present application. The video generating device is used to execute the video generating method applied to the client provided in any of the above embodiments, such as Fig.12 As shown, the video generating device comprises:

[0181] The receiving module 1210 is used to receive the storyboard text fragment and the first shooting parameter input by the user through the first information input interface; the generating module 1220 is used to generate a first image generation instruction including the storyboard text fragment and the first shooting parameter; the sending module 1230 is used to send the first image generation instruction to the server, so that the server generates a corresponding video based on the storyboard text fragment and the first shooting parameter; the receiving module 1210 is also used to receive the corresponding video sent by the server.

[0182] In some embodiments, the receiving module 1210 is also used to receive the storyboard description information input by the user through the second information input interface before receiving the storyboard text segment and the first shooting parameter input by the user through the first information input interface; the video generating device also includes: a sending module 1230, used to send a text generation instruction to the server, the text generation instruction including the description information; the text generation instruction is used to instruct the server to generate the storyboard text according to the storyboard description information.

[0183] In some embodiments, the video generating device further includes: a determining module, which is used to determine a set of alternative storyboard images based on each first image after sending a first image generating instruction to a server; an interface display module, which is used to display a user interaction interface, the user interaction interface is used to display predetermined interactive elements, and the predetermined interactive elements are used to trigger the selection of an image from the set of alternative storyboard images; the determining module is also used to determine a target storyboard image in response to a selection operation performed on an image in the set of alternative storyboard images; and the sending module 1230 is also used to send an image selection instruction to the server, the image selection instruction is used to indicate a target storyboard image.

[0184] In some embodiments, the generating module 1220 is further used to generate a second image generating instruction based on the received detail description information and the second shooting parameters through a third information input interface before displaying the user interaction interface, and the detail description information is used to describe the detail features of the storyboard image constituent elements through natural language; the sending module 1230 is further used to send the second image generating instruction to the server; the video generating device also includes: an updating module, used to update the set of alternative storyboard images according to each second image increment.

[0185] In some embodiments, the sending module 1230 is further used to send at least one third image submitted by the client in advance to the server before displaying the user interaction interface; the updating module is further used to update the set of alternative storyboard images according to each third image increment.

[0186] In some embodiments, the corresponding video of the target storyboard image is a video generated based on the post-processed storyboard image obtained after image editing of the target storyboard image; the video generation device also includes: an access module, which is used to access the post-processed storyboard image of the server after sending an image selection instruction to the server; the determination module is also used to obtain at least one specified storyboard image in response to a selection instruction for at least one of the accessed storyboard images; the receiving module 1210 is also used to receive motion parameters for each specified storyboard image, the motion parameters are used to indicate the motion characteristics of the elements contained in each specified storyboard image; the sending module 1230 is also used to send a video generation instruction to the server, the video generation instruction is used to indicate each specified storyboard image and the motion parameters.

[0187] In some embodiments, the selected image is used on the server side to generate a corresponding storyboard video; the access module is used to access each storyboard video on the server side after sending a video generation instruction to the server side; the generation module 1220 is also used to generate a video processing instruction in response to receiving a selection instruction for at least one specified storyboard video in each storyboard video, the video processing instruction is used to indicate at least one specified storyboard video; the sending module 1230 is also used to send the video processing instruction to the server side.

[0188] The video generation device applied to the client provided in the embodiment of the present application and the video generation method applied to the client provided in the embodiment of the present application are based on the same inventive concept and have the same beneficial effects as the methods adopted, run or implemented therein.

[0189] Corresponding to the above-mentioned embodiment of the video display method, the present application also provides an embodiment of a video display device. Fig.13 FIG. 1 is a schematic diagram of the structure of a video display device according to an embodiment of the present application. The video display device is used to execute the video display method provided in any of the above embodiments, such as Fig.13 As shown, the video display device includes:

[0190] The display module 1310 is used to display the operation interface of the predetermined application, which includes the interactive entry element of the member page; the opening module 1320 is used to respond to the operation instruction for the interactive entry element, open the member page and play the predetermined video, and the predetermined video is a video generated by any of the above-mentioned video generation methods applied to the server.

[0191] The video display device provided in the embodiment of the present application and the video display method provided in the embodiment of the present application are based on the same inventive concept and have the same beneficial effects as the methods adopted, operated or implemented therein.

[0192] It should be clear that the present application is not limited to the specific configurations and processes described in the above embodiments and shown in the figures. For the convenience and brevity of description, a detailed description of the known method is omitted here, and the implementation process of the functions and effects of each module in the above device is specifically detailed in the implementation process of the corresponding steps in the above method, which will not be repeated here.

[0193] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can refer to the partial description of the method embodiments. The device embodiments described above are merely illustrative, wherein the modules described as separate components may or may not be physically separated, and the components illustrated as modules may or may not be physical modules, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the present application. A person of ordinary skill in the art may understand and implement it without creative work.

[0194] Some embodiments of the present application also provide an electronic device corresponding to the video generation method or video display method provided in the aforementioned implementation manner, so as to execute the aforementioned video generation method or video display method.

[0195] Fig.14 1 is a hardware structure diagram of an electronic device according to an exemplary embodiment, and the electronic device includes: a communication interface 1401, a processor 1402, a memory 1403 and a bus 1404; wherein the communication interface 1401, the processor 1402 and the memory 1403 communicate with each other through the bus 1404. The processor 1402 can execute the video generation method or video display method described above by reading and executing the machine executable instructions corresponding to the control logic of the video generation method or the video display method in the memory 1403. The specific content of the method is referred to the above embodiment and will not be repeated here.

[0196] The memory 1403 mentioned in the embodiment of the present application can be any electronic, magnetic, optical or other physical storage device, and can contain storage information, such as executable instructions, data, etc. Specifically, the memory 1403 can be RAM (Random Access Memory), flash memory, storage drive (such as hard disk drive), any type of storage disk (such as optical disk, DVD, etc.), or similar storage medium, or a combination thereof. The communication connection between the system network element and at least one other network element is realized through at least one communication interface 1401 (which can be wired or wireless), and the Internet, wide area network, local area network, metropolitan area network, etc. can be used.

[0197] The bus 1404 may be an ISA bus, a PCI bus or an EISA bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc. The memory 1403 is used to store programs, and the processor 1402 executes the programs after receiving the execution instruction.

[0198] Processor 1402 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by an integrated logic circuit of hardware in processor 1402 or an instruction in software form. The above-mentioned processor 1402 can be a general-purpose processor, including a network processor (Network Processor, referred to as NP), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a readily available programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware controls, etc. The disclosed methods, steps and logic block diagrams in the embodiments of the present application can be implemented or executed. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in conjunction with the embodiments of the present application can be directly embodied as a hardware decoding processor to execute, or the hardware and software modules in the decoding processor are combined to execute.

[0199] The electronic device provided in the embodiment of the present application and the video generation method or video display method provided in the embodiment of the present application are based on the same inventive concept and have the same beneficial effects as the methods adopted, operated or implemented therein.

[0200] The present application also provides a computer-readable storage medium corresponding to the video generation method or video display method provided in the above-mentioned embodiment. Fig.15 As shown, the computer-readable storage medium shown is a CD 1510, on which a computer program (ie, a program product) is stored. When the computer program is run by a processor, it executes the video generation method or video display method provided by any of the aforementioned embodiments.

[0201] It should be noted that examples of computer-readable storage media may also include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other optical or magnetic storage media, which are not listed here one by one.

[0202] The computer-readable storage medium provided in the above-mentioned embodiments of the present application and the video generation method or video display method provided in the embodiments of the present application are based on the same inventive concept and have the same beneficial effects as the methods adopted, run or implemented by the application programs stored therein.

[0203] An embodiment of the present application also provides a computer program product corresponding to the video generation method or video display method provided in the aforementioned embodiment. The computer program product includes a computer program, and the computer program is executed by a processor to implement the video generation method or video display method provided in the aforementioned embodiment.

[0204] The computer program product provided by the above-mentioned embodiments of the present application and the video generation method or video display method provided by the embodiments of the present application are based on the same inventive concept and have the same beneficial effects as the methods adopted, run or implemented by the application programs stored therein.

[0205] Those skilled in the art will readily appreciate other embodiments of the present application after considering the specification and practicing the invention disclosed herein. The present application is intended to cover any modification, use or adaptation of the present application, which follows the general principles of the present application and includes common knowledge or customary techniques in the art that are not disclosed in the present application. The specification and examples are intended to be exemplary only, and the true scope and spirit of the present application are indicated by the claims.

[0206] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0207] The above description is only a preferred embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application.

Claims

1. A video generation method, characterized in that: Applied to the server, the method includes: Receiving a first image generation instruction from a client, wherein the first image generation instruction includes a storyboard text segment obtained in advance from the storyboard text and a first shooting parameter; Using an image generation model, performing image generation processing on the storyboard text segment and the first shooting parameter to obtain at least one first image; A corresponding video is generated based on each first image.

2. The method according to claim 1, characterized in that Before receiving the first image generation instruction from the client, the method further includes: Receiving a text generation instruction from the client, wherein the text generation instruction includes storyboard description information; In response to the text generation instruction, a storyboard text is generated based on the storyboard description information.

3. The method according to claim 1, characterized in that The generating a corresponding video based on each first image includes: Determine a set of candidate storyboard images according to each of the first images; Receiving an image selection instruction from the client, wherein the image selection instruction is used to indicate at least one target storyboard image in the set of candidate storyboard images; Generate corresponding videos according to each target storyboard image.

4. The method according to claim 3, wherein: Before receiving the image selection instruction from the client, the method further includes: receiving a second image generation instruction from the client, wherein the second image generation instruction includes detail description information and description information of a second shooting parameter, wherein the detail description information is used to describe detail features of elements constituting the storyboard image in natural language; Using the image generation model, perform image generation processing on the detail description information and the description information of the second shooting parameter to obtain at least one second image; The set of candidate storyboard images is updated using each second image increment.

5. A video generation method, characterized in that: Applied to a client, the method comprises: Receiving, via the first information input interface, a storyboard text segment and a first shooting parameter input by a user; Generate a first image generation instruction including the storyboard text segment and the first shooting parameter; Sending the first image generation instruction to a server, so that the server generates a corresponding video based on the storyboard text segment and the first shooting parameter; Receive the corresponding video sent by the server.

6. A video display method, characterized in that: The method comprises: Displaying the operation interface of the predetermined application, wherein the operation interface includes the interactive entry element of the member page; In response to an operation instruction for the interactive entry element, the member page is opened and a predetermined video is played, wherein the predetermined video is a video generated according to any one of the methods of claims 1-4.

7. A video generation system, characterized in that: The system includes a client and a server; The client is used to receive a storyboard text segment and a first shooting parameter in the storyboard text, generate a first image generation instruction according to the storyboard text segment and the first shooting parameter, and send the first image generation instruction to the server; The server is used to obtain the storyboard text fragment and the first shooting parameter from the first image generation instruction, perform image generation processing on the storyboard text fragment and the first shooting parameter using an image generation model, obtain at least one first image, and generate a corresponding video based on each first image.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: The processor executes the program to implement the method according to any one of claims 1 to 4, claim 5 or claim 6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: The program is executed by a processor to implement the method according to any one of claims 1 to 4, claim 5 or claim 6.

10. A computer program product, comprising a computer program, characterized in that The computer program is executed by a processor to implement the method of any one of claims 1 to 4, claim 5 or claim 6.