Video generation method and apparatus, computer-readable storage medium, and program product

By receiving and processing multiple reference images and their sequence information, target prompt information for the video content logic is generated, which solves the problem of poor controllability in video generation in the prior art and achieves video generation effects that better meet user expectations.

WO2026066116A1PCT designated stage Publication Date: 2026-04-02BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-05-12
Publication Date
2026-04-02

AI Technical Summary

Technical Problem

In existing technologies, the controllability of video generation based on user-input text is poor, resulting in generated videos that do not meet user expectations.

Method used

By receiving multiple reference images and their sequence information, the system identifies objects and scenes, generates target prompts for video content logic, and generates videos based on these prompts. This includes multimodal fusion and audio processing to improve the controllability and accuracy of video generation.

Benefits of technology

It improves the controllability and accuracy of video generation, enhances the user experience, and makes the generated videos more in line with user expectations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025094238_02042026_PF_FP_ABST
    Figure CN2025094238_02042026_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to the technical field of computers, and relates to a video generation method and apparatus, a computer-readable storage medium, and a program product. The video generation method comprises: receiving input content from a user, wherein the input content comprises a plurality of reference images and sequence information of the reference images; on the basis of the sequence information, and changing process information of objects and scenes in the plurality of reference images, generating target prompt information having video content logic; and on the basis of the target prompt information, generating a video comprising the plurality of reference images.
Need to check novelty before this filing date? Find Prior Art

Description

Video generation method and device, computer-readable storage medium, and program product

[0001] Cross-reference to Related Applications

[0002] This application is based on and claims priority to the application with the Chinese application number 202411354682.6 and the filing date of September 26, 2024, the disclosure of which is hereby incorporated by reference in its entirety into this application. TECHNICAL FIELD

[0003] The present disclosure relates to the technical field of computer, and particularly relates to a video generation method and device, a computer-readable storage medium, and a program product. BACKGROUND

[0004] With the development of artificial intelligence technology, relevant applications have penetrated into various aspects of our life, such as intelligent video generation.

[0005] In the related art, an agent feeds back a video corresponding to a text to a user according to a text input by the user. SUMMARY

[0006] This summary is provided to introduce a selection of concepts, which are further described below in the detailed description. This summary is not intended to identify key or essential features of the claimed technology, nor is it intended to be used to limit the scope of the claimed technology.

[0007] According to a first aspect of some embodiments of the present disclosure, a video generation method is provided, comprising:

[0008] receiving input content from a user, wherein the input content includes multiple reference images and sequence information of the multiple reference images;

[0009] generating target prompt information with video content logic according to the sequence information and change process information of objects and scenes in the multiple reference images;

[0010] generating a video including the multiple reference images according to the target prompt information.

[0011] In some embodiments, generating target prompt information with video content logic according to the sequence information and change process information of objects and scenes in the multiple reference images comprises:

[0012] identifying objects and scenes in each reference image;

[0013] performing semantic understanding on the multiple reference images according to the sequence information to obtain change process information of the objects and change process information of the scenes.

[0014] generate target prompt information with video content logic according to the object, the scene and the change process information.

[0015] In some embodiments, the semantic understanding of the plurality of reference images according to the sequence information to obtain the change process information of the object and the scene comprises:

[0016] extracting object features of the object and scene features of the scene in each reference image;

[0017] determining first difference information between the object features of the same object and second difference information between the scene features of the same scene in adjacent two reference images according to the sequence information;

[0018] determining the change process information of the object according to the first difference information;

[0019] determining the change process information of the scene according to the second difference information.

[0020] In some embodiments, the semantic understanding of the plurality of reference images according to the sequence information to obtain the change process information of the object and the scene comprises:

[0021] multimodal fusion of the object and the scene in the plurality of reference images, the sequence information and the plurality of reference images;

[0022] generating the change process information of the object and the change process information of the scene according to the result of the multimodal fusion.

[0023] In some embodiments, the multimodal fusion of the object and the scene in the plurality of reference images, the sequence information and the plurality of reference images comprises:

[0024] extracting text features from the object and the scene in the plurality of reference images and the sequence information;

[0025] extracting visual features from the plurality of reference images;

[0026] fusing the text features and the visual features to obtain fusion features as the result of the multimodal fusion.

[0027] In some embodiments, the generation of target prompt information with video content logic according to the object, the scene and the change process information comprises:

[0028] generating first picture description information describing a video picture of the object according to the object and the change process information of the object;

[0029] generating second picture description information describing a video picture of the scene according to the scene and change process information of the scene;

[0030] generating target prompt information with video content logic according to the first picture description information and the second picture description information.

[0031] In some embodiments, generating a video including the plurality of reference images according to the target prompt information comprises:

[0032] determining a time length of a video segment based on the plurality of reference images according to the change process information;

[0033] generating a video including the plurality of reference images according to the target prompt information and the time length of the video segment.

[0034] In some embodiments, generating a video including the plurality of reference images according to the target prompt information and the time length of the video segment comprises:

[0035] determining a visual effect of a video segment based on the plurality of reference images according to the change process information;

[0036] generating a video including the plurality of reference images according to the target prompt information, the time length of the video segment and the visual effect.

[0037] In some embodiments, the input content further comprises at least one segment of reference audio, and generating a video including the plurality of reference images according to the target prompt information comprises:

[0038] determining an audio effect of a video segment based on the plurality of reference images according to the change process information;

[0039] audio processing the at least one segment of reference audio according to the audio effect;

[0040] generating a video including the plurality of reference images and the at least one segment of reference audio after the audio processing according to the target prompt information.

[0041] In some embodiments, generating a video including the plurality of reference images according to the target prompt information comprises:

[0042] generating at least one segment of reference audio according to the sequence information and the plurality of reference images;

[0043] generating a video including the plurality of reference images and the at least one segment of reference audio according to the target prompt information.

[0044] In some embodiments, the change process information of the object includes action change information and shape change information; and / or

[0045] The change process information of the scene includes scene angle change information and switching information between different scenes.

[0046] In some embodiments, the plurality of reference images include a first frame image, a last frame image and at least one intermediate frame image between the first frame image and the last frame image of the video to be generated.

[0047] According to a second aspect of some embodiments of the present disclosure, a video generation apparatus is provided, comprising:

[0048] A receiving module configured to receive input content from a user, wherein the input content includes a plurality of reference images and sequence information of the plurality of reference images;

[0049] A first generating module configured to generate target prompt information with video content logic according to the sequence information and change process information of objects and scenes in the plurality of reference images;

[0050] A second generating module configured to generate a video including the plurality of reference images according to the target prompt information.

[0051] According to a third aspect of some embodiments of the present disclosure, a video generation apparatus is provided, comprising: a memory; and a processor coupled to the memory, the processor being configured to execute a video generation method of any of the embodiments described in the present disclosure based on instructions stored in the memory.

[0052] According to a fourth aspect of some embodiments of the present disclosure, a computer readable storage medium is provided, having stored thereon a computer program, which, when executed by a processor, performs a video generation method of any of the embodiments described in the present disclosure.

[0053] According to a fifth aspect of some embodiments of the present disclosure, a computer program product is provided, which, when executed on a computer, causes the computer to implement a video generation method of any of the embodiments.

[0054] According to a sixth aspect of some embodiments of the present disclosure, a computer program is provided, comprising: instructions which, when executed by a processor, cause the processor to perform a video generation method of any of the embodiments described in the present disclosure.

[0055] Other features, aspects, and advantages of the present disclosure will become apparent from the following detailed description of the exemplary embodiments with reference to the accompanying drawings. BRIEF DESCRIPTION OF DRAWINGS

[0056] Preferred embodiments of the present disclosure are explained hereinafter with reference to the accompanying drawings. The accompanying drawings, which are included to provide a further understanding of the present disclosure and are incorporated in and constitute a part of this specification, illustrate embodiments of the present disclosure and together with the detailed description serve to explain the present disclosure. It is to be understood that the drawings are only schematic and are used only for providing a conceptual understanding of the present disclosure. In the drawings:

[0057] FIG. 1 is a flowchart illustrating a method of generating a video according to some embodiments of the present disclosure;

[0058] FIG. 2 is a flowchart illustrating a method of generating a video according to some other embodiments of the present disclosure;

[0059] FIG. 3 is a flowchart illustrating a method of generating a video according to some other embodiments of the present disclosure;

[0060] FIG. 4 is a flowchart illustrating a method of generating a video according to some other embodiments of the present disclosure;

[0061] FIG. 5 is a flowchart illustrating a method of generating a video according to some other embodiments of the present disclosure;

[0062] FIG. 6 is a flowchart illustrating a method of generating a video according to some other embodiments of the present disclosure;

[0063] FIG. 7 is a flowchart illustrating a method of generating a video according to some other embodiments of the present disclosure;

[0064] FIG. 8 is a flowchart illustrating a method of generating a video according to some other embodiments of the present disclosure;

[0065] FIG. 9 is a flowchart illustrating a method of generating a video according to some other embodiments of the present disclosure;

[0066] FIG. 10A is a diagram illustrating an interface for generating a video according to some embodiments of the present disclosure;

[0067] FIG. 10B is a diagram illustrating an interface for generating a video according to some other embodiments of the present disclosure;

[0068] FIG. 11 is a block diagram illustrating a video generating apparatus according to some embodiments of the present disclosure;

[0069] FIG. 12 is a block diagram illustrating a video generating apparatus according to some other embodiments of the present disclosure;

[0070] FIG. 13 illustrates a block diagram of an electronic device according to some embodiments of the present disclosure.

[0071] It should be understood that the dimensions of the various parts shown in the drawings are not necessarily to scale. Identical or similar components are identified throughout the various figures with identical or similar reference numerals. Therefore, when a component is identified in one figure, it can not be further discussed in subsequent figures. DETAILED DESCRIPTION

[0072] The technical solutions in the embodiments of the present disclosure will be described clearly and completely in combination with the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only some of the embodiments of the present disclosure, but not all the embodiments. The description of the embodiments is actually only illustrative, but not as any limitation on the present disclosure and its application or use. It should be understood that the present disclosure can be implemented in various forms, and should not be interpreted as being limited to the embodiments set forth herein.

[0073] It should be understood that the various steps in the method embodiments of the present disclosure can be performed in different orders and / or in parallel. In addition, the method embodiments can include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect. Unless otherwise specifically stated, the relative arrangement of components and steps, numerical expressions, and numerical values set forth in these embodiments should be interpreted as merely exemplary, not limiting the scope of the present disclosure.

[0074] The term "comprise" and variations thereof used in the present disclosure means an open term that includes at least the recited elements, but does not exclude other elements. In other words, the term "comprising" means "including, but not limited to". In addition, the term "comprise" and variations thereof used in the present disclosure means an open term that includes at least the recited elements, but does not exclude other elements. In other words, the term "comprising" means "including, but not limited to". Therefore, comprising and including are synonymous. The term "based on" means "at least partially based on".

[0075] Throughout the specification, the terms "one embodiment", "some embodiments" or "embodiments" mean that the specific features, structures or characteristics described in connection with the embodiments are included in at least one embodiment of the present invention. For example, the term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Moreover, the appearance of the phrase "in one embodiment", "in some embodiments" or "in embodiments" at various places in the specification does not necessarily refer to the same embodiment, but can refer to different embodiments.

[0076] It should be noted that the terms "first", "second", and the like in the present disclosure are only used to distinguish different devices, modules or units, and are not intended to limit the order or interdependence of the functions performed by these devices, modules or units. Unless otherwise specified, the terms "first", "second", and the like are not intended to imply a given order or any other manner of given order in time, space, ranking, or any other manner.

[0077] It should be noted that the terms "one", "multiple" in the present disclosure are illustrative rather than limiting, and those skilled in the art should understand that unless otherwise explicitly specified in the context, it should be understood as "one or more".

[0078] The names of the messages or information exchanged between the plurality of devices in the embodiments of the present disclosure are only for illustrative purposes, and are not intended to limit the scope of the messages or information.

[0079] The embodiments of the present disclosure will be described in detail below in conjunction with the accompanying drawings, but the present disclosure is not limited to these specific embodiments. The following specific embodiments can be combined with each other, and the same or similar concepts or processes can not be described in some embodiments. In addition, in one or more embodiments, specific features, structures or characteristics can be combined by any suitable manner from the present disclosure which is clear to those skilled in the art.

[0080] In the related art, a video is generated according to a text input by a user, and controllability of the video generation is poor.

[0081] The present disclosure provides a technical solution, which can improve the controllability of video generation, improve the usability of generated video, and improve user experience.

[0082] FIG. 1 is a flow diagram illustrating a video generation method according to some embodiments of the present disclosure.

[0083] As shown in FIG. 1, the video generation method includes: step S10, receiving input content from a user, wherein the input content includes a plurality of reference images and sequence information of the plurality of reference images; step S20, generating target prompt information with video content logic according to the sequence information and change process information of objects and scenes in the plurality of reference images; and step S30, generating a video including the plurality of reference images according to the target prompt information. The video generation method can be performed by an intelligent agent.

[0084] The sequence information of the plurality of reference images, for example, includes position information of the plurality of reference images as image frames in a video to be generated.

[0085] The object in each reference image can be a person or an object, and the scene can be an environment in which the object is located. The change process information of the object describes the change of the object between different reference images (for example, adjacent reference images). The change information of the scene describes the change of the scene between different reference images.

[0086] The video content logic can also be referred to as video story logic, and can include at least one of a story theme, background information, character information, and plot development of the video, for example. The target prompt information with the video content logic can describe detailed content of a story that occurs in the video to be generated.

[0087] In the above embodiment, in the video generation process, the plurality of reference images input by the user and the sequence information of the plurality of reference images are considered, the prompt information with the video content logic is generated first, and then the video including the plurality of reference images input by the user is generated according to the prompt information. In the entire video generation process, the reference images are used as part of the video, and the sequence information is configured by the user, so that the video generation process is more controllable, and the generated video is more in line with the user's expectation, thereby improving the controllability and accuracy of video generation, improving the usability of the generated video, and improving the user experience.

[0088] The video generation method in other embodiments of the present disclosure will be described in detail below with reference to FIGS. 2-9.

[0089] FIG. 2 is a flow diagram illustrating a video generation method according to some embodiments of the present disclosure. FIG. 2 differs from FIG. 1 in that steps S21-S23 in FIG. 2 are an implementation of step S20 in FIG. 1. Only the differences between FIG. 2 and FIG. 1 will be described below, and the same parts will not be described again.

[0090] As shown in FIG. 2, in step S21, the object and the scene in each reference image are identified.

[0091] In step S22, according to the sequence information, the plurality of reference images are semantically understood to obtain the change process information of the object and the change process information of the scene.

[0092] In step S23, target prompt information with video content logic is generated according to the object, the scene, and the change process information.

[0093] For example, the identification of the object and the scene in each reference image can be achieved by object recognition.

[0094] In this embodiment, the order of the plurality of reference images to some extent reflects the video content logic, such as the development of the plot, of the video to be generated. In combination with the order information of the plurality of reference images, semantic understanding of the plurality of reference images can accurately obtain the change process information of the object and the change process information of the scene representing the development of the plot, thereby further guiding accurate generation of the target prompt information with the video content logic, so as to further improve the controllability and accuracy of video generation, and further improve the user experience.

[0095] FIG. 3 is a flow diagram illustrating a video generation method according to some embodiments of the present disclosure. FIG. 3 differs from FIG. 2 in that steps S220-S226 in FIG. 3 are an implementation of step S22 in FIG. 2. Hereinafter, only the differences between FIG. 3 and FIG. 2 will be described, and the same parts will not be described again.

[0096] As shown in FIG. 3, in step S220, the object features of the object and the scene features of the scene in each reference image are extracted.

[0097] In step S222, according to the order information, the first difference information between the object features of the same object in the adjacent two reference images and the second difference information between the scene features of the same scene are determined.

[0098] In step S224, the change process information of the object is determined according to the first difference information.

[0099] In step S226, the change process information of the scene is determined according to the second difference information.

[0100] For example, taking a cat as an example, the change process information of the cat can be determined by analyzing the difference changes in the actions and expressions of the cat in the adjacent two reference images.

[0101] For example, taking a coffee shop as an example, the change process information of the coffee shop can be determined by analyzing the difference changes in the existence and perspective of the coffee shop in the adjacent two reference images.

[0102] In this embodiment, the difference changes of the same object and the difference changes of the same scene in the adjacent two reference images are analyzed based on the order information, and then the change process information of the object and the scene is determined according to the difference changes, which is simple and convenient, can improve the efficiency of video generation, and improve the user experience.

[0103] FIG. 4 is a flowchart illustrating a video generation method according to some embodiments of the present disclosure. FIG. 4 differs from FIG. 2 in that steps S221 and S223 in FIG. 4 are an implementation of step S22 in FIG. 2. Only the differences between FIG. 4 and FIG. 2 will be described below, and the same parts will not be described again.

[0104] As shown in FIG. 4, in step S221, the object and the scene in the plurality of reference images, the sequence information, and the plurality of reference images are multi-modal fused.

[0105] In step S223, the change process information of the object and the change process information of the scene are generated according to the result of the multi-modal fusion.

[0106] The object, the scene, and the sequence information are text information, and the plurality of reference images are image information. The multi-modal fusion of the text information and the image information can be implemented by using a machine learning model, for example. The result of the multi-modal fusion can be multi-modal fusion features at a feature level, such as a feature vector, or text description information for describing the visual features of the object and the scene in the plurality of reference images.

[0107] In this embodiment, through multi-modal fusion, the change process of the object and the scene in the generated video caused by the plurality of reference images can be learned more deeply, so that the change process information of the object and the change process information of the scene can be generated more accurately, the accuracy of video generation is further improved, and the user experience is further improved.

[0108] FIG. 5 is a flowchart illustrating a video generation method according to some embodiments of the present disclosure. FIG. 5 differs from FIG. 4 in that steps S2211 and S2213 in FIG. 5 are an implementation of step S221 in FIG. 4. Only the differences between FIG. 5 and FIG. 4 will be described below, and the same parts will not be described again.

[0109] As shown in FIG. 5, in step S2211, text features are extracted from the object and the scene in the plurality of reference images and the sequence information.

[0110] In step S2212, visual features are extracted from the plurality of reference images.

[0111] In step S2213, the text features and the visual features are fused to obtain fusion features as the result of the multi-modal fusion.

[0112] For example, the text features and the visual features can both be represented by feature vectors.

[0113] In this embodiment, the accuracy of multi-modal fusion can be improved through multi-modal feature level fusion of text features and visual features, so as to further improve the accuracy of video generation and enhance user experience.

[0114] FIG. 6 is a flow diagram illustrating a video generation method according to some embodiments of the present disclosure. FIG. 6 differs from FIG. 2 in that steps S231-S233 in FIG. 6 are an implementation of step S23 in FIG. 2. Only the differences between FIG. 6 and FIG. 2 will be described below, and the same parts will not be described again.

[0115] As shown in FIG. 6, in step S231, first picture description information describing a video picture of the object is generated according to the object and the change process information of the object.

[0116] In step S232, second picture description information describing a video picture of the scene is generated according to the scene and the change process information of the scene.

[0117] In step S233, target prompt information with video content logic is generated according to the first picture description information and the second picture description information.

[0118] For example, taking the object as a kitten, the first picture description information can include description information of changes in the kitten's actions, expressions, etc. over time.

[0119] For example, taking the scene as a school playground, the second picture description information can include description information of changes in the size, shape, light and shadow effects, etc. of the school playground under different perspectives.

[0120] In this embodiment, the first picture description information of the video picture of the object and the second picture description information of the video picture of the scene are respectively generated based on the change process information of the object and the change process information of the scene, and then the target prompt information is generated based on the picture description information, which can improve the accuracy of the target prompt information, thereby improving the accuracy of video generation and enhancing user experience.

[0121] FIG. 7 is a flow diagram illustrating a video generation method according to some embodiments of the present disclosure. FIG. 7 differs from FIG. 2 in that steps S31-S33 in FIG. 7 are an implementation of step S30 in FIG. 2. Only the differences between FIG. 7 and FIG. 2 will be described below, and the same parts will not be described again.

[0122] As shown in FIG. 7, in step S31, the duration of a video segment based on the plurality of reference images is determined according to the change process information.

[0123] In step S33, a video including the plurality of reference images is generated according to the target prompt information and the time length of the video segment.

[0124] For example, a longer time length can be determined when the change process information indicates that the change process is complex, and a shorter time length can be determined when the change process information indicates that the change process is simple.

[0125] In this embodiment, the time length of the video segment is determined based on the change process information, and then the video is generated according to the target prompt information and the time length of the video segment, which further improves the controllability and accuracy of video generation and enhances user experience.

[0126] In some embodiments, a visual effect of the video segment based on the plurality of reference images can be determined according to the change process information; and then a video including the plurality of reference images is generated according to the target prompt information, the time length of the video segment, and the visual effect.

[0127] For example, the visual effect can be the light and shade degree of color, a dynamic effect, a switching mode, etc.

[0128] In this embodiment, in the video generation process, the visual effect is also determined based on the change process information, which can further improve the controllability and accuracy of video generation and enhance user experience.

[0129] FIG. 8 is a flow diagram illustrating a video generation method according to some other embodiments of the present disclosure. FIG. 8 differs from FIG. 2 in that steps S32-S36 in FIG. 8 are an implementation of step S30 in FIG. 2. Hereinafter, only the differences between FIG. 8 and FIG. 2 will be described, and the same parts will not be described again.

[0130] As shown in FIG. 8, in step S32, an audio effect of a video segment based on the plurality of reference images is determined according to the change process information.

[0131] In step S34, the at least one segment of reference audio is subjected to audio processing according to the audio effect.

[0132] In step S36, a video including the plurality of reference images and the at least one segment of reference audio subjected to the audio processing is generated according to the target prompt information.

[0133] For example, the audio effect can include effects in terms of sound quality, spatial sense, pitch change, etc.

[0134] In this embodiment, the user can input reference audio, and in the video generation process, the audio effect is also determined in the change process information to process the user input reference audio, so that the generated video is more consistent with the actual situation, and the controllability, accuracy and usability of video generation can be further improved, and the user experience is improved.

[0135] FIG. 9 is a flow diagram illustrating a video generation method according to some embodiments of the present disclosure. FIG. 9 differs from FIG. 2 in that steps S32' to S34' in FIG. 9 are an implementation of step S30 in FIG. 2. Only the differences between FIG. 9 and FIG. 2 will be described below, and the same parts will not be described again.

[0136] As shown in FIG. 9, in step S32', at least one piece of reference audio is generated according to the sequence information and the plurality of reference images.

[0137] In step S34', a video including the plurality of reference images and the at least one piece of reference audio is generated according to the target prompt information.

[0138] In this embodiment, reference audio can be generated according to the sequence information of the plurality of reference images and the plurality of reference images, and more user-demanding audio can be generated in the case where the user does not provide audio, and then a video is generated, further improving the accuracy and usability of video generation and improving the user experience.

[0139] In some embodiments, the change process information of the object includes action change information and shape change information. For example, the action change information can include position change, posture change, gesture change, etc. For example, the shape change information can include shape change, size change, color or texture change, etc.

[0140] In some embodiments, the change process information of the scene includes scene angle change information and switching information between different scenes. The scene angle change information may, for example, include view angle height change, view angle distance change, etc. The switching information between different scenes may, for example, include the process of switching from a first scene to a second scene, etc.

[0141] In some embodiments, the plurality of reference images includes a first frame image of the video to be generated, a last frame image, and at least one intermediate frame image between the first frame image and the last frame image. In this embodiment, through the first frame, the last frame and the intermediate frame, the process of switching from the first frame picture to the intermediate frame picture and then to the last frame picture is described, so that the video generation process is more controllable, further improving the controllability and accuracy of video generation, and improving the user experience.

[0142] In some embodiments, the input content of the user can further include initial prompt information input by the user, which represents the demand or expected information of the user for video generation. The initial prompt information can be used to generate target prompt information throughout the whole process of video generation.

[0143] The video generation process in some embodiments of the present disclosure will be described below in conjunction with FIG. 10A and FIG. 10B.

[0144] FIG. 10A is a schematic diagram of an interface for video generation according to some embodiments of the present disclosure.

[0145] FIG. 10B is a schematic diagram of an interface for video generation according to some other embodiments of the present disclosure.

[0146] As shown in FIG. 10A, the user enters an interactive interface for interacting with an intelligent agent, and can perform settings for video generation by operating controls 10 of the interactive interface, such as inputting multiple reference images, inputting initial prompt information, and inputting audio, etc.

[0147] For example, the control 10 is a button with an identifier “video generation”, and the user can call up a video generation panel as shown in FIG. 10B by clicking the button. The video generation panel shown in FIG. 10B includes one or more controls, such as control 101, control 102, etc. The one or more controls include a control for the user to upload one or more reference images, and can further include a control for the user to set some attributes (such as aspect ratio, playback speed, etc.) of the video to be generated.

[0148] For the sequence information of the multiple reference images, the one or more controls can be identified as different frame identifiers, so that the frame position of the reference image in the video to be generated can be defined by uploading the reference image using the control with the corresponding frame identifier. For example, the one or more controls include multiple image upload buttons, which are differentiated by texts such as “first frame image”, “last frame image”, etc. to indicate the frame position of the reference image corresponding to the different image upload buttons in the video to be generated. For another example, the order between the multiple reference images can also be determined according to the order of image uploading, and the frame position of the reference image in the video to be generated can be set using other controls.

[0149] Taking the image upload button of the control 101 for the user to upload the reference image as an example, the user clicks the image upload button to call up the local or cloud photo album of the user, and selects and checks one image in the photo album as the reference image for uploading.

[0150] As shown in FIGS. 10A and 10B, the interactive interface further includes a control 20 and a control 30. For example, the control 20 is an input box, and the control 30 is a send button. The user can input a message indicating video generation in the input box in FIG. 10A and click the send button to trigger the agent to call up the video generation panel shown in FIG. 10B.

[0151] After the user completes the settings in the video generation panel, the user can click the send button shown in FIG. 10B to trigger the agent to execute the video generation method in any embodiment of the present disclosure to generate a video.

[0152] For example, taking the control 20 as an input box as an example, the user can also input original prompt information in the input box, where the original prompt information represents the user's demand or expectation for video generation. For example, the original prompt information can be a complete narrative text, including but not limited to information such as characters, scenes, and behaviors.

[0153] The above is the video generation method provided by some embodiments of the present disclosure. Next, the video generation apparatus in some embodiments of the present disclosure will be described in conjunction with FIG. 11.

[0154] FIG. 11 is a block diagram illustrating a video generation apparatus according to some embodiments of the present disclosure.

[0155] As shown in FIG. 11, the video generation apparatus 11 includes a receiving module 111, a first generation module 112, and a second generation module 113.

[0156] The receiving module 111 is configured to receive input content from a user, where the input content includes a plurality of reference images and sequence information of the plurality of reference images.

[0157] The first generation module 112 is configured to generate target prompt information having video content logic according to the sequence information and change process information of objects and scenes in the plurality of reference images.

[0158] The second generation module 113 is configured to generate a video including the plurality of reference images according to the target prompt information.

[0159] The video generation apparatus 11 can be used to execute steps S10-S30 of FIG. 1. In some embodiments, the video generation apparatus 11 can also execute any step shown in FIGS. 2-9.

[0160] It should be noted that the above-mentioned various modules are only logical modules according to the specific functions they implement, and are not used to limit the specific implementation manner, for example, they can be implemented in software, hardware or a combination of software and hardware. In actual implementation, the above-mentioned various modules can be implemented as independent physical entities, or can also be implemented by a single entity (for example, a processor (CPU or DSP, etc.), an integrated circuit, etc.). In addition, the above-mentioned various modules are indicated by dashed lines in the drawings, indicating that these modules can not actually exist, and the operations / functions they implement can be implemented by the processing circuit itself.

[0161] FIG. 12 is a block diagram illustrating a video generation apparatus according to some embodiments of the present disclosure.

[0162] As shown in FIG. 12, the video generation apparatus 12 includes a memory 121 and a processor 122 coupled to the memory 121, the processor 122 being configured to perform the video generation method according to any one of the preceding embodiments based on instructions stored in the memory 121.

[0163] The memory 121 is configured to store one or more computer-readable instructions. The memory 121 can include any combination of various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory, including but not limited to random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), read-only memory (ROM), flash memory. The memory 121 may, for example, store an operating system, an application program, a boot loader, a database, and other programs, and can also store various application programs and various data.

[0164] The processor 122 is configured to run the computer-readable instructions to implement the video generation method according to any one of the preceding embodiments. For specific implementation of each step of the video generation method, please refer to the above-mentioned embodiments, and the repeated parts will not be described here.

[0165] The processor 122 and the memory 121 can directly or indirectly communicate with each other. For example, the processor 122 and the memory 121 can communicate through a network. The network can include a wireless network, a wired network, and / or any combination of a wireless network and a wired network. The processor 122 and the memory 121 can also communicate with each other through a system bus, and the present disclosure does not limit the communication between the processor 122 and the memory 121.

[0166] It should be noted that the components of the video generation apparatus 12 shown in FIG. 12 are only exemplary and not limiting, and the video generation apparatus 12 can also have other components according to actual application needs. The processor 122 can perform the desired functions with other components in the video generation apparatus 12.

[0167] The video generation apparatus can be implemented in software, firmware and / or hardware, and can be integrated in an electronic device in which a related application is installed.

[0168] The above is the video generation apparatus in some embodiments of the disclosure.

[0169] FIG. 13 illustrates a block diagram of an electronic device according to some embodiments of the disclosure.

[0170] The electronic device 13 illustrated in FIG. 13 can be a computer system having a dedicated hardware structure, and can perform a corresponding function when a related application is installed.

[0171] The electronic device includes, but is not limited to, a mobile terminal such as a smartphone, a notebook computer, a Personal Digital Assistant (PDA), a Tablet Personal Computer (Tablet PC), a PMP (Portable Multimedia Player), a car terminal (e.g., a car navigation terminal), a wearable device, and the like, and a stationary terminal such as a digital TV, a desktop computer, and the like.

[0172] As illustrated in FIG. 13, a central processing unit (CPU) 131 performs various processes according to a program stored in a read-only memory (ROM) 132 or a program loaded from a storage portion 138 to a random access memory (RAM) 133. In the RAM 133, data required when the CPU 131 performs various processes and the like is stored as necessary. The central processing unit is merely exemplary, and can be other types of processors such as various processors described above. The ROM 132, the RAM 133, and the storage portion 138 can be various forms of computer readable storage media. It is noted that although the ROM 132, the RAM 133, and the storage portion 138 are illustrated separately in FIG. 13, one or more of them can be combined, or located in the same or different memory or storage module.

[0173] The CPU 131, the ROM 132, and the RAM 133 are connected to each other via a bus 134. An input / output interface 135 is also connected to the bus 134.

[0174] The following components are connected to the input / output interface 135: an input portion 136, such as a touch panel, a touch pad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, and the like; an output portion 137, including a display, such as a cathode ray tube (CRT), a liquid crystal display (LCD), a speaker, a vibrator, and the like; a storage portion 138, including a hard disk, a magnetic tape, and the like; and a communication portion 139, including a network interface card, such as a LAN card, a modem, and the like. The communication portion 139 allows communication processing to be performed via a network, such as the Internet. It is easily understood that, although the respective devices or modules in the electronic device 13 are shown in FIG. 13 as communicating through the bus 134, they can also communicate through a network or other means, where the network can include a wireless network, a wired network, and / or any combination of a wireless network and a wired network.

[0175] The driver 1310 is also connected to the input / output interface 135 as necessary. A removable medium 1311, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, and the like, is attached to the driver 1310 as necessary, so that a computer program read therefrom is installed in the storage portion 138 as necessary.

[0176] In the case where the above series of processes are implemented by software, the program constituting the software can be installed from a network or a storage medium 1311 such as a removable medium.

[0177] According to an embodiment of the present disclosure, the processes described above with reference to the flowcharts can be implemented as a computer software program. For example, some embodiments of the present disclosure include a computer program product that, when run on a computer, causes the computer to implement the video generation method described in any of the embodiments. The computer program product includes a computer program carried on a computer-readable medium, which contains program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through the communication portion 139, or installed from the storage portion 138, or installed from the ROM 132. When the computer program is executed by the CPU 131, the video generation method of the embodiment of the present disclosure is executed.

[0178] Note that, in the context of the present disclosure, the computer-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device.

[0179] The computer-readable medium can be a computer-readable storage medium, or a computer-readable signal medium, or any combination of the two.

[0180] Computer-readable storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium which can be used to store the desired information in a computer-readable format. In some embodiments, computer-readable storage media can include non-transitory computer-readable storage media. In some embodiments, computer-readable storage media can include transitory computer- readable storage media.

[0181] Computer-readable signal media can include a propagated data signal with computer- readable program code embodied therein. For example, a propagated signal can be an electromagnetic signal, an optical signal, and / or the like. Such a propagated signal can take a wide variety of forms, including, but not limited to, electro-magnetic signals, optical signals, and / or the like. Computer-readable signal media can include computer-readable instructions embodied thereon. The computer-readable instructions can be downloaded to a computer from computer-readable signal media and / or stored on and / or transmitted to computer-readable storage media as described above. Computer-readable signal media can be any computer-readable medium that is not a computer-readable storage medium.

[0182] The computer-readable medium can be included inside the electronic device; alternatively, the computer-readable medium can exist as a separate entity transportable to the electronic device.

[0183] In some embodiments, a computer program product is provided which, when run on a computer, causes the computer to implement the video generation method of any of the above embodiments.

[0184] In some embodiments, a computer program is provided which comprises instructions which, when executed on a processor, cause the processor to carry out the video generation method of any of the above embodiments. The instructions can be embodied in one or more computer programs.

[0185] Computer program code for carrying out operations of the present disclosure can be written in any one or more of a variety of programming languages or combinations of languages including an object oriented programming language such as Java, Smalltalk, C++ or the like and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).

[0186] The computer program instructions can also be loaded onto a computer or other programmable information processing apparatus to cause a series of operations to be performed on the computer or other programmable information processing apparatus to produce a computer implemented process such that the instructions which execute on the computer or other programmable information processing apparatus implement the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0187] The functions described above can be implemented in at least part by one or more hardware logic components. For example, and without limitation, illustrative hardware logic components that can be used include Field-programmable Gate Arrays (FPGAs), Program-specific Integrated Circuits (ASICs), Program-specific Standard Products (ASSPs), System-on-a-chip systems (SOCs), Complex Programmable Logic Devices (CPLDs), etc.

[0188] While certain aspects of the present disclosure have been described with reference to one or more particular embodiments thereof, those skilled in the art will understand that many alternative embodiments can be made therefrom. In general, embodiments of the present disclosure are applicable to any suitable electronic device, system, or architecture. In addition, unless otherwise indicated, the functions performed by the various components described herein can be implemented using electronic components, software, firmware, or any suitable combination thereof. Furthermore, from the disclosure herein, one skilled in the art will recognize that the various embodiments described herein can be implemented in a computer program product tangibly embodied in a machine readable storage medium (e.g., memory component) including many instructions that can be used to control the behavior of a computer or other machine.

Claims

1. A method of video generation, the method comprising: The method comprises: receiving input content from a user, wherein the input content comprises a plurality of reference images and sequence information of the plurality of reference images; generating target prompt information with video content logic according to the sequence information and change process information of objects and scenes in the plurality of reference images; generating a video comprising the plurality of reference images according to the target prompt information.

2. The video generation method of claim 1, wherein, Generating target prompt information with video content logic according to the sequence information and change process information of objects and scenes in the plurality of reference images comprises: identifying objects and scenes in each reference image; performing semantic understanding on the plurality of reference images according to the sequence information to obtain change process information of the objects and change process information of the scenes; generating target prompt information with video content logic according to the objects, the scenes and the change process information.

3. The video generation method of claim 2, wherein, Performing semantic understanding on the plurality of reference images according to the sequence information to obtain change process information of the objects and change process information of the scenes comprises: extracting object features of objects and scene features of scenes in each reference image; determining first difference information between object features of the same objects and second difference information between scene features of the same scenes in adjacent two reference images according to the sequence information; determining change process information of the objects according to the first difference information; determining change process information of the scenes according to the second difference information.

4. The video generation method of claim 2, wherein, Performing semantic understanding on the plurality of reference images according to the sequence information to obtain change process information of the objects and change process information of the scenes comprises: performing multi-modal fusion on the objects and scenes in the plurality of reference images, the sequence information and the plurality of reference images; generating change process information of the objects and change process information of the scenes according to a result of the multi-modal fusion.

5. The video generation method of claim 4, wherein, Performing multi-modal fusion on the objects and scenes in the plurality of reference images, the sequence information and the plurality of reference images comprises: extracting text features from the objects and scenes in the plurality of reference images and the sequence information; extracting visual features from the plurality of reference images; fusing the text features and the visual features to obtain fusion features as a result of the multi-modal fusion.

6. The video generation method of any of claims 2-5, Generating target prompt information with video content logic according to the objects, the scenes and the change process information comprises: generating first picture description information describing video pictures of the objects according to the objects and the change process information of the objects; generating second picture description information describing video pictures of the scenes according to the scenes and the change process information of the scenes; generating target prompt information with video content logic according to the first picture description information and the second picture description information.

7. The video generation method of any of claims 2-6, Generating a video comprising the plurality of reference images according to the target prompt information comprises: determining a time length of a video segment based on the plurality of reference images according to the change process information; generating a video comprising the plurality of reference images according to the target prompt information and the time length of the video segment.

8. The video generation method of claim 7, wherein, According to the target prompt information and the time length of the video segment, the video including the multiple reference images is generated. According to the change process information, a visual effect of the video segment based on the multiple reference images is determined. According to the target prompt information, the time length of the video segment and the visual effect, the video including the multiple reference images is generated.

9. The video generation method of any of claims 2-8, The input content further includes at least one piece of reference audio, and according to the target prompt information, the video including the multiple reference images is generated. According to the change process information, an audio effect of the video segment based on the multiple reference images is determined. According to the audio effect, the at least one piece of reference audio is audio processed. According to the target prompt information, the video including the multiple reference images and the at least one piece of audio processed reference audio is generated.

10. The video generation method of any of claims 2-8, According to the target prompt information, the video including the multiple reference images is generated. According to the sequence information and the multiple reference images, at least one piece of reference audio is generated. According to the target prompt information, the video including the multiple reference images and the at least one piece of reference audio is generated.

11. The video generation method of claim 3, wherein, the change process information of the object includes action change information and shape change information; and / or the change process information of the scene includes scene angle change information and switching information between different scenes.

12. The video generation method of any of claims 1-11, wherein, The multiple reference images include a first frame image, a last frame image and at least one intermediate frame image between the first frame image and the last frame image of the video to be generated.

13. A video generating apparatus, comprising: including: a receiving module configured to receive input content from a user, wherein the input content includes multiple reference images and sequence information of the multiple reference images; a first generation module configured to generate target prompt information with video content logic according to the sequence information and change process information of objects and scenes in the multiple reference images; a second generation module configured to generate a video including the multiple reference images according to the target prompt information.

14. A video generating apparatus characterized by comprising: including: a memory; and a processor coupled to the memory, the processor being configured to execute a video generation method according to any one of claims 1 to 12 based on instructions stored in the memory.

15. A computer-readable storage medium, characterized in that, a computer program having computer program instructions stored therein, which instructions, when executed by a processor, implement a video generation method according to any one of claims 1 to 12.

16. A computer program product, characterised in that, a computer program which, when executed by a processor, implements a video generation method according to any one of claims 1 to 12.

17. A computer program comprising: instructions which, when executed by a processor, cause the processor to perform a video generation method according to any one of claims 1 to 12.

Citation Information

Patent Citations

  • Multi-lens video recording method and related equipment

    CN114285963A

  • Video generation method and server

    CN116233491A

  • Multi-lens video recording method and related equipment

    CN116405776A

  • Creating cinematic video from multi-view capture data

    US20210227195A1