Video generation method, apparatus, device, and storage medium
By encoding text prompts and multiple reference information, a target video is generated, which solves the problem that existing video generation schemes cannot maintain diversity and motion smoothness, and achieves accurate restoration and diversity of object instances in the video.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
- Filing Date
- 2024-12-02
- Publication Date
- 2026-05-12
AI Technical Summary
Existing image generation schemes based on multiple object instances cannot be directly transferred to video generation scenarios, resulting in a loss of diversity in video content and smoothness of motion.
The target video is generated by encoding text prompts and multiple reference information separately. This includes encoding the text prompts to obtain text prompt features, and visually and semantically encoding multiple reference information separately. The visual features and text semantic features are then fused to generate the target video.
The generated video accurately reproduces the attributes of each reference object instance, exhibiting diversity and smooth motion, and avoids dependence on region constraints.
Smart Images

Figure CN119676472B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of computer technology, and in particular to a video generation method, apparatus, device, and storage medium. Background Technology
[0002] With the development of computer technology, video generation schemes have emerged that generate video content containing specific object instances using image guidance. In practical applications, users may want to generate video content that includes multiple instances of specific objects. Therefore, how to implement the above schemes to meet user needs is a research direction.
[0003] Currently, there are image generation schemes based on multiple object instances. These schemes typically impose region constraints on each object instance when generating the image to avoid confusion regarding the object instance's attributes.
[0004] However, because video content is dynamic, the positions of object instances change over time, and region constraints affect the diversity and motion smoothness of the generated video content, image generation schemes based on multiple images cannot be directly transferred to video generation scenarios. Therefore, a new video generation scheme is urgently needed. Summary of the Invention
[0005] This disclosure provides a video generation method, apparatus, device, and storage medium. This solution debinds features from different reference information to avoid confusion between them, enabling the generated video to accurately reproduce the attributes of each reference object instance. Furthermore, since this solution does not rely on region constraints, the generated video exhibits diversity and smooth motion.
[0006] According to one aspect of the embodiments of this disclosure, a video generation method is provided, the method comprising:
[0007] The text prompt information is encoded to obtain text prompt features;
[0008] Multiple reference information are encoded separately to obtain multiple semantic visual features. Each reference information includes a reference image and descriptive text. The reference image includes a reference object instance, and the descriptive text is used to describe the corresponding reference object instance.
[0009] The target video is generated based on the text prompt features and the multiple semantic visual features.
[0010] According to another aspect of the present disclosure, a video generation apparatus is provided, the apparatus comprising:
[0011] The first encoding unit is configured to encode the text prompt information to obtain text prompt features;
[0012] The second encoding unit is configured to encode multiple reference information respectively to obtain multiple semantic visual features. Each reference information includes a reference image and descriptive text. The reference image includes a reference object instance, and the descriptive text is used to describe the corresponding reference object instance.
[0013] The generation unit is configured to generate a target video based on the text prompt features and the plurality of semantic visual features.
[0014] In some embodiments, the second encoding unit is configured to, for any reference information, perform visual encoding on a reference image in the reference information to obtain visual features of the reference object instance; perform semantic encoding on descriptive text in the reference information to obtain textual semantic features of the descriptive text; and fuse the visual features and the textual semantic features to obtain semantic visual features of the reference information.
[0015] In some embodiments, the first encoding unit is configured to align the visual features and the text semantic features; and to fuse the aligned visual features and the text semantic features based on a cross-attention mechanism.
[0016] In some embodiments, the generation unit is configured to fuse the video features of the diffusion model with the text prompt features to obtain intermediate fused features, wherein the video features of the diffusion model are hidden layer features in the video generation process; to fuse the intermediate fused features with the plurality of semantic visual features in sequence to obtain video features; and to decode the video features to obtain the target video.
[0017] In some embodiments, the generation unit is configured to fuse the video features of the diffusion model and the text prompt features through a cross-attention mechanism to obtain the intermediate fused features.
[0018] In some embodiments, the generation unit is configured to sequentially acquire multiple semantic visual features through a feature injection module; for each semantic visual feature acquired, the semantic visual feature is fused with the intermediate fusion feature to obtain an intermediate video feature; and in response to the completion of fusion, the video feature is obtained.
[0019] In some embodiments, the first encoding unit is further configured to acquire input text prompt information; perform compliance verification on the text prompt information; and display prompt information if the text prompt information fails the verification.
[0020] According to another aspect of the embodiments of this disclosure, an electronic device is provided, the electronic device comprising:
[0021] One or more processors;
[0022] Memory used to store the executable program code of the processor;
[0023] The processor is configured to execute the program code to implement the video generation method described above.
[0024] According to another aspect of the present disclosure, a computer-readable storage medium is provided that, when the instructions in the computer-readable storage medium are executed by a processor of an electronic device, enables the electronic device to perform the video generation method described above.
[0025] According to another aspect of the present disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the above-described video generation method.
[0026] This disclosure provides a video generation scheme that, by encoding multiple pieces of reference information separately, binds the features of the reference object instance corresponding to each piece of reference information with the features of the descriptive text, and debindes the features of different reference information to avoid confusion between them. This allows the generated video to accurately reproduce the attributes of each reference object instance. Furthermore, since the above scheme does not rely on region constraints, the generated video exhibits diversity and smooth motion.
[0027] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description
[0028] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure, and are not intended to unduly limit this disclosure.
[0029] Figure 1 This is a schematic diagram illustrating the implementation environment of a video generation method according to an exemplary embodiment.
[0030] Figure 2 This is a flowchart illustrating a video generation method according to an exemplary embodiment.
[0031] Figure 3 This is a flowchart illustrating another video generation method according to an exemplary embodiment.
[0032] Figure 4 This is a schematic diagram of a process for encoding reference information according to an exemplary embodiment.
[0033] Figure 5 This is a flowchart illustrating a video generation method according to an exemplary embodiment.
[0034] Figure 6 This is a block diagram illustrating a video generation apparatus according to an exemplary embodiment.
[0035] Figure 7 This is a block diagram illustrating an electronic device according to an exemplary embodiment. Detailed Implementation
[0036] To enable those skilled in the art to better understand the technical solutions of this disclosure, the technical solutions in the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings.
[0037] It should be noted that the terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.
[0038] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.), and signals involved in this disclosure are all authorized by the user or fully authorized by all parties, and the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, the images involved in this disclosure were obtained with full authorization.
[0039] Figure 1 This is a schematic diagram illustrating an implementation environment for a video generation method according to an exemplary embodiment. See also... Figure 1 The implementation environment specifically includes: terminal 101 and server 102. Terminal 101 can be connected to server 102 via wireless network or wired network.
[0040] Terminal 101 can be at least one of the following devices: smartphone, smartwatch, desktop computer, laptop, MP3 player (Moving Picture Experts Group Audio Layer III), MP4 player (Moving Picture Experts Group Audio Layer IV), and laptop computer. An application can be installed and run on terminal 101, which generates video based on user input. This application is associated with server 102, which provides background services to terminal 101.
[0041] Terminal 101 can refer to one of a plurality of terminals, and this embodiment uses terminal 101 as an example. Those skilled in the art will know that the number of terminals can be more or less. For example, there can be several terminals, or dozens or hundreds of terminals, or more. This embodiment does not limit the number of terminals or the type of devices.
[0042] Server 102 can be at least one of a single server, multiple servers, a cloud computing platform, and a virtualization center. Optionally, the number of servers may be more or fewer, and this disclosure does not limit this. Of course, server 102 may also include other functional servers to provide more comprehensive and diversified services. In some embodiments, server 102 undertakes the main computing work, and terminal 101 undertakes the secondary computing work; or, server 102 undertakes the secondary computing work, and terminal 101 undertakes the main computing work; or, server 102 and terminal 101 collaborate on computing using a distributed computing architecture. Server 102 can be connected to terminal 101 and other terminals via a wireless network or a wired network. Optionally, the number of servers may be more or fewer, and this disclosure does not limit this.
[0043] Figure 2 This is a flowchart illustrating a video generation method according to an exemplary embodiment, such as... Figure 2 As shown, the method is performed by an electronic device and includes the following steps:
[0044] In step S201, the text prompt information is encoded to obtain text prompt features.
[0045] In this embodiment of the disclosure, the text prompt information is information input by the user and serves as a prompt during the video generation process. The text prompt information includes at least one prompt word. By encoding the at least one prompt word, text prompt features can be obtained.
[0046] In step S202, multiple reference information are encoded to obtain multiple semantic visual features.
[0047] In this embodiment of the disclosure, each piece of reference information includes a pair of reference images and descriptive text. The reference image includes a reference object instance, and the descriptive text describes the corresponding reference object instance. The reference object instance represents a concept contained in the video to be generated, such as a person, animal, building, landscape, or style.
[0048] In step S203, the target video is generated based on text prompt features and multiple semantic visual features.
[0049] In this embodiment of the disclosure, multiple semantic visual features can guide the generation of video content, while text prompt features can yield video content with different styles, characters, or scenes. Therefore, based on text prompt features and multiple semantic visual features, a target video conforming to the aforementioned features can be generated.
[0050] This disclosure provides a video generation scheme that, by encoding multiple pieces of reference information separately, binds the features of the reference object instance corresponding to each piece of reference information with the features of the descriptive text, and debindes the features of different reference information to avoid confusion between them. This allows the generated video to accurately reproduce the attributes of each reference object instance. Furthermore, since the above scheme does not rely on region constraints, the generated video exhibits diversity and smooth motion.
[0051] In some embodiments, multiple reference information are encoded to obtain multiple semantic visual features, including:
[0052] For any reference information, the reference image in the reference information is visually encoded to obtain the visual features of the reference object instance;
[0053] Semantic encoding is performed on the descriptive text in the reference information to obtain the textual semantic features of the descriptive text;
[0054] By integrating visual features and textual semantic features, the semantic visual features of the reference information are obtained.
[0055] In this embodiment of the disclosure, the operation can effectively integrate the visual content in the reference information with the semantic information of the descriptive text, accurately extract the semantic visual features of the reference information, thereby enabling the electronic device to better understand the reference object instance, providing a more accurate and comprehensive basis for subsequent video generation tasks, and improving the processing effect.
[0056] In some embodiments, visual features and textual semantic features are fused to obtain semantic visual features of the reference information, including:
[0057] Align visual features with textual semantic features;
[0058] Based on the cross-attention mechanism, the aligned visual features and textual semantic features are fused together.
[0059] In this embodiment of the disclosure, by first aligning visual features and textual semantic features, and then fusing them based on a cross-attention mechanism, the relevant information of images and text can be combined more accurately and comprehensively, thereby providing more accurate and effective feature basis for subsequent video generation and improving processing performance.
[0060] In some embodiments, a target video is generated based on text prompt features and multiple semantic visual features, including:
[0061] The video features of the diffusion model are fused with the text prompt features to obtain intermediate fused features. The video features of the diffusion model are the hidden features in the video generation process.
[0062] The intermediate fusion features are sequentially fused with multiple semantic visual features to obtain video features;
[0063] The target video is obtained by decoding the video features.
[0064] In this embodiment of the disclosure, by sequentially fusing the hidden video features of the diffusion model with text prompt features and multiple semantic visual features when generating the video, and then decoding the result to obtain the target video, the diffusion model's multiple iterations can fully integrate multi-faceted feature information, effectively improving the accuracy, richness, and relevance of the target video generation to user needs.
[0065] In some embodiments, the video features of the diffusion model are fused with the text prompt features to obtain intermediate fused features, including:
[0066] By using a cross-attention mechanism, video features and text prompt features from the diffusion model are fused to obtain intermediate fused features.
[0067] In this embodiment of the disclosure, the intermediate fusion feature is obtained by fusing video features and text prompt features of the diffusion model through the cross-attention mechanism. This enables electronic devices to accurately capture the correlation and interaction information between the two, effectively combining relevant characteristics, thereby laying a good foundation for the subsequent generation of target videos that are more in line with the requirements and of higher quality.
[0068] In some embodiments, intermediate fusion features are sequentially fused with multiple semantic visual features to obtain video features, including:
[0069] Multiple semantic visual features are obtained sequentially through the feature injection module;
[0070] For each semantic visual feature acquired, the semantic visual feature is fused with the intermediate fusion feature to obtain the intermediate video feature;
[0071] Once the fusion is complete, the video features are obtained.
[0072] In this embodiment of the disclosure, by introducing a feature injection module into the diffusion model, multiple semantic visual features can be effectively acquired and integrated, ensuring that the features are deeply integrated with visual features and text prompt features. When an electronic device operates according to this process, video features with high integration and better meeting the requirements can be generated, thereby improving the quality and effect of target video generation.
[0073] In some embodiments, the method further includes:
[0074] Get the input text prompt information;
[0075] Perform compliance verification on text prompts;
[0076] If the text prompt fails validation, a prompt message will be displayed.
[0077] In this embodiment of the disclosure, validating the text prompt information can effectively filter out compliant content, avoid errors in subsequent processes or invalid results due to non-compliant input, improve the accuracy and stability of the overall video generation process, and optimize the user interaction experience and ensure the effectiveness of system operation by promptly prompting the user to re-enter the information.
[0078] The above Figure 2 The diagram shown is a flowchart of a video generation method according to this disclosure. The video generation scheme provided by this disclosure will be further elaborated below. Figure 3 This is a flowchart illustrating another video generation method according to an exemplary embodiment, see [link to flowchart]. Figure 3 This method is performed by an electronic device and includes the following steps:
[0079] In step S301, the input text prompt information is obtained.
[0080] In this embodiment of the disclosure, text prompts are used to guide and instruct the generation of video content. Simply put, text prompts use words to describe various aspects of the video to be generated, such as the video's style, scene, and context.
[0081] For example, text prompts can include the video's theme, such as "Generate a video about a sunrise at the beach." This theme clearly indicates the main scene and general content direction of the video. Text prompts can also describe the video's style, such as "Show a city night scene in an oil painting style." Here, "oil painting style" is the text prompt indicating the style, giving the generated video an oil painting texture; aspects such as brushstrokes and color saturation will be generated according to the characteristics of oil painting. Text prompts can also include the video's plot development. For example, "Generate a video that begins with a little girl getting lost in the forest, then she meets a talking elf who guides her home." Such text prompts can construct a complete plot chain, including characters (the little girl and the elf), the scene (the forest), and the beginning, development, and ending of the plot.
[0082] In some embodiments, the electronic device can validate the text prompt information entered by the user. Accordingly, the electronic device acquires the entered text prompt information. Then, the electronic device performs compliance validation on the text prompt information. If the text prompt information fails validation, the electronic device displays a prompt message to inform the user that the currently entered text prompt information is non-compliant and requests the user to re-enter it. If the text prompt information passes validation, the electronic device can either execute subsequent processes or display that the input is compliant. Validating the text prompt information effectively filters out compliant content, avoiding errors or invalid results in subsequent processes due to non-compliant input, improving the accuracy and stability of the overall video generation process. Simultaneously, by promptly prompting the user to re-enter, it optimizes the user interaction experience and ensures the effectiveness of system operation.
[0083] In some embodiments, users can input the above-mentioned text prompt information through an input box, or by voice input, or by selecting a template from preset prompt information templates that closely matches their needs, and then modify and supplement it. This disclosure does not limit the input method for text prompt information.
[0084] In step S302, the text prompt information is encoded to obtain text prompt features.
[0085] In this embodiment of the disclosure, since text prompts are usually descriptions in natural language, it is difficult for electronic devices to directly generate videos based on such natural language text. Therefore, the text prompts are encoded and converted into text prompt features that can be understood and processed by a computer.
[0086] Alternatively, the electronic device can encode the text prompt information using a text feature encoder. This encoder extracts text prompt words from the prompt information and then extracts these words into a complete text feature, i.e., a text prompt feature.
[0087] Optionally, the text feature encoder can encode the text prompt using word vector encoding. For example, the text feature encoder can convert each word in the text prompt into a vector. Suppose the text prompt is "Generate a happy birthday party video". The text feature encoder converts words such as "happy", "birthday party", and "video" into corresponding word vectors. These word vectors are points in a high-dimensional space, and the vector dimension may be hundreds of dimensions. For example, the vector of the word "happy" might be represented as [0.1, 0.2, -0.3, ...], where these values represent the position of the word in this high-dimensional space. Word vectors can capture the semantic information of words; for example, "happy" and "joyful" are semantically similar, so their word vectors will be relatively close in space.
[0088] Optionally, the text feature encoder can also be a deep learning encoding model, such as BERT, etc., and this disclosure does not limit this.
[0089] In step S303, multiple pieces of reference information are obtained. Each piece of reference information includes a reference image and descriptive text. The reference image includes a reference object instance, and the descriptive text is used to describe the corresponding reference object instance.
[0090] In this embodiment of the disclosure, reference information is used to assist in video generation. Each piece of reference information consists of two parts: a reference image and descriptive text.
[0091] Reference images are used to visually represent content. For example, a reference image might be a photograph of a cat, clearly showing its appearance, fur color, posture, and other visual features. Each cat in the photograph is a reference object instance, meaning a reference object instance is an individual object with actual form and characteristics.
[0092] The descriptive text is used to describe the corresponding reference object instance in words. Continuing with the example of a photo including a cat, the descriptive text might read "This is a white cat." Such descriptive text accurately conveys the information about the reference object instance.
[0093] In step S304, multiple reference information are encoded to obtain multiple semantic visual features.
[0094] In this embodiment of the disclosure, each piece of reference information is encoded individually to obtain its semantic visual features. This effectively distinguishes the semantic visual features corresponding to different pieces of reference information.
[0095] In some embodiments, semantic visual features are visual features with semantic information. Accordingly, for any reference information, the electronic device first performs visual encoding on the reference image in the reference information to obtain the visual features of the reference object instance. Then, the electronic device performs semantic encoding on the descriptive text in the reference information to obtain the textual semantic features of the descriptive text. Finally, the electronic device fuses the visual features and the textual semantic features to obtain the semantic visual features of the reference information. This operation effectively fuses the visual content in the reference information with the semantic information of the descriptive text, accurately extracting the semantic visual features of the reference information. This allows the electronic device to better understand the reference object instance, providing a more accurate and comprehensive basis for subsequent video generation tasks and improving processing performance.
[0096] Visual coding is an operation that transforms the visual information of an image into a form that computers can understand and process. Visual coding uses specific algorithms and techniques to analyze the visual attributes of each pixel in a reference image, such as color, brightness, and texture, and then represents this information using a specific encoding method.
[0097] For example, for a reference image containing a red apple, visual encoding can extract visual features such as the apple's shape (approximately round), color (specific hue of red), and surface texture (smoothness, etc.). These features constitute the visual features of the reference object instance (this apple).
[0098] Semantic encoding is used to understand the semantic information conveyed by the descriptive text, such as analyzing the meaning of words in the text, the relationships between words, and the overall meaning of sentences.
[0099] For example, continuing with the descriptive text containing the reference image of a red apple, semantic encoding can extract the semantics of the fruit category represented by the word "apple," the semantics describing the appearance of the apple such as "bright color" and "smooth," and the semantics of the location involved in "placed on a wooden cutting board." These constitute the textual semantic features of the descriptive text.
[0100] In some embodiments, the electronic device fuses visual features and textual semantic features using a cross-attention mechanism. Specifically, the electronic device first aligns the visual features and textual semantic features. Then, based on the cross-attention mechanism, the electronic device fuses the aligned visual features and textual semantic features. By first aligning the visual features and textual semantic features and then fusing them based on the cross-attention mechanism, the relevant information of the image and text can be combined more accurately and comprehensively, thereby providing more accurate and effective feature basis for subsequent video generation and improving processing performance.
[0101] In this process, the electronic device acquires visual features from a reference image and textual semantic features from descriptive text. These features may initially have different dimensions. Alignment operations allow the visual features to be aligned with the textual features in terms of dimensions.
[0102] For example, matching certain key elements in visual features (such as numerical values corresponding to color features) with the semantic codes corresponding to words describing color in textual semantic features allows elements in visual features and elements in textual semantic features to be compared with each other in a common "space" or "scale," ensuring that the object features seen from a visual perspective and the object features described in words can be accurately matched.
[0103] Cross-attention is a special type of attention mechanism. It involves mutual attention between data from two different modalities (visual and textual modalities in this embodiment). Cross-attention allows visual features to "attention" to relevant parts of textual semantic features, and vice versa, thereby better capturing the correlation and interaction information between the two. Optionally, visual features and textual semantic features are jointly injected into the cross-attention module, thus injecting each textual semantic feature into a corresponding visual feature to obtain visual features with textual semantics.
[0104] In some embodiments, an electronic device may encode multiple reference information separately using a multi-instance feature encoder. This multi-instance feature encoder includes a visual encoder and a text encoder. The visual encoder, such as the CLIP (Contrastive Language-Image Pretraining) model and multilayer perceptual resampling, is used to extract visual features of the reference image. The text encoder, such as the T5 (Text-To-Text Transfer Transformer) model, is used to extract textual semantic features describing the text.
[0105] For example, see Figure 4 As shown, Figure 4 This is a schematic diagram illustrating a process for encoding reference information according to an exemplary embodiment. For example... Figure 4 As shown, the reference image from the reference information is input into the visual encoder to obtain visual features. The descriptive text from the reference information is input into the text encoder to obtain semantic text features. The visual features and semantic text features are then input into the cross-attention module to obtain semantic visual features.
[0106] In step S305, the target video is generated based on text prompt features and multiple semantic visual features.
[0107] In this embodiment of the disclosure, the electronic device injects text prompt features and multiple semantic visual features into a diffusion model. Based on the diffusion model, starting from noise, the various features are fused to obtain the target video.
[0108] In some embodiments, the electronic device first fuses the video features of the diffusion model with the text prompt features to obtain intermediate fused features. The video features of the diffusion model are the hidden features in the video generation process. Then, the electronic device sequentially fuses the intermediate fused features with multiple semantic visual features to obtain video features. Finally, the electronic device decodes the video features to obtain the target video. The diffusion model can generate the target video through multiple iterations. For any given iteration, the video features obtained after the previous iteration are the video features of the aforementioned diffusion model. By sequentially fusing the hidden video features of the diffusion model with the text prompt features and multiple semantic visual features during video generation, and then decoding to obtain the target video, the multiple iterations of the diffusion model can fully integrate multi-faceted feature information, effectively improving the accuracy, richness, and relevance to user needs in target video generation.
[0109] In some embodiments, electronic devices can use a cross-attention mechanism to fuse video features and text prompt features from a diffusion model to obtain intermediate fused features. This intermediate fused feature, obtained by fusing video features and text prompt features from a diffusion model using a cross-attention mechanism, enables the electronic device to accurately capture the correlation and interaction information between the two, effectively combining relevant characteristics, thereby laying a solid foundation for subsequently generating more demanding and higher-quality target videos.
[0110] In some embodiments, a novel feature injection module is introduced into the diffusion model. This module acquires and injects multiple semantic visual features into the diffusion model, ensuring that the features of the diffusion model are effectively integrated with visual and textual prompt features. Accordingly, the electronic device acquires multiple semantic visual features sequentially through the feature injection module. For each acquired semantic visual feature, the electronic device fuses it with intermediate fusion features to obtain intermediate video features. Finally, upon completion of the fusion, the video features are obtained. The core of the feature injection module is the introduction of a new cross-attention module into the traditional Transformer Block. This allows the diffusion model to adaptively learn the ability to distinguish different concepts based on the semantics of the conceptual features, ensuring that each concept in the generated video is accurately reproduced. By introducing a feature injection module into the diffusion model, multiple semantic visual features can be effectively acquired and integrated, ensuring deep fusion between these features and visual and textual prompt features. When the electronic device operates according to this process, it can generate video features with high fusion and better meet the requirements, thereby improving the quality and effect of the target video generation.
[0111] It should be noted that, in order to make the video generation scheme provided in this disclosure easier to understand, please refer to... Figure 5 As shown, Figure 5 This is a flowchart illustrating a video generation method according to an exemplary embodiment. Figure 5 As shown, this video generation scheme is implemented through a video generation model, which consists of three parts: an input layer, a processing layer, and an output layer. The input layer is used to input text prompts and multiple reference information. The processing layer includes a text feature encoder, a multi-instance feature encoder, and a diffusion model. The multi-instance feature encoder includes a visual encoder and a text encoder. The diffusion model includes a video diffusion model and a feature injection module. The output layer outputs video features and decodes these features using a latent variable decoder to generate the target video.
[0112] This disclosure provides a video generation scheme that, by encoding multiple pieces of reference information separately, binds the features of the reference object instance corresponding to each piece of reference information with the features of the descriptive text, and debindes the features of different reference information to avoid confusion between them. This allows the generated video to accurately reproduce the attributes of each reference object instance. Furthermore, since the above scheme does not rely on region constraints, the generated video exhibits diversity and smooth motion.
[0113] Figure 6 This is a block diagram illustrating a video generation apparatus according to an exemplary embodiment. Figure 6 As shown, the device includes a first encoding unit 601, a second encoding unit 602, and a generation unit 603.
[0114] The first encoding unit 601 is configured to encode the text prompt information to obtain text prompt features;
[0115] The second encoding unit 602 is configured to encode multiple reference information respectively to obtain multiple semantic visual features. Each reference information includes a reference image and descriptive text. The reference image includes a reference object instance, and the descriptive text is used to describe the corresponding reference object instance.
[0116] The generation unit 603 is configured to generate a target video based on the text prompt features and the plurality of semantic visual features.
[0117] In some embodiments, the second encoding unit 602 is configured to, for any reference information, perform visual encoding on a reference image in the reference information to obtain visual features of a reference object instance; perform semantic encoding on descriptive text in the reference information to obtain textual semantic features of the descriptive text; and fuse visual features and textual semantic features to obtain semantic visual features of the reference information.
[0118] In some embodiments, the first encoding unit 601 is configured to align visual features and textual semantic features; and to fuse the aligned visual features and textual semantic features based on a cross-attention mechanism.
[0119] In some embodiments, the generation unit 603 is configured to fuse the video features of the diffusion model with the text prompt features to obtain intermediate fused features, wherein the video features of the diffusion model are the hidden layer features in the video generation process; to fuse the intermediate fused features with multiple semantic visual features in sequence to obtain video features; and to decode the video features to obtain the target video.
[0120] In some embodiments, the generation unit 603 is configured to fuse video features and text prompt features of a diffusion model through a cross-attention mechanism to obtain intermediate fused features.
[0121] In some embodiments, the generation unit 603 is configured to sequentially acquire multiple semantic visual features through a feature injection module; for each semantic visual feature acquired, the semantic visual feature is fused with an intermediate fusion feature to obtain an intermediate video feature; and in response to the completion of fusion, a video feature is obtained.
[0122] In some embodiments, the first encoding unit 601 is further configured to acquire input text prompt information; perform compliance verification on the text prompt information; and display prompt information if the text prompt information fails the verification.
[0123] This disclosure provides a video generation apparatus that, by encoding multiple pieces of reference information separately, binds the features of the reference object instance corresponding to each piece of reference information with the features of the descriptive text, and debindes the features of different reference information to avoid confusion between them. This allows the generated video to accurately reproduce the attributes of each reference object instance. Furthermore, since the above scheme does not rely on region constraints, the generated video exhibits diversity and smooth motion.
[0124] It should be noted that the video generation apparatus provided in the above embodiments is only an example of the division of the above functional units. In practical applications, the above functions can be assigned to different functional units as needed, that is, the internal structure of the electronic device can be divided into different functional units to complete all or part of the functions described above. In addition, the video generation apparatus and the video generation method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.
[0125] Regarding the video generation apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.
[0126] In the embodiments of this disclosure, the electronic device can be a terminal or a server. When the electronic device is a terminal, the terminal acts as the execution subject to implement the technical solutions provided in the embodiments of this disclosure; when the electronic device is a server, the server acts as the execution subject to implement the technical solutions provided in the embodiments of this disclosure; or, the technical solutions provided in this disclosure can be implemented through interaction between the terminal and the server. This disclosure does not limit the scope of the embodiments.
[0127] Figure 7 This is a block diagram illustrating an electronic device according to an exemplary embodiment. Typically, the electronic device 700 includes a processor 701 and a memory 702.
[0128] Processor 701 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. Processor 701 may be implemented using at least one hardware form selected from DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). Processor 701 may also include a main processor and a coprocessor. The main processor, also known as a CPU (Central Processing Unit), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, processor 701 may integrate a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content required to be displayed on the screen. In some embodiments, processor 701 may also include an AI (Artificial Intelligence) processor, which is used to handle computational operations related to machine learning.
[0129] The memory 702 may include one or more computer-readable storage media, which may be non-transitory. The memory 702 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In some embodiments, the non-transitory computer-readable storage media in the memory 702 are used to store at least one program code, which is executed by the processor 701 to implement the video generation method provided in the method embodiments of this disclosure.
[0130] In some embodiments, the electronic device 700 may optionally include a peripheral device interface 703 and at least one peripheral device. The processor 701, memory 702, and peripheral device interface 703 can be connected via a bus or signal line. Each peripheral device can be connected to the peripheral device interface 703 via a bus, signal line, or circuit board. Specifically, the peripheral device includes at least one of the following: a radio frequency circuit 704, a display screen 705, a camera assembly 706, an audio circuit 707, and a power supply 708.
[0131] Peripheral device interface 703 can be used to connect at least one I / O (Input / Output) related peripheral device to processor 701 and memory 702. In some embodiments, processor 701, memory 702 and peripheral device interface 703 are integrated on the same chip or circuit board; in some other embodiments, any one or two of processor 701, memory 702 and peripheral device interface 703 can be implemented on separate chips or circuit boards, which is not limited in this embodiment.
[0132] The radio frequency (RF) circuit 704 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The RF circuit 704 communicates with communication networks and other communication devices via electromagnetic signals. The RF circuit 704 converts electrical signals into electromagnetic signals for transmission, or converts received electromagnetic signals back into electrical signals. Optionally, the RF circuit 704 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, etc. The RF circuit 704 can communicate with other electronic devices through at least one wireless communication protocol. This wireless communication protocol includes, but is not limited to: metropolitan area networks (MANs), various generations of mobile communication networks (2G, 3G, 4G, and 5G), wireless local area networks (WLANs), and / or WiFi (Wireless Fidelity) networks. In some embodiments, the RF circuit 704 may also include circuitry related to NFC (Near Field Communication), which is not limited in this disclosure.
[0133] Display screen 705 is used to display a UI (User Interface). This UI may include graphics, text, icons, videos, and any combination thereof. When display screen 705 is a touch display screen, it also has the ability to collect touch signals on or above its surface. These touch signals can be input as control signals to processor 701 for processing. In this case, display screen 705 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there may be one display screen 705, which serves as the front panel of electronic device 700; in other embodiments, there may be at least two display screens 705, respectively disposed on different surfaces of electronic device 700 or in a folded design; in still other embodiments, display screen 705 may be a flexible display screen, disposed on a curved or folded surface of electronic device 700. Furthermore, display screen 705 may be configured as a non-rectangular, irregular shape, i.e., a non-rectangular screen. Display screen 705 may be made of materials such as LCD (Liquid Crystal Display) or OLED (Organic Light-Emitting Diode).
[0134] Camera assembly 706 is used to acquire images or videos. Optionally, camera assembly 706 includes a front-facing camera and a rear-facing camera. Typically, the front-facing camera is located on the front panel of the electronic device, and the rear-facing camera is located on the back of the electronic device. In some embodiments, there are at least two rear-facing cameras, which are any one of a main camera, a depth-sensing camera, a wide-angle camera, and a telephoto camera, to achieve background blurring by fusion of the main camera and the depth-sensing camera, panoramic shooting by fusion of the main camera and the wide-angle camera, VR (Virtual Reality) shooting, or other fusion shooting functions. In some embodiments, camera assembly 706 may also include a flash. The flash can be a single-color temperature flash or a dual-color temperature flash. A dual-color temperature flash refers to a combination of a warm-light flash and a cool-light flash, which can be used for light compensation at different color temperatures.
[0135] The audio circuit 707 may include a microphone and a speaker. The microphone is used to collect sound waves from the user and the environment, converting the sound waves into electrical signals that are input to the processor 701 for processing, or input to the radio frequency circuit 704 for voice communication. For stereo sound acquisition or noise reduction purposes, multiple microphones may be used, each located in a different part of the electronic device 700. The microphone may also be an array microphone or an omnidirectional microphone. The speaker is used to convert the electrical signals from the processor 701 or the radio frequency circuit 704 into sound waves. The speaker may be a conventional diaphragm speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can convert electrical signals not only into audible sound waves but also into inaudible sound waves for purposes such as distance measurement. In some embodiments, the audio circuit 707 may also include a headphone jack.
[0136] Power supply 708 is used to supply power to the various components in electronic device 700. Power supply 708 can be AC power, DC power, a disposable battery, or a rechargeable battery. When power supply 708 includes a rechargeable battery, the rechargeable battery can support wired charging or wireless charging. The rechargeable battery can also be used to support fast charging technology.
[0137] Those skilled in the art will understand that Figure 7 The structure shown does not constitute a limitation on the electronic device 700, and may include more or fewer components than shown, or combine certain components, or use different component arrangements.
[0138] In an exemplary embodiment, a computer-readable storage medium including instructions is also provided, such as a memory 702 including instructions, which can be executed by a processor 701 of an electronic device 700 to complete the video generation method described above. Optionally, the computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.
[0139] A computer program product includes a computer program that, when executed by a processor, implements the video generation method described above.
[0140] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the following claims.
[0141] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.
Claims
1. A video generation method, characterized in that, The method includes: The text prompt information is encoded to obtain text prompt features, which are used to describe various aspects of the video to be generated; Multiple reference information are obtained, each of which includes a reference image and descriptive text. The reference image includes a reference object instance, and the descriptive text is used to describe the corresponding reference object instance. The reference object instance is an individual object with actual form and characteristics. For any reference information, the reference image in the reference information is visually encoded to obtain the visual features of the reference object instance, and the descriptive text in the reference information is semantically encoded to obtain the text semantic features of the descriptive text. The visual features and the text semantic features are fused to obtain the semantic visual features of the reference information. Encode multiple reference information separately to obtain multiple semantic visual features; The target video is generated based on the text prompt features and the multiple semantic visual features.
2. The video generation method according to claim 1, characterized in that, The process of fusing the visual features and the textual semantic features to obtain the semantic visual features of the reference information includes: Align the visual features and the text semantic features; Based on the cross-attention mechanism, the aligned visual features and the text semantic features are fused together.
3. The video generation method according to claim 1, characterized in that, The process of generating a target video based on the text prompt features and the multiple semantic visual features includes: The video features of the diffusion model are fused with the text prompt features to obtain intermediate fused features, wherein the video features of the diffusion model are the hidden layer features in the video generation process; The intermediate fusion features are sequentially fused with the multiple semantic visual features to obtain video features; The video features are decoded to obtain the target video.
4. The video generation method according to claim 3, characterized in that, The process of fusing the video features of the diffusion model with the text prompt features to obtain intermediate fused features includes: The intermediate fused features are obtained by fusing the video features of the diffusion model and the text prompt features through a cross-attention mechanism.
5. The video generation method according to claim 3, characterized in that, The step of sequentially fusing the intermediate fusion features with the plurality of semantic visual features to obtain video features includes: Multiple semantic visual features are obtained sequentially through the feature injection module; For each semantic visual feature acquired, the semantic visual feature is fused with the intermediate fusion feature to obtain intermediate video features; Upon completion of the fusion process, the video features are obtained.
6. The video generation method according to any one of claims 1-5, characterized in that, The method further includes: Get the input text prompt information; Perform compliance verification on the text prompt information; If the text prompt fails verification, a prompt message will be displayed.
7. A video generation apparatus, characterized in that, The device includes: The first encoding unit is configured to encode text prompt information to obtain text prompt features, wherein the text prompt information is used to describe various aspects of the video to be generated; The second encoding unit is configured to acquire multiple pieces of reference information, each piece of reference information including a reference image and descriptive text. The reference image includes a reference object instance, and the descriptive text describes the corresponding reference object instance, which is an individual object with actual form and characteristics. For any piece of reference information, the reference image in the reference information is visually encoded to obtain the visual features of the reference object instance, and the descriptive text in the reference information is semantically encoded to obtain the textual semantic features of the descriptive text. The visual features and the textual semantic features are fused to obtain the semantic visual features of the reference information. Multiple pieces of reference information are encoded separately to obtain multiple semantic visual features. The generation unit is configured to generate a target video based on the text prompt features and the plurality of semantic visual features.
8. An electronic device, characterized in that, The electronic device includes: One or more processors; Memory used to store the executable program code of the processor; The processor is configured to execute the program code to implement the video generation method as described in any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, When the instructions in the computer-readable storage medium are executed by the processor of the electronic device, the electronic device is able to perform the video generation method as described in any one of claims 1 to 6.
10. A computer program product comprising a computer program that, when executed by a processor, implements the video generation method as described in any one of claims 1 to 6.