Video generation method and device, equipment and storage medium

By filtering out features with low similarity to the video description text from the style reference image, and generating videos with local and global style information, the content leakage problem is solved and video generation with consistent content and stylized effects is achieved.

CN120343299APending Publication Date: 2025-07-18BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411756556.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-02
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

Existing video generation technology injects all hidden features of the style reference image when generating video, resulting in content leakage in non-style elements.

Method used

By filtering out features with low similarity to the video description text from multiple image area features of the style reference image, and generating target videos based on local and global style information, avoiding content leakage and retaining style features.

Benefits of technology

The generated video content is consistent with the video description text, the style is consistent with the style reference image, and the content leakage is avoided, and the stylization effect is good.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120343299A_ABST
    Figure CN120343299A_ABST
Patent Text Reader

Abstract

The invention provides a video generation method and device, equipment and a storage medium, and belongs to the technical field of computers. The method comprises the steps of determining a plurality of target features from a plurality of image region features of a style reference image based on a video description text; determining local style information of the style reference image based on the plurality of target features and the video description text; and generating a target video based on the video description text, the local style information and the global style information of the style reference image. According to the method, the image region features with relatively high similarity with the video description text are discarded from the plurality of image region features of the style reference image, and the image region features with relatively low similarity with the video description text are left, so that the texture features in the style reference image can be extracted, and content leakage can be avoided. And the global style information of the style reference image and the video description text are combined, so that a target video which meets the content described by the video description text and has a relatively good stylization effect can be generated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technologies, and particularly to a video generation method, apparatus, device, and storage medium. Background Art

[0002] Text-to-video refers to a technical solution for generating a video from text. In this solution, the content, structure, action pattern, etc. of the video are determined by the text. For example, in Video-Composer, this method generates a video based on multiple control conditions, including control conditions for controlling the video style, such as a style reference image. However, in the above solution, all hidden layer features of the style reference image are injected into the video generation model when generating the video, resulting in non-style elements in the style reference image appearing in the generated video, thus causing a serious content leakage problem. Therefore, there is an urgent need for a new video generation solution. Summary of the Invention

[0003] The present disclosure provides a video generation method, apparatus, device, and storage medium. Through this method, a target video that meets the content described in the video description text and has a good stylization effect can be generated.

[0004] According to one aspect of the embodiments of the present disclosure, a video generation method is provided, the method including:

[0005] Based on a video description text, determining a plurality of target features from a plurality of image region features of a style reference image, the similarity between the plurality of target features and the video description text being lower than that of other image region features among the plurality of image region features;

[0006] Based on the plurality of target features, determining local style information of the style reference image;

[0007] Based on the video description text, the local style information, and global style information of the style reference image, generating a target video, the content of the target video being consistent with the content of the video description text, and the style of the target video being consistent with the style reference image.

[0008] According to another aspect of the embodiments of the present disclosure, a video generation apparatus is provided, the apparatus including:

[0009] A first determination unit configured to determine a plurality of target features from a plurality of image region features of a style reference image based on a video description text, the similarity between the plurality of target features and the video description text being lower than that of other image region features among the plurality of image region features;

[0010] A second determination unit configured to determine local style information of the style reference image based on the plurality of target features;

[0011] A generation unit, configured to generate a target video based on the video description text, the local style information, and the global style information of the style reference image, where the content of the target video is consistent with the content of the video description text, and the style of the target video is consistent with the style reference image.

[0012] In some embodiments, the first determination unit is configured to extract the plurality of image region features from the style reference image based on a multimodal training model; determine the plurality of target features according to a target ratio based on the similarity between each image region feature and the video description text, where the target ratio is the proportion of the plurality of target features in the plurality of image region features.

[0013] In some embodiments, the local style information of the style reference image is extracted by a local projection module; the apparatus further includes: a training unit, configured to: obtain a plurality of sample vocabulary vectors; learn the plurality of sample vocabulary vectors based on the local projection module; the second determination unit is configured to process the plurality of target features based on a self-attention mechanism through the local projection module to obtain the local style information of the style reference image.

[0014] In some embodiments, the generation unit is configured to splice the local style information and the global style information of the style reference image to obtain the style information of the style reference image; inject the style information into a video diffusion model based on a style cross-attention mechanism; input the video description text into the video diffusion model based on a text cross-attention mechanism; and generate the target video starting from a noise video through the video diffusion model.

[0015] In some embodiments, the generation unit is configured to input the video description text, the local style information, and the global style information of the style reference image into a video diffusion model; adjust the weights of the temporal processing layer in the video diffusion model based on a motion adaptation module; and generate the target video based on the adjusted video diffusion model.

[0016] In some embodiments, the generation unit is further configured to obtain a reference video, where the reference video is used to provide layout and structure information; convert the reference video into a grayscale thumbnail video; input the video description text, the local style information, and the global style information of the style reference image into the video diffusion model to obtain video hidden layer features; splice the features of the grayscale thumbnail video and the video hidden layer features to obtain intermediate video features; and generate the target video based on the intermediate video features.

[0017] In some embodiments, the global style information of the style reference image is extracted by a global projection module;

[0018] The apparatus further includes:

[0019] An acquisition unit configured to acquire a set of sample paired images, the set of sample paired images including a plurality of paired image groups, and each paired image group including two images with the same style but different contents;

[0020] A feature extraction unit configured to, for any paired image group, input the first image and the second image in the paired image group into a feature extraction network respectively to obtain the image feature of the first image and the image feature of the second image;

[0021] A feature processing unit configured to input the image feature of the first image and the image feature of the second image into a multi-layer perceptron network respectively to obtain the global style feature of the first image and the style feature of the second image;

[0022] A training unit configured to train the global projection module based on the global style feature of the first image and the global style feature of the second image.

[0023] According to another aspect of the embodiments of the present disclosure, there is provided an electronic device, which includes:

[0024] One or more processors;

[0025] A memory for storing executable program code of the processor;

[0026] Wherein, the processor is configured to execute the program code to implement the above video generation method.

[0027] According to another aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium, when the instructions in the computer-readable storage medium are executed by a processor of an electronic device, enabling the electronic device to execute the above video generation method.

[0028] According to another aspect of the embodiments of the present disclosure, there is provided a computer program product, including a computer program, and when the computer program is executed by a processor, implementing the above video generation method.

[0029] Embodiments of the present disclosure provide a video generation solution. By discarding image region features with a relatively high similarity to the video description text from multiple image region features of a style reference image, the remaining image region features with a relatively low similarity to the video description text are obtained, so that both texture features in the style reference image can be extracted and content leakage can be avoided. Combining the global style information of the style reference image and the video description text can generate a target video that meets the content described in the video description text and has a good stylization effect.

[0030] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and do not limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure, and do not constitute an improper limitation of the present disclosure.

[0032] Figure 1 is a schematic diagram of an implementation environment of a video generation method shown according to an exemplary embodiment.

[0033] Figure 2 is a flowchart of a video generation method shown according to an exemplary embodiment.

[0034] Figure 3 is a flowchart of another video generation method shown according to an exemplary embodiment.

[0035] Figure 4 is a schematic diagram of a paired image group provided according to an exemplary embodiment.

[0036] Figure 5 is a video generation flowchart provided according to an exemplary embodiment.

[0037] Figure 6 is a block diagram of a video generation device shown according to an exemplary embodiment.

[0038] Figure 7 is a block diagram of an electronic device shown according to an exemplary embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0039] To enable those of ordinary skill in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings.

[0040] It should be noted that the terms "first", "second", etc. in the description, claims and the above drawings of the present disclosure are used to distinguish similar objects and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0041] It should be noted that the information (including but not limited to user equipment information, user personal information, etc.), data (including but not limited to data for analysis, stored data, displayed data, etc.) and signals involved in the present disclosure are all authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data need to comply with relevant laws, regulations and standards of relevant countries and regions. For example, the images involved in the present disclosure are all obtained under full authorization.

[0042] Figure 1 It is a schematic diagram of an implementation environment of a video generation method shown according to an exemplary embodiment. Refer to Figure 1 In this implementation environment, it specifically includes: a terminal 101 and a server 102. The terminal 101 can be connected to the server 102 through a wireless network or a wired network.

[0043] The terminal 101 can be at least one of devices such as a smart phone, a smart watch, a desktop computer, a laptop computer, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 (Moving Picture Experts Group Audio Layer IV) player, and a laptop portable computer. An application can be installed and run on the terminal 101, and this application is used to generate a corresponding video based on the video description text input by the user. This application is associated with the server 102, and the server 102 provides background services to the terminal 101.

[0044] The terminal 101 can generally refer to one of multiple terminals, and the terminal 101 is used as an example in this embodiment. Those skilled in the art can know that the number of the above terminals can be more or less. For example, the above terminals can be several, or the above terminals can be dozens or hundreds, or a larger number. The present disclosure embodiments do not limit the number and device types of the terminals.

[0045] The server 102 is at least one of a single server, multiple servers, a cloud computing platform, and a virtualization center. Optionally, the number of the above-mentioned servers may be more or less, and the embodiments of the present disclosure do not limit this. Of course, the server 102 may further include other functional servers to provide more comprehensive and diverse services. In some embodiments, the server 102 undertakes the main computing work, and the terminal 101 undertakes the secondary computing work; or, the server 102 undertakes the secondary computing work, and the terminal 101 undertakes the main computing work; or, the server 102 and the terminal 101 adopt a distributed computing architecture for collaborative computing. The server 102 can be connected to the terminal 101 and other terminals through a wireless network or a wired network. Optionally, the number of the above-mentioned servers may be more or less, and the embodiments of the present disclosure do not limit this.

[0046] Figure 2 is a flowchart of a video generation method shown according to an exemplary embodiment, as Figure 2 shown, and this method is executed by an electronic device and includes the following steps:

[0047] In step S201, based on the video description text, multiple target features are determined from multiple image region features of the style reference image.

[0048] In the embodiments of the present disclosure, the similarity between the multiple target features and the video description text is lower than that of other image region features among the multiple image region features. That is, from the multiple image region features, some image region features with a higher similarity to the video description text are removed, and the remaining ones are the target features.

[0049] In step S202, based on the multiple target features, the local style information of the style reference image is determined.

[0050] In the embodiments of the present disclosure, the target features include the texture features of the style reference image. The texture features can be learned from the multiple target features, and the above-mentioned texture features are used as the local style information of the style reference image.

[0051] In step S203, based on the video description text, the local style information, and the global style information of the style reference image, a target video is generated.

[0052] In the embodiments of the present disclosure, since the target video is generated based on the video description text, the local style information, and the global style information of the style reference image, the content of the target video is consistent with the content of the video description text, and the style of the target video is consistent with that of the style reference image. That is, the target video has the image style of the style reference image.

[0053] The embodiments of the present disclosure provide a video generation solution. By discarding the image region features with a high similarity to the video description text from multiple image region features of the style reference image, the remaining image region features with a low similarity to the video description text are obtained, so that both the texture features in the style reference image can be extracted and content leakage can be avoided. Combining the global style information of the style reference image and the video description text, a target video that meets the content described in the video description text and has a good stylization effect can be generated.

[0054] In some embodiments, based on the video description text, determining multiple target features from multiple image region features of the style reference image includes:

[0055] Extracting multiple image region features from the style reference image based on a multimodal training model;

[0056] Based on the similarity between each image region feature and the video description text, determining multiple target features according to a target ratio, where the target ratio is the proportion of the multiple target features in the multiple image region features.

[0057] In the embodiments of the present disclosure, by extracting image region features based on a multimodal training model and screening out multiple target features according to the target ratio determined based on the similarity to the video description text, it helps to accurately extract the features whose correlation with the video description text meets specific requirements, and both texture features can be extracted and content leakage can be avoided.

[0058] In some embodiments, the local style information of the style reference image is extracted through a local projection module; the method further includes:

[0059] Obtaining multiple sample vocabulary vectors;

[0060] Learning multiple sample vocabulary vectors based on the local projection module;

[0061] Based on the multiple target features and the video description text, determining the local style information of the style reference image includes:

[0062] Through the local projection module, processing the multiple target features based on the self-attention mechanism to obtain the local style information of the style reference image.

[0063] In the embodiments of the present disclosure, the electronic device, with the help of the self-attention mechanism, learns sample vocabulary vectors through the local projection module and processes the learned content together with the target features, which can more accurately and effectively determine the local style information of the style reference image.

[0064] In some embodiments, based on the video description text, the local style information, and the global style information of the style reference image, generating a target video includes:

[0065] Splice the local style information and the global style information of the style reference image to obtain the style information of the style reference image;

[0066] Based on the style cross-attention mechanism, inject the style information into the video diffusion model;

[0067] Based on the text cross-attention mechanism, input the video description text into the video diffusion model;

[0068] Generate the target video starting from the noise video through the video diffusion model.

[0069] In the embodiments of the present disclosure, by integrating the video description text, the local style information, and the global style information of the style reference image, and injecting relevant information into the video diffusion model through operations such as splicing, the style cross-attention mechanism, and the text cross-attention mechanism, the target video is generated starting from the noise video, so that the target video meets the requirements of the above information, thereby improving the quality and effect of the generated video.

[0070] In some embodiments, generating the target video based on the video description text, the local style information, and the global style information of the style reference image includes:

[0071] Input the video description text, the local style information, and the global style information of the style reference image into the video diffusion model;

[0072] Adjust the weights of the temporal processing layer in the video diffusion model based on the motion adaptation module;

[0073] Generate the target video based on the adjusted video diffusion model.

[0074] In the embodiments of the present disclosure, by integrating the video description text, the local style information, and the global style information of the style reference image, and adjusting the weights of the temporal processing layer of the video diffusion model using the motion adaptation module, the video diffusion model can generate a target video that integrates multiple elements and adapts to the motion characteristics, improving the motion and stylization degree of the target video.

[0075] In some embodiments, the method further includes:

[0076] Obtain a reference video, where the reference video is used to provide layout and structure information;

[0077] Convert the reference video into a grayscale thumbnail video;

[0078] Generating the target video based on the video description text, the local style information, and the global style information of the style reference image includes:

[0079] Input the video description text, local style information, and global style information of the style reference image into the video diffusion model to obtain video hidden layer features;

[0080] Concatenate the features of the grayscale thumbnail video with the video hidden layer features to obtain intermediate video features;

[0081] Generate a target video based on the intermediate video features.

[0082] In the embodiments of the present disclosure, the electronic device controls the generation of video content by adopting the controlNet method, and uses the grayscale thumbnail video converted from the reference video as a guide to ensure temporal stability. It can effectively utilize the layout and structure information of the reference video to accurately control the presentation of the generated video, improving the quality and effect of the generated video.

[0083] In some embodiments, the global style information of the style reference image is extracted through a global projection module;

[0084] The method further includes:

[0085] Obtain a sample paired image set, which includes multiple paired image groups, and each paired image group includes two images with the same style but different contents;

[0086] For any paired image group, input the first image and the second image in the paired image group into the feature extraction network respectively to obtain the image features of the first image and the image features of the second image;

[0087] Input the image features of the first image and the image features of the second image into the multi-layer perceptron network respectively to obtain the global style features of the first image and the style features of the second image;

[0088] Train the global projection module based on the global style features of the first image and the global style features of the second image.

[0089] In the embodiments of the present disclosure, through specific training steps, the sample paired image set is used, and the paired images are processed by the feature extraction network and the multi-layer perceptron network to obtain their style features, so as to effectively extract the global style information of the style reference image, which helps subsequent related applications to better utilize the global style information to generate videos with specific styles.

[0090] The above Figure 2 shows a flowchart of a video generation method of the present disclosure. Next, the image processing solution provided by the present disclosure will be further elaborated. Figure 3 is a flowchart of another video generation method shown according to an exemplary embodiment. Refer to Figure 3 , this method is executed by an electronic device and includes the following steps:

[0091] In step S301, the global style information of the style reference image is obtained.

[0092] In the embodiments of the present disclosure, the style reference image is an image with a specific visual style. For example, Van Gogh's painting "The Starry Night" has unique color combinations (such as blue and yellow as the main colors), brushstroke styles (swirling brushstrokes), and composition methods (the layout of the sky and the village), etc. These characteristics together constitute the style of this painting.

[0093] The global style information refers to the features that can represent the style of the entire image. The global style information is the information that describes the image style as a whole. Optionally, the global style information includes the overall pattern of color distribution (such as the proportion of warm colors, the distribution range of cold colors, etc.), the overall characteristics of the texture (whether it is smooth, rough, or regular texture, etc.), and the overall rules of the shape and layout of the objects (such as whether the objects in the picture are symmetrically distributed or scattered).

[0094] In some embodiments, the global style information of the style reference image is extracted by a global projection module. Optionally, the training steps of the global projection module include: First, the electronic device obtains a sample paired image set, which includes multiple paired image groups, and each paired image group includes two images with the same style but different contents. Then, taking the training process of one paired image group as an example, for any paired image group, the electronic device inputs the first image and the second image in the paired image group into the feature extraction network respectively to obtain the image features of the first image and the image features of the second image. Then, the electronic device inputs the image features of the first image and the image features of the second image into the multi-layer perceptron network respectively to obtain the global style features of the first image and the global style features of the second image. Finally, the electronic device trains the global projection module based on the global style features of the first image and the global style features of the second image. By using the sample paired image set through specific training steps, and processing the paired images with the feature extraction network and the multi-layer perceptron network to obtain their style features, the effective extraction of the global style information of the style reference image is realized, which helps subsequent related applications to better utilize the global style information to generate videos with specific styles.

[0095] For example, the electronic device uses the model illusion model to generate multiple paired image groups. Among them, each paired image group includes at least two images. Refer to Figure 4 as shown Figure 4 is a schematic diagram of a paired image group provided according to an exemplary embodiment. As Figure 4As shown, an example is given where a paired image group includes two images. The two images in the paired image group are basically rearrangements of pixels, that is, they have the same style but different contents. For ease of explanation, the two images in a paired image group are referred to as Image A and Image B. Input Image A into a feature extraction network, such as a pre-trained CLIP (Contrastive Language-Image Pretraining) network, to obtain the image features of Image A. These image features include both content features and style features. Then, input the image features of Image A into a multi-layer perceptron network, that is, a few layers of MLP (Multilayer Perceptron) networks can be trained to obtain the global style features of Image A. The global style features of Image B are obtained in the same way and will not be elaborated here. Finally, the electronic device can use triplet loss to train the global projection module. The training objective is that the distance between the global style features of Image A and Image B is less than the distance between the global style features of Image A and the images in other paired image groups.

[0096] Optionally, refer to the following formula:

[0097] F image_embed = CLIPImageEncoder(A)

[0098] F global = MLPLayer(F image_embed )

[0099] where F image_embed represents the image features of Image A, and F global represents the global style features of Image A.

[0100] In some embodiments, the electronic device can obtain global style information related to colors through a color histogram. The electronic device can count the frequency of each color (in different color spaces, such as RGB, HSV, etc.) in the image through a color histogram. For example, in the RGB color space, for a landscape photo, by calculating the color histogram, the pixel distribution of each of the red, green, and blue channels can be obtained. If the proportion of pixels in the high-brightness area of the red channel is relatively high, it may mean that the overall photo is warmer in tone. By comparing the color histograms of the reference image and the target image, the global style information of the reference image in terms of colors can be obtained, such as the overall color brightness and the richness of colors.

[0101] In some embodiments, the electronic device can also obtain global style information related to color through color moments. Among them, color moments include the first moment (mean), the second moment (variance), and the third moment (skewness). The first color moment can represent the average intensity of the image color. For example, for an image dominated by gray tones, the values of the first moment of its RGB channels may be relatively close and relatively low. The second color moment describes the distribution range of colors. The larger the variance, the more dispersed the color distribution. The third color moment reflects the symmetry of the color distribution. These color moments can provide global style information in terms of color when combined.

[0102] In some embodiments, the electronic device can also obtain global style information related to texture through the gray-level co-occurrence matrix (GLCM). Among them, for gray-scale images, the GLCM can describe the spatial correlation characteristics of pixel gray levels in the image. The GLCM can characterize the texture by calculating the probability of two pixels appearing simultaneously at a certain distance and direction. For example, in an image with regular stripe textures, in the stripe direction, the probability that adjacent pixels have similar gray values is relatively high. Through the GLCM, overall characteristics of such textures can be obtained, such as the thickness of the texture (by changing the calculation distance parameter), the direction of the texture (by analyzing the GLCM in different directions), and other global texture style information.

[0103] In some embodiments, the electronic device can also obtain global style information related to texture by filtering methods, such as filtering the image using different filters (such as Gaussian filters, Laplacian filters, etc.). For example, a Gaussian filter can smooth the image. By observing the result after filtering, the roughness of the image texture can be understood. If the image changes little after Gaussian filtering, it indicates that the texture is relatively smooth; conversely, if the change is obvious, the texture may be relatively rough. These overall characteristics after filtering can be used as global style information in terms of texture.

[0104] In step S302, based on the video description text, multiple target features are determined from multiple image region features of the style reference image.

[0105] In the embodiments of the present disclosure, the video description text is used to describe the video to be generated, such as a literal description of aspects such as video content, atmosphere, emotion, etc. The video description text can convey many key information of the video. For example, if the video shows a quiet seaside sunset scene, the text may describe the golden sunlight shining on the sea level, the waves gently lapping on the beach, the gorgeous sunset glow in the sky, etc. The multiple image region features of the style reference image are the features of multiple image blocks after dividing the style reference image into image blocks, which are called multiple image region features. Optionally, the electronic device inputs the style reference image into the CLIP model to obtain multiple image region features of the style reference image. Then, according to the similarity between the image region features and the video description text, multiple target features are filtered out from the multiple image region features. Among them, the similarity between the multiple target features and the video description text is lower than that of other image region features in the multiple image region features.

[0106] In some embodiments, the electronic device filters multiple image region features according to a target ratio. Correspondingly, the electronic device extracts multiple image region features from the style reference image based on a multimodal pre-trained model. Then, the electronic device determines multiple target features according to the target ratio based on the similarity between each image region feature and the video description text, and the target ratio is the proportion of the multiple target features in the multiple image region features. By extracting image region features based on a multimodal training model and filtering out multiple target features according to the target ratio determined based on the similarity with the video description text, it helps to accurately extract features whose correlation with the video description text meets specific requirements, and can extract texture features and avoid content leakage.

[0107] Optionally, the electronic device can measure the similarity between each image region feature and the video description text based on a distance calculation method of feature vectors, such as cosine similarity, etc. Correspondingly, the electronic device can convert the image region features into a quantifiable feature vector form, and at the same time perform corresponding quantization processing on the video description text (such as converting the text into a vector through word vectors, etc.), and then calculate the distance or similarity value between the two vectors.

[0108] In step S303, based on the multiple target features, the local style information of the style reference image is determined.

[0109] In the embodiments of the present disclosure, the local style information includes the texture information of the style reference image, such as the oil painting brushstroke information of Van Gogh's starry sky.

[0110] In some embodiments, the local style information of the style reference image is extracted by a local projection module. The electronic device can determine the local style information based on the self-attention mechanism. The electronic device creates multiple learnable tokens, and obtains the above-mentioned multiple vocabulary vectors by learning the above tokens. The embodiments of the present disclosure do not limit this. Correspondingly, during training, the electronic device obtains multiple sample vocabulary vectors. Then, the electronic device learns the multiple sample vocabulary vectors based on the local projection module. During inference, the electronic device processes multiple target features based on the self-attention mechanism through the above local projection module to obtain the local style information of the style reference image.

[0111] It should be noted that by extracting the local style information of the style reference image, the electronic device can solve the problem that local texture information cannot be obtained from the global style information. For example, the electronic device first passes the image I through CLIP to obtain patch features, that is, image region features. Then, the electronic device discards 90% of the patches with high similarity according to the similarity between the video description text and the patch features to obtain F_fil_patch. Then, since the electronic device can use Q-Former to learn N tokens during training, which are represented as queries, during inference, the queries and F_fil_patch can be concatenated together to perform self-attention.

[0112] Correspondingly, the above process is shown in the following formula.

[0113] F image_embed ,F patch =CLIPImageEncoder(I)

[0114] F fil_patch =Filter(F patch )

[0115] F local =SELF-ATTENTION(CONCAT(query,F fil_path ))

[0116] Among them, F image_embed ,F patch represents the patch features obtained by passing the image I through CLIP, that is, the image region features. F fil_patch represents F_fil_patch obtained by discarding 90% of the patches with high similarity. F local represents the local style information. SELF-ATTENTION() represents the self-attention mechanism. CONCAT( ) represents concatenation.

[0117] In step S304, the video description text, the local style information, and the global style information of the style reference image are input into the video diffusion model to obtain video hidden layer features.

[0118] In the embodiments of the present disclosure, the electronic device may splice the local style information and the global style information together to obtain the style information of the style reference image. Among them, the style information of the style reference image can be expressed as F_style. Then, under the guidance of the video description text, based on the video diffusion model, starting from the noise video, referring to the above style information, the electronic device gradually generates a video. Among them, the video diffusion model can obtain video hidden layer features during the iterative process of generating the video.

[0119] For example, after receiving the above three inputs, the video diffusion model can perform semantic analysis on the video description text to extract the key semantic information therein; perform operations such as feature fusion and transformation on the local style information and the global style information, and integrate and process these style information from different sources and different levels so that they can cooperate better with the information conveyed by the video description text.

[0120] It should be noted that the electronic device can train the above video diffusion model based on the style cross-attention mechanism and the text cross-attention mechanism. Correspondingly, the video diffusion model can learn the style information and the video description text during the training process to gradually generate a video, and thus output the video. Correspondingly, as shown in the following formula.

[0121] F out = TCA(F in , F text ) + SCA(F in , F style )

[0122] Among them, F out represents the output of the video diffusion model. TCA() represents text cross-attention. F in represents the input of the video diffusion model. SCA() represents style cross-attention. F text represents the video description text. F style represents the style information.

[0123] In step S305, based on the motion adaptation module, the weights of the temporal processing layer in the video diffusion model are adjusted.

[0124] In the embodiments of the present disclosure, the electronic device enhances the motion degree and stylization degree of the generated video by adjusting the weights of the temporal processing layer in the video diffusion model.

[0125] Optionally, a motion adaptation module is trained, which is used to adjust the weights \(W\in\{W Q , W K , W V \}\) of the temporal processing layer in the video diffusion model. The adjustment method is shown in the following formula.

[0126]

[0127] Wherein, represents the adjusted weight. The two A matrices are learnable parameters. When \(\alpha = 1\), training is performed using a completely static video. When \(\alpha = 0\), training is performed using a normal video. Then, when inferring, setting \(\alpha=-1\) can enhance the motion ability more. Moreover, since we use a realistic video for training when \(\alpha = 1\), setting \(\alpha=-1\) can also make the stylization degree stronger. It should be noted that the above formula decomposes the weight matrix into the product of two smaller matrices, and only trains these two small matrices while fixing the weights of the original pre-trained model. The advantage of doing this is to significantly reduce the number of training parameters, speed up the training, and reduce the video memory occupancy.

[0128] In step S306, the features of the grayscale thumbnail video are concatenated with the video hidden layer features to obtain intermediate video features.

[0129] In the embodiments of the present disclosure, the electronic device can control the content of the generated video in the way of controlNet (control network). To ensure the temporal stability during the control process, the embodiments of the present disclosure use video thumbnails as a guide. Correspondingly, the electronic device acquires a reference video, which is used to provide layout and structure information. Then, the electronic device converts the reference video into a grayscale thumbnail video. Among them, the thumbnail not only smooths the details, gives space for style changes, but also gives sufficient layout and object content hints. And considering that the color of the input reference video may affect the result of style control, the input video is changed into a grayscale thumbnail video. By adopting the controlNet method to control the generated video content and using the grayscale thumbnail video converted from the reference video as a guide to ensure temporal stability, the electronic device can effectively utilize the layout and structure information of the reference video to accurately control the presentation of the generated video and improve the quality and effect of the generated video.

[0130] In step S307, based on the intermediate video features, a target video is generated through the adjusted video diffusion model.

[0131] In an embodiment of the present disclosure, an electronic device may inject style information into a video diffusion model based on a style cross-attention mechanism. The electronic device may input video description text into the video diffusion model based on a text cross-attention mechanism. Correspondingly, the video diffusion model learns style information and video description text based on intermediate video features through multiple iterations to gradually generate a video and obtain a target video.

[0132] Optionally, as shown in the following formula.

[0133] F out = TCA(F in , F text ) + SCA(F in , F style )

[0134] where F out represents the output of the video diffusion model. TCA( ) represents text cross-attention. F in represents the input of the video diffusion model. SCA( ) represents style cross-attention. F text represents the video description text. F style represents the style information. It should be noted that during multiple iterations of the video diffusion model, the feature input for this iteration is the intermediate video feature output by the model at the end of the previous iteration.

[0135] It should be noted that to make the video generation solution provided in the embodiment of the present disclosure easier to understand, refer to Figure 5 shown in Figure 5 which is a video generation flowchart provided according to an exemplary embodiment. As shown in Figure 5As shown in the figure, it can be divided into four steps. Step 1: Train a global projection module. Input the images in the paired image group into CLIP for processing, and then input the output of CLIP into the global projection module to obtain the global style features of the images. Based on these global style features, model training is carried out. Step 2: During inference, input the style reference image into CLIP for processing to obtain multiple image region features. After randomly discarding 90% according to the similarity with the text, and then processing through Q-Former, local style information is obtained. Concatenate the local style information and the global style information to obtain the style information. Based on the video diffusion model, process the video description text and the style information to generate a stylized video, that is, the target video. Step 3: During inference, the weights of the temporal processing layer in the video diffusion model can also be adjusted through a motion adaptation module. Among them, the motion adaptation module trains the parameter matrix in the motion adaptation module through a static video or a motion video. Step 4: During inference, control can also be performed through a control network. Correspondingly, input the reference video into the video diffusion model, and during the process of generating the video, the layout and structure of the reference video can be learned.

[0136] The embodiment of the present disclosure provides a video generation scheme. By discarding the image region features with a higher similarity to the video description text from the multiple image region features of the style reference image, the remaining image region features with a lower similarity to the video description text can be obtained, so that both the texture features in the style reference image can be extracted and content leakage can be avoided. Combining the global style information of the style reference image and the video description text, a target video that meets the content described in the video description text and has a good stylized effect can be generated.

[0137] Figure 6 is a block diagram of a video generation device shown according to an exemplary embodiment. As Figure 6 shown, the device includes: a first determination unit 601, a second determination unit 602, and a generation unit 603.

[0138] The first determination unit 601 is configured to determine multiple target features from the multiple image region features of the style reference image based on the video description text, and the similarity of the multiple target features to the video description text is lower than that of other image region features in the multiple image region features;

[0139] The second determination unit 602 is configured to determine the local style information of the style reference image based on the multiple target features;

[0140] The generation unit 603 is configured to generate a target video based on the video description text, the local style information, and the global style information of the style reference image. The content of the target video is consistent with the content of the video description text, and the style of the target video is consistent with the style reference image.

[0141] In some embodiments, the first determination unit 601 is configured to extract multiple image region features from a style reference image based on a multimodal training model; determine multiple target features according to a target ratio based on the similarity between each image region feature and the video description text, where the target ratio is the proportion of the multiple target features in the multiple image region features.

[0142] In some embodiments, the local style information of the style reference image is extracted through a local projection module; the apparatus further includes: a training unit configured to: obtain multiple sample vocabulary vectors; learn the multiple sample vocabulary vectors based on the local projection module; a second determination unit 602 configured to process the multiple target features through the local projection module based on a self-attention mechanism to obtain the local style information of the style reference image.

[0143] In some embodiments, the generation unit 603 is configured to splice the local style information and the global style information of the style reference image to obtain the style information of the style reference image; inject the style information into a video diffusion model based on a style cross-attention mechanism; input the video description text into the video diffusion model based on a text cross-attention mechanism; start from a noise video through the video diffusion model to generate a target video.

[0144] In some embodiments, the generation unit 603 is configured to input the video description text, the local style information, and the global style information of the style reference image into the video diffusion model; adjust the weights of the temporal processing layer in the video diffusion model based on a motion adaptation module; generate a target video based on the adjusted video diffusion model.

[0145] In some embodiments, the generation unit 603 is further configured to obtain a reference video, where the reference video is used to provide layout and structure information; convert the reference video into a grayscale thumbnail video; input the video description text, the local style information, and the global style information of the style reference image into the video diffusion model to obtain video hidden layer features; splice the features of the grayscale thumbnail video and the video hidden layer features to obtain intermediate video features; generate a target video based on the intermediate video features.

[0146] In some embodiments, the global style information of the style reference image is extracted through a global projection module;

[0147] The apparatus further includes:

[0148] An acquisition unit configured to acquire a set of sample paired images, where the set of sample paired images includes multiple paired image groups, and each paired image group includes two images with the same style but different contents;

[0149] A feature extraction unit, configured to, for any paired image group, input the first image and the second image in the paired image group into a feature extraction network respectively to obtain the image features of the first image and the image features of the second image;

[0150] A feature processing unit, configured to input the image features of the first image and the image features of the second image into a multi-layer perceptron network respectively to obtain the global style features of the first image and the style features of the second image;

[0151] A training unit, configured to train a global projection module based on the global style features of the first image and the global style features of the second image.

[0152] An embodiment of the present disclosure provides a video generation device. By discarding the image region features with a relatively high similarity to the video description text from the multiple image region features of the style reference image, the remaining image region features with a relatively low similarity to the video description text are obtained, so that both the texture features in the style reference image can be extracted and content leakage can be avoided. Combining the global style information of the style reference image and the video description text, a target video that meets the content described in the video description text and has a good stylization effect can be generated.

[0153] It should be noted that, for the video generation device provided in the above embodiment, only the above-mentioned division of each functional unit is used for illustration. In actual applications, the above functions can be allocated to different functional units as needed, that is, the internal structure of the electronic device is divided into different functional units to complete all or part of the functions described above. In addition, the video generation device provided in the above embodiment and the embodiment of the video generation method belong to the same concept, and the specific implementation process is detailed in the method embodiment, which will not be elaborated here.

[0154] Regarding the video generation device in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method related thereto, and will not be elaborated here.

[0155] In the embodiment of the present disclosure, the electronic device can be a terminal or a server. When the electronic device is a terminal, the terminal is used as the execution subject to implement the technical solution provided in the embodiment of the present disclosure; when the electronic device is a server, the server is used as the execution subject to implement the technical solution provided in the embodiment of the present disclosure; or, the technical solution provided in the present disclosure is implemented through the interaction between the terminal and the server. The embodiment of the present disclosure does not limit this.

[0156] Figure 7 It is a block diagram of an electronic device shown according to an exemplary embodiment. Generally, the electronic device 700 includes: a processor 701 and a memory 702.

[0157] The processor 701 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. The processor 701 may be implemented in at least one of the following hardware forms: DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). The processor 701 may also include a main processor and a coprocessor. The main processor is used to process data in the wake state and is also referred to as the CPU (Central Processing Unit); the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, the processor 701 may be integrated with a GPU (Graphics Processing Unit), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 701 may further include an AI (Artificial Intelligence) processor, and the AI processor is used to process computational operations related to machine learning.

[0158] The memory 702 may include one or more computer-readable storage media, and the computer-readable storage media may be non-transitory. The memory 702 may further include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash storage devices. In some embodiments, the non-transitory computer-readable storage media in the memory 702 is used to store at least one program code, and the at least one program code is used to be executed by the processor 701 to implement the video generation method provided in the method embodiments of the present disclosure.

[0159] In some embodiments, the electronic device 700 may further optionally include: a peripheral device interface 703 and at least one peripheral device. The processor 701, the memory 702, and the peripheral device interface 703 may be connected through a bus or signal lines. Each peripheral device may be connected to the peripheral device interface 703 through a bus, signal lines, or a circuit board. Specifically, the peripheral devices include at least one of the following: a radio frequency circuit 704, a display screen 705, a camera module 706, an audio circuit 707, and a power supply 708.

[0160] The peripheral device interface 703 can be used to connect at least one I / O (Input / Output) related peripheral device to the processor 701 and the memory 702. In some embodiments, the processor 701, the memory 702, and the peripheral device interface 703 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 701, the memory 702, and the peripheral device interface 703 can be implemented on a separate chip or circuit board, and this embodiment does not limit this.

[0161] The radio frequency circuit 704 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The radio frequency circuit 704 communicates with the communication network and other communication devices through electromagnetic signals. The radio frequency circuit 704 converts an electrical signal into an electromagnetic signal for transmission, or converts the received electromagnetic signal into an electrical signal. Optionally, the radio frequency circuit 704 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a subscriber identity module card, and so on. The radio frequency circuit 704 can communicate with other electronic devices through at least one wireless communication protocol. The wireless communication protocol includes but is not limited to: metropolitan area network, each generation of mobile communication network (2G, 3G, 4G, and 5G), wireless local area network, and / or WiFi (Wireless Fidelity) network. In some embodiments, the radio frequency circuit 704 may further include a circuit related to NFC (Near Field Communication), and this disclosure does not limit this.

[0162] The display screen 705 is used to display the UI (User Interface). The UI may include graphics, text, icons, videos, and any combination thereof. When the display screen 705 is a touch display screen, the display screen 705 also has the ability to collect touch signals on or above the surface of the display screen 705. The touch signals can be input as control signals to the processor 701 for processing. At this time, the display screen 705 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there may be one display screen 705, which is provided on the front panel of the electronic device 700; in other embodiments, there may be at least two display screens 705, which are respectively provided on different surfaces of the electronic device 700 or are in a foldable design; in still other embodiments, the display screen 705 may be a flexible display screen, which is provided on the curved surface or the folding surface of the electronic device 700. Even more, the display screen 705 can also be set to an irregular non-rectangular shape, that is, a special-shaped screen. The display screen 705 can be prepared using materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).

[0163] The camera module 706 is used to capture images or videos. Optionally, the camera module 706 includes a front camera and a rear camera. Generally, the front camera is provided on the front panel of the electronic device, and the rear camera is provided on the back of the electronic device. In some embodiments, there are at least two rear cameras, which are any one of a main camera, a depth-of-field camera, a wide-angle camera, and a telephoto camera, so as to implement functions such as background blurring by fusing the main camera and the depth-of-field camera, panoramic shooting by fusing the main camera and the wide-angle camera, and VR (Virtual Reality) shooting function or other fused shooting functions. In some embodiments, the camera module 706 may also include a flash. The flash can be a single-color temperature flash or a dual-color temperature flash. A dual-color temperature flash refers to a combination of a warm light flash and a cold light flash, which can be used for light compensation under different color temperatures.

[0164] The audio circuit 707 may include a microphone and a speaker. The microphone is used to collect sound waves of the user and the environment, and convert the sound waves into electrical signals for input to the processor 701 for processing, or input to the radio frequency circuit 704 to achieve voice communication. For the purpose of stereo collection or noise reduction, there may be multiple microphones, which are respectively arranged at different parts of the electronic device 700. The microphone may also be an array microphone or an omnidirectional collection microphone. The speaker is used to convert the electrical signal from the processor 701 or the radio frequency circuit 704 into sound waves. The speaker may be a traditional thin film speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can not only convert the electrical signal into sound waves audible to humans, but also convert the electrical signal into sound waves inaudible to humans for uses such as ranging. In some embodiments, the audio circuit 707 may further include a headphone jack.

[0165] The power supply 708 is used to supply power to each component in the electronic device 700. The power supply 708 may be alternating current, direct current, a primary battery or a rechargeable battery. When the power supply 708 includes a rechargeable battery, the rechargeable battery may support wired charging or wireless charging. The rechargeable battery may also be used to support fast charging technology.

[0166] Those skilled in the art can understand that Figure 7 the structure shown in

[0167] does not limit the electronic device 700, and may include more or fewer components than shown in the figure, or combine certain components, or adopt different component arrangements.

[0168] A computer program product includes a computer program, and when the computer program is executed by a processor, the above video generation method is implemented.

[0169] After considering the specification and practicing the invention disclosed herein, those skilled in the art will readily conceive of other embodiments of the present disclosure. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure, which follow the general principles of the present disclosure and include common general knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and embodiments are only regarded as exemplary, and the true scope and spirit of the present disclosure are pointed out by the following claims.

[0170] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is limited only by the appended claims.

Claims

1. A video generation method, characterized in that, The method includes: Based on the video description text, determining multiple target features from multiple image region features of the style reference image, where the similarity between the multiple target features and the video description text is lower than that of other image region features among the multiple image region features; Based on the multiple target features, determining the local style information of the style reference image; Based on the video description text, the local style information, and the global style information of the style reference image, generating a target video, where the content of the target video is consistent with the content of the video description text, and the style of the target video is consistent with the style reference image.

2. The video generation method according to claim 1, wherein The determining multiple target features from multiple image region features of the style reference image based on the video description text includes: Extracting the multiple image region features from the style reference image based on a multimodal training model; Determining the multiple target features according to a target ratio based on the similarity between each image region feature and the video description text, where the target ratio is the proportion of the multiple target features in the multiple image region features.

3. The video generation method according to claim 1, wherein The local style information of the style reference image is obtained by extracting through a local projection module; the method further includes: Obtaining multiple sample vocabulary vectors; Learning the multiple sample vocabulary vectors based on the local projection module; The determining the local style information of the style reference image based on the multiple target features includes: Processing the multiple target features through the local projection module based on the self-attention mechanism to obtain the local style information of the style reference image.

4. The video generation method according to claim 1, wherein The generating a target video based on the video description text, the local style information, and the global style information of the style reference image includes: Concatenating the local style information and the global style information of the style reference image to obtain the style information of the style reference image; Injecting the style information into a video diffusion model based on a style cross-attention mechanism; Inputting the video description text into the video diffusion model based on a text cross-attention mechanism; Generating the target video starting from a noise video through the video diffusion model.

5. The video generation method according to claim 1, wherein The generating a target video based on the video description text, the local style information, and the global style information of the style reference image includes: Inputting the video description text, the local style information, and the global style information of the style reference image into a video diffusion model; Adjusting the weights of the temporal processing layer in the video diffusion model based on a motion adaptation module; Generating the target video based on the adjusted video diffusion model.

6. The video generation method according to claim 1, wherein The method further includes: Obtaining a reference video, where the reference video is used to provide layout and structure information; Converting the reference video into a grayscale thumbnail video; The generating a target video based on the video description text, the local style information, and the global style information of the style reference image includes: Inputting the video description text, the local style information, and the global style information of the style reference image into a video diffusion model to obtain video hidden layer features; Concatenate the features of the grayscale thumbnail video with the video hidden layer features to obtain intermediate video features; Generate the target video based on the intermediate video features.

7. The video generation method according to any one of claims 1-6, characterized in that, The global style information of the style reference image is extracted by a global projection module; The method further includes: Obtain a sample paired image set, the sample paired image set includes multiple paired image groups, and each paired image group includes two images with the same style but different contents; For any paired image group, input the first image and the second image in the paired image group into the feature extraction network respectively to obtain the image features of the first image and the image features of the second image; Input the image features of the first image and the image features of the second image into a multi-layer perceptron network respectively to obtain the global style features of the first image and the style features of the second image; Train the global projection module based on the global style features of the first image and the global style features of the second image.

8. A video generation device, characterized in that, The device includes: A first determination unit configured to determine multiple target features from multiple image region features of a style reference image based on a video description text, and the similarity between the multiple target features and the video description text is lower than that of other image region features in the multiple image region features; A second determination unit configured to determine the local style information of the style reference image based on the multiple target features; A generation unit configured to generate a target video based on the video description text, the local style information, and the global style information of the style reference image, the content of the target video is consistent with the content of the video description text, and the style of the target video is consistent with the style reference image.

9. An electronic device, characterized in that, The electronic device includes: One or more processors; A memory for storing program code executable by the processor; Wherein, the processor is configured to execute the program code to implement the video generation method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, When the instructions in the computer-readable storage medium are executed by the processor of the electronic device, the electronic device is enabled to execute the video generation method according to any one of claims 1 to 7.

11. A computer program product, including a computer program, which when executed by a processor implements the video generation method according to any one of claims 1 to 7.