Video generation method, device, computer program, and electronic device

By integrating content description text with a reference video for feature extraction, the method addresses the issue of frame jitter in AI-generated videos, achieving enhanced consistency and quality in the generated content.

JP2026504980APending Publication Date: 2026-02-10TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2025542396
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-07-26
Filing Date
2024-05-22
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

Existing AI models generate videos with poor quality due to weak consistency between consecutive frames, resulting in noticeable jitter, as the generation of each image is independent and lacks coherent action continuity.

Method used

A video generation method that combines content description text with a content reference video to perform feature extraction, using text semantic features and video reference features to generate a target video that maintains consistent action and improves quality.

Benefits of technology

The method enhances video consistency and quality by aligning the actions in the generated video with the reference content, resulting in higher accuracy and improved video generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026504980000001_ABST
    Figure 2026504980000001_ABST
Patent Text Reader

Abstract

This application discloses a video generation method, an apparatus, a storage medium, and an electronic device. The method includes the steps of: obtaining a content description text and a content reference video, where the content description text includes information for describing target content represented by a target video to be generated, and the content reference video includes action reference information related to the target content; performing feature extraction on the content description text to obtain text semantic features, where the text semantic features are used to characterize the semantic information of the content description text; performing feature extraction on the content reference video to obtain video reference features, where the video reference features are used to characterize the action reference information in the content reference video; and generating a target video based on the text semantic features and the video reference features. This application can improve the quality of the generated target video.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present application relates to the field of artificial intelligence, and in particular to the field of computer vision techniques.

[0002] This application claims priority from a Chinese patent application filed with the China Patent Office on July 26, 2023, bearing application number 2023109234935 and entitled "Video Acquisition Method, Apparatus, and Storage Medium, and Electronic Device," the entire contents of which are incorporated herein by reference. [Background technology]

[0003] In video generation scenarios, artificial intelligence (AI) models are typically used to generate a sequence of images based on input content description text, and then this sequence of images is used to form a continuous video.

[0004] In the above method, the content description text is specifically used for generating images, not directly used for generating videos, and the generation processes of each image are independent of each other, so the consistency of each generated image is likely to be weak. Correspondingly, in the video composed of each image, obvious jitter occurs between consecutive frames, which affects the quality of the generated video, that is, there may be a problem that the quality of the generated video is poor. Summary of the Invention [Problem to be solved by the invention]

[0005] The embodiments of the present application provide a video generation method, an apparatus, a storage medium, and an electronic device that can improve the quality of the generated video. [Means for solving the problem]

[0006] According to one aspect of an embodiment of the present application, there is provided a video generation method executed by an electronic device, comprising: obtaining a content description text and a content reference video, wherein the content description text includes information for describing a target content represented by a target video to be generated, and the content reference video includes action reference information related to the target content; performing feature extraction on the content description text to obtain text semantic features for characterizing semantic information of the content description text, and performing feature extraction on the content-referred video to obtain video-referred features for characterizing the action-referred information in the content-referred video; generating a target video based on the text semantic features and the video reference features.

[0007] According to another aspect of an embodiment of the present application, there is provided a video production device, comprising: a first acquisition unit for acquiring a content description text and a content reference video, wherein the content description text includes information for describing a target content represented by a target video to be generated, and the content reference video includes action reference information related to the target content; an extraction unit performing feature extraction on the content description text to obtain text semantic features for characterizing semantic information of the content description text, and performing feature extraction on the content-referenced video to obtain video-referenced features for characterizing the action-referenced information in the content-referenced video; a generation unit for generating a target video based on the text semantic features and the video reference features.

[0008] According to yet another aspect of an embodiment of the present application, there is provided a computer-readable storage medium including a program stored thereon, the program performing the above video generation method when executed by an electronic device.

[0009] According to yet another aspect of an embodiment of the present application, there is provided a computer program product or a computer program including computer instructions stored in a computer-readable storage medium, wherein a processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, thereby causing the computer device to perform the above-described video generation method.

[0010] According to yet another aspect of an embodiment of the present application, there is further provided an electronic device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the above-mentioned video generation method by means of the computer program.

[0011] In an embodiment of the present application, a content description text and a content reference video are obtained, the content description text includes information for describing target content expressed by a desired target video to be generated, and the content reference video includes action reference information related to the target content, and feature extraction is performed on the content description text to obtain text semantic features for characterizing semantic information of the content description text, and feature extraction is performed on the content reference video to obtain video reference features for characterizing the action reference information in the content reference video, and a target video is generated based on the text semantic features and the video reference features. Since the embodiment of the present application uses the content description text to describe the target content expressed by the desired target video to be generated and the video-consistent content reference video to provide a reference manner for the target content, the action information in the consistent content reference video can be fully referenced in the process of generating the target video, thereby improving the consistency of actions in the generated video content, i.e., improving the consistency of the generated target video, thereby achieving the objective of obtaining a higher-quality output video and achieving the technical effect of improving the accuracy of video acquisition. [Brief explanation of the drawings]

[0012] The drawings set forth herein provide a further understanding of the present application, constitute a part of the present application, and the illustrative examples and descriptions thereof should not be construed as undue limitations on the present application, but should be used to interpret the present application.

[0013] [Figure 1] 1 is a schematic diagram of an application environment of a selectable video generation method according to an embodiment of the present application; [Figure 2] 1 is a flowchart of a selectable video generation method according to an embodiment of the present application. [Figure 3] 1 is a schematic diagram of a selectable video generation method according to an embodiment of the present application; [Figure 4] FIG. 1 is a schematic diagram of another selectable video generation method according to an embodiment of the present application. [Figure 5] FIG. 1 is a schematic diagram of another selectable video generation method according to an embodiment of the present application. [Figure 6] FIG. 1 is a schematic diagram of another selectable video generation method according to an embodiment of the present application. [Figure 7] FIG. 1 is a schematic diagram of another selectable video generation method according to an embodiment of the present application. [Figure 8] FIG. 1 is a schematic diagram of another selectable video generation method according to an embodiment of the present application. [Figure 9] FIG. 1 is a schematic diagram of another selectable video generation method according to an embodiment of the present application. [Figure 10] 1 is a schematic diagram of a selectable video generation device according to an embodiment of the present application; [Figure 11] 1 is a schematic diagram illustrating the configuration of a selectable electronic device according to an embodiment of the present application. DETAILED DESCRIPTION OF THE INVENTION

[0014] In order to help those skilled in the art better understand the technical solutions of the present application, the technical solutions in the embodiments of the present application will be described below clearly and completely with reference to the drawings of the embodiments of the present application, and obviously, the described embodiments are not all embodiments, but only some embodiments of the present application. All other embodiments obtained by those skilled in the art based on the embodiments of the present application without creative work belong to the protection scope of the present application.

[0015] It should be noted that terms such as "first," "second," etc. in the specification, claims, and drawings of this application are used to distinguish between similar objects, not to describe a particular order or chronology. Such terms may be substituted for one another where appropriate, such that the embodiments of this application described herein may be performed in orders other than those illustrated or described herein. Furthermore, the terms "comprise," "have," and any variations thereof are intended to cover non-exclusive inclusions; for example, a process, method, system, product, or apparatus comprising a series of steps or units is not necessarily limited to the explicitly recited steps or units, but may include other steps or units not explicitly recited or inherent in the process, method, product, or apparatus.

[0016] According to one aspect of the embodiments of the present application, a video generation method is provided, and optionally, as an alternative embodiment, the video generation method can be applied to, but is not limited to, the environment shown in Figure 1, which may include, but is not limited to, a user device 102 and a server 112. The user device 102 may include a display 104, a processor 106, and a memory 108, and the server 112 may include a database 114 and a processing engine 116.

[0017] The specific process is as follows: Step S102, the user device 102 obtains the content description text and the content reference video. In steps S104-S106, the user device 102 transmits the content description text and the content reference video to the server 112 via the network 110. In steps S108-S112, the server 112, via the processing engine 116, performs feature extraction on the content description text to obtain text semantic features, performs feature extraction on the content reference video to obtain video reference features, and further generates a target video based on the text semantic features and the video reference features. In steps S114-S116, the server 112 transmits the target video to the user device 102 via the network 110, and the user device 102 displays the target video on the display 104 via the processor 106, and stores the target video in the memory 108.

[0018] In addition to the example shown in FIG. 1 , the above steps may be completed independently by the user device or the server, or may be completed jointly and cooperatively by the user device and the server. For example, the user device 102 may perform steps S108-S112 above, thereby reducing the processing burden on the server 112. The user device 102 may include, but is not limited to, a handheld device (e.g., a mobile phone), a laptop computer, a tablet computer, a desktop computer, an in-vehicle device, a smart TV, etc. This application does not limit the specific form of the user device 102. The server 112 may be a single server, a server cluster consisting of multiple servers, or a cloud server.

[0019] Optionally, as an alternative embodiment, as shown in FIG. 2, a video generation method can be performed by an electronic device such as a user device or a server shown in FIG. 1, and specifically includes the following steps: Step S202: Obtain a content description text and a content reference video, where the content description text includes information for describing the target content represented by the target video to be generated, and the content reference video includes action reference information related to the target content. Step S204: perform feature extraction on the content description text to obtain text semantic features for characterizing the semantic information of the content description text; perform feature extraction on the content reference video to obtain video reference features for characterizing the action reference information in the content reference video. Step S206: Generate a target video based on the text semantic features and the video reference features.

[0020] Optionally, in this embodiment, the above video generation method can be applied to, but is not limited to, application scenarios of artificial intelligence generated content (abbreviated as AIGC). AIGC is an artificial intelligence technology that can generate new content, audio, images, etc., such as AI-generated images and AI-generated videos. This embodiment combines content description text and content reference video to support users in inputting content reference video and combining it with content description text to generate video content, thereby effectively improving the controllability and generation quality of the generated video content.

[0021] Optionally, in this embodiment, the content description text for describing the target content expressed by the target video to be generated relates to information such as objects, scenarios, actions, emotions, etc. in the video.

[0022] To further illustrate, the content description text may provide a general description of the entire video, including basic information such as the theme, scenario, time, and location of the video, e.g., "A red car is driving on the screen," or may describe the actions or behaviors of objects (e.g., people or objects) in the video, e.g., "A puppy is chasing a ball," or may evaluate the emotional color of the video content, e.g., "A heartwarming family moment is shown," or may describe the scenario or environment shown in the video, e.g., "A beautiful beach is shown in the video," or may explain the chronological order of events or changes occurring in the video, e.g., "At first, the sun rises slowly, and then a spectacular sunset follows."

[0023] Optionally, in this embodiment, in order to improve the correlation between the input content description text and the output target video, the video generation process can combine a content reference video provided as a reference to the target content, for example, the content description text is "a puppy chasing a ball", and the content reference video shows a series of actions of a kitten chasing a ball, and by combining the content description text and the content reference video, a target video showing a series of actions of a puppy chasing a ball can finally be generated, and the series of actions shown in the target video may be similar or identical to, but not limited to, the series of actions shown in the content reference video.

[0024] Optionally, in this embodiment, the text semantic features are used to characterize the semantic information of the content description text, i.e., the semantic information of the content description text for describing the target content, and may be, but are not limited to, the expression method of the meaning or information conveyed by the content description text. For example, the vocabulary selection, word meanings, and parts of speech in the text may reflect the meaning and connotation of the text, the sentence structure, grammatical rules, and inter-word relationships in the text may reflect the logical and semantic connections between sentences, the context in the text, the tone and emotional color of the text, the theme, topic, and related professional knowledge of the text.

[0025] Optionally, in this embodiment, the video reference features are used to characterize key information that the content-referred video provides reference to the target content, i.e., action reference information in the content-referred video. In order to improve the consistency of the target video, the video reference features may be, but are not limited to, features that satisfy dynamic property conditions. And / or the key information may correspond to, but are not limited to, key content in the target content. The key content may be, but are not limited to, dynamic content.

[0026] 3 , the content-referred video 302 and the content description text 304 may be combined to generate a target video 308 showing a puppy dancing, and the video-referred features 306 obtained by performing feature extraction on the content-referred video 302 may be, but are not limited to, dance action features in the content-referred video 302. The dance action features in the content-referred video 302 are extracted as the video-referred features 306 because, first, the dance action features satisfy the dynamic property condition, and second, the word "dancing" in the content description text 304 "A puppy dancing" also belongs to dynamic content and "dancing" also corresponds to dance action features, so the dance action features in the content-referred video 302 are extracted as the video-referred features 306.

[0027] Optionally, in this embodiment, the target video is generated based on the text semantic features and the video reference features, for example, the text semantic features provide the desired video theme, content, and keywords to be generated, and at least one video element is obtained based on the video theme, content, and keywords, and the action reference features provided by the video reference features are used to guide the generation of key video elements among the video elements to obtain the target video, where the text semantic features ensure the consistency between the target video and the desired target content to be generated, and the video reference features ensure the video quality of the target video.

[0028] Furthermore, by using content description text to describe the target content expressed by the desired video to be generated, and by using a content reference video with video consistency to provide a reference to the target content, the consistency of the generated video content is improved, and further, the relevance between the input content description text and the output target video is improved, thereby achieving the technical effect of improving the accuracy of video generation and enabling the generation of higher quality videos.

[0029] To further illustrate, optionally, based on the scenario shown in FIG. 3 , and continuing to generate a content description text 304 and a content reference video 302, for example as shown in FIG. 4 , the content description text 304 includes information for describing the target content represented by the desired video to be generated, the content reference video 302 includes action reference information that provides a reference to the target content, feature extraction is performed on the content description text 304 to obtain text semantic features 402, which are used to characterize the semantic information used by the content description text 304 to describe the target content, feature extraction is performed on the content reference video 302 to obtain video reference features 306, which are used to characterize the action reference information used by the content reference video 302 to provide a reference to the target content, for example, the key information that provides a reference for generating the video for ``dancing'' in the target content is the dancing action information in the content reference video 302, and a target video 308 is generated based on the text semantic features 402 and the video reference features 306.

[0030] According to an embodiment provided by the present application, a content description text and a content reference video are obtained, the content description text includes information for describing target content expressed by a desired target video to be generated, and the content reference video includes action reference information related to the target content, feature extraction is performed on the content description text to obtain text semantic features for characterizing the semantic information of the content description text, feature extraction is performed on the content reference video to obtain video reference features for characterizing the action reference information in the content reference video, and a target video is generated based on the text semantic features and the video reference features. Since the content description text is used to describe the target content expressed by the desired target video to be generated, and the video-consistent content reference video is used to provide a reference manner for the target content, the action information in the consistent content reference video can be fully referenced in the process of generating the target video, thereby improving the consistency of actions in the generated video content, i.e., improving the consistency of the generated target video, thereby achieving the objective of obtaining a higher quality output video and achieving the technical effect of improving the accuracy of video acquisition.

[0031] As an alternative solution, generating a target video based on text semantic features and video reference features includes the following steps. Step S1-1: determine at least one video element in the target video based on text semantic features, where the at least one video element includes a first subject object. Step S1-2: Determine the posture change situation of the first subject object in the target video based on the video reference features. Step S1-3: generating a target video based on at least one video element and the posture change situation of the first subject object in the target video;

[0032] Optionally, in this embodiment, text semantic features are utilized to determine at least one video element to be displayed in the target video. For example, if the content description text is "a puppy is dancing", the text semantic features obtained by performing feature extraction on the content description text can show that the desired video to be generated needs a puppy as a main object (first main object) and a soccer field as a video background, and both the puppy and the soccer field can be understood as video elements.

[0033] Optionally, in this embodiment, the pose change situation may refer to, but is not limited to, a change in the posture, position, or shape of an object (e.g., an object or a human body) in (target video) space, such as a displacement change (a change in the position of an object in space, which may be a translation, rotation, or shear movement along a straight or curved path), a posture change (a change in the posture of a part or the whole body when the object is stationary or moving, which may be, for example, bending, stretching, twisting, etc.), a shape change (a change in the appearance of an object, which may be, for example, a size change, deformation, or expansion caused by compression or stretching, etc.), etc.

[0034] It should be noted that the pose change situation is usually an important attribute for determining whether a video is consistent, that is, whether the video display is consistent is closely related to the pose change situation. In this embodiment, to further improve the consistency of the target video, the video reference features are used to determine the pose change situation of the first subject object in the target video, so that the pose change situation in the target video is more adapted to the characteristics of the video itself, and the quality of the generated target video is improved.

[0035] To further illustrate, optionally, for example, by determining the position distribution and element shape of at least one video element in each video frame in the target video, and dynamically and specifically adjusting the position distribution and element shape of the first subject object in each video frame in the target video based on the posture change situation of the first subject object in the target video, the final display effect of the target video is not limited to a collection of multiple frame images, but has posture changes that are more compatible with the video characteristics, i.e., a more consistent target video can be displayed.

[0036] According to an embodiment provided by the present application, at least one video element in a target video is determined based on text semantic features, and the at least one video element includes a first subject object; a posture change situation of the first subject object in the target video is determined based on video reference features; and a target video is generated based on the at least one video element and the posture change situation of the first subject object in the target video, thereby achieving the purpose of generating a more consistent target video and realizing the technical effect of improving the video quality of the target video.

[0037] In an optional embodiment, performing feature extraction on the content-referenced video to obtain video-referenced features comprises: performing feature extraction on a second subject object in the content-referred video to obtain object representation features, the object representation features being for characterizing a posture change situation of the second subject object in the content-referred video, and the video-referred features including the object representation features; The step of determining a posture change situation of the first subject object in the target video based on the video reference features includes: The method includes determining a posture change situation of a first subject object in a target video based on object representation features, wherein the posture change situation of the first subject object in the target video corresponds to the posture change situation of a second subject object in the content reference video.

[0038] Optionally, in this embodiment, the posture change situation of the first subject object in the target video corresponds to the posture change situation of the second subject object in the content-referred video, for example, the posture change situation of the person (second subject object) while dancing in the content-referred video 302 shown in Figure 4 corresponds to the posture change situation of the puppy (first subject object) while dancing in the target video 308.

[0039] It can be understood that improving the video quality of the target video by matching the posture change situations between the existing content-referred video and the desired target video to be generated means that the posture change situations of the first subject object in the target video are restored based on the posture change situations of the second subject object in the content-referred video, or that the posture change situations of the first subject object in the target video may be a series of target actions performed by the first subject object in the target video, and this series of target actions may be identical to or similar to a series of actions performed by the second subject object in the content-referred video.

[0040] According to an embodiment provided by the present application, feature extraction is performed on a second subject object in the content-referred video to obtain object representation features, the object representation features are used to characterize the posture change situation of the second subject object in the content-referred video, the video reference features include object representation features, and based on the object representation features, the posture change situation of the first subject object in the target video is determined, and the posture change situation of the first subject object in the target video corresponds to the posture change situation of the second subject object in the content-referred video, thereby achieving the purpose of matching the posture change situations between the existing content-referred video and the desired target video to be generated, and achieving the technical effect of improving the video quality of the target video, and the series of actions performed by the first subject object in the target video refer to the series of actions performed by the second subject object in the content-referred video.

[0041] As an optional embodiment, performing feature extraction on the second subject object in the content-referred video to obtain object representation features includes: Step S2-1: perform feature extraction on at least two target video frames including a second subject object in the content-referred video to obtain at least two object static features, which are used to characterize the position and shape of the second subject object in the target video frames. S2-2, based on timing relationship information between at least two target video frames, integrate and process at least two object static features to obtain object dynamic features, the object dynamic features are used to characterize the posture change situation of the second subject object in the content-reference video, and the object expression features include the object dynamic features.

[0042] It should be noted that the object static features are used to characterize the position and shape of the second subject object in the target video frame, and the position and shape itself is a static attribute, which is usually sufficient as the basis for generating an image, but when used as the basis for generating a video, it lacks dynamic attributes, because the consistency of the video is usually determined by dynamic attributes, and without dynamic attributes, it is naturally impossible to generate high-quality video.

[0043] Furthermore, in this embodiment, the object dynamic features are used to characterize the posture change situation of the second subject object in the content-referencing video, that is, this embodiment does not use the object static features as the basis for generating the video as they are, but uses the object static features to obtain object dynamic features, and generates high-quality video based on the dynamic attributes contained in the object dynamic features.

[0044] To further illustrate, optionally, feature extraction is performed on at least two target video frames including a second subject object 504 in the content-referred video 502 to obtain at least two object static features 506, which are used to characterize the position and shape of the second subject object 504 in the target video frames, and an ordering and integration process is performed on the at least two object static features 506 using timing relationship information 508 between the at least two target video frames to obtain object dynamic features 510, which are used to characterize the posture change situation of the second subject object 504 in the content-referred video 502, for example, as shown in FIG.

[0045] According to an embodiment provided by the present application, feature extraction is performed on at least two target video frames including a second subject object in the content-referred video to obtain at least two object static features, the object static features are used to characterize the position and shape of the second subject object in the target video frames, and timing relationship information between the at least two target video frames is used to integrate and process the at least two object static features to obtain object dynamic features, the object dynamic features are used to characterize the posture change situation of the second subject object in the content-referred video, and the object expression features include object dynamic features, thereby achieving the purpose of using the object static features to obtain object dynamic features and generating high-quality video based on the dynamic attributes contained in the object dynamic features, and realizing the technical effect of improving the video quality of the target video.

[0046] As an optional solution, the step of performing feature extraction on at least two target video frames including the second subject object in the content-referencing video to obtain at least two object static features includes at least one of the following: Step S3-1: perform keypoint extraction on a second subject object in at least two target video frames to obtain at least two keypoint features, the keypoint features being used to characterize the positions of the keypoints of the second subject object in the target video frames, and the object static features include the keypoint features. Step S3-2: perform key line extraction on the second subject object in at least two target video frames to obtain at least two key line features, the key line features are used to characterize the positions of the key lines of the second subject object in the target video frames, and the object static features include the key line features. Step S3-3: performing contour extraction on the second subject object in at least two target video frames to obtain at least two contour features, the contour features being used to characterize the shape position of the contour of the second subject object in the target video frames, and the object static features including the contour features; Step S3-4: performing edge extraction on the second subject object in the at least two target video frames to obtain at least two first object features, and the object static features include the first object features. Step S3-5: performing depth extraction on a second subject object in at least two target video frames to obtain at least two second object features, and the object static features include the second object features. Step S3-6: performing white model extraction on the second subject object in the at least two target video frames to obtain at least two third object features, and the object static features include the third object features.

[0047] Optionally, in this embodiment, keypoint extraction refers to, but is not limited to, automatically detecting and identifying important feature points in an image. Such feature points usually have important structure, texture, or shape information. For example, feature detection algorithms such as Harris corner detection, SIFT (Scale Invariant Feature Transform), and SURF (Super Fast Robust Features) can be used to find keypoints in an image. The above algorithms can determine keypoints based on features such as the local structure of the image, gradient direction, and scale change.

[0048] Optionally, in this embodiment, key line extraction refers to, but is not limited to, extracting lines with important visual information from an image. The lines may typically be major contours, boundaries, or other important linear structures in the image. Visually prominent and important lines can be selected based on criteria such as line length, curvature, and straight line fit. Curvature calculation, straight line fitting algorithm, etc. can be used to evaluate the quality and importance of the lines.

[0049] Optionally, in this embodiment, contour extraction refers to, but is not limited to, extracting the boundary contour of an object from an image. A contour can be regarded as a line segment connecting discontinuous points on the surface of an object, and can characterize the shape and structure information of the object. Based on a binary image, a contour extraction algorithm can be used to detect and extract the boundary contour of the object, specifically, a contour detection algorithm based on edge connection (e.g., Moore-Neighbor algorithm, kNN detection algorithm, etc.) and a region growing algorithm based on pixel connection are used.

[0050] Optionally, in this embodiment, edge extraction can be used to detect and extract edge information of objects in an image, but is not limited to this. An edge usually represents a sudden change or discontinuity in brightness, color, texture, etc. in an image, and is the boundary between objects or between an object and a background. Examples of implementation methods include Canny edge detection (obtaining high-quality edge results through a multi-stage process, such as Gaussian filtering, image gradient calculation, non-maximum suppression, double thresholding, and edge connection), Sobel operator (performing convolution operations in the horizontal and vertical directions of an image to obtain two gradient images, and combining these two gradient images to obtain an edge strength image), Laplacian operator (performing Laplacian filtering on an image to obtain an edge image), and self-adaptive thresholding (dividing an image into small regions and selecting a threshold for each region based on local statistical characteristics).

[0051] Optionally, in this embodiment, depth extraction refers to, but is not limited to, the process of obtaining depth information from an image or a scenario. Depth information represents the distance between different points in an object or a scenario and a camera. Examples of implementation methods include facial depth extraction (using an infrared camera, structured light, or time-of-flight sensor to obtain depth information of a face region), binocular stereo vision (estimating the depth of a scenario through images from two perspectives and inferring the distance of an object based on the disparity between the left-eye and right-eye images, i.e., the offset between corresponding pixels), 3D reconstruction (recovering the geometric structure of a 3D scenario through multiple images or video sequences and using technologies such as multi-view geometry, structured light projection, and optical flow estimation to perform depth extraction and stereo reconstruction), and deep learning methods (using structures such as deep convolutional neural networks to learn and predict depth information directly from a single image).

[0052] Optionally, in this embodiment, a white model refers to a prototype model of a building, product, sculpture, etc., used as a reference for creating a formal product or structure, and may be, but is not limited to, a model made of materials such as wood, clay, polymer, etc.

[0053] In addition, various subject object extraction methods are provided, and according to the actual application scenario and processing object, the corresponding extraction method can be flexibly used to meet the needs of the corresponding scenario and processing object, so that the object static features of the second subject object in the target video frame can be accurately determined.

[0054] In an optional embodiment, generating a target video based on the text semantic features and the video reference features comprises: A step of inputting the text semantic features and video reference features into a video generation model to obtain a target video output by the video generation model, wherein the video generation model is a neural network model for generating videos obtained by training based on multiple video sample data.

[0055] In addition, in order to improve the efficiency of generating the target video, a video generation model is used to generate the corresponding target video, and the video generation model may be, but is not limited to, a video diffusion model that generates video data by gradually removing noise from random noise.

[0056] 6, optionally, feature extraction is performed on the acquired content description text 602-1 and content reference video 602-2 using encoder A and encoder B to obtain text semantic features 604-1 and video reference features 604-2. The text semantic features 604-1 and video reference features 604-2 are input to a video generation model 606 and processed by the video generation model 606 to obtain video features, which are then converted into video data of a target video 608 by a decoder.

[0057] According to the embodiments provided by the present application, text semantic features and video reference features are input into a video generation model to obtain a target video output by the video generation model, and the video generation model is a neural network model for generating videos obtained by training based on multiple video sample data, thereby achieving the purpose of using the video generation model to generate a desired target video and achieving the technical effect of improving the generation efficiency of the target video.

[0058] As an alternative embodiment, before inputting the text semantic features and video reference features into the video generation model, the method further includes the following steps: Step S4-1: Obtain an image generation model. The image generation model is a neural network model for generating an image obtained by training based on a plurality of image sample data. Step S4-2: Adjust the image generation model to obtain an initial video generation model, which is composed of a convolutional layer and an attention layer capable of processing temporal dimension information. Step S4-3: Train an initial video generation model based on multiple video sample data to obtain a video generation model.

[0059] However, directly training a video generation model to generate a desired video typically requires a large amount of training data and high computational power to support it, and the training quality of the video generation model cannot be reasonably guaranteed due to both cost and efficiency considerations.

[0060] In this embodiment, a more mature image generation model is adjusted, and the adjusted image generation model (video generation model) can be applied to the video generation process. At the same time, before making any improvements, an appropriate amount of image training samples is first used to train the image generation model, and after obtaining the trained image generation model, an appropriate amount of video training samples is used to train the adjusted image generation model.

[0061] To further illustrate, optionally, assuming that the convolutional layer of the image generation model is 3x3, the adjustment extends the original convolutional layer to a 1x3x3 convolutional layer, and adds a temporal attention layer to the video generation model to improve the model's understanding of sequence frames and the stability of generating consecutive frames of video, and adapt to the video generation process.

[0062] According to an embodiment provided by the present application, an image generation model is obtained, the image generation model being a neural network model for generating images obtained by training based on a plurality of image sample data, and the image generation model is adjusted to obtain an initial video generation model, the initial video generation model being composed of a convolution layer and an attention layer capable of processing temporal dimension information. The initial video generation model is trained using a plurality of video sample data to obtain the video generation model, thereby achieving the purpose of adjusting and obtaining a video generation model based on a more mature image generation model and achieving the technical effect of ensuring the training quality of the video generation model within a reasonable range.

[0063] In an optional embodiment, the step of inputting the text semantic features and the video reference features into a video generation model to obtain a target video output by the video generation model comprises: Invoking a single graphics processor unit to execute a video generation model and process the input text semantic features and video reference features to obtain a target video output by the video generation model; After invoking the single graphics processor unit to execute a video generation model and process the input text semantic features and video reference features to obtain a target video output by the video generation model, the method further comprises: The method includes inserting the relevant video frames into a video frame sequence corresponding to the target video to obtain a new video, wherein the length of the video corresponding to the new video is greater than the length of the video corresponding to the target video.

[0064] Optionally, in this embodiment, a graphics processing unit (abbreviated as GPU) is a type of dedicated hardware for processing graphics and parallel computing tasks. GPUs have many cores and high-speed memory, and can process multiple data operations simultaneously and perform large-scale computing tasks in parallel. GPUs excel at processing intensive computing tasks such as large data sets, matrix operations, image processing, and simulations.

[0065] Optionally, in this embodiment, inserting related video frames may be understood as, but is not limited to, adding additional frames to the target video. Adding additional frames to the target video to change the frame rate of the target video not only increases the video length of the output new video, but also makes the movement of the new video smoother. Specifically, adding additional frames to the target video may be achieved by, but is not limited to, linear interpolation, optical flow estimation, frame blending, etc.

[0066] Among them, linear interpolation distributes pixel values ​​evenly over time by interpolating between adjacent frames. Optical flow estimation is based on pixel motion characteristics and estimates pixel values ​​of intermediate frames by analyzing the optical flow between two frames. Blending frames takes into account the motion and deformation of objects in a video sequence and generates an interpolated frame by sampling and blending multiple adjacent frames.

[0067] In order to improve the consistency of the target video, the target video may be processed by a single graphics processor unit, but is not limited to this. However, the performance of a single graphics processor unit is limited, and it can only output a short target video, which naturally limits the amount of video data that can be output. A short target video usually cannot meet user needs. In addition, this embodiment uses a frame interpolation method to compensate for the above drawback of not being able to meet user needs.

[0068] According to the embodiment provided by the present application, a single graphics processor unit is invoked to execute a video generation model, and the input text semantic features and video reference features are processed to obtain a target video output by the video generation model. Related video frames are inserted into the video frame sequence corresponding to the target video to obtain a new video, and the length of the video corresponding to the new video is greater than the length of the video corresponding to the target video. That is, the frame interpolation method is used to compensate for the short length of the video generated by invoking a single graphics processor unit, thereby achieving the technical effect of improving the consistency of the target video.

[0069] As an alternative embodiment, for ease of understanding, the above video generation method is applied to a scenario in which a person video is generated based on text. In existing technical solutions, videos generated using two-dimensional (2D) painting techniques often have obvious inter-frame jitter and require a lot of computation time. Three-dimensional (3D) video generation methods rely on huge amounts of data and computing power, which makes them very costly, difficult to implement, and results in low stability.

[0070] This embodiment extends the 2D diffusion model to a 3D video generation network and combines it with the Controlnet control model, significantly reducing the data and computing power costs required for training the 3D video generation model and achieving the goal of being able to run it on a single data card. At the same time, the reuse of the 2D diffusion model ensures the quality of the generated video, and the extension of the temporal attention module effectively maintains video stability. Compared to previous solutions, this embodiment overcomes the shortcomings of both effectiveness and efficiency, generating high-quality video content with very low training and inference costs, making it widely applicable to various categories of video generation tasks. It should be noted that this embodiment can also be applied to other types of video content besides people. This is merely an example and is not limiting.

[0071] Optionally, this embodiment builds a video generation model based on an extension of a 2D logical data model (referred to as Latent Diffusion Model, or LDM for short), which can significantly reduce the data and computing power costs required for model training and simultaneously effectively reuse the generation functions of existing large-scale 2D paint models for video creation. At the same time, this embodiment also integrates a control function into the video generation model to extract pose motions from a reference video to control the content of the generated video, thereby effectively improving the stability and controllability of the generated video content. Furthermore, this embodiment combines the above video generation technology with a frame interpolation algorithm to extract keyframes to generate an initial video with a low frame rate, and then uses a neural network-based frame interpolation algorithm to increase the video frame rate, thereby significantly improving the efficiency and fluency of video generation.

[0072] Optionally, this embodiment integrates a video generation model and a posture control signal to support user input of content-referred video, which is then combined with content description text to generate video content, thereby effectively improving the controllability and generation quality of the generated video content. By extracting skeletal animation sequences from differentiated content-referred videos and combining them with diverse content description text, this embodiment achieves precise control over video content and can output video content with different styles, characters, and scenarios. This allows this embodiment to output high-quality video at the minute level using a single GPU. This efficient, low-cost video generation solution has great potential for application in video and film production, game promotional videos, and animation production, and can meet diverse and customized video production needs.

[0073] To further illustrate, optionally, in this embodiment, the initial input is a content-referencing video and a content-description text description provided by a user, as shown in FIG. 7. The content-referencing video is used to extract posture and motion to control the human motion in the generated video, and the content-description text is used to control specific content, such as the human image, style, and background, in the generated video. The extracted posture and motion and text semantic information are integrated into a video diffusion model and used to generate the initial video. The initial video can then be subjected to post-processing methods, such as video frame interpolation and super-resolution reconstruction, to generate the final video effect.

[0074] This embodiment can be used to generate video content related to a person. The input content-referred video should include the person's movements, and pose and movement are extracted to control the content of the generated video. Specifically, this embodiment uses the openpose algorithm to extract human body keypoints (18 keypoints including the head and torso) from each frame of the content-referred video, visualizes the keypoints as skeletal connections, and obtains a transformed pose and movement diagram. The transformed pose and movement diagram is used as input for the ControlNet to guide the generation of the video content. Due to the limited video memory of a single GPU, there is an upper limit to the number of video frames that can be generated at one time (a single A100 GPU can generate 70 frames of video at a time). To extend the length of the generated video, this embodiment can extract an 8-fps keyframe sequence from the content-referred video for movement sequence extraction, improving processing efficiency and the length of the generated video.

[0075] To further illustrate, optionally, an image 802 of a current video frame is determined from a person video (i.e., a content-referenced video including a person), and key points of the human body are extracted from the image 802, and the key points are visualized as skeletal connectivity relationships to obtain a transformed posture-motion diagram 804, for example as shown in FIG. 8 .

[0076] Optionally, in this embodiment, as shown in FIG. 9, the video diffusion model is composed of four parts: a latent variable decoder, a text encoder, a video diffusion model, and a posture control module. First, a text encoder is used to extract text semantic features from the input content description text, and a posture control module is used to extract video reference features, i.e., action features, from the content reference video. Specifically, this embodiment follows the stable diffusion model, using a CLIP Text Encoder for text encoding and a ControlNet as the posture control module. Next, the text semantic features and the video reference features are simultaneously input into the video diffusion model for calculation, and latent variable features of the generated video are obtained. The latent variable features are finally converted into video content through processing by the latent variable decoder.

[0077] The video diffusion model inherits the basic network structure of the traditional 2D diffusion model (LDM). This structure uses U-Net as the basis and processes input features using networks such as 2D convolutional layers, upsampling layers, downsampling layers, and self-attention mechanisms. This traditional 2D network structure can only generate a single image. To directly generate video content, this embodiment extends the 2D diffusion model into a 3D video generation model. Specifically, this embodiment extends the original 3x3 convolutional layers in U-Net to 1x3x3 convolutional layers, and adds a temporal attention layer to the network to improve the model's understanding of sequential frames and the stability of generating consecutive frames of video. Before deployment, this embodiment uses a small amount of video data to train the network. During the training process, the parameters of the original 2D U-Net are fixed, and only the newly added temporal attention module is trained. Such a design can make full use of traditional 2D model resources to generate videos with different styles and concepts, while the original 2D diffusion model parameters are fixed during model training, and only a very small amount of data training is required to give the 3D video generation model video generation capabilities.

[0078] To further enhance the control capabilities over video content, this embodiment also incorporates a posture control module into the video generation system to output highly controllable video content. Specifically, this embodiment uses a Controlnet model, which can realize motion control of image content drawn in 2D AI painting through input of skeletal movements. In this system, this embodiment inputs 8fps skeletal sequence frames extracted from content-reference video into the posture control module, which processes the skeletal sequence frame by frame and stitches together features from multiple frames to guide the calculation of the video generation model. The output of this model is the latent variable features of the video content, which are finally converted into a low-frame-rate (8fps) initial video via a latent variable decoder.

[0079] Optionally, in order to further improve the smoothness of the video, the embodiment extends the frame rate of the generated initial video to 32 fps through video interpolation. Specifically, the embodiment first decomposes the generated initial video into single frames, and then uses a neural network-based FILM algorithm to perform a 4x frame interpolation process on the frame sequence, thereby merging the newly generated frames into a high frame rate video, which can greatly improve the smoothness.

[0080] In the embodiment provided by this application, the 2D diffusion model is extended to a 3D video generation model and combined with the Controlnet control model, significantly reducing the data and computational costs required to train the 3D video generation model and enabling it to be run on a single data card. At the same time, this reuse of the 2D diffusion model ensures the quality of the generated video, and the extension of the temporal attention module effectively ensures video stability. Furthermore, compared to previous solutions, this embodiment simultaneously compensates for the shortcomings of both effectiveness and efficiency, enabling the generation of high-quality video content with very low training and inference costs, and can be widely used for various categories of video generation tasks.

[0081] In specific embodiments of the present application, where relevant data such as user information is involved, it is understood that when the above examples of the present application are applied to a particular product or technology, user permission or consent is required, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions.

[0082] Although the above-described method embodiments are all expressed as a combination of a series of operations for ease of explanation, those skilled in the art should understand that the present application is not limited to the order of operations described, because certain steps may be performed in other orders or simultaneously according to the present application. Next, those skilled in the art should also recognize that all the embodiments described herein are preferred embodiments, and related operations and modules are not necessarily essential to the present application.

[0083] According to another aspect of the embodiment of the present application, there is further provided a video generation device for implementing the above video generation method. As shown in Figure 10, the device includes: a first acquisition unit 1002 for acquiring a content description text and a content reference video, the content description text including information for describing a target content represented by a target video to be generated, and the content reference video including action reference information related to the target content; an extraction unit 1004, which performs feature extraction on the content description text to obtain text semantic features for characterizing semantic information of the content description text, and performs feature extraction on the content-referenced video to obtain video-referenced features for characterizing action-referenced information in the content-referenced video; and a generating unit 1006 for generating a target video based on the text semantic features and the video reference features.

[0084] For a specific embodiment, reference can be made to the example given in the video generation method above, which will not be described in detail here.

[0085] In an alternative embodiment, the generating unit 1006: a first determining module for determining at least one video element in the target video based on the text semantic features, the at least one video element including a first subject object; a second determining module for determining a pose change situation of the first subject object in the target video based on the video reference features; an acquisition module for acquiring the target video based on at least one video element and a pose change of the first subject object in the target video.

[0086] For a specific embodiment, reference can be made to the example given in the video generating device above, which will not be described in detail here.

[0087] In an alternative embodiment, the extraction unit 1004 comprises: an extraction module that performs feature extraction on a second subject object in the content-referred video to obtain object representation features, the object representation features being used to characterize a posture change situation of the second subject object in the content-referred video, and the video-referred features including the object representation features; The generating unit 1006 and a third determination module for determining a posture change situation of the first subject object in the target video based on the object representation features, wherein the posture change situation of the first subject object in the target video corresponds to a posture change situation of the second subject object in the content reference video.

[0088] For a specific embodiment, reference can be made to the example given in the video generating device above, which will not be described in detail here.

[0089] In an alternative embodiment, the extraction module: an extraction sub-module for performing feature extraction on at least two target video frames including a second subject object in the content-referred video to obtain at least two object static features, the object static features being used to characterize the position and shape of the second subject object in the target video frames; and a processing sub-module that integrates and processes at least two object static features to obtain object dynamic features based on timing relationship information between at least two target video frames, wherein the object dynamic features are used to characterize the posture change situation of the second subject object in the content-referred video, and the object expression features include the object dynamic features.

[0090] For a specific embodiment, reference can be made to the example given in the video generating device above, which will not be described in detail here.

[0091] In an alternative embodiment, the extraction sub-module: a first extraction subunit for performing keypoint extraction on a second subject object in at least two target video frames to obtain at least two keypoint features, the keypoint features being used to characterize keypoint positions of the second subject object in the target video frames, and the object static features including the keypoint features; a second extraction subunit for performing key line extraction on a second subject object in at least two target video frames to obtain at least two key line features, the key line features being used to characterize key line positions of the second subject object in the target video frames, and the object static features including the key line features; a third extraction subunit for performing contour extraction on a second subject object in at least two target video frames to obtain at least two contour features, the contour features being used to characterize a shape position of the contour of the second subject object in the target video frames, and the object static features including the contour features; a fourth extraction subunit for performing edge extraction on a second subject object in at least two target video frames to obtain at least two first object features, where the object static features include the first object features; a fifth extraction subunit for performing depth extraction on a second subject object in at least two target video frames to obtain at least two second object features, where the object static features include the second object features; a sixth extraction subunit for performing white model extraction on a second subject object in at least two target video frames to obtain at least two third object features, wherein the object static features include the third object features; Contains at least one of the following:

[0092] For a specific embodiment, reference can be made to the example given in the video generating device above, which will not be described in detail here.

[0093] In an alternative embodiment, the generating unit 1006: The video generation model includes an input module that inputs the text semantic features and the video reference features into a video generation model to obtain a target video to be output by the video generation model, the video generation model being a neural network model for generating videos obtained by training based on a plurality of video sample data.

[0094] For a specific embodiment, reference can be made to the example given in the video generating device above, which will not be described in detail here.

[0095] In an alternative embodiment, the device further comprises: a second acquisition unit for acquiring an image generation model before inputting the text semantic features and the video reference features into the video generation model, the image generation model being a neural network model for generating images obtained by training based on a plurality of image sample data; A determination unit for adjusting the image generation model to obtain an initial video generation model, the initial video generation model being composed of a convolution layer and an attention layer capable of processing time dimension information; a training unit for obtaining a video generation model by training an initial video generation model based on a plurality of video sample data.

[0096] For a specific embodiment, reference can be made to the example given in the video generating device above, which will not be described in detail here.

[0097] In an alternative embodiment, the input module comprises: an input sub-module that invokes a single graphics processor unit to execute a video generation model to process the input text semantic features and video reference features and obtain a target video output by the video generation model; The apparatus further comprises: The frame interpolation sub-module includes: a frame interpolation sub-module that invokes a single graphics processor unit to execute a video generation model to process text semantic features and video reference features, obtains a target video output by the video generation model, and then inserts related video frames into a video frame sequence corresponding to the target video to obtain a new video, wherein the length of the video corresponding to the new video is greater than the length of the video corresponding to the target video.

[0098] For a specific embodiment, reference can be made to the example given in the video generating device above, which will not be described in detail here.

[0099] According to another aspect of the embodiment of the present application, there is further provided an electronic device for implementing the above-mentioned video generation method, which may be, but is not limited to, the user device 102 or the server 112 shown in Fig. 1. This embodiment is described by taking the electronic device as the user device 102 as an example, and further, as shown in Fig. 11, the electronic device includes a memory 1102 and a processor 1104, the memory 1102 stores a computer program, and the processor 1104 is configured to execute the steps of any of the above-mentioned method embodiments according to the computer program.

[0100] Optionally, in this embodiment, the electronic device may be located in at least one of a plurality of network devices of a computer network.

[0101] Optionally, those skilled in the art will understand that the configuration shown in Figure 11 is merely exemplary and does not limit the configuration of the electronic device described above. For example, the electronic device may include more or fewer components (such as a network interface) than those shown in Figure 11, or may have a different configuration than that shown in Figure 11.

[0102] The memory 1102 may store software programs and modules, such as program instructions / modules corresponding to the video generation method and apparatus of the present application. The processor 1104 executes the software programs and modules stored in the memory 1102 to perform various functional applications and data processing, thereby realizing the video generation method. The memory 1102 may include high-speed random access memory or non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some examples, the memory 1102 may further include memory located remotely from the processor 1104, and these remote memories may be connected to the electronic device via a network. Examples of such networks include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof. The memory 1102 may specifically be used to store information such as, but not limited to, content description text, content reference videos, and target videos. 11, the memory 1102 may include, but is not limited to, the first acquisition unit 1002, the extraction unit 1004, and the generation unit 1006 in the video generation device. It may also include, but is not limited to, other module units in the video generation device. In this embodiment, they will not be described in detail.

[0103] Optionally, the transmission device 1106 transmits and receives data via a network. Specific examples of the network include a wired network and a wireless network. In one example, the transmission device 1106 includes a network interface controller (NIC) that can connect to other network devices or routers via a network cable to communicate with the Internet or a local area network. In one example, the transmission device 1106 includes a radio frequency (RF) module for wirelessly communicating with the Internet.

[0104] The electronic device further includes a display 1108 for displaying information such as the content description text, content reference video, and target video, and a connection bus 1110 for connecting each modular component in the electronic device.

[0105] In another embodiment, the user device or server may be a node in a distributed system, which may be a blockchain system, which may be formed by connecting the plurality of nodes via network communication. A peer-to-peer network may be formed among the nodes, and any type of computing device, such as a server or a user device, may become a node in the blockchain system by joining the peer-to-peer network.

[0106] According to one aspect of the present application, a computer program product is provided that includes computer programs / instructions, the computer programs / instructions including program code for performing the method illustrated in the flowcharts. In such an embodiment, the computer program may be downloaded and installed from a network via a communication unit and / or from removable media. When the computer program is executed by a central processing unit, various functions defined in the system of the present application are performed.

[0107] The numbers of the examples in the present application above are for illustrative purposes only and do not represent the superiority or inferiority of the examples.

[0108] It should be noted that the computer system of the electronic device is merely an example and does not limit the functions and scope of use of the embodiments of the present application.

[0109] A computer system includes a central processing unit (CPU) that performs various appropriate operations and processes based on programs stored in read-only memory (ROM) or programs loaded from the storage unit into random access memory (RAM). The random access memory also stores various programs and data required for the operation of the system. The CPU, ROM, and RAM are interconnected via a bus. An input / output (I / O) interface is also connected to the bus.

[0110] The following components are connected to the I / O interface: an input unit including a keyboard, mouse, etc.; an output unit including a cathode ray tube (CRT), a liquid crystal display (LCD), and speakers, etc.; a storage unit including a hard disk, etc.; and a communication unit including a network interface card such as a LAN card or modem. The communication unit performs communication processing via a network such as the Internet. Drives are connected to the I / O interface as needed. Removable media such as magnetic disks, optical disks, magneto-optical disks, and semiconductor memories are loaded into the drives as needed, and computer programs read from them are installed in the storage unit as needed.

[0111] In particular, according to an embodiment of the present application, the procedures illustrated in each method flowchart are implemented as a computer software program. For example, an embodiment of the present application includes a computer program product including a computer program stored on a computer-readable medium, the computer program including program code for executing the method illustrated in the flowchart. In such an embodiment, the computer program may be downloaded and installed from a network via a communication unit and / or from removable media. When the computer program is executed by a central processing unit, various functions defined in the system of the present application are performed.

[0112] According to one aspect of the present application, a computer-readable storage medium is provided, and a processor of a computer device reads computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, thereby causing the computer device to perform the method according to the various selectable implementations described above.

[0113] Optionally, in this embodiment, those skilled in the art can understand that all or part of the steps in each method of the above embodiment can be completed by instructing relevant hardware of an electronic device through a program. The program can be stored in a computer-readable storage medium, and the storage medium can include a flash memory disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, an optical disk, etc.

[0114] The numbers of the examples in the present application above are for illustrative purposes only and do not represent the superiority or inferiority of the examples.

[0115] The integrated units in the above embodiments can be realized in the form of software functional units and stored in the above computer-readable storage medium when sold or used as an independent product. Based on this understanding, the technical solutions of the present application, or the portions contributing to existing technology, or all or part of the technical solutions, can be embodied in the form of a software product. The computer software product is stored in a storage medium. The computer software product includes several instructions for causing one or more computer devices (such as personal computers, servers, or network devices) to execute all or part of the steps of the methods described in each embodiment of the present application.

[0116] In the above embodiments of the present application, the description of each embodiment has its own emphasis, and for the parts not described in detail in one embodiment, reference can be made to the relevant descriptions of other embodiments.

[0117] It should be understood that in some embodiments provided by the present application, the disclosed user device can be realized in other ways. Furthermore, the device embodiments described above are merely exemplary. For example, the division of the above units is merely a division of logical functions. In actual implementation, other division methods may exist. For example, multiple units or components may be combined or integrated into other systems, or some features may be ignored or not implemented. On the other hand, the shown or discussed mutual couplings or direct couplings or communication connections may be indirect couplings or communication connections via some interfaces, units, or modules, and may be electrical or other forms.

[0118] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, and may be located in one place or distributed across multiple network units. According to actual needs, some or all of the units may be selected to achieve the objectives of the technical solutions of this embodiment.

[0119] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, may exist physically alone, or two or more units may be integrated into one unit. The integrated unit may be realized in the form of hardware or in the form of a software functional unit.

[0120] The above are only preferred embodiments of the present application, and those skilled in the art may make some improvements and modifications without departing from the principles of the present application, and these improvements and modifications should also fall within the protection scope of the present application.

Claims

1. 1. A video generation method implemented by an electronic device, comprising: obtaining a content description text and a content reference video, wherein the content description text includes information for describing a target content represented by a target video to be generated, and the content reference video includes action reference information related to the target content; performing feature extraction on the content description text to obtain text semantic features, wherein the text semantic features are used to characterize semantic information of the content description text; performing feature extraction on the content-referred video to obtain video-referred features, the video-referred features being used to characterize the action-referred information in the content-referred video; generating the target video based on the text semantic features and the video reference features; A method comprising:

2. generating the target video based on the text semantic features and the video reference features, determining at least one video element in the target video based on the text semantic features, wherein the at least one video element includes a first subject object; determining a pose change situation of the first subject object in the target video based on the video reference features; generating the target video based on the at least one video element and a pose change situation of the first subject object in the target video; 10. The method of claim 1, comprising:

3. performing feature extraction on the content-referenced video to obtain video-referenced features; performing feature extraction on a second subject object in the content-referred video to obtain object representation features, the object representation features being used to characterize a pose change situation of the second subject object in the content-referred video, and the video-referred features including the object representation features; Including, The step of determining a posture change situation of the first subject object in the target video based on the video reference features includes: determining a posture change situation of the first subject object in the target video based on the object representation features, wherein the posture change situation of the first subject object in the target video corresponds to a posture change situation of the second subject object in the content reference video; 3. The method of claim 2, comprising:

4. performing feature extraction on a second subject object in the content-referred video to obtain object representation features; performing feature extraction on at least two target video frames in the content-referred video, the target video frames including the second subject object, to obtain at least two object static features, the object static features being used to characterize a positional morphology of the second subject object in the target video frames; synthesizing the at least two object static features to obtain object dynamic features based on timing relationship information between the at least two target video frames, the object dynamic features being used to characterize a posture change situation of the second subject object in the content-referred video, and the object representation features including the object dynamic features; 4. The method of claim 3, comprising:

5. The step of performing feature extraction on at least two target video frames in the content-referred video including the second subject object to obtain at least two object static features includes: performing keypoint extraction on the second subject object in the at least two target video frames to obtain at least two keypoint features, the keypoint features being used to characterize keypoint positions of the second subject object in the target video frames, and the object static features including the keypoint features; performing keyline extraction on the second subject object in the at least two target video frames to obtain at least two keyline features, the keyline features being used to characterize keyline positions of the second subject object in the target video frames, and the object static features including the keyline features; performing contour extraction on the second subject object in the at least two target video frames to obtain at least two contour features, the contour features being used to characterize a morphological position of a contour of the second subject object in the target video frames, and the object static features include the contour features; performing edge extraction on the second subject object in the at least two target video frames to obtain at least two first object features, wherein the object static features include the first object features; performing depth extraction on the second subject object in the at least two target video frames to obtain at least two second object features, wherein the object static features include the second object features; performing white model extraction on the second subject object in the at least two target video frames to obtain at least two third object features, wherein the object static features include the third object features; The method of claim 4, comprising at least one of:

6. generating the target video based on the text semantic features and the video reference features, inputting the text semantic features and the video reference features into a video generation model to obtain the target video output by the video generation model, wherein the video generation model is a neural network model for generating videos obtained by training based on a plurality of video sample data; 6. The method of any one of claims 1 to 5, comprising:

7. Before inputting the text semantic features and the video reference features into a video generation model, the method further comprises: obtaining an image generation model, the image generation model being a neural network model for generating an image obtained by training based on a plurality of image sample data; A step of obtaining an initial video generation model for the image generation model, the initial video generation model being composed of a convolution layer and an attention layer capable of processing time dimension information; training the initial video generation model based on the plurality of video sample data to obtain the video generation model; 7. The method of claim 6, comprising:

8. inputting the text semantic features and the video reference features into a video generation model to obtain the target video output by the video generation model, invoking a single graphics processor unit to execute the video generation model to process the text semantic features and the video reference features to obtain the target video output by the video generation model; Including, After invoking the single graphics processor unit to execute the video generation model to process the text semantic features and the video reference features and obtain the target video output by the video generation model, the method further comprises: inserting related video frames into a video frame sequence corresponding to the target video to obtain a new video, wherein a video length corresponding to the new video is greater than a video length corresponding to the target video; 8. The method of claim 6 or 7, comprising:

9. 1. A video generation device, comprising: a first acquisition unit for acquiring a content description text and a content reference video, the content description text including information for describing a target content represented by a target video to be generated, and the content reference video including action reference information related to the target content; an extraction unit performing feature extraction on the content description text to obtain text semantic features for characterizing semantic information of the content description text, and performing feature extraction on the content-referenced video to obtain video-referenced features for characterizing the action-referenced information in the content-referenced video; a generation unit for generating the target video based on the text semantic features and the video reference features; 10. A video generation device comprising:

10. A computer-readable storage medium containing a program stored thereon, A computer-readable storage medium that, when executed by an electronic device, performs the method of any one of claims 1 to 8.

11. A computer program product comprising computer programs / instructions, A computer program product, the computer program / instructions of which, when executed by a processor, implement the steps of the method according to any one of claims 1 to 8.

12. An electronic device including a memory and a processor, An electronic device, the memory storing a computer program, the processor being configured to execute the method of any one of claims 1 to 8 by means of the computer program.

Citation Information

Patent Citations

  • Depicting Humans in Text-Defined Outfits

    US20210272341A1