Video generation method and device, computer program product and electronic equipment

The process of event description text and video frame characteristics through timing position encoding and timing cross attention mechanisms is solved, and the problem of difficult event sequence and duration control in multi-event videos is achieved, and high-quality multi-event video generation is achieved.

CN120017929APending Publication Date: 2025-05-16NETEASE (HANGZHOU) NETWORK CO LTD
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202510168668.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-14
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The prior art is difficult to accurately control the event sequence and duration of multi-event videos, resulting in the disordered sequence of events in the generated videos, the duration does not meet expectations, and the unnatural transition, which affects the generation quality of multi-event videos.

Method used

By receiving the event description text sequence, the reference frame feature vector of the video to be generated is determined, and the timing position encoding process is performed based on the start and end time information, the video frame and text features are fused, and the target video is generated using the timing cross attention mechanism.

Benefits of technology

Accurate time control of multi-event videos is achieved, ensuring the correct sequence of events and reasonable duration, and generating coherent and smooth videos, with natural transitions between events.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120017929A_ABST
    Figure CN120017929A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of video processing, and relates to a video generation method and device, a computer program product and electronic equipment. The video generation method comprises the steps that an event description text sequence is received, a reference frame feature vector of a video to be generated is determined according to the event description text sequence, and the event description text sequence comprises text description information of multiple events and corresponding starting and ending time information; performing time sequence position coding processing on the reference frame feature vector according to the starting and ending time information to obtain a video frame time sequence fusion feature, and performing time sequence position coding processing on the text description information to obtain a text time sequence fusion feature; based on a time sequence cross attention mechanism, fusing the video frame time sequence fusion feature and the text time sequence fusion feature to obtain a target fusion feature; and generating a dynamic image according to the target fusion feature corresponding to each reference frame feature vector to obtain a target video. The quality of the generated multi-event video can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of video processing technology, and more specifically, to a video generating method, a video generating device, a computer program product, and an electronic device. Background Art

[0002] With the development of computer and video processing technology, deep learning-based video generation has been widely used in various technical fields. For example, games and short video production. At present, high-quality video images can be produced by generating a single scene or event video through a single text or image description. However, for complex scenes of multi-event video generation, it is difficult to accurately control the start time and duration of each event, which leads to the disorder of the event sequence in the generated video, the duration does not meet expectations, and the transition is not smooth, which affects the generation quality of multi-event videos to a certain extent.

[0003] It should be noted that the information disclosed in the above background technology section is only used to enhance the understanding of the background of the present disclosure, and therefore may include information that does not constitute the prior art known to ordinary technicians in the field. Summary of the invention

[0004] The purpose of the present disclosure is to provide a video generation method and apparatus, a computer program product and an electronic device, thereby improving the quality of generated multi-event videos at least to a certain extent.

[0005] Other features and advantages of the present disclosure will become apparent from the following detailed description, or may be learned in part by the practice of the present disclosure.

[0006] According to one aspect of the present disclosure, a video generation method is provided, including: receiving an event description text sequence, and determining a reference frame feature vector of a to-be-generated video according to the event description text sequence, wherein the event description text sequence includes text description information of a plurality of events and corresponding start and end time information; performing temporal position encoding processing on the reference frame feature vector according to the start and end time information to obtain a video frame temporal fusion feature, and performing temporal position encoding processing on the text description information to obtain a text temporal fusion feature; based on a temporal cross-attention mechanism, fusing the video frame temporal fusion feature and the text temporal fusion feature to obtain a target fusion feature; and performing dynamic image generation according to the target fusion feature corresponding to each reference frame feature vector to obtain a target video.

[0007] In an exemplary embodiment of the present disclosure, a reference frame feature vector of a video to be generated is determined based on an event description text sequence, including: determining a reference image frame of the video to be generated based on text description information of multiple events and corresponding start time information, wherein each event corresponds to at least one reference image frame; and numerically encoding the reference image frame to obtain a reference frame feature vector.

[0008] In an exemplary embodiment of the present disclosure, a reference frame feature vector is temporally encoded based on start and end time information to obtain a video frame temporal fusion feature, including: for each event, the occurrence time of the reference image frame corresponding to the event is standardized based on the start and end time information corresponding to the event to obtain a standard occurrence time; a feature rotation angle is calculated based on the standard occurrence time and a reference angular velocity; and a reference frame feature vector corresponding to the standard occurrence time is rotated based on the feature rotation angle to obtain a video frame temporal fusion feature.

[0009] In an exemplary embodiment of the present disclosure, for each event, the occurrence time of the reference image frame corresponding to the event is standardized according to the start and end time information corresponding to the event to obtain a standard occurrence time, including: determining the time period position corresponding to the event according to the start and end time information corresponding to the event; and standardizing the occurrence time of the reference image frame corresponding to the event based on the start and end time information corresponding to the event, the time period position and a preset standard event time length to obtain a standard occurrence time.

[0010] In an exemplary embodiment of the present disclosure, a reference frame feature vector corresponding to the standard occurrence moment is rotated according to a feature rotation angle to obtain a video frame temporal fusion feature, including: splitting the reference frame feature vector corresponding to the standard occurrence moment into multiple dimensional components; dividing the multiple dimensional components into multiple component groups in the manner of each adjacent dimension as a group; for each component group, rotating the component group in a two-dimensional plane based on the feature rotation angle to obtain a time fusion component group; and determining the video frame temporal fusion feature according to the time fusion component groups corresponding to each component group.

[0011] In an exemplary embodiment of the present disclosure, text description information is subjected to temporal position encoding processing to obtain text temporal fusion features, including: encoding the text description information to obtain text embedding information; determining the text time information of the text embedding information in the video to be generated according to the start and end time information of the event corresponding to the text description information; encoding the text time information into the text embedding information to obtain text temporal fusion features.

[0012] In an exemplary embodiment of the present disclosure, based on the temporal cross-attention mechanism, the video frame temporal fusion features and the text temporal fusion features are fused to obtain the target fusion features, including: mapping the video frame temporal fusion features to the query space to obtain the query vector; mapping the text temporal fusion features to the key-value space to obtain the key vector, and mapping the text temporal fusion features to the key-value space to obtain the value vector, wherein the mapping matrices used to obtain the key vector and the value vector are different; performing attention calculation according to the query vector, the key vector and the value vector to obtain the target fusion features.

[0013] In an exemplary embodiment of the present disclosure, based on the temporal cross-attention mechanism, the video frame temporal fusion features and the text temporal fusion features are fused to obtain the target fusion features, and also include: receiving scene switching information, and obtaining the scene embedding vector corresponding to the scene switching information according to the scene switching type; performing temporal position encoding processing on the scene embedding vector according to the start and end time information to obtain the scene temporal fusion features; fusing the scene temporal fusion features and the text temporal fusion features to obtain the updated text temporal fusion features; based on the temporal cross-attention mechanism, the video frame temporal fusion features are fused with the updated text temporal fusion features to obtain the target fusion features.

[0014] In an exemplary embodiment of the present disclosure, before performing temporal position encoding processing on the scene embedding vector according to the start and end time information to obtain the scene temporal fusion feature, the method also includes: performing learning training on video frame samples annotated with different types of scene switching information to obtain scene embedding vectors corresponding to different scene switching types, and the scene embedding vectors of different scene switching types are learnable.

[0015] In an exemplary embodiment of the present disclosure, dynamic image generation is performed according to target fusion features corresponding to feature vectors of each reference frame to obtain a target video, including: according to the target fusion features corresponding to the feature vectors of each reference frame, content generation is performed through a diffusion model to obtain a target video; wherein, based on the target fusion features fused with updated text temporal fusion features, scene transition frames are generated in the target video.

[0016] According to one aspect of the present disclosure, a video generation device is provided, including: an information acquisition module, used to receive an event description text sequence, and determine a reference frame feature vector of a to-be-generated video according to the event description text sequence, wherein the event description text sequence includes text description information of a plurality of events and corresponding start and end time information; an encoding processing module, used to perform temporal position encoding processing on the reference frame feature vector according to the start and end time information to obtain a video frame temporal fusion feature, and perform temporal position encoding processing on the text description information to obtain a text temporal fusion feature; an attention module, used to fuse the video frame temporal fusion feature and the text temporal fusion feature based on a temporal cross-attention mechanism to obtain a target fusion feature; and a content generation module, used to generate dynamic images according to the target fusion feature corresponding to each reference frame feature vector to obtain a target video.

[0017] According to one aspect of the present disclosure, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, any one of the above methods is implemented.

[0018] According to one aspect of the present disclosure, an electronic device is provided, comprising: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to perform any one of the above methods by executing the executable instructions.

[0019] The video generation method in the exemplary embodiment of the present disclosure receives an event description text sequence, and determines the reference frame feature vector of the video to be generated according to the event description text sequence, wherein the event description text sequence includes text description information of multiple events and corresponding start and end time information; performs temporal position encoding processing on the reference frame feature vector according to the start and end time information to obtain the video frame temporal fusion feature, and performs temporal position encoding processing on the text description information to obtain the text temporal fusion feature; based on the temporal cross attention mechanism, the video frame temporal fusion feature and the text temporal fusion feature are fused to obtain the target fusion feature; and dynamic image generation is performed according to the target fusion feature corresponding to each reference frame feature vector to obtain the target video. On the one hand, the process can make the content and timing of the video frame consistent with the text description based on the received event description text sequence, by performing temporal position encoding on the video frame feature vector and the text description information, so as to realize the time control at the frame level, thereby generating a multi-event video with precise time control. On the other hand, dynamic image generation is performed according to the target fusion feature corresponding to each reference frame feature vector, which fully integrates the text description and video frame feature vector of multiple events, so that the target video is coherent and smooth, and the transition of each event is natural.

[0020] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] The above and other objects, features and advantages of the exemplary embodiments of the present disclosure will become readily understood by reading the following detailed description with reference to the accompanying drawings, in which several embodiments of the present disclosure are shown in an exemplary and non-limiting manner.

[0022] Figure 1 An application environment according to an exemplary embodiment of the present disclosure is shown.

[0023] Figure 2 A flowchart of a video generating method according to an exemplary embodiment of the present disclosure is shown.

[0024] Figure 3 A flow chart of obtaining temporal fusion features of video frames according to an exemplary embodiment of the present disclosure is shown.

[0025] Figure 4 A flowchart of an implementation method for acquiring temporal fusion features of video frames according to an exemplary embodiment of the present disclosure is shown.

[0026] Figure 5 A flowchart of acquiring text temporal fusion features according to an exemplary embodiment of the present disclosure is shown.

[0027] Figure 6 A flow chart of acquiring target fusion features according to an exemplary embodiment of the present disclosure is shown.

[0028] Figure 7 A flowchart of yet another target fusion feature according to an exemplary embodiment of the present disclosure is shown.

[0029] Figure 8 A schematic diagram showing the composition of a video generating device according to an exemplary embodiment of the present disclosure is shown.

[0030] Fig. 9 A block diagram of an electronic device according to an exemplary embodiment of the present disclosure is shown.

[0031] In the drawings, the same or corresponding reference numerals represent the same or corresponding parts. DETAILED DESCRIPTION

[0032] The exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, the exemplary embodiments can be implemented in a variety of forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that the present disclosure will be more comprehensive and complete and the concepts of the exemplary embodiments will be fully conveyed to those skilled in the art. The same reference numerals in the figures represent the same or similar structures, and thus their detailed description will be omitted.

[0033] In addition, the described features, structures or characteristics may be combined in one or more embodiments in any suitable manner. In the following description, many specific details are provided to provide a full understanding of the embodiments of the present disclosure. However, those skilled in the art will appreciate that the technical solutions of the present disclosure may be practiced without one or more of the specific details, or other methods, components, devices, steps, etc. may be adopted. In other cases, known structures, methods, devices, implementations or operations are not shown or described in detail to avoid blurring the various aspects of the present disclosure.

[0034] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities may be implemented in software form, or these functional entities or parts of functional entities may be implemented in one or more software hardened modules, or these functional entities may be implemented in different networks and / or processor devices and / or microcontroller devices.

[0035] At present, it is difficult to accurately control the start time and duration of each event in the complex scenes of multi-event video generation, which leads to the disorder of event sequence in the generated video and the duration not meeting expectations. In addition, the separately generated video clips are simply spliced ​​together, resulting in unclear scene transitions and abrupt switches, which affects the generation quality of multi-event videos to a certain extent.

[0036] Based on this, an exemplary embodiment of the present disclosure provides a video generation method, which makes the content and timing of the video frame consistent with the text description by temporal position encoding of the text description and the video frame feature vector, and combines the temporal cross-attention mechanism to achieve frame-level time control, thereby generating a multi-event video with precise time control, making the target video coherent and smooth, and the transition of each event is natural.

[0037] It should be noted that the video generation method of the exemplary embodiment of the present disclosure can be applied to the fields of advertising creation, image production, virtual reality, animation production, etc., without limitation thereto.

[0038] The video generation method provided by the exemplary embodiment of the present disclosure can be applied to Figure 1 In the application environment shown, the terminal 101 communicates with the server 102 via a network. The data storage system can store data that the server 102 needs to process. The data storage system can be integrated on the server 102, or placed on the cloud or other network servers.

[0039] In an exemplary embodiment, the video generation method provided by the exemplary embodiment of the present disclosure may be executed by the server 102, and the corresponding video generation device is disposed in the server 102. Correspondingly, in this mode of execution by the server 102, the server 102 may start to execute the steps in the technical solution of the exemplary embodiment of the present disclosure in response to a trigger command, wherein the trigger command may be sent by a terminal used by a user, or may be triggered locally by the server in response to some automation events.

[0040] The server 102 may be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The server 102 may perform background tasks.

[0041] Furthermore, in another exemplary embodiment, the terminal 101 may also have similar functions to the server 102 , so as to execute the video generating method provided by the exemplary embodiment of the present disclosure.

[0042] The terminal 101 may be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, an IoT device, and a portable wearable device. The IoT device may be a smart speaker, a smart TV, a smart air conditioner, and a smart vehicle-mounted device, etc. The portable wearable device may be a smart watch, a smart bracelet, a head-mounted device, etc. The terminal 101 may also be referred to as a mobile terminal, a terminal device, a mobile device, etc. The exemplary embodiments of the present disclosure do not limit the type of the terminal 101.

[0043] In addition, the technical solution of the exemplary embodiment of the present disclosure can also be executed collaboratively by the terminal 101 and the server 102. In this manner of collaborative execution by the terminal 101 and the server 102, some steps in the technical solution provided by the exemplary embodiment of the present disclosure are executed by the terminal 101, while other steps are executed by the server 102. It should be noted that in this manner of collaborative execution by the terminal 101 and the server 102, the steps respectively executed by the terminal 101 and the server 102 can be dynamically adjusted according to actual conditions, and there is no special limitation on this.

[0044] The terminal 101 and the server 102 may be connected directly or indirectly via wireless communication, and the exemplary embodiments of the present disclosure do not impose any special limitation thereto.

[0045] like Figure 2 FIG. 1 is a flowchart of a video generation method according to an exemplary embodiment of the present disclosure. The video generation method includes steps S210 to S240, which are specifically as follows:

[0046] In step S210, an event description text sequence is received, and a reference frame feature vector of a video to be generated is determined according to the event description text sequence, wherein the event description text sequence includes text description information of multiple events and corresponding start and end time information.

[0047] In the exemplary embodiments of the present disclosure, an event refers to a certain thing or an action that occurs at an instant, and an event description text is a detailed description of the action or situation in which the event occurs. For example, "a person first walks to the window", "opens the window", and "looks out" are all event description texts. The event description text sequence can be provided by the user, that is, the user can specify the text description information of multiple events and the start and end time information of each event. For example, "a person first walks to the window, then opens the window, and finally looks out", the start and end time information of the event "a person first walks to the window" is 0 to 3 seconds, the start and end time information of the event "opens the window" is 3 to 5 seconds, and the start and end time information of the event "looks out" is 5 to 8 seconds. Based on this, the time period of each event on the timeline of the video to be generated can be determined according to the event description sequence to ensure that the subsequent generation process is carried out in the time order set by the user.

[0048] Among them, the reference frame feature vector refers to the frame feature representation corresponding to the reference image frame corresponding to the video to be generated. The reference image frame is the potential representation of the current video corresponding to each time step when the diffusion model is used for time processing. Each reference image frame has its own corresponding reference frame feature vector. The exemplary embodiment of the present disclosure is explained by taking the processing of a reference frame feature vector as an example. Subsequently, in each time step, it is necessary to perform cross-attention calculation on the current video potential representation and the text temporal fusion feature.

[0049] In an exemplary embodiment, the reference frame feature vector may be a numerical representation of each reference image frame of the video to be generated after being processed by the deep learning model. Specifically, determining the reference frame feature vector of the video to be generated according to the event description text sequence may include:

[0050] First, the reference image frame of the video to be generated is determined according to the text description information of multiple events and the corresponding start time information, and each event corresponds to at least one reference image frame; then the reference image frame is numerically encoded to obtain a reference frame feature vector.

[0051] Among them, digital coding means that the image itself is composed of pixels, and the "reference frame feature vector" converts the pixel information into a high-dimensional numerical vector, which can be understood as a highly generalized and abstract representation of the image content, including the visual information of the reference image frame, such as the shape, texture, color, etc. of the object. The processing by the deep learning model can be completed by a pre-trained deep learning model (for example, a convolutional neural network), which is trained to identify and extract important features in the image.

[0052] By extracting a high-dimensional vector of the reference frame feature vector, the dimension can reach hundreds or thousands of dimensions, which can capture the complex information of the reference image frame and provide comprehensive information input for subsequent processing.

[0053] Based on the event description text sequence, the correspondence between events and time can also be obtained, so that subsequent event description texts, reference frame feature vectors, etc. can be time-aligned.

[0054] In step S220, the reference frame feature vector is subjected to temporal position coding according to the start and end time information to obtain the video frame temporal fusion feature, and the text description information is subjected to temporal position coding to obtain the text temporal fusion feature.

[0055] In an exemplary embodiment of the present disclosure, the temporal position coding process is to perform a rotation operation on the feature vector to encode the time information. The temporal position coding process for the reference frame feature vector refers to adding time perception capability to each reference image frame and adding time perception capability to each text description information, that is, by temporal position coding, the time information is respectively integrated into the reference frame feature and the text description, so that when the image is generated, the temporal correspondence between the text description and the image content can be established according to the different time point information.

[0056] In step S230, based on the temporal cross-attention mechanism, the video frame temporal fusion features and the text temporal fusion features are fused to obtain the target fusion features.

[0057] In an exemplary embodiment of the present disclosure, the temporal cross-attention mechanism is mainly used to process the interaction between two temporal sequences (video frame temporal fusion features, text temporal fusion features), and dynamically focus on important time points by calculating the correlation between one sequence (query sequence) and another sequence (key-value sequence). When generating each frame of the video, the exemplary embodiment of the present disclosure can pay more attention to the relevant text description information according to the current moment. For example, when generating the picture of the event "walking to the window", it focuses more on understanding the meaning of the keywords "walk" and "window". For example, when the picture of the 2nd second of the video is being generated (corresponding to the event "walking to the window"), the temporal cross-attention mechanism is used to guide the model to pay more attention to the text description information describing "walking to the window" rather than the text description information describing "opening the window" or "looking out".

[0058] Among them, the temporal cross-attention processing process can be performed based on a temporal cross-attention model, such as a deep learning model based on a Transformer architecture, for example, DiT (Diffusion Transformer, a diffusion model combined with a Transformer architecture), each layer of which is responsible for fusing the temporal fusion features of the video frame and the temporal fusion features of the text.

[0059] In step S240, dynamic image generation is performed according to the target fusion features corresponding to the feature vectors of each reference frame to obtain a target video.

[0060] In an exemplary embodiment of the present disclosure, after obtaining the target fusion features of the feature vectors of each reference frame, a diffusion model such as DiT can be used to generate dynamic images to obtain a target video. In the denoising process, the text description information is used to control the generation process through a cross-attention mechanism, so that the content of the generated video frame is consistent with the text description, and the content and timing of the video frame are consistent with the text description.

[0061] In the exemplary embodiment of the present disclosure, the video generation method can, on the one hand, be based on the received event description text sequence, and by encoding the temporal position of the video frame feature vector and the text description information, make the content and timing of the video frame consistent with the text description, and realize the time control at the frame level, thereby generating a multi-event video with precise time control. On the other hand, dynamic image generation is performed according to the target fusion features corresponding to the feature vectors of each reference frame, which fully integrates the text description of multiple events and the video frame feature vectors, so that the target video is coherent and smooth, and the transition of each event is natural.

[0062] Each step is described in detail below.

[0063] In an exemplary embodiment, if Figure 3 As shown, the reference frame feature vector is subjected to temporal position encoding processing according to the start and end time information, and the obtained video frame temporal fusion features may include:

[0064] Step S310: For each event, the occurrence time of the reference image frame corresponding to the event is standardized according to the start and end time information corresponding to the event to obtain the standard occurrence time.

[0065] Taking into account that different event durations may vary, in order to avoid deviations caused by different event durations, the occurrence time of each reference image frame of the event is standardized, that is, no matter how long the actual duration of the event is, it is mapped to a preset standard event time length.

[0066] Specifically, for each event, the occurrence time of the reference image frame corresponding to the event is standardized according to the start and end time information corresponding to the event, and obtaining the standard occurrence time may include:

[0067] First, determine the time period corresponding to the event based on the start and end time information corresponding to the event; then, based on the start and end time information corresponding to the event, the time period position and the preset standard event time length, standardize the occurrence time of the reference image frame corresponding to the event to obtain the standard occurrence time.

[0068] The time period position refers to the event order of the event in the video to be generated, that is, the nth time period of an event in the video to be generated. The preset standard event time length can be set according to the time length requirement of the generated video, for example, set to 8 seconds, 12 seconds, etc. If a longer video needs to be generated, the standard event time length can be increased, and there is no limit on its specific value.

[0069] Taking the event of the nth time period of the video to be generated as an example, for any occurrence time t of the event (corresponding to the reference image frame), the process of standardizing the occurrence time t of the reference image frame corresponding to the event based on the start and end time information corresponding to the event, the time period position and the preset standard event time length can be expressed by Formula 1:

[0070] t'=((tT start ) / ( T end -T start ))*L+(n-1)*L Formula 1

[0071] Among them, t' is the standard occurrence time corresponding to t, n is the time period position, L is the preset standard event time length, T start is the start time of the event, T end is the end time of the event, expressed by (tT start ) / (T end -T start ) can determine the relative proportion of the current time t in the event duration.

[0072] The occurrence time of the reference image frame corresponding to each event is standardized in a similar manner to obtain a standard occurrence time.

[0073] Step S320: Calculate the characteristic rotation angle based on the standard occurrence time and the reference angular velocity.

[0074] The feature rotation angle is used to rotate the feature vector in the future. The reference angular velocity determines the encoding frequency of the time information. The time encoding frequency is set according to the needs. The feature rotation angle can be calculated by formula 2:

[0075] θ t =t'*ω Formula 2

[0076] Among them, θ t is the characteristic rotation angle, ω is the reference angular velocity, and the larger the standard occurrence time, the larger the characteristic rotation angle obtained.

[0077] Step S330: rotating the reference frame feature vector corresponding to the standard occurrence time according to the feature rotation angle to obtain the video frame temporal fusion feature.

[0078] Rotating the "feature vector" means integrating time information into each dimension of the feature vector to adjust certain directions of the "feature vector". The exemplary embodiment of the present disclosure rotates the reference frame feature vector corresponding to its time through different standard occurrence moments, and different moments correspond to different rotation angles, so that the obtained video frame time sequence fusion feature contains time sequence information.

[0079] Specifically, Figure 4 A flowchart of an implementation method for obtaining the temporal fusion features of video frames is shown, as Figure 4 As shown, the reference frame feature vector corresponding to the standard occurrence time is rotated according to the feature rotation angle, and the video frame temporal fusion feature obtained may include:

[0080] Step S410: split the reference frame feature vector corresponding to the standard occurrence moment into multiple dimensional components.

[0081] The reference frame feature vector is split into multiple dimensional components. If the reference frame feature vector is a d-dimensional feature vector, it is split according to the dimension to obtain v1, v2, …, v 2i ,v 2i+1 Dimensional components, i = 0, ..., d / 2-1, v 2i is the 2ith component of the reference frame feature vector, v 2i+1 is the 2i+1th component of the reference frame feature vector.

[0082] Step S420: Divide the multiple dimensional components into multiple component groups in such a way that each adjacent dimension is a group.

[0083] The components of each adjacent dimension are grouped together, that is, every two components (v 2i ,v 2i+1 ) as a group to obtain multiple component groups, so that two adjacent dimensions in the reference frame feature vector can be processed each time subsequently.

[0084] Step S430: For each component group, the component group is rotated in a two-dimensional plane based on a characteristic rotation angle to obtain a time-fused component group.

[0085] By rotating a vector in a two-dimensional plane, the modulus of the vector can be kept unchanged and only the direction of the vector is changed. That is, the numerical size information of the vector is retained while the time information is embedded.

[0086] Specifically, the exemplary embodiment of the present disclosure is based on a component group (v 2i ,v 2i+1 ) is rotated in the two-dimensional plane to convert it to a new coordinate representation. 2i ,v 2i+1 ), rotate it by formula 3 and formula 4 to get the time fusion component group:

[0087] vortated 2i =v 2i *cos(θ t )- v 2i+1 *sin(θ t ) Formula 3

[0088] vortated 2i+1 =v 2i *sin(θ t )+v 2i+1 *cos(θ t ) Formula 4

[0089] Among them, vortated 2i is the component v 2i The corresponding time fusion component, vortated 2i+1 is the component v 2i+1 The corresponding time fusion component is equivalent to the vector (v 2i ,v 2i+1 ) is rotated counterclockwise by θ t angle, and obtain the time fusion component group (vortated 2i ,vortated 2i+1 ). It should be understood that the dimension of the rotated vector is the same as the original vector.

[0090] Step S440: Determine the temporal fusion features of the video frames according to the time fusion component groups corresponding to the component groups.

[0091] Since the standard occurrence time of feature vectors of different reference frames is different, the corresponding feature rotation angles are different, and the video frame temporal fusion features determined by the time fusion component groups corresponding to each component group contain temporal information. At the same time, the relative position relationship between vectors can be changed through rotation operations, which is conducive to the subsequent model to understand the differences between different time points during video generation, capture the time sequence of events, and thus better handle dynamic image generation tasks with time dependence.

[0092] In an exemplary embodiment, if Figure 5 As shown, the text description information is subjected to temporal position encoding processing, and the text temporal fusion features obtained may include:

[0093] Step S510: Encode the text description information to obtain text embedding information.

[0094] For the text description information of each event, a pre-trained text encoder can be used for encoding, and the text description information can be converted into a high-dimensional embedding vector (i.e., text embedding information) to capture the semantic information of the event. For example, T5 (Text-to-Text Transfer Transformer, a natural language processing model based on the Transformer architecture) can be used for encoding. One or more text embedding information can be obtained for each event.

[0095] Step S520: Determine the text time information of the text embedded information in the video to be generated according to the start and end time information of the event corresponding to the text description information.

[0096] According to the start and end time information of the event in the video to be generated, the corresponding time information can be embedded into the text feature. Specifically, the text time information refers to the start and end time of the text embedded information of the event in the video to be generated.

[0097] Step S530: Encode the text time information into the text embedding information to obtain the text time sequence fusion feature.

[0098] Similarly, in accordance with the method of standardizing the occurrence time of the reference image frame corresponding to the event, the text time information is also standardized, and then the feature rotation angle is calculated based on the standardization result and the reference angular velocity, so that the text is embedded in the information according to the feature rotation angle to obtain the text time fusion. It should be understood that the above-mentioned method of rotating the reference frame feature vector corresponding to the standard occurrence time according to the feature rotation angle to obtain the video frame time fusion feature is also applicable to the time position encoding of the text description information, that is, according to the obtained feature rotation angle θ t (t'*ω) performs a rotation operation on each component of the text embedding information, which will not be described again.

[0099] The text temporal fusion features obtained through temporal position encoding can include "semantics of events" + "temporal position information". For example, the text temporal fusion features can reflect the event or time point t' to which the text embedding information belongs. This makes it easier for the model to know the position of the text embedding information on the entire video timeline during image generation, and accurately match it to the corresponding video frame during cross-attention calculation.

[0100] It is worth noting that if the temporal position encoding is not performed, the text embedding information only represents the semantics of the event, such as "open the window". At the 3rd or 5th second, it is difficult to determine whether the text embedding information is related to the picture of this frame (at the 3rd or 5th second). On the contrary, by encoding the temporal information into the text embedding vector, the text embedding vector will carry time-related information such as "when the event occurred" and "the position in the standard event time length", which will help the model automatically match the current frame picture with the text embedding information of the event when the picture is generated, and realize accurate timing control of the generated video.

[0101] In an exemplary embodiment, if Figure 6 , based on the temporal cross attention mechanism, the video frame temporal fusion features and the text temporal fusion features are fused, and the target fusion features may include:

[0102] Step S610: Map the temporal fusion features of the video frames to the query space to obtain a query vector.

[0103] In the temporal cross-attention mechanism based on the Transformer architecture, three sets of vectors need to be paid special attention to: the query vector Query (Q), the key vector Key (K) and the value vector Value (V). In the exemplary embodiment of the present disclosure, the features on the video side are used as the source of the query vector, that is, the temporal fusion features of the video frame are mapped to the query space through features to obtain the query vector. That is to say, when the subsequent video is generated, when processing the tth frame, its Query (Q) carries the time label of the t frame.

[0104] Among them, using the defined query matrix (WQW ∧ Q) is mapped, and the query matrix is ​​obtained in the same way as the conventional method. Similarly, the subsequent key vector (WKW ∧ K) and the value vector (WVW ∧ V) The mapping matrix (key matrix, value matrix) used is also the same as the conventional method and will not be described again.

[0105] Step S620: Map the text temporal fusion features to the key-value space to obtain a key vector, and map the text temporal fusion features to the key-value space to obtain a value vector, wherein the mapping matrices used to obtain the key vector and the value vector are different.

[0106] The dimensions of the key vector and the value vector are the same as those of the query vector, so that attention calculation and weighted fusion can be performed later. That is, in the case of the same input source, the key vector and the value vector will each play different roles. The key vector is used to calculate the attention weight, and the value vector is used for weighted summation.

[0107] Step S630: Perform attention calculation based on the query vector, key vector and value vector to obtain the target fusion feature.

[0108] Since the sources of the query vector, key vector and value vector are all features processed by temporal position encoding, when processing the t-th frame, its Query (Q) carries the time label of the t-th frame and can be matched with the "text description information of events near the t-th frame" in the Key / Value.

[0109] Specifically, Formula 5 can be used to calculate the relevance score (attention weight) of Query (Q) relative to each Key vector:

[0110] A=softmax((Q*K^T) / sqrt(d k )) Formula 5

[0111] Among them, A is the attention weight, d kis the dimension of the Key, sqrt() is the square root operation, sqrt(d k ) is used to prevent the calculated value from being too large.

[0112] Then, the value vector is weighted and summed using the attention weight to obtain the target fusion feature. After the temporal attention calculation, the text description information of the corresponding time or event is embedded in the "frame feature", and the target fusion feature that integrates time and text is obtained, so that the subsequent model can learn to "match the t-th frame with the text period with similar description" to form the correct temporal connection and content association.

[0113] For example, when the model processes the third frame (Query(Q) carries "t=3"), it will form a high similarity with the feature of the middle position of the Key="open window" event between the 3rd and 5th seconds, and thus pay attention to and use the text description information of this event. On the contrary, if the description information of an event is only valid between the 7th and 8th seconds, and the similarity with "Query(Q) carries "t=3" is low, it will not be overly concerned, thus avoiding problems such as out-of-order events or premature appearance of content in the generated video.

[0114] The exemplary embodiment of the present disclosure fuses the video frame timing fusion features and the text timing fusion features through the temporal cross-attention mechanism, realizes precise time control of the text description information of multiple events and multiple scenes, and focuses on the matching text vector when generating the video frame of the corresponding moment. At the same time, the events (scenes) are smoothly arranged in the time dimension, so that the output dynamic image matches the time both logically and visually. At the same time, it should be understood that a kind of "soft" time attention is achieved by controlling the rotation angle. Since the time information has been marked for each event and video frame through the temporal position coding, that is, the model can understand the sequence and relative relationship between different time points through coding, and then at a certain time point, the model will focus more on the event description information related to the time point, and will pay appropriate attention to the information of the adjacent time to ensure the natural and smooth transition of the picture.

[0115] In an exemplary embodiment, it is also possible to generate a natural and controllable scene transition effect in response to a scene switching point specified by a user. Figure 7 As shown, based on the temporal cross attention mechanism, the video frame temporal fusion features and the text temporal fusion features are fused to obtain the target fusion features, which may include:

[0116] Step S710: receiving scene switching information, and acquiring a scene embedding vector corresponding to the scene switching information according to the scene switching type.

[0117] The scene switching information refers to the description information of the scene switching point, such as "at the 5th second, the scene switches from the indoor scene to the outdoor scene". It should be understood that the scene switching information provided by the user may specify the scene switching time point or the time interval, and there is no limitation on this.

[0118] The scene embedding vectors corresponding to different scene switching types are learned in advance, and then when the scene switching information is obtained, the scene embedding vector can be obtained according to the scene switching type of the scene switching information. The attribute information such as the shape and color of the scene embedding vector can be changed to represent different screen connection methods (such as smooth connection, fast switching, etc.).

[0119] In an exemplary embodiment, before performing temporal position encoding processing on the scene embedding vector according to the start and end time information to obtain the scene temporal fusion feature, the method further includes:

[0120] By performing learning and training on video frame samples annotated with different types of scene switching information, scene embedding vectors corresponding to different scene switching types are obtained, and the scene embedding vectors of different scene switching types are learnable.

[0121] Specifically, video frame samples are first obtained, that is, video clips containing various scene switches such as "indoor scene→outdoor scene" and "outdoor scene→indoor scene", and split into frame sequences. Each frame carries a corresponding global timestamp (such as the 1st second, the 2nd second,..., the 10th second), and the scene switching time is marked (such as 5 seconds).

[0122] Secondly, the scene embedding vector is input into the model to be trained. The scene embedding vector is integrated at the scene switching time point or several frames before and after it. The specific methods include: one method is to insert a "switch token" (the vector representation is the scene embedding vector) in the sequence corresponding to the text description information. Another method is to add the scene embedding vector to the cross-attention input at the scene switching moment (such as 5 seconds) to let the model know that "the scene needs to be switched here". It should be understood that the scene embedding vector is also obtained by performing temporal position encoding processing on the scene switching information input by the user, and also has temporal position information.

[0123] Therefore, the model generates images by taking the video frames, timestamp information, text description information (if any) and scene embedding vectors of multiple events as input. Among them, reconstruction loss or denoising loss can be constructed to measure the difference in vision or features between the "generated frame" and the "real video frame", such as the Euclidean norm, the KL (Kullback-Leibler) loss of the diffusion process or other perceptual losses, without limitation. When the transition frame near a certain frame is not ideal (such as abrupt switching), the loss will increase, and parameters can be adjusted based on this. The gradient can be back-propagated to the various learning parameters of the model, including but not limited to the weight parameters inside the diffusion model, the parameters involved in the temporal position encoding process, and the parameters of the scene embedding vector. With the continuous input of training samples, the model learns how to adjust the scene embedding vector under different switching models, and the value of the scene embedding vector will converge to a vector representation suitable for generating suitable transition frames under different switching models, that is, the scene embedding vector corresponding to different scene switching types is obtained.

[0124] Based on this, in actual reasoning, when the user specifies "scene switching occurs at the 5th second", the corresponding scene embedding vector is obtained to achieve screen transition. Since the scene embedding vector is not written statically, compared with "hard splicing" or "manually specified switching method", it is learned through video frame samples containing a large amount of scene switching information, which can flexibly adapt to the scene switching requirements proposed by the user, and is conducive to generating screen switching that meets user needs and has a more natural transition.

[0125] Step S720: Perform temporal position encoding processing on the scene embedding vector according to the start and end time information to obtain the scene temporal fusion feature.

[0126] The method of temporal position encoding of the scene embedding vector is similar to the process of temporal position encoding of the reference frame feature vector according to the start and end time information. However, it should be noted that the scene switching information provided by the user may only contain one time point, such as "at the 5th second, transition from the indoor scene to the outdoor scene". In the process of temporal position encoding, since the start time and the end time are involved, the degenerate interval method can be used, that is, the start and end values ​​are set to be the same, such as T start =T end=5 seconds, but there will be a division by zero problem. In actual implementation, a minimum value can be set for processing. Another processing method is to choose the cell expansion method to diffuse the single point "5 seconds" to [5-δ,5+δ] (for example, [4.5,5.5]). Based on this, there will be no division by zero problem, and the model can perform progressive scene switching within the interval. For example, when the diffusion model is processed, the switching can be completed from the 4.5th second to the 5.5th second. In another processing method, if the user only provides "switch to the outdoor scene after opening the window" and does not include the specific scene switching time, the scene switching time can be automatically inferred, for example, the time to switch the scene is determined based on context analysis.

[0127] Step S730: Fusing the scene temporal fusion feature and the text temporal fusion feature to obtain an updated text temporal fusion feature.

[0128] The scene temporal fusion features and the text temporal fusion features can be connected to form updated text temporal fusion features, so that the model knows at which moment / moments it needs to switch to the next scene when the image is generated.

[0129] Step S740: Based on the temporal cross-attention mechanism, the video frame temporal fusion feature is fused with the updated text temporal fusion feature to obtain the target fusion feature.

[0130] The fusion method of this step is similar to the above-mentioned temporal cross-attention mechanism based on which the video frame temporal fusion features and the text temporal fusion features are fused to obtain the target fusion features, and will not be repeated here.

[0131] By integrating the scene switching information required by the user into the target fusion feature, when generating dynamic images, it will first determine "whether the current moment corresponds to event A or event B" based on the current moment, then determine "whether there is a scene embedding vector that needs to be inserted at the current moment", and finally decide which feature vectors to fuse (such as video frame timing fusion features and text timing fusion features, or video timing fusion features and updated text timing fusion features). Since all features are the same and have been processed by temporal position encoding, they can match the moment of the reference image frame (Query) at the current moment during the actual cross-attention calculation.

[0132] Among them, for scene switching, due to the temporal position encoding, when generating dynamic images, the model can know at which second the specified scene switching point is reached, and the Key / Value on the text side also contains the time tag (temporal fusion), so that the attention can be partially directed to the semantics of the next scene. On the contrary, if temporal position encoding is not performed, the model can only obtain the text description information of the "indoor scene" and the text description information of the "outdoor scene", but cannot know at which second the scene switching occurs, which may cause switching confusion or the picture will always stay indoors.

[0133] Based on the aforementioned exemplary embodiment, dynamic image generation is performed according to the target fusion features corresponding to the feature vectors of each reference frame to obtain a target video, including:

[0134] According to the target fusion features corresponding to the feature vectors of each reference frame, content generation is performed through a diffusion model to obtain a target video; wherein, based on the target fusion features fused with updated text temporal fusion features, scene transition frames are generated in the target video.

[0135] Among them, the scene transition frame is generated by the model during the iterative processing. Taking the diffusion model DiT as an example, the model will perform multiple denoising (generation) steps on the frames of the entire video. When processing to 4.5 to 5.5 seconds, it detects that there is a scene switching point at the 5th second (scene A switches to scene B). Then, through attention calculation, at the 5th second, it will partially focus on scene A and partially focus on scene B, and generate intermediate frames (scene A→scene B). In other words, the transition frame is not fixedly written through the interpolation script, but the model learns during the training process that when the "temporal position is close to the scene switching point" + "text temporal fusion features are embedded with switching instructions", a progressive scene connection picture is generated. This is because the "video frame temporal fusion features" and "text temporal fusion features" play a guiding role in the cross-attention mechanism, indicating the model near which frame it needs to mix the pictures of the previous and next scenes to generate a transition picture.

[0136] The exemplary embodiments of the present disclosure can generate dynamic images with scene transition effects by allowing users to customize scene switching time points, making the generated videos richer and more diverse.

[0137] The video generation method in the exemplary embodiment of the present disclosure receives an event description text sequence, and determines the reference frame feature vector of the video to be generated according to the event description text sequence, wherein the event description text sequence includes text description information of multiple events and corresponding start and end time information; performs temporal position encoding processing on the reference frame feature vector according to the start and end time information to obtain the video frame temporal fusion feature, and performs temporal position encoding processing on the text description information to obtain the text temporal fusion feature; based on the temporal cross attention mechanism, the video frame temporal fusion feature and the text temporal fusion feature are fused to obtain the target fusion feature; and dynamic image generation is performed according to the target fusion feature corresponding to each reference frame feature vector to obtain the target video. On the one hand, the process can make the content and timing of the video frame consistent with the text description based on the received event description text sequence, by performing temporal position encoding on the video frame feature vector and the text description information, so as to realize the time control at the frame level, thereby generating a multi-event video with precise time control. On the other hand, dynamic image generation is performed according to the target fusion feature corresponding to each reference frame feature vector, which fully integrates the text description and video frame feature vector of multiple events, so that the target video is coherent and smooth, and the transition of each event is natural. In the exemplary embodiment of the present disclosure, users only need to input a description of multiple events and specify the start and end time of each event to generate a multi-event video with precise time control. The video can be applied to advertising creation, image production, virtual reality, animation production and other fields to improve the efficiency of video generation.

[0138] In an exemplary embodiment of the present disclosure, a video generating device is also provided. Figure 8 As shown, the video generation device 800 may include an information acquisition module 810, an encoding processing module 820, an attention module 830 and a content generation module 840. Specifically:

[0139] The information acquisition module 810 is used to receive an event description text sequence, and determine the reference frame feature vector of the video to be generated based on the event description text sequence, wherein the event description text sequence includes text description information of multiple events and corresponding start and end time information; the encoding processing module 820 is used to perform temporal position encoding processing on the reference frame feature vector according to the start and end time information to obtain video frame temporal fusion features, and perform temporal position encoding processing on the text description information to obtain text temporal fusion features; the attention module 830 is used to fuse the video frame temporal fusion features and the text temporal fusion features based on a temporal cross-attention mechanism to obtain a target fusion feature; the content generation module 840 is used to generate dynamic images based on the target fusion features corresponding to each reference frame feature vector to obtain a target video.

[0140] In an exemplary embodiment of the present disclosure, the information acquisition module 810 is configured to perform: determining the reference image frame of the video to be generated based on the text description information and corresponding start time information of multiple events, each event corresponding to at least one reference image frame; and numerically encoding the reference image frame to obtain a reference frame feature vector.

[0141] In an exemplary embodiment of the present disclosure, the encoding processing module 820 is configured to perform: for each event, based on the start and end time information corresponding to the event, standardizing the occurrence time of the reference image frame corresponding to the event to obtain the standard occurrence time; calculating the characteristic rotation angle based on the standard occurrence time and the reference angular velocity; rotating the reference frame feature vector corresponding to the standard occurrence time according to the characteristic rotation angle to obtain the video frame timing fusion feature.

[0142] In an exemplary embodiment of the present disclosure, the encoding processing module 820 is configured to perform: determining the time period position corresponding to the event based on the start and end time information corresponding to the event; and standardizing the occurrence time of the reference image frame corresponding to the event based on the start and end time information corresponding to the event, the time period position and the preset standard event time length to obtain the standard occurrence time.

[0143] In an exemplary embodiment of the present disclosure, the encoding processing module 820 is configured to perform: splitting the reference frame feature vector corresponding to the standard occurrence moment into multiple dimensional components; dividing the multiple dimensional components into multiple component groups in the manner of each adjacent dimension as a group; for each component group, rotating the component group in a two-dimensional plane based on the feature rotation angle to obtain a time fusion component group; determining the video frame timing fusion feature according to the time fusion component groups corresponding to each component group.

[0144] In an exemplary embodiment of the present disclosure, the encoding processing module 820 is configured to perform: encoding processing on the text description information to obtain text embedding information; determining the text time information of the text embedding information in the video to be generated based on the start and end time information of the event corresponding to the text description information; encoding the text time information into the text embedding information to obtain text timing fusion features.

[0145] In an exemplary embodiment of the present disclosure, the attention module 830 is configured to perform: mapping the temporal fusion features of the video frames to the query space to obtain a query vector; mapping the temporal fusion features of the text to the key-value space to obtain a key vector, and mapping the text temporal fusion features to the key-value space to obtain a value vector, wherein the mapping matrices used to obtain the key vector and the value vector are different; performing attention calculation based on the query vector, the key vector and the value vector to obtain the target fusion feature.

[0146] In an exemplary embodiment of the present disclosure, the attention module 830 is further configured to execute: receiving scene switching information, and obtaining a scene embedding vector corresponding to the scene switching information according to the scene switching type; performing temporal position encoding processing on the scene embedding vector according to the start and end time information to obtain a scene temporal fusion feature; fusing the scene temporal fusion feature with the text temporal fusion feature to obtain an updated text temporal fusion feature; based on the temporal cross-attention mechanism, fusing the video frame temporal fusion feature with the updated text temporal fusion feature to obtain a target fusion feature.

[0147] In an exemplary embodiment of the present disclosure, the information acquisition module 810 is further configured to execute: before performing temporal position encoding processing on the scene embedding vector according to the start and end time information to obtain the scene temporal fusion feature, learning and training are performed on video frame samples annotated with different types of scene switching information to obtain scene embedding vectors corresponding to different scene switching types, and the scene embedding vectors of different scene switching types are learnable.

[0148] In an exemplary embodiment of the present disclosure, the content generation module 840 is configured to execute: based on the target fusion features corresponding to the feature vectors of each reference frame, content generation is performed through a diffusion model to obtain a target video; wherein, based on the target fusion features fused with updated text temporal fusion features, scene transition frames are generated in the target video.

[0149] Since the details of the functional modules of the video generating device of the exemplary embodiment of the present disclosure have been described in the exemplary embodiment of the video generating method described above, they will not be repeated here.

[0150] It should be noted that although several modules or units of the video generation device are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided into multiple modules or units to be embodied.

[0151] The exemplary embodiments of the present disclosure further provide a computer program product, which includes a computer program, and when the computer program is executed by a processor, the above-mentioned video generation method is implemented.

[0152] In one embodiment, the computer program product may be a tangible product containing a computer program, such as a computer-readable storage medium storing a computer program. The readable storage medium may be a storage medium based on electrical, magnetic, optical, electromagnetic, infrared, or other signals, including but not limited to: random access memory (RAM), read-only memory (ROM), magnetic tape, floppy disk, flash memory (Flash), mechanical hard disk (HDD), solid-state drive (SSD), and the like. Exemplarily, the computer program product may be implemented as a non-volatile storage medium storing a computer program, such as a read-only memory, a NAND flash memory, and the like.

[0153] In one embodiment, the computer program product may be an intangible product including a computer program. Exemplarily, the computer program product may be implemented as a virtual digital product, such as a digital file storing an executable file, an installation package, etc. of the computer program.

[0154] The code of the computer program can be written in one or more programming languages. Programming languages ​​such as C language, Java, C++, etc. The program code can be executed entirely on the user computing device, or partially on the user computing device, or as a separate software package, or partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user computing device through any type of network, such as a local area network (LAN), a wide area network (WAN), etc., or can be connected to an external computing device (e.g., an Internet connection provided by an operator).

[0155] The computer program may be carried or transmitted via electrical, magnetic, optical, electromagnetic, infrared, or other signals. The electronic device may convert the signal carrying the computer program into a digital signal, thereby running the computer program. When the computer program is run on an electronic device, its code is used to enable the electronic device to execute (more specifically, the processor of the electronic device may execute) the method steps of various exemplary embodiments of the present disclosure, such as the above-mentioned video generation method, which includes the following steps:

[0156] An event description text sequence is received, and a reference frame feature vector of a to-be-generated video is determined based on the event description text sequence, wherein the event description text sequence includes text description information of multiple events and corresponding start and end time information; the reference frame feature vector is temporally encoded based on the start and end time information to obtain a video frame temporal fusion feature, and the text description information is temporally encoded to obtain a text temporal fusion feature; based on a temporal cross-attention mechanism, the video frame temporal fusion feature and the text temporal fusion feature are fused to obtain a target fusion feature; dynamic images are generated based on the target fusion features corresponding to each reference frame feature vector to obtain a target video.

[0157] In addition, in an exemplary embodiment of the present disclosure, an electronic device capable of implementing the above method is also provided. It will be appreciated by those skilled in the art that various aspects of the present disclosure may be implemented as a system, method, or program product. Therefore, various aspects of the present disclosure may be specifically implemented in the following forms, namely: a complete hardware embodiment, a complete software embodiment (including firmware, microcode, etc.), or an embodiment combining hardware and software aspects, which may be collectively referred to herein as a "circuit", "module", or "system".

[0158] Refer to the following Fig. 9 hereinafter describes an electronic device 900 according to such an embodiment of the present disclosure. Fig. 9 The electronic device 900 shown is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.

[0159] like Fig. 9 As shown, the electronic device 900 is in the form of a general computing device. The components of the electronic device 900 may include, but are not limited to: the at least one processing unit 910, the at least one storage unit 920, a bus 930 connecting different system components (including the storage unit 920 and the processing unit 910), and a display unit 940.

[0160] The storage unit stores a program code, which can be executed by the processing unit 910, so that the processing unit 910 performs the steps according to various exemplary embodiments of the present disclosure described in the above “Exemplary Method” section of this specification. For example, the processing unit 910 can perform the following steps:

[0161] An event description text sequence is received, and a reference frame feature vector of a to-be-generated video is determined based on the event description text sequence, wherein the event description text sequence includes text description information of multiple events and corresponding start and end time information; the reference frame feature vector is temporally encoded based on the start and end time information to obtain a video frame temporal fusion feature, and the text description information is temporally encoded to obtain a text temporal fusion feature; based on a temporal cross-attention mechanism, the video frame temporal fusion feature and the text temporal fusion feature are fused to obtain a target fusion feature; dynamic images are generated based on the target fusion features corresponding to each reference frame feature vector to obtain a target video.

[0162] The storage unit 920 may include a readable medium in the form of a volatile storage unit, such as a random access memory unit (RAM) 921 and / or a cache memory unit 922 , and may further include a read-only memory unit (ROM) 923 .

[0163] The storage unit 920 may also include a program / utility 924 having a set (at least one) of program modules 925, such program modules 925 including but not limited to: an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment.

[0164] Bus 930 may represent one or more of several types of bus structures, including a memory unit bus or memory unit controller, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of a variety of bus architectures.

[0165] The electronic device 900 may also communicate with one or more external devices 1000 (e.g., keyboards, pointing devices, Bluetooth devices, etc.), may also communicate with one or more devices that enable a user to interact with the electronic device 900, and / or communicate with any device that enables the electronic device 900 to communicate with one or more other computing devices (e.g., routers, modems, etc.). Such communication may be performed via an input / output (I / O) interface 950. Furthermore, the electronic device 900 may also communicate with one or more networks (e.g., local area networks (LANs), wide area networks (WANs), and / or public networks, such as the Internet) via a network adapter 960. As shown, the network adapter 960 communicates with other modules of the electronic device 900 via a bus 930. It should be understood that, although not shown in the figure, other hardware and / or software modules may be used in conjunction with the electronic device 900, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.

[0166] Through the description of the above embodiments, it is easy for those skilled in the art to understand that the example embodiments described here can be implemented by software, or by software combined with necessary hardware. Therefore, the technical solution according to the embodiment of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, and includes several instructions to enable a computing device (which can be a personal computer, a server, a terminal device, or a network device, etc.) to execute the method according to the embodiment of the present disclosure.

[0167] In addition, the above-mentioned figures are only schematic illustrations of the processes included in the method according to the exemplary embodiments of the present disclosure, and are not intended to be limiting. It is easy to understand that the processes shown in the above-mentioned figures do not indicate or limit the time sequence of these processes. In addition, it is also easy to understand that these processes can be performed synchronously or asynchronously, for example, in multiple modules.

[0168] Those skilled in the art will readily appreciate other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary technical means in the art that are not disclosed in the present disclosure. The description and examples are to be considered exemplary only, and the true scope and spirit of the present disclosure are indicated by the claims.

Claims

1. A video generation method, characterized in that: include: Receive an event description text sequence, and determine a reference frame feature vector of a video to be generated according to the event description text sequence, wherein the event description text sequence includes text description information of multiple events and corresponding start and end time information; Performing temporal position coding processing on the reference frame feature vector according to the start and end time information to obtain a video frame temporal fusion feature, and performing temporal position coding processing on the text description information to obtain a text temporal fusion feature; Based on the temporal cross attention mechanism, the video frame temporal fusion feature and the text temporal fusion feature are fused to obtain the target fusion feature; Dynamic images are generated according to the target fusion features corresponding to the feature vectors of each reference frame to obtain a target video.

2. The method according to claim 1, characterized in that The step of determining a reference frame feature vector of a video to be generated according to the event description text sequence comprises: Determine, according to the text description information of the plurality of events and the corresponding start time information, a reference image frame of the video to be generated, wherein each of the events corresponds to at least one reference image frame; The reference image frame is numerically encoded to obtain the reference frame feature vector.

3. The method according to claim 1, characterized in that: The step of performing temporal position encoding processing on the reference frame feature vector according to the start and end time information to obtain a video frame temporal fusion feature includes: For each of the events, the occurrence time of the reference image frame corresponding to the event is standardized according to the start and end time information corresponding to the event to obtain a standard occurrence time; Calculate a characteristic rotation angle based on the standard occurrence time and a reference angular velocity; The reference frame feature vector corresponding to the standard occurrence moment is rotated according to the feature rotation angle to obtain the video frame temporal fusion feature.

4. The method according to claim 3, characterized in that For each of the events, the occurrence time of the reference image frame corresponding to the event is standardized according to the start and end time information corresponding to the event to obtain the standard occurrence time, including: Determine the time period corresponding to the event according to the start and end time information corresponding to the event; Based on the start and end time information corresponding to the event, the time period position and the preset standard event time length, the occurrence time of the reference image frame corresponding to the event is standardized to obtain the standard occurrence time.

5. The method according to claim 3, characterized in that: The step of rotating the reference frame feature vector corresponding to the standard occurrence time according to the feature rotation angle to obtain the video frame temporal fusion feature includes: Splitting the reference frame feature vector corresponding to the standard occurrence moment into multiple dimensional components; Dividing the multiple dimensional components into multiple component groups in a manner where each adjacent dimension is a group; For each of the component groups, rotating the component group in a two-dimensional plane based on the characteristic rotation angle to obtain a time-fused component group; The video frame temporal fusion feature is determined according to the time fusion component group corresponding to each of the component groups.

6. The method according to claim 1, characterized in that The step of performing temporal position encoding processing on the text description information to obtain a text temporal fusion feature includes: Encoding the text description information to obtain text embedding information; Determining text time information of the text embedded information in the video to be generated according to the start and end time information of the event corresponding to the text description information; The text time information is encoded into the text embedding information to obtain the text temporal fusion feature.

7. The method according to claim 1, characterized in that The method of fusing the video frame temporal fusion feature and the text temporal fusion feature based on the temporal cross attention mechanism to obtain the target fusion feature includes: Mapping the video frame temporal fusion features to a query space to obtain a query vector; Mapping the text temporal fusion feature to a key value space to obtain a key vector, and mapping the text temporal fusion feature to a key value space to obtain a value vector, wherein the mapping matrices used to obtain the key vector and the value vector are different; Attention calculation is performed according to the query vector, the key vector and the value vector to obtain the target fusion feature.

8. The method according to claim 1, characterized in that The method further comprises: fusing the video frame temporal fusion feature and the text temporal fusion feature based on the temporal cross attention mechanism to obtain a target fusion feature; Receive scene switching information, and obtain a scene embedding vector corresponding to the scene switching information according to the scene switching type; Performing temporal position encoding processing on the scene embedding vector according to the start and end time information to obtain a scene temporal fusion feature; The scene temporal fusion feature and the text temporal fusion feature are fused to obtain an updated text temporal fusion feature; Based on the temporal cross-attention mechanism, the video frame temporal fusion feature is fused with the updated text temporal fusion feature to obtain the target fusion feature.

9. The method according to claim 8, characterized in that Before performing temporal position encoding processing on the scene embedding vector according to the start and end time information to obtain the scene temporal fusion feature, the method further includes: By performing learning and training on video frame samples annotated with different types of scene switching information, scene embedding vectors corresponding to different scene switching types are obtained, and the scene embedding vectors of different scene switching types are learnable.

10. The method according to claim 8, characterized in that The step of generating a dynamic image according to the target fusion features corresponding to the feature vectors of each reference frame to obtain a target video includes: According to the target fusion features corresponding to the feature vectors of each reference frame, content generation is performed through a diffusion model to obtain a target video; Wherein, based on the target fusion feature fused with the updated text temporal fusion feature, a scene transition frame is generated in the target video.

11. A video generating device, characterized in that: The device comprises: An information acquisition module, used to receive an event description text sequence, and determine a reference frame feature vector of a video to be generated according to the event description text sequence, wherein the event description text sequence includes text description information of multiple events and corresponding start and end time information; A coding processing module, used for performing temporal position coding processing on the reference frame feature vector according to the start and end time information to obtain a video frame temporal fusion feature, and performing temporal position coding processing on the text description information to obtain a text temporal fusion feature; An attention module, used for fusing the video frame temporal fusion feature and the text temporal fusion feature based on a temporal cross attention mechanism to obtain a target fusion feature; The content generation module is used to generate dynamic images according to the target fusion features corresponding to the feature vectors of each reference frame to obtain a target video.

12. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method according to any one of claims 1 to 10 is implemented.

13. An electronic device, characterized in that: include: processor; as well as A memory, configured to store executable instructions of the processor; The processor is configured to perform the method of any one of claims 1 to 10 by executing the executable instructions.

Citation Information

Cited By

  • Method for obtaining frame sequence from text and related device

    CN120279470A

  • A method and related device for obtaining frame sequence from text

    CN120279470B

  • Video generation method and device, storage medium and program product

    CN120416623A

  • Video generation method, device, storage medium and program product

    CN120416623B

  • Video text method and device based on time sequence motion perception and electronic equipment

    CN120853083A