Video dataset construction, model training, video generation method and device
By constructing lighting trajectories and training datasets in 3D lighting grids, generating and annotating video datasets, and combining frozen denoising and external transformer network training models, the problem of poor lighting effects is solved, achieving better lighting control and video generation effects.
Patent Information
- Application Number
- CN202411535402.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-30
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2044-10-30
AI Technical Summary
The lack of clearly labeled training datasets that show lighting changes makes it impossible to generate videos that express lighting effects well, affecting the aesthetic quality of video generation models.
Multiple lighting trajectories are constructed in the 3D lighting grid, image sets of whiteboards and 3D models are generated, text information is annotated, a video dataset is constructed, and a text-to-video generation model is trained through a frozen denoising network and an external transformer network, fusing lighting features to control lighting effects.
The generated video can better express the lighting effect, improve the user experience, and meet the user's needs for lighting control.
Smart Images

Figure CN119450026B_ABST
Abstract
Description
Technical Field
[0001] One or more embodiments of this specification relate to the field of machine vision technology, and in particular, to a video dataset construction method, a model training method, and a video generation method and device. Background Art
[0002] Lighting plays a crucial role in ensuring the naturalness of generated videos and has a significant impact on the aesthetic quality of the generated videos. However, the lack of clearly labeled training datasets that demonstrate lighting variations and well-defined lighting information has hindered the development of video generation models that can generate videos with well-expressed lighting effects.
[0003] Therefore, it is necessary to provide a video dataset construction method to obtain a video dataset that can better express lighting information, and then train a video generation model that can generate videos with good lighting effects. Summary of the Invention
[0004] In order to obtain a video dataset that can better express lighting information, and then train a video generation model that can generate videos that express lighting effects well, one or more embodiments of this specification provide a video dataset construction method, a model training method, a video generation method and an apparatus.
[0005] In a first aspect, one or more embodiments of the present specification provide a method for constructing a video dataset, the method comprising: constructing multiple lighting trajectories in a 3D lighting grid; the 3D lighting grid is a three-dimensional grid in which grid points are used as lighting positions; the lighting trajectory is a trajectory connected by multiple grid points; obtaining a first image set and at least one second image set corresponding to each lighting trajectory; the first image set includes whiteboard images generated by point light sources being located at each lighting position of the corresponding lighting trajectory to illuminate a whiteboard in a blank space; the second image set includes model images generated by the point light sources being located at each lighting position of the corresponding lighting trajectory to illuminate a 3D model; the second image set corresponds one-to-one to the 3D model; generating a first video corresponding to each first image set and a second video corresponding to each second image set; annotating text information for each second video; constructing a video dataset for training a text video generation model based on all first videos and all second videos annotated with text information; the video dataset includes multiple training samples, each training sample including a first video corresponding to a lighting trajectory and a second video annotated with text information of a 3D model corresponding to the lighting trajectory, and both the lighting trajectory and the 3D model correspond one-to-one to the training samples.
[0006] In one possible implementation, the first lighting trajectory is any one of the multiple lighting trajectories; obtaining the first image set corresponding to the first lighting trajectory includes: controlling the point light source to move in sequence at each lighting position of the first lighting trajectory, and after the point light source moves to any lighting position to illuminate the whiteboard in the blank space, generating a frame of whiteboard image through virtual engine rendering; all whiteboard images corresponding to the first lighting trajectory are constructed into a first image set corresponding to the first lighting trajectory in the order of generation time.
[0007] In one possible implementation, the first lighting trajectory is any one of the multiple lighting trajectories; the third image set is any one of the at least one second image set corresponding to the first lighting trajectory; the first 3D model is the 3D model corresponding to the third image set; obtaining the third image set corresponding to the first lighting trajectory includes: controlling the point light source to move sequentially on each lighting position of the first lighting trajectory, and after the point light source moves to any lighting position to illuminate the first 3D model, generating a frame of model image through virtual engine rendering; constructing all model images corresponding to the first lighting trajectory and the first 3D model into a third image set corresponding to the first lighting trajectory and the first 3D model in chronological order of generation time.
[0008] In one possible implementation, text information is annotated for each second video, including: generating initial text information for each second video through an image language pre-training model; performing title enhancement processing on the initial text information of each second video to generate enhanced text information for each second video; and annotating corresponding enhanced text information for each second video.
[0009] In a possible implementation, the 3D model is a 3D character model.
[0010] In a second aspect, one or more embodiments of the present specification further provide a model training method based on a video dataset, the model training method comprising: constructing a frozen denoising network and an external transformer network; the frozen denoising network has the same architecture as a pre-trained text video generation model, and is used to freeze the network weight parameters of the video denoising network in the pre-trained text video generation model; the video denoising network comprises a variational self-attention encoder for feature extraction and a preset number of diffusion transformer blocks; the external transformer network comprises the same variational self-attention encoder as the video denoising network, and the preset number of illumination encoder blocks, wherein the illumination encoder block is a self-attention layer. Transformer architecture; based on the training samples and objective function in the video dataset, the pre-trained text video generation model, the external transformer network and the frozen denoising network are trained to generate a trained text video generation model and an external transformer network; the objective function is a decoupling loss function of the predicted noise of the video denoising network and the predicted noise of the frozen denoising network; wherein, in any training iteration process, the illumination features extracted from each illumination encoder block are input into the corresponding diffusion transformer block, and after feature fusion with the video features extracted from the diffusion transformer block, are input into the multi-head attention layer of the diffusion transformer block.
[0011] In one possible implementation, the pre-trained text video generation model, the external transformer network, and the frozen denoising network are trained based on the training samples and objective function in the video dataset, including: inputting the second video with text information labeled in the training samples in the video dataset into the pre-trained text video generation model and the frozen denoising network, respectively, and inputting the first video in the corresponding training sample into the external transformer network, and performing model training based on the objective function.
[0012] In a possible implementation, the lighting feature and the video feature are fused by element-by-element addition.
[0013] In a third aspect, one or more embodiments of the present specification also provide a video generation method based on an external transformer network and a text video generation model trained and generated by the model training method of the second aspect, the video generation method comprising: selecting a target lighting trajectory required to generate a target video from a plurality of lighting trajectories in a 3D lighting grid; inputting a first video corresponding to the target lighting trajectory into the external transformer network, and inputting target text information used to generate the target video into the text video generation model to generate the target video; wherein, in the process of generating the target video, the lighting features extracted from each lighting encoder block are input into the corresponding diffusion transformer block, and after feature fusion with the video features extracted by the diffusion transformer block, are input into the multi-head attention layer of the diffusion transformer block.
[0014] In a fourth aspect, one or more embodiments of the present specification further provide a video data set construction device, the device comprising: a first construction module, configured to construct a plurality of lighting tracks in a 3D lighting grid; the 3D lighting grid being a three-dimensional grid in which grid points are used as lighting positions; the lighting track being a track formed by connecting a plurality of the grid points; an acquisition module, configured to acquire a first image set and at least one second image set corresponding to each lighting track; the first image set comprising whiteboard images corresponding to whiteboards in blank spaces illuminated by point light sources at respective lighting positions of the corresponding lighting tracks; the second image set comprising images corresponding to 3D images illuminated by the point light sources at respective lighting positions of the corresponding lighting tracks; The model image is generated corresponding to the model; the second image set corresponds to the 3D model one-to-one; a first generation module is used to generate a first video corresponding to each first image set and a second video corresponding to each second image set; a labeling module is used to label text information for each second video; a second construction module is used to construct a video data set for training the text video generation model based on all first videos and all second videos labeled with text information; the video data set includes multiple training samples, each training sample includes a first video corresponding to a lighting trajectory, and a second video labeled with text information of a 3D model corresponding to the lighting trajectory, and the lighting trajectory and the 3D model both correspond to the training samples one-to-one.
[0015] In one possible implementation, the first lighting trajectory is any one of the multiple lighting trajectories; the acquisition module is used to obtain a first image set corresponding to the first lighting trajectory, including: the acquisition module is used to: control the point light source to move in sequence on each lighting position of the first lighting trajectory, after the point light source moves to any lighting position to illuminate the whiteboard in the blank space, generate a frame of whiteboard image through virtual engine rendering; all whiteboard images corresponding to the first lighting trajectory are constructed into a first image set corresponding to the first lighting trajectory in the order of generation time.
[0016] In one possible implementation, the first lighting trajectory is any one of the multiple lighting trajectories; the third image set is any one of the at least one second image set corresponding to the first lighting trajectory; the first 3D model is the 3D model corresponding to the third image set; the acquisition module is used to acquire the third image set corresponding to the first lighting trajectory, including: the acquisition module is used to: control the point light source to move in sequence on each lighting position of the first lighting trajectory, after the point light source moves to any lighting position to illuminate the first 3D model, generate a frame of model image through virtual engine rendering; all model images corresponding to the first lighting trajectory and the first 3D model are constructed into a third image set corresponding to the first lighting trajectory and the first 3D model in order of generation time.
[0017] In one possible implementation, the annotation module is used to annotate text information for each second video, including: the annotation module is used to: generate initial text information for each second video through an image language pre-training model; perform title enhancement processing on the initial text information of each second video to generate enhanced text information for each second video; and annotate corresponding enhanced text information for each second video.
[0018] In a possible implementation, the 3D model is a 3D character model.
[0019] In a fifth aspect, one or more embodiments of the present specification further provide a model training device based on a video dataset, wherein the video dataset is constructed by the video dataset construction device of the fourth aspect, and the model training device comprises: a third construction module for constructing a frozen denoising network and an external transformer network; the frozen denoising network has the same architecture as the pre-trained text video generation model, and is used to freeze the network weight parameters of the video denoising network in the pre-trained text video generation model; the video denoising network comprises a variational self-decomposition encoder for feature extraction and a preset number of diffusion transformer blocks; the external transformer network comprises the same variational self-decomposition encoder as the video denoising network, and the preset number of illumination encoder blocks, The illumination encoder block is a transformer architecture including a self-attention layer; a training module is used to train the pre-trained text video generation model, the external transformer network, and the frozen denoising network based on training samples and an objective function in the video dataset to generate a trained text video generation model and an external transformer network; the objective function is a decoupling loss function of the predicted noise of the video denoising network and the predicted noise of the frozen denoising network; wherein, during any training iteration, the illumination features extracted from each illumination encoder block are input into the corresponding diffusion transformer block, and after feature fusion with the video features extracted from the diffusion transformer block, are input into the multi-head attention layer of the diffusion transformer block.
[0020] In one possible implementation, the training module is used to train the pre-trained text video generation model, the external transformer network and the frozen denoising network based on the training samples and objective function in the video dataset, including: the training module is used to input the second video with text information annotated in the training sample in the video dataset into the pre-trained text video generation model and the frozen denoising network respectively, and input the first video in the corresponding training sample into the external transformer network, and perform model training based on the objective function.
[0021] In a possible implementation, the training module is configured to perform feature fusion on the lighting feature and the video feature by element-by-element addition.
[0022] In a sixth aspect, one or more embodiments of the present specification also provide a video generation device based on an external transformer network and a text video generation model trained and generated by the model training device of the fifth aspect, the video generation device comprising: a selection module for selecting a target lighting trajectory required to generate a target video from a plurality of lighting trajectories in a 3D lighting grid; a second generation module for inputting a first video corresponding to the target lighting trajectory into the external transformer network, and inputting the target text information used to generate the target video into the text video generation model to generate the target video; wherein, in the process of generating the target video, the lighting features extracted from each lighting encoder block are input into the corresponding diffusion transformer block, and after feature fusion with the video features extracted by the diffusion transformer block, are input into the multi-head attention layer of the diffusion transformer block.
[0023] In the seventh aspect, one or more embodiments of this specification also provide an electronic device, which includes a memory and a processor; the memory is used to store a computer program product; the processor is used to execute the computer program product stored in the memory, and when the computer program product is executed, the method of the above-mentioned first aspect, second aspect or third aspect is implemented.
[0024] In an eighth aspect, one or more embodiments of this specification further provide a computer-readable storage medium, wherein the computer-readable storage medium stores computer program instructions, and when the computer program instructions are executed, the method of the first aspect, the second aspect or the third aspect described above is implemented.
[0025] In summary, in order to obtain a video dataset that can better express lighting information, and then train a text video generation model to generate a video that expresses lighting effects well, one or more embodiments of this specification provide a video dataset construction method, a model training method, a video generation method, and an apparatus. In this video dataset construction method, a large number of lighting trajectories are constructed in a 3D lighting network, and then a video of a whiteboard image representing the lighting information and a model video representing a 3D model (such as a 3D character model) corresponding to each lighting trajectory are generated. Afterwards, a video dataset for training the text video generation model is generated based on the video representing the lighting information and the video representing the model information.
[0026] When training a text generation model based on the above video dataset, the lighting information can be integrated into the text video generation model. In the subsequent video generation process, the lighting direction and motion trajectory can be controlled so that the generated video can better express the lighting effect, better meet user needs, and provide a better user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] In order to more clearly illustrate the technical solutions of one or more embodiments of the present specification, the drawings needed to be used in the embodiment description will be briefly introduced as follows. Obviously, the drawings in the following description are only some embodiments of one or more embodiments of the present specification, and other drawings can be obtained by those skilled in the art without creative effort.
[0028] Figure 1 A flowchart of a video dataset construction method provided by one or more embodiments of the present specification is shown in the figure.
[0029] Figure 2 An application scenario diagram provided by one or more embodiments of the present specification is shown in the figure.
[0030] Figure 3 A flowchart of a model training method provided by one or more embodiments of the present specification is shown in the figure.
[0031] Figure 4 Another application scenario diagram provided by one or more embodiments of the present specification is shown in the figure.
[0032] Figure 5 A flowchart of a video generation method provided by one or more embodiments of the present specification is shown in the figure.
[0033] Figure 6 A structural block diagram of a video dataset construction device provided by one or more embodiments of the present specification is shown in the figure.
[0034] Figure 7 A structural block diagram of a model training device provided by one or more embodiments of the present specification is shown in the figure.
[0035] Figure 8 A structural block diagram of a video generation device provided by one or more embodiments of the present specification is shown in the figure.
[0036] Figure 9 A structural block diagram of an electronic device provided by one or more embodiments of the present specification is shown in the figure. DETAILED DESCRIPTION
[0037] One or more embodiments of the present specification will be further described in detail below with the drawings and embodiments. Through these descriptions, the features and advantages of one or more embodiments of the present specification will become clearer and more apparent.
[0038] The word "exemplary" is used herein to mean "serving as an example, instance, or illustration." Any implementation described herein as "exemplary" is not necessarily to be construed as preferred or advantageous over other implementations. Unless specifically indicated otherwise, the drawings shown in the figures are not necessarily drawn to scale.
[0039] In addition, the technical features involved in different implementations of one or more embodiments of this specification described below can be combined with each other as long as they do not conflict with each other.
[0040] To facilitate understanding, the application scenarios of the technical solutions provided by one or more embodiments of this specification are first described below.
[0041] Lighting is an essential element in video generation, a factor that determines a video's overall aesthetic quality and a crucial means of conveying emotion, highlighting characters, and directing viewers' attention. However, the scarcity of clearly labeled training datasets that demonstrate lighting variations and well-defined lighting information has hindered the development of video generation models capable of generating videos with well-expressed lighting effects. Alternatively, it has been impossible to effectively control lighting during the video generation process.
[0042] In order to obtain a video with good lighting effects, one or more embodiments of this specification provide a video dataset construction method, a model training method, a video generation method and a device.
[0043] See also Figure 1 , Figure 1 This is a flow chart of a method for constructing a video data set provided in one or more embodiments of this specification. This method can be applied to a server or terminal device, etc. The following takes the terminal device as an example to introduce the content of the embodiment. Figure 1 As shown, the method may include the following steps:
[0044] Step S102: construct multiple lighting tracks in the 3D lighting grid.
[0045] The 3D lighting grid is a three-dimensional grid whose grid points are used as lighting positions. The lighting track is a track formed by connecting multiple grid points in the 3D lighting grid.
[0046] The size of the 3D lighting grid and the spacing between grid points in the 3D lighting grid can be set according to the needs of the actual application scenario. For example, the size of the 3D lighting grid can be set to 160cm (centimeter) × 160cm × 160cm. The grid points in the 3D lighting grid can be set to a uniform spacing mode, and the spacing between grid points can be set to 5cm. Figure 2 As shown, the 3D lighting grid can be set to Figure 2 The three-dimensional grid shown.
[0047] The lighting tracks constructed within the 3D lighting grid can be configured based on the needs of the actual application scenario. For example, the lighting tracks can include horizontal lighting tracks along the horizontal direction, vertical lighting tracks along the vertical direction, and diagonal lighting tracks along diagonal corners of the 3D lighting grid. The number of lighting tracks can also be configured based on the needs of the actual application scenario.
[0048] Step S104: Acquire a first image set and at least one second image set corresponding to each lighting track.
[0049] The first image set includes images of a whiteboard generated by illuminating a whiteboard in a blank space with a point light source positioned at each lighting position along a corresponding lighting trajectory. For example, if the first lighting trajectory is any one of the multiple constructed lighting trajectories, and there are 10 grid points along the first lighting trajectory, then the first image set corresponding to the first lighting trajectory includes images of the whiteboard generated by illuminating the whiteboard with a point light source positioned at each of the 10 grid points. No other objects exist in the blank space except the whiteboard.
[0050] The second image set includes model images generated by illuminating the 3D model with point light sources located at respective lighting positions along the corresponding lighting trajectory. The second image set corresponds one-to-one with the 3D model. For example, the 3D model may be a 3D human figure. It should be understood that the 3D model may also be a 3D model of an animal or other object, and this is not limited in this embodiment.
[0051] Optionally, the number of 3D models can be set based on the needs of the actual application scenario. The number of second image sets corresponding to each lighting trajectory is the same as the number of 3D models. For example, when constructing a video dataset, 65 3D character models can be set. Then, the number of second image sets corresponding to each lighting trajectory is 65.
[0052] Exemplarily, taking the first lighting trajectory and any 3D model as an example, the corresponding second image set includes model images generated after the point light source is located at each of the above 10 grid points and illuminates the 3D model.
[0053] There are multiple ways to obtain the first image set and at least one second image set corresponding to each lighting track. The following takes the first lighting track as an example to illustrate the way to obtain the first image set and at least one second image set corresponding to each lighting track.
[0054] In some optional embodiments, obtaining the first image set corresponding to the first lighting trajectory can be achieved as follows: controlling a point light source to sequentially move to each lighting position of the first lighting trajectory; after the point light source moves to any lighting position of the first lighting trajectory to illuminate a whiteboard in a blank space, a frame of whiteboard image is generated by rendering using a virtual engine; and constructing all whiteboard images corresponding to the first lighting trajectory into a first image set corresponding to the first lighting trajectory in chronological order of generation. In this manner, the first image set corresponding to the first lighting trajectory can be directly and quickly obtained, and the efficiency of obtaining the first image set is relatively high.
[0055] In some optional embodiments, obtaining the first image set corresponding to the first lighting trajectory can also be achieved in the following manner: controlling the point light source to move to each lighting position of the 3D lighting grid, and after the point light source moves to any lighting position of the 3D lighting grid to illuminate the whiteboard in the blank space, a frame of whiteboard image is generated by virtual engine rendering, and the position information of the corresponding grid point is marked in each frame of the whiteboard image; then, based on the position information of the grid point, the whiteboard image corresponding to each grid point in the first lighting trajectory is selected from all the whiteboard images generated above, and then all the selected whiteboard images are constructed into the first image set corresponding to the first lighting trajectory according to the trajectory order of the first lighting trajectory. In this way, there is no need to repeatedly render and generate the same whiteboard image for the same lighting position, which can save the cost of rendering resources.
[0056] In some optional embodiments, obtaining the first image set corresponding to the first lighting trajectory can also be achieved in the following manner: controlling the point light source to move to each lighting position of each constructed lighting trajectory respectively, after the point light source moves to any lighting position of the lighting trajectory to illuminate the whiteboard in the blank space, a frame of whiteboard image is generated by virtual engine rendering, and the position information of the corresponding grid point is marked in each frame of the whiteboard image; then, based on the position information of the grid point, all the generated whiteboard images are deduplicated, that is, only one frame of whiteboard images with the same lighting position is retained; then, the whiteboard image corresponding to each grid point in the first lighting trajectory is selected from all the deduplicated whiteboard images, and then, according to the trajectory order of the first lighting trajectory, all the selected whiteboard images are constructed into the first image set corresponding to the first lighting trajectory. In this way, there is no need to store a large number of identical whiteboard images, and storage resources can be optimized.
[0057] Similarly, various methods can be used to obtain any second image set corresponding to each lighting trajectory. Assume that the first lighting trajectory is any one of the multiple constructed lighting trajectories. The third image set is any one of the at least one second image set corresponding to the first lighting trajectory. The first 3D model is the 3D model corresponding to the third image set. Below, using the first lighting trajectory, the first 3D model, and the third image set as examples, we will illustrate methods for obtaining the second image set corresponding to each lighting trajectory.
[0058] In some optional embodiments, obtaining the third image set corresponding to the first lighting trajectory can be achieved as follows: controlling a point light source to move sequentially along each lighting position of the first lighting trajectory; after the point light source moves to any lighting position along the first lighting trajectory to illuminate the first 3D model, generating a frame of model image through virtual engine rendering; and constructing all model images corresponding to the first lighting trajectory and the first 3D model into a third image set corresponding to the first lighting trajectory and the first 3D model in chronological order of generation. In this manner, the third image set corresponding to the first lighting trajectory can be directly and quickly obtained, and the efficiency of obtaining the third image set is relatively high.
[0059] In some optional embodiments, obtaining the third image set corresponding to the first lighting trajectory can also be achieved in the following manner: controlling the point light source to move to each lighting position of the 3D lighting grid, and after the point light source moves to any lighting position of the 3D lighting grid to illuminate the first 3D model, a frame of model image is generated by virtual engine rendering, and the position information of the corresponding grid point is marked in each frame of the model image; then, based on the position information of the grid point, the model image corresponding to each grid point in the first lighting trajectory is selected from all the model images generated above, and then, according to the trajectory order of the first lighting trajectory, all the selected model images are constructed into a third image set corresponding to the first lighting trajectory and the first 3D model. In this way, there is no need to repeatedly render and generate the same model image for the same lighting position, which can save the cost of rendering resources.
[0060] In some optional embodiments, obtaining the third image set corresponding to the first lighting trajectory can also be achieved in the following manner: controlling the point light source to move to each lighting position of each constructed lighting trajectory respectively, after the point light source moves to any lighting position of the lighting trajectory to illuminate the first 3D model, a frame of model image is generated by virtual engine rendering, and the position information of the corresponding grid point is marked in each frame of the model image; then, based on the position information of the grid point, all the generated model images are deduplicated, that is, only one frame of model images with the same lighting position is retained; then, the model image corresponding to each grid point in the first lighting trajectory is selected from all the deduplicated model images, and then, according to the trajectory order of the first lighting trajectory, all the selected model images are constructed into a third image set corresponding to the first lighting trajectory and the first 3D model. In this way, there is no need to store a large number of identical template images, and storage resources can be optimized.
[0061] Step S106: Generate a first video corresponding to each first image set and a second video corresponding to each second image set.
[0062] After obtaining the first image set and at least one second image set corresponding to each lighting track, a video corresponding to each first image set is generated, denoted as a first video; and a video corresponding to each second image set is generated, denoted as a second video. Furthermore, each first video is associated with the corresponding lighting track. Each second video is associated with the corresponding lighting track and 3D model.
[0063] Step S108: annotate text information for each second video.
[0064] In some optional embodiments, the initial text information of each second video can be generated by a bootstrapping language-image pre-training (BLIP) model, and then the corresponding initial text information is annotated for each second video, that is, the initial text information of each second video is annotated on the second video as the text information.
[0065] In some optional embodiments, initial text information for each second video can be generated using an image-language pre-trained model. Then, title enhancement processing is performed on the initial text information for each second video to generate enhanced text information for each second video. Finally, the corresponding enhanced text information is annotated for each second video. That is, the enhanced text information for each second video is annotated as text information on the second video. This approach results in more accurate text information annotated for the second video.
[0066] Step S110: construct a video dataset for training a text-to-video generation model based on all first videos and all second videos annotated with text information.
[0067] The video dataset includes multiple training samples, each training sample includes a first video corresponding to a lighting trajectory and a second video with annotated text information of a 3D model corresponding to the lighting trajectory, and both the lighting trajectory and the 3D model correspond one-to-one to the training sample.
[0068] For example, if there are 65 3D character models, the first lighting trajectory can correspond to 65 training samples. Each training sample includes a first video corresponding to the first lighting trajectory and a second video corresponding to one of the 65 second videos annotated with text information corresponding to the first lighting trajectory. Different training samples correspond to different second videos of 3D character models.
[0069] In the video dataset construction method provided in the above embodiments, a large number of lighting trajectories are constructed in a 3D lighting network. For each lighting trajectory, a video of a whiteboard image representing the lighting information and a model video representing a 3D model (e.g., a 3D person model) are generated. Subsequently, a video dataset for training a text video generation model is generated based on the videos representing the lighting information and the model information.
[0070] When training a text generation model based on the above video dataset, the lighting information can be integrated into the text video generation model. In the subsequent video generation process, the lighting direction and motion trajectory can be controlled so that the generated video can better express the lighting effect, better meet user needs, and provide a better user experience.
[0071] It can be understood that the above embodiments are only examples and can be modified in actual implementation. Those skilled in the art can understand that the modification methods of the above embodiments without creative work fall within the protection scope of one or more embodiments of this specification and will not be repeated in the embodiments.
[0072] Based on the same inventive concept, one or more embodiments of this specification further provide a model training method based on a video dataset, wherein the video dataset can be constructed according to the video dataset construction method provided in the above embodiments.
[0073] See also Figure 3 , Figure 3 This is a flow chart of a model training method based on a video dataset provided in one or more embodiments of this specification. This method can be applied to a server or terminal device, etc. The following takes the terminal device as an example to introduce the content of the embodiment. Figure 3 As shown, the method may include the following steps:
[0074] Step S202: construct a frozen denoising network and an external transformer network.
[0075] The frozen denoising network has the same architecture as the pre-trained text-video generation model and is used to freeze the network weight parameters of the video denoising network in the pre-trained text-video generation model. In other words, the pre-trained text-video generation model can include a video denoising network. The frozen denoising network can be set to the same architecture as the pre-trained text-video generation model, and then the network weight parameters of the video denoising network in that architecture are frozen.
[0076] The video denoising network in the pre-trained text-to-video generation model can include a variational encoder for feature extraction and a preset number of diffusion transformer blocks. The preset number can be set based on the actual application scenario. For example, the preset number can be set to 32, meaning that the video denoising network can include 32 diffusion transformer blocks. Each diffusion transformer block includes a multi-head attention layer.
[0077] The outer transformer network consists of the same variational self-attention encoder as the video denoising network, and a preset number of light encoder blocks, which are transformer architectures including self-attention layers.
[0078] For example, the architecture of the video denoising network and the external transformer network can be found in Figure 4 .
[0079] Step S204: Based on the training samples and the objective function in the video data set, the pre-trained text-video generation model, the external transformer network, and the frozen denoising network are trained to generate a trained text-video generation model and an external transformer network.
[0080] The objective function is a decoupling loss function between the predicted noise of the video denoising network and the predicted noise of the frozen denoising network.
[0081] Exemplarily, the objective function may include the following:
[0082]
[0083] L total =L denoise +βL dis .
[0084] Among them, L dis represents the decoupling loss, L denoise represents the denoising loss, σ(·) represents the standard deviation, and μ(·) represents the mean. represents the predicted noise of the frozen denoising network at time step t. represents the predicted noise of the video denoising network with time step t. N represents the total step size. c t Indicates text condition. β is set to 3.0.
[0085] like Figure 4 As shown in the figure, during any training iteration, the lighting features extracted from each lighting encoder block in the external transformer network are input into the corresponding diffusion transformer block in the video denoising network. After being fused with the video features extracted from the diffusion transformer block, they are input into the multi-head attention layer of the diffusion transformer block. This allows the lighting information to be integrated into the video.
[0086] In some optional embodiments, the lighting features and the video features may be fused by element-by-element addition.
[0087] In some optional embodiments, when executing step S204, the second video with text information annotated in the training sample in the video dataset is input into the pre-trained text video generation model and the frozen denoising network respectively, and the first video in the corresponding training sample is input into the external transformer network, and the model training is performed based on the objective function.
[0088] The first training sample is assumed to be any training sample in a video dataset. During training, the second video annotated with text information included in the first training sample is input into a pre-trained text-video generation model and a frozen denoising network, respectively. Furthermore, the first video included in the first training sample is input into an external transformer network. The predicted noise of the video denoising network and the predicted noise of the frozen denoising network are then calculated. The predicted noise of the video denoising network and the predicted noise of the frozen denoising network are used to optimize the objective function and adjust the model parameters.
[0089] When the objective function is optimized to the minimum loss, the model training is completed, and the trained text-video generation model and external transformer network can be obtained.
[0090] In the model training method provided by the above embodiment, on the basis of the pre-trained text video generation model, a frozen denoising network with the same architecture as the pre-trained text video generation model is constructed, and the network weight parameters of the video denoising network are frozen in the frozen denoising network. An external transformer network that shares the same variable self-encoder and the same number of transformer blocks as the video denoising network is also constructed. During the model training process, the lighting features are integrated into the video features through the external transformer network, and the pre-trained text video generation model and the external transformer network (together as a control branch) are established through the frozen denoising network, and the decoupling between the two branches of the frozen denoising network (as a frozen branch) avoids the overfitting of the model, thereby training to obtain a more accurate text video generation model and an external transformer network that can control the lighting direction and lighting motion trajectory. Subsequently, a video that better meets user needs can be generated, and the user experience is better.
[0091] Based on the same inventive concept, one or more embodiments of this specification further provide a video generation method. In this video generation method, the text video generation model trained by the model training method provided in the above embodiments and an external transformer network are used to generate the target video required by the user.
[0092] See also Figure 5 , Figure 5 This is a flow chart of a video generation method provided in one or more embodiments of this specification. This method can be applied to a server or terminal device, etc. The following takes the terminal device as an example to introduce the content of the embodiment. Figure 5 As shown, the method may include the following steps:
[0093] Step S302: Select a target lighting trajectory required for generating a target video from multiple lighting trajectories in the 3D lighting grid.
[0094] The specific contents of the 3D lighting grid and the lighting trajectory can be found in the above embodiments and will not be repeated here.
[0095] In some optional embodiments, a target lighting track required for generating a target video may be selected from a plurality of lighting tracks in a 3D lighting grid according to user input.
[0096] Step S304: input the first video corresponding to the target lighting trajectory into an external transformer network, and input the target text information used to generate the target video into a text video generation model to generate the target video.
[0097] In the process of generating the target video, the lighting features extracted from each lighting encoder block in the external transformer network are input into the corresponding diffusion transformer block of the video denoising network in the text video generation model, and are fused with the video features extracted by the diffusion transformer block and then input into the multi-head attention layer of the diffusion transformer block.
[0098] The method for obtaining the first video corresponding to the target lighting trajectory can refer to the content of the aforementioned embodiment and will not be repeated here.
[0099] In some optional embodiments, the target text information may be obtained based on user input, or may be pre-stored in the system and retrieved from the system, which is not limited in the embodiments of this specification.
[0100] In the video generation method provided in the above embodiment, lighting information that meets user needs is integrated during the video generation process, so that the generated video can better express the lighting effect and provide a better user experience.
[0101] Based on the same inventive concept, one or more embodiments of this specification also provide a video dataset construction device, a model training device, and a video generation device. Since the principles of the problems solved by these devices are similar to those of the aforementioned methods, the implementation of the devices can refer to the implementation of the aforementioned methods, and the repeated parts will not be repeated.
[0102] See also Figure 6 , Figure 6 This is a structural block diagram of a video data set construction device provided in one or more embodiments of this specification. Figure 6 As shown, the video data set construction device 600 may include: a first construction module 601, an acquisition module 602, a first generation module 603, a labeling module 604, and a second construction module 605.
[0103] A first construction module 601 is configured to construct a plurality of lighting trajectories in a 3D lighting grid; the 3D lighting grid is a three-dimensional grid whose grid points serve as lighting positions; and the lighting trajectories are trajectories formed by connecting a plurality of the grid points;
[0104] Acquisition module 602 is configured to acquire a first image set and at least one second image set corresponding to each lighting trajectory; the first image set includes whiteboard images generated by point light sources located at respective lighting positions of the corresponding lighting trajectory illuminating a whiteboard in a blank space; the second image set includes model images generated by the point light sources located at respective lighting positions of the corresponding lighting trajectory illuminating a 3D model; the second image sets correspond one-to-one to the 3D models;
[0105] A first generating module 603 is configured to generate a first video corresponding to each first image set and a second video corresponding to each second image set;
[0106] Annotation module 604, configured to annotate text information for each second video;
[0107] The second construction module 605 is used to construct a video dataset for training a text video generation model based on all the first videos and all the second videos annotated with text information; the video dataset includes multiple training samples, each training sample includes a first video corresponding to a lighting trajectory, and a second video annotated with text information of a 3D model corresponding to the lighting trajectory, and the lighting trajectory and the 3D model have a one-to-one correspondence with the training sample.
[0108] In one possible implementation, the first lighting trajectory is any one of the multiple lighting trajectories; the acquisition module 602 is used to obtain a first image set corresponding to the first lighting trajectory, including: the acquisition module 602 is used to: control the point light source to move in sequence on each lighting position of the first lighting trajectory, after the point light source moves to any lighting position to illuminate the whiteboard in the blank space, generate a frame of whiteboard image through virtual engine rendering; all whiteboard images corresponding to the first lighting trajectory are constructed into a first image set corresponding to the first lighting trajectory in the order of generation time.
[0109] In one possible implementation, the first lighting trajectory is any one of the multiple lighting trajectories; the third image set is any one of the at least one second image set corresponding to the first lighting trajectory; the first 3D model is the 3D model corresponding to the third image set; the acquisition module 602 is used to acquire the third image set corresponding to the first lighting trajectory, including: the acquisition module 602 is used to: control the point light source to move in sequence on each lighting position of the first lighting trajectory, after the point light source moves to any lighting position to illuminate the first 3D model, generate a frame of model image through virtual engine rendering; all model images corresponding to the first lighting trajectory and the first 3D model are constructed into a third image set corresponding to the first lighting trajectory and the first 3D model in order of generation time.
[0110] In one possible implementation, the annotation module 604 is used to annotate text information for each second video, including: the annotation module 604 is used to: generate initial text information for each second video through an image language pre-training model; perform title enhancement processing on the initial text information of each second video to generate enhanced text information for each second video; and annotate corresponding enhanced text information for each second video.
[0111] In a possible implementation, the 3D model is a 3D character model.
[0112] See also Figure 7 , Figure 7 This is a structural block diagram of a model training device provided in one or more embodiments of this specification. Figure 7 As shown, the model training device 700 may include: a third construction module 701 and a training module 702.
[0113] A third construction module 701 is configured to construct a frozen denoising network and an external transformer network; the frozen denoising network has the same architecture as the pre-trained text-video generation model and is configured to freeze network weight parameters of the video denoising network in the pre-trained text-video generation model; the video denoising network includes a variational self-decomposition encoder for feature extraction and a preset number of diffusion transformer blocks; the external transformer network includes the same variational self-decomposition encoder as the video denoising network and the preset number of illumination encoder blocks, wherein the illumination encoder blocks are transformer architectures including self-attention layers;
[0114] A training module 702 is configured to train the pre-trained text-video generation model, the external transformer network, and the frozen denoising network based on training samples and an objective function in the video dataset, thereby generating a trained text-video generation model and external transformer network. The objective function is a decoupling loss function between the predicted noise of the video denoising network and the predicted noise of the frozen denoising network. During any training iteration, the illumination features extracted from each illumination encoder block are input into the corresponding diffusion transformer block, where they are fused with the video features extracted from the diffusion transformer block and then input into the multi-head attention layer of the diffusion transformer block. The video dataset is constructed based on the aforementioned video dataset construction device 600.
[0115] In one possible implementation, the training module 702 is used to train the pre-trained text video generation model, the external transformer network and the frozen denoising network based on the training samples and the objective function in the video data set, including: the training module is used to: input the second video with text information annotated in the training sample in the video data set into the pre-trained text video generation model and the frozen denoising network respectively, and input the first video in the corresponding training sample into the external transformer network, and perform model training based on the objective function.
[0116] In a possible implementation, the training module 702 is configured to perform feature fusion on the lighting feature and the video feature by element-by-element addition.
[0117] See also Figure 8 , Figure 8 This is a structural block diagram of a video generation device provided in one or more embodiments of this specification. Figure 8 As shown, the video generation device 800 may include: a selection module 801 and a second generation module 802.
[0118] A selection module 801 is configured to select a target lighting track required for generating a target video from a plurality of lighting tracks in a 3D lighting grid;
[0119] The second generation module 802 is configured to input the first video corresponding to the target lighting trajectory into the external transformer network and input the target text information used to generate the target video into the text video generation model to generate the target video. During the generation of the target video, the lighting features extracted from each lighting encoder block are input into the corresponding diffusion transformer block, where they are fused with the video features extracted by the diffusion transformer block and then input into the multi-head attention layer of the diffusion transformer block. The external transformer network and text video generation model are trained and generated by the aforementioned model training device 700.
[0120] See also Figure 9 , Figure 9 This is a structural block diagram of an electronic device provided in one or more embodiments of this specification. Figure 9 As shown, the electronic device 900 may include a processor 901 and a memory 902; the memory 902 may be coupled to the processor 901. Figure 9 is exemplary; other types of structures may also be used to supplement or replace this structure to implement telecommunication functions or other functions.
[0121] In one possible implementation, the functions of the video dataset construction device 600 , the model training device 700 , and the video generation device 800 may be integrated into the processor 901 .
[0122] In another possible implementation, the video dataset construction device 600, the model training device 700, and the video generation device 800 can be configured separately from the processor 901. For example, the video dataset construction device 600, the model training device 700, and the video generation device 800 can be configured as a chip connected to the processor 901, and corresponding processing is implemented through the control of the processor 901.
[0123] In addition, in some optional implementations, the electronic device 900 may further include: a communication module, an input unit, an audio processor, a display, a power supply, etc. It is worth noting that the electronic device 900 does not necessarily have to include Figure 9 In addition, the electronic device 900 may also include all components shown in Figure 9For components not shown, reference may be made to the prior art.
[0124] In some optional implementations, the processor 901 is sometimes also referred to as a controller or operation control, and may include a microprocessor or other processor device and / or logic device, which receives input and controls the operation of various components of the electronic device 900.
[0125] Memory 902 may be, for example, one or more of a cache, flash memory, hard drive, removable media, volatile memory, non-volatile memory, or other suitable devices. It may store information related to the video dataset construction device 600, model training device 700, and video generation device 800, as well as programs that execute the relevant information. Processor 901 may execute the programs stored in memory 902 to implement information storage or processing.
[0126] The input unit can provide input to the processor 901. The input unit can be, for example, a keypad or a touch input device. The power supply can be used to provide power to the electronic device 900. The display can be used to display objects such as images and text. The display can be, for example, an LCD display, but is not limited thereto.
[0127] The memory 902 may be a solid-state memory, such as a read-only memory (ROM), a random access memory (RAM), a SIM card, or the like. Alternatively, it may be a memory that retains information even when power is off, can be selectively erased, and is provided with more data. Examples of such memory are sometimes referred to as EPROMs. The memory 902 may also be some other type of device. The memory 902 includes a buffer memory (sometimes referred to as a buffer). The memory 902 may include an application / function storage unit for storing application programs and function programs or processes for executing electronic device 900 operations via the processor 901.
[0128] The memory 902 may also include a data storage unit for storing data, such as contacts, digital data, pictures, sounds, and / or any other data used by the electronic device. The driver storage unit of the memory 902 may include various driver programs for the computer device for communication functions and / or for executing other functions of the computer device (such as a messaging application, a contact book application, etc.).
[0129] The communication module is a transmitter / receiver that sends and receives signals via an antenna. The communication module (transmitter / receiver) is coupled to the processor 901 to provide input signals and receive output signals, which may be the same as in a conventional mobile communication terminal.
[0130] Based on different communication technologies, multiple communication modules can be provided in the same computer device, such as a cellular network module, a Bluetooth module, and / or a wireless local area network module. The communication module (transmitter / receiver) is also coupled to a speaker and a microphone via an audio processor to provide audio output via the speaker and receive audio input from the microphone, thereby implementing common telecommunication functions. The audio processor may include any suitable buffer, decoder, amplifier, etc. In addition, the audio processor is also coupled to the processor 901, thereby enabling recording of the device via the microphone and playback of stored audio via the speaker.
[0131] One or more embodiments of this specification also provide a computer-readable storage medium that can implement all the steps of the video dataset construction method, model training method, and video generation method in the above-mentioned embodiments. The computer-readable storage medium stores a computer program that, when executed by a processor, implements all the steps of the video dataset construction method, model training method, and video generation method in the above-mentioned embodiments.
[0132] Although one or more embodiments of this specification provide method operation steps as described in the embodiments or flowcharts, more or fewer operation steps may be included based on conventional or non-creative work. The order of steps listed in the embodiments is only one way of executing the steps among many steps and does not represent the only execution order. When an actual device or client product is executed, it can be executed in sequence or in parallel according to the method shown in the embodiments or the drawings (for example, in a parallel processor or multi-threaded processing environment).
[0133] Those skilled in the art will appreciate that the embodiments of this specification may be provided as methods, devices (systems), or computer program products. Therefore, the embodiments of this specification may take the form of entirely hardware embodiments, entirely software embodiments, or embodiments combining software and hardware. Furthermore, one or more embodiments of this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0134] The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks. Figure 1 The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks. Figure 1 The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks.
[0135] The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks. Figure 1 The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks. Figure 1 The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks.
[0136] The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks. Figure 1 The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks. Figure 1 The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks.
[0137] The embodiments in the specification are described progressively, and the same or similar parts among the embodiments can be mutually referred to. Each embodiment focuses on the difference from other embodiments. In particular, the device and system embodiments are described simply because they are basically similar to the method embodiments, and the relevant parts can be referred to the description of the method embodiments.
[0138] In this document, relational terms such as first and second and the like can be used solely to distinguish one entity or action from another entity or action without necessarily requiring or implying any actual such relationship or order between such entities or actions. Moreover, the terms "comprises", "comprising", or any other variation thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can include other elements not expressly listed or inherent to such process, method, article, or apparatus. Any process, method, article, or apparatus that comprises a list of elements can comprise only those elements but can also comprise other elements not expressly listed or inherent to such process, method, article, or apparatus.
[0139] It should be noted that the features of one or more embodiments and embodiments described in the specification above can be combined with each other and the application is not limited to any combination or permutation of the principles described above; the application is limited only by the claims that follow, including any equivalent claims submitted during prosecution.
[0140] Finally, it should be noted that the above-mentioned embodiments merely illustrate the technical solutions of the one or more embodiments, instead of limiting the one or more embodiments; no matter how specific the above-mentioned embodiments are, the ordinary skill in the art should understand: the technical solutions recorded in the above-mentioned embodiments can be modified, or some or all of the technical features can be substituted with equivalent features; if these modifications or substitutions do not deviate from the scope of the technical solutions of the one or more embodiments, they should be contained in the scope of the one or more embodiments.
[0141] The one or more embodiments of the present specification are described above in combination with optional embodiments, but these embodiments are only exemplary and serve only to illustrate. On this basis, various substitutions and improvements can be made to the one or more embodiments of the present specification, and these all fall within the protection scope of the one or more embodiments of the present specification.
Claims
1. A method for constructing a video dataset, characterized in that: The method comprises: Constructing a plurality of lighting tracks in a 3D lighting grid; the 3D lighting grid is a three-dimensional grid whose grid points serve as lighting positions; the lighting tracks are tracks formed by connecting a plurality of the grid points; Acquire a first image set and at least one second image set corresponding to each lighting trajectory; the first image set includes whiteboard images generated by point light sources located at each lighting position of the corresponding lighting trajectory illuminating a whiteboard in a blank space; the second image set includes model images generated by the point light sources located at each lighting position of the corresponding lighting trajectory illuminating a 3D model; the second image sets correspond one-to-one to the 3D models; generating a first video corresponding to each first image set and a second video corresponding to each second image set; Annotating text information for each second video; Based on all first videos and all second videos annotated with text information, a video dataset for training a text video generation model is constructed; the video dataset includes multiple training samples, each training sample includes a first video corresponding to a lighting trajectory, and a second video annotated with text information of a 3D model corresponding to the lighting trajectory, and both the lighting trajectory and the 3D model have a one-to-one correspondence with the training samples.
2. The method according to claim 1, wherein The first lighting track is any one of the multiple lighting tracks; Acquiring a first image set corresponding to a first lighting trajectory includes: Controlling the point light source to move sequentially to each lighting position of the first lighting trajectory, after the point light source moves to any lighting position to illuminate the whiteboard in the blank space, generating a frame of whiteboard image through virtual engine rendering; All whiteboard images corresponding to the first lighting trajectory are constructed into a first image set corresponding to the first lighting trajectory in order of generation time.
3. The method according to claim 1, wherein The first lighting track is any one of the multiple lighting tracks; the third image set is any one of the at least one second image set corresponding to the first lighting track; The first 3D model is a 3D model corresponding to the third image set; Acquiring a third image set corresponding to the first lighting trajectory, including: Controlling the point light source to move sequentially to each lighting position of the first lighting trajectory, and after the point light source moves to any lighting position to illuminate the first 3D model, generating a frame of the model image through virtual engine rendering; All model images corresponding to the first lighting trajectory and the first 3D model are constructed into a third image set corresponding to the first lighting trajectory and the first 3D model in order of generation time.
4. The method according to claim 1, wherein Annotate each second video with text information, including: Generate initial text information for each second video through an image-language pre-training model; Performing title enhancement processing on the initial text information of each second video to generate enhanced text information of each second video; Each second video is labeled with corresponding enhanced text information.
5. The method according to any one of claims 1 to 4, wherein: The 3D model is a 3D character model.
6. A model training method based on a video dataset, characterized in that: The video dataset is constructed by the method according to any one of claims 1 to 5, the method comprising: Constructing a frozen denoising network and an external transformer network; the frozen denoising network has the same architecture as a pre-trained text-video generation model and is used to freeze the network weight parameters of the video denoising network in the pre-trained text-video generation model; the video denoising network includes a variational self-decomposition encoder for feature extraction and a preset number of diffusion transformer blocks; the external transformer network includes the same variational self-decomposition encoder as the video denoising network and the preset number of illumination encoder blocks, each illumination encoder block having a transformer architecture including a self-attention layer; Based on the training samples and the objective function in the video dataset, the pre-trained text-video generation model, the external transformer network, and the frozen denoising network are trained to generate a trained text-video generation model and an external transformer network; the objective function is a decoupling loss function of the predicted noise of the video denoising network and the predicted noise of the frozen denoising network; Among them, in any training iteration process, the lighting features extracted from each lighting encoder block are input into the corresponding diffusion transformer block, and are fused with the video features extracted from the diffusion transformer block and then input into the multi-head attention layer of the diffusion transformer block.
7. The method according to claim 6, wherein Training the pre-trained text-to-video generation model, the external transformer network, and the frozen denoising network based on training samples and an objective function in the video dataset includes: The second video with text information annotated in the training sample in the video dataset is input into the pre-trained text video generation model and the frozen denoising network respectively, and the first video in the corresponding training sample is input into the external transformer network, and the model training is performed based on the objective function.
8. The method according to claim 6, wherein The lighting features and the video features are fused by element-by-element addition.
9. A video generation method based on an external transformer network and a text-to-video generation model trained and generated by the method according to any one of claims 6 to 8, characterized in that: The method comprises: Selecting a target lighting track required to generate a target video from a plurality of lighting tracks in the 3D lighting grid; Inputting a first video corresponding to the target lighting trajectory into the external transformer network, and inputting target text information for generating a target video into the text video generation model to generate the target video; In the process of generating the target video, the lighting features extracted from each lighting encoder block are input into the corresponding diffusion transformer block, and are fused with the video features extracted by the diffusion transformer block and then input into the multi-head attention layer of the diffusion transformer block.
10. A video data set construction device, characterized in that: The device comprises: A first construction module is configured to construct a plurality of lighting tracks in a 3D lighting grid; the 3D lighting grid is a three-dimensional grid whose grid points serve as lighting positions; and the lighting tracks are tracks formed by connecting a plurality of the grid points; an acquisition module, configured to acquire a first image set and at least one second image set corresponding to each lighting trajectory; the first image set comprising whiteboard images generated by point light sources located at respective lighting positions of the corresponding lighting trajectory illuminating a whiteboard in a blank space; and the second image set comprising model images generated by the point light sources located at respective lighting positions of the corresponding lighting trajectory illuminating a 3D model; the second image sets corresponding to the 3D models. A first generating module is configured to generate a first video corresponding to each first image set and a second video corresponding to each second image set; a marking module, configured to mark text information for each second video; The second construction module is used to construct a video dataset for training a text video generation model based on all the first videos and all the second videos annotated with text information; the video dataset includes multiple training samples, each training sample includes a first video corresponding to a lighting trajectory, and a second video annotated with text information of a 3D model corresponding to the lighting trajectory, and the lighting trajectory and the 3D model have a one-to-one correspondence with the training samples.
11. The device according to claim 10, wherein The first lighting track is any one of the multiple lighting tracks; The acquisition module is used to acquire a first image set corresponding to the first lighting trajectory, including: The acquisition module is used to: Controlling the point light source to move sequentially to each lighting position of the first lighting trajectory, after the point light source moves to any lighting position to illuminate the whiteboard in the blank space, generating a frame of whiteboard image through virtual engine rendering; All whiteboard images corresponding to the first lighting trajectory are constructed into a first image set corresponding to the first lighting trajectory in order of generation time.
12. The device according to claim 10, wherein The first lighting track is any one of the multiple lighting tracks; the third image set is any one of the at least one second image set corresponding to the first lighting track; The first 3D model is a 3D model corresponding to the third image set; The acquisition module is used to acquire a third image set corresponding to the first lighting trajectory, including: The acquisition module is used to: Controlling the point light source to move sequentially to each lighting position of the first lighting trajectory, and after the point light source moves to any lighting position to illuminate the first 3D model, generating a frame of the model image through virtual engine rendering; All model images corresponding to the first lighting trajectory and the first 3D model are constructed into a third image set corresponding to the first lighting trajectory and the first 3D model in order of generation time.
13. The device according to claim 10, wherein The annotation module is used to annotate text information for each second video, including: The annotation module is used to: Generate initial text information for each second video through an image-language pre-training model; Performing title enhancement processing on the initial text information of each second video to generate enhanced text information of each second video; Each second video is labeled with corresponding enhanced text information.
14. The device according to any one of claims 10 to 13, characterized in that The 3D model is a 3D character model.
15. A model training device based on a video dataset, characterized in that: The video dataset is constructed by the video dataset construction device according to any one of claims 10 to 14, and the model training device comprises: a third construction module, configured to construct a frozen denoising network and an external transformer network; the frozen denoising network having the same architecture as the pre-trained text-video generation model, and configured to freeze network weight parameters of the video denoising network in the pre-trained text-video generation model; the video denoising network comprising a variational self-decomposition encoder for feature extraction and a preset number of diffusion transformer blocks; the external transformer network comprising the same variational self-decomposition encoder as the video denoising network and the preset number of illumination encoder blocks, wherein the illumination encoder blocks are transformer architectures including self-attention layers; A training module is configured to train the pre-trained text-video generation model, the external transformer network, and the frozen denoising network based on training samples in the video dataset and an objective function to generate a trained text-video generation model and an external transformer network; the objective function is a decoupling loss function of the predicted noise of the video denoising network and the predicted noise of the frozen denoising network; Among them, in any training iteration process, the lighting features extracted from each lighting encoder block are input into the corresponding diffusion transformer block, and are fused with the video features extracted from the diffusion transformer block and then input into the multi-head attention layer of the diffusion transformer block.
16. The model training device according to claim 15, wherein: The training module is used to train the pre-trained text-video generation model, the external transformer network, and the frozen denoising network based on the training samples and the objective function in the video dataset, including: The training module is used to: The second video with text information annotated in the training sample in the video dataset is input into the pre-trained text video generation model and the frozen denoising network respectively, and the first video in the corresponding training sample is input into the external transformer network, and the model training is performed based on the objective function.
17. The model training device according to claim 15, wherein: The training module is used to perform feature fusion on the lighting feature and the video feature by element-by-element addition.
18. A video generation device based on an external transformer network and a text-to-video generation model trained and generated by the model training device according to any one of claims 15 to 17, characterized in that: The video generating device comprises: A selection module, configured to select a target lighting track required to generate a target video from a plurality of lighting tracks in the 3D lighting grid; a second generation module, configured to input the first video corresponding to the target lighting trajectory into the external transformer network, and input target text information for generating the target video into the text video generation model to generate the target video; In the process of generating the target video, the lighting features extracted from each lighting encoder block are input into the corresponding diffusion transformer block, and are fused with the video features extracted by the diffusion transformer block and then input into the multi-head attention layer of the diffusion transformer block.
19. An electronic device, characterized in that: The electronic device comprises: a memory for storing a computer program product; The processor is configured to execute the computer program product stored in the memory, and when the computer program product is executed, the method described in any one of claims 1 to 5, 6 to 8, and 9 above is implemented.
20. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer program instructions, which, when executed, implement the method described in any one of claims 1-5, 6-8, and 9 above.