Video generation method and apparatus, electronic device, storage medium, and product
By employing convolutional layers in the encoding and decoding modules of the video generation model, combined with single-process and multi-process processing methods, the problem of high device memory consumption is solved, and efficient and high-quality video generation is achieved.
Patent Information
- Application Number
- CN202411869323.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-18
- Publication Date
- 2026-01-13
- Estimated Expiration
- 2044-12-18
AI Technical Summary
Existing video generation models cause excessive memory consumption when processing large-scale video frame data, limiting their application on resource-constrained devices and affecting the efficiency and quality of video generation.
By acquiring video description text and random noise, and utilizing the convolutional layers in the encoding and decoding modules, combined with the target processing method corresponding to the current process type, the processing of the video generation model is controlled, including convolutional processing methods under single-process and multi-process conditions, thereby reducing memory usage and improving processing efficiency.
It effectively reduces the memory footprint of the video generation model, improves the efficiency and quality of video generation, and enables the generation of high-resolution videos on resource-constrained devices.
Smart Images

Figure CN119743644B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image processing technology, and in particular to a video generation method, apparatus, electronic device, storage medium, and product. Background Technology
[0002] With the rapid development of artificial intelligence technology, Artificial Intelligence Generated Content (AIGC) is being used for video generation. For example, AIGC can be used to generate high-resolution, long-duration videos with rich detail and dynamic effects.
[0003] With the continuous improvement of video resolution and the exponential growth of video data volume, the aforementioned video generation process requires processing massive amounts of video frame data. However, existing video generation models, when processing large-scale video frame data, are prone to causing excessive memory consumption on devices where the video generation models are deployed. This limits the application of video generation models on resource-constrained devices. For example, when performing video generation processing on mobile terminal devices, model performance is limited, resulting in excessively long video generation times and affecting the processing efficiency and quality of video generation. Summary of the Invention
[0004] This invention provides a video generation method, apparatus, electronic device, storage medium, and product. By controlling the processing of the video generation model through a target processing method corresponding to the current process type, the efficiency and quality of video generation are improved.
[0005] According to one aspect of the present invention, a video generation method is provided, the method comprising:
[0006] Obtain the video description text and random noise;
[0007] Based on the target processing method corresponding to the current process type, the video generation model is controlled to process the input video description text and random noise so that the video generation model outputs the target video corresponding to the video description text.
[0008] The video generation model includes an encoding module and a decoding module. The encoding module encodes and processes the video description text and random noise, and inputs the processing result into the decoding module. The decoding module includes at least one convolutional layer, which includes a convolutional kernel and a processing matrix corresponding to the convolutional kernel. The processing matrix is constructed based on the processing parameters corresponding to the input and output channels of the convolutional kernel.
[0009] According to another aspect of the present invention, a video generation apparatus is provided, the apparatus comprising:
[0010] The text and noise acquisition module is used to acquire video description text and random noise;
[0011] The target video determination module is used to control the video generation model to process the input video description text and random noise based on the target processing method corresponding to the current process type, so that the video generation model outputs a target video corresponding to the video description text;
[0012] The video generation model includes an encoding module and a decoding module. The encoding module encodes and processes the video description text and random noise, and inputs the processing result into the decoding module. The decoding module includes at least one convolutional layer, which includes a convolutional kernel and a processing matrix corresponding to the convolutional kernel. The processing matrix is constructed based on the processing parameters corresponding to the input and output channels of the convolutional kernel.
[0013] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:
[0014] At least one processor; and
[0015] A memory that is communicatively connected to at least one processor; wherein,
[0016] The memory stores a computer program that can be executed by at least one processor, such that the at least one processor is able to perform the video generation method of any embodiment of the present invention.
[0017] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the video generation method of any embodiment of the present invention.
[0018] According to another aspect of the present invention, a computer program product is provided, comprising a computer program, characterized in that the computer program, when executed by a processor, implements a video generation method as described in any embodiment of the present invention.
[0019] The technical solution of this invention acquires video description text and random noise, and controls a video generation model to process the input video description text and random noise according to the target processing method corresponding to the current process type, so that the video generation model outputs a target video corresponding to the video description text. The video generation model includes an encoding module and a decoding module. The encoding module encodes the video description text and random noise and inputs the processing result to the decoding module, so that the decoding module obtains the target video based on the processing result. The decoding module includes at least one convolutional layer, and the at least one convolutional layer includes a convolutional kernel and a processing matrix corresponding to the convolutional kernel. The processing matrix is constructed based on the processing parameters corresponding to the input and output channels of the convolutional kernel. The processing matrix can accelerate the convolutional processing of corresponding features and reduce the memory usage of the device hosting the video generation model. This invention solves the problem in the prior art where, when performing video generation processing based on existing video generation models, the device's memory usage is too high, resulting in limited model processing performance and low video generation efficiency and quality. By controlling the processing of the video generation model by the target processing method corresponding to the current process type, the system usage during model processing is reduced, and the efficiency and quality of video generation are improved.
[0020] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0021] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0022] Figure 1 This is a flowchart of a video generation method provided in an embodiment of the present invention;
[0023] Figure 2 This is a flowchart of a video generation method provided in an embodiment of the present invention;
[0024] Figure 3 This is an example diagram of the processing matrix provided in this embodiment of the invention processing sub-noise;
[0025] Figure 4 This is an example diagram of performing three-dimensional convolution processing on the noise to be processed and the column-segmented three-dimensional convolution tensor provided in the embodiments of the present invention;
[0026] Figure 5 This is an example diagram of the output result and row-segmented three-dimensional convolution tensor for three-dimensional convolution processing provided by the embodiments of the present invention;
[0027] Figure 6 This is a flowchart of a video generation model training method provided in an embodiment of the present invention;
[0028] Figure 7 This is a schematic diagram of the structure of a video generation device provided in an embodiment of the present invention;
[0029] Figure 8 This is a schematic diagram of the structure of an electronic device that implements the video generation method of this invention. Detailed Implementation
[0030] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0031] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0032] Example 1
[0033] Figure 1 This is a flowchart of a video generation method provided in Embodiment 1 of the present invention. This embodiment is applicable to single-process or multi-process scenarios, where a target processing method corresponding to the single-process or multi-process is used to control the video generation model to output a target video based on the input video description text and random noise. This method can be executed by a video generation device, which can be implemented in hardware and / or software, and can be configured in electronic devices such as mobile phones, computers, or servers. Figure 1 As shown, the method includes:
[0034] S110, Obtain video description text and random noise.
[0035] The video description text can be used to summarize or explain in detail the content of the video to be generated. The video description text determines the content of the video to be generated. For example, if the video description text is "a puppy is running," then the subsequently generated video will also be a video of a puppy running. The random noise can be random Gaussian noise that follows a Gaussian distribution.
[0036] Specifically, in video generation scenarios, the video content information, including scenes, characters, actions, and dialogues, can be determined based on actual needs. Video description text is then determined based on this content information. Random noise following a preset Gaussian distribution is then acquired to generate a video corresponding to the video description text, based on the video description text and the random noise.
[0037] S120. Based on the target processing method corresponding to the current process type, control the video generation model to process the input video description text and random noise, so that the video generation model outputs the target video corresponding to the video description text.
[0038] The current process type can be determined based on the number of Graphics Processing Units (GPUs) on the device deploying the video generation model. For example, if the device deploying the video generation model has one GPU, the current process type can be single-process. If the device deploying the video generation model has multiple GPUs, the current process type can be multi-process. Single-process can be understood as only one process executing during the video generation model's processing of the input video description text and random noise. Multi-process can be understood as multiple processes executing in parallel during the video generation model's processing of the input video description text and random noise, thereby increasing the speed of convolution processing and improving the output speed of the target video.
[0039] The target processing method can be based on determining the corresponding processing matrix according to the number of input and output channels of the convolution kernel, and then processing the noise to be processed based on the processing matrix. The number of input channels of the convolution kernel can be understood as the number of channels of the input data. The number of input channels is used to characterize the number of feature dimensions that the convolution kernel needs to process. Each input channel can be regarded as a response to different features. The number of output channels can be used to determine the number of output feature maps after the convolution operation.
[0040] The video generation model includes an encoding module and a decoding module. The encoding module encodes and processes the video description text and random noise, and inputs the processing result into the decoding module. The decoding module includes at least one convolutional layer, which includes a convolutional kernel and a processing matrix corresponding to the convolutional kernel. The processing matrix is constructed based on the processing parameters corresponding to the input and output channels of the convolutional kernel.
[0041] The processing results can be used to characterize the features output by the encoding module of the video generation model. In the decoding module, convolutional layers are used to perform weighted summation of local connectivity opinions on the input features of the convolutional layer using convolutional kernels to obtain the corresponding output features. The convolutional kernel is the basic unit in the convolutional layer. The shape of the convolutional kernel is determined by the number of input channels, the number of output channels, and the spatial dimension of the convolutional kernel. Optionally, the shape of the convolutional kernel can be (Cin, Cout, Kd, Kh, Kw), where Cin represents the number of input channels, Cout represents the number of output channels, Kd represents the size of the convolutional kernel in the time dimension, Kh represents the size of the convolutional kernel in the height dimension, and Kw represents the size of the convolutional kernel in the width dimension. The processing matrix can be understood as the parameter matrix associated with the convolutional kernel. Optionally, the processing matrix can be a matrix determined based on the number of input channels and the number of output channels of the convolutional kernel. The processing parameters can be used to determine how to extract input features and how to obtain output features. The target video can be a video corresponding to the video description text.
[0042] Specifically, in the case of a single-process model, the encoding module in the video generation model encodes the video description text and random noise separately, obtaining the processing results corresponding to the video description text and the random noise. The processing result corresponding to the video description text is then input into the decoding module for decoding, yielding the text features corresponding to the video description text.
[0043] The processing result corresponding to the random noise is input into the decoding module. In a scenario where the current process type is determined to be single-process, the target processing method corresponding to single-process is determined. Based on the target processing method and the number of input and output channels of the convolutional kernel in the convolutional layer, the processing matrix corresponding to the convolutional kernel is determined.
[0044] The input random noise is processed by convolution using the processing matrix of the convolution kernel in at least one convolutional layer to obtain output features corresponding to the random noise. Based on the output features corresponding to the random noise and the text features corresponding to the video description text, the target video corresponding to the video description text is determined.
[0045] The technical solution of this embodiment acquires video description text and random noise, and controls the video generation model to process the input video description text and random noise according to the target processing method corresponding to the current process type, so that the video generation model outputs a target video corresponding to the video description text. The video generation model includes an encoding module and a decoding module. The encoding module encodes the video description text and random noise, and inputs the processing result to the decoding module, so that the decoding module obtains the target video based on the processing result. The decoding module includes at least one convolutional layer, and the at least one convolutional layer includes a convolutional kernel and a processing matrix corresponding to the convolutional kernel. The processing matrix is constructed based on the processing parameters corresponding to the input and output channels of the convolutional kernel. The processing matrix can accelerate the convolutional processing of corresponding features and reduce the memory usage of the device hosting the video generation model. This invention solves the problem in the prior art where, when performing video generation processing based on existing video generation models, the device's memory is excessively occupied, resulting in limited model processing performance and low video generation efficiency and quality. By controlling the processing of the video generation model by the target processing method corresponding to the current process type, the system usage during model processing is reduced, and the efficiency and quality of video generation are improved.
[0046] Example 2
[0047] Figure 2 This is a flowchart of a video generation method provided in Embodiment 2 of the present invention. This embodiment refines the step of "controlling the video generation model to process the input video description text and random noise, so that the video generation model outputs a target video corresponding to the video description text" based on the above embodiments. For specific implementation details, please refer to the technical solution of this embodiment. Technical terms that are the same as or corresponding to those in the above embodiments will not be repeated here. Figure 2 As shown, the method includes:
[0048] S210, Obtain video description text and random noise.
[0049] S220. Based on the encoding module in the video generation model, random noise is encoded to obtain noise to be processed. Based on the encoding module, the video description text features are extracted to obtain the text features to be processed corresponding to the video description text.
[0050] Here, random noise can be understood as random Gaussian noise that follows a Gaussian distribution. The noise to be processed can be the noise features corresponding to the random noise obtained after encoding the random noise. The text features to be processed can be understood as the text features corresponding to the video description text obtained after feature extraction processing of the video description text.
[0051] Specifically, the random noise is encoded using the encoding module in the video generation model to obtain the noise to be processed corresponding to the random noise. The video description text is then processed using the encoding module in the video generation model to obtain the text features to be processed corresponding to the video description text. These noise and text features are then input into the decoding module of the video generation model to obtain the target video corresponding to the video description text.
[0052] S230. Based on the target processing method corresponding to the current process type, control the current convolutional layer in the decoding module of the video generation model to process the noise to be processed, obtain the noise output result, input the noise output result to the next convolutional layer, and use the next convolutional layer as the current convolutional layer, until the last convolutional layer outputs the target noise feature.
[0053] The noise output can be the noise features obtained after the current convolutional layer processes the noise to be processed. The target noise features are obtained by sequentially processing the noise to be processed through at least one convolutional layer. That is, the target noise features can be understood as the features output by all convolutional layers of the decoding module after sequentially processing the noise to be processed.
[0054] Specifically, when the current process type is single-process, the target processing method corresponding to single-process is determined. Using the target processing method, the number of input and output channels of the convolutional kernel in the current convolutional layer is processed to obtain a processing matrix. The processing matrix is then segmented to obtain at least two segmentation matrices. Based on the at least two segmentation matrices and the processing parameters corresponding to the spatial dimensions of the convolutional kernel, convolution processing is performed on the noise to be processed, yielding the noise output result. It should be noted that the processing parameters corresponding to the spatial dimensions of the convolutional kernel are processed in the normal manner for convolution processing of the noise to be processed. The noise output result is input into the next convolutional layer, and the next convolutional layer is used as the current convolutional layer. The noise output result is then processed using the processing matrix corresponding to the current convolutional layer to obtain the corresponding output features. The above process is repeated until the last convolutional layer outputs the target noise features.
[0055] In this embodiment of the invention, when the current process type is a single process, the target processing method is as follows: based on the number of input channels and the number of output channels of the convolution kernel in the current convolutional layer, determine the processing matrix of the convolution kernel, wherein the processing matrix is an m×n matrix, where m is the number of input channels and n is the number of output channels; divide the noise to be processed into sub-noise corresponding to the number of input channels, process the sub-noise based on the current convolutional layer, and determine the noise output result corresponding to the current convolutional layer based on the intermediate processing results and the processing matrix.
[0056] Here, sub-noise can be noise features obtained by segmenting the noise to be processed based on the number of input channels. Intermediate processing results can be obtained by processing the noise to be processed based on the processing matrix corresponding to the convolution kernel and other processing parameters. Noise output results can be noise features obtained by concatenating the sub-noise after processing the sub-noise according to the processing matrix and other processing parameters of the convolution kernel.
[0057] Specifically, in the case of a single-process operation, the shape of the convolutional kernel in the current convolutional layer is (Cin, Cout, Kd, Kh, Kw), where Cin represents the number of input channels, Cout represents the number of output channels, Kd represents the size of the kernel in the time dimension, Kh represents the size of the kernel in the height dimension, and Kw represents the size of the kernel in the width dimension. The processing matrix is determined based on the number of input and output channels of the kernel in the current convolutional layer. In this m×n processing matrix, row n represents the number of input channels, and column m represents the number of output channels.
[0058] The convolutional kernel slides along the time, height, and width dimensions. Let's illustrate this with the shape of the noise to be processed as (N, Cin, Din, Hin, Win), where N represents the number of samples, Cin represents the number of input channels, Din represents the magnitude of the noise in the time dimension, Hin represents the magnitude of the noise in the width dimension, and Win represents the magnitude of the noise in the height dimension. By sliding the convolutional kernel along these dimensions, the first part of the noise to be processed is obtained. The shape of this first part is (N, Cin, Kd, Kh, Kw). It should be noted that if Din = Kd, Hin = Kh, and Win = Kw, then the obtained first part of the noise is the noise to be processed. This first part of the noise is then segmented to obtain the corresponding sub-noise. Based on the shape of each sub-noise, the sub-noise matrix is determined.
[0059] The input and output channels in the processing matrix are segmented to obtain at least four segmentation matrices. Matrix multiplication is then performed on the sub-noise matrices using these four segmentation matrices to obtain at least four output features corresponding to the noise to be processed. These four output features are then concatenated according to a preset concatenation order to obtain intermediate processing results. The convolutional kernel continues to slide along the time, height, and width dimensions until it covers the entire data range of the noise to be processed in these dimensions, obtaining at least one intermediate processing result. Based on at least one intermediate processing result, the noise output result corresponding to the current convolutional layer is obtained.
[0060] For example, see Figure 3 , Figure 3 This is an example diagram of processing sub-noise based on the processing matrix. Figure 3 The rectangle on the left, containing X0 and X1, represents an N×Cin matrix determined by the number of samples N and the number of input channels Cin in the shape (N,Cin,Kd,Kh,Kw) of the noise to be processed. The rows of this matrix represent the number of samples N, and the columns represent the number of input channels Cin. The matrix corresponding to the noise to be processed is then divided into sub-noise matrices corresponding to the number of channels, namely the X0 matrix and the X1 matrix. Figure 3 The rectangle on the right containing W0, W1, W2, and W3 represents the processing matrix. The processing matrix is a Cin × Cout matrix, where Cin corresponds to m (mentioned above), representing the number of input channels, and Cout corresponds to n (mentioned above), representing the number of output channels. The processing matrix is divided into at least four partition matrices corresponding to the number of output and input channels. For example, Figure 3The processing matrix is divided according to the number of input channels (Cin) and output channels (Cout), resulting in four segmentation matrices: W0, W1, W2, and W3. The sub-noise matrices X0 and X1 corresponding to the sub-noise are then multiplied with the segmentation matrices W0, W1, W2, and W3 to obtain the multiplication results X0W0, X1W2, X0W1, and X1W3. These multiplication results are then subjected to feature concatenation to obtain the processing result corresponding to the noise portion to be processed (N, Cin, Kd, Kh, Kw), i.e., the intermediate processing result Y = Concat(X0W0 + X1W2, X0W1 + X1W3), where Concat represents feature concatenation. The shape of the intermediate processing result is (N, Cout, Kd, Kh, Kw). The convolutional kernel continues to slide along the time, height, and width dimensions until it covers the entire data range corresponding to the noise to be processed in these dimensions. During this sliding process, the steps described above for processing the noise portion of the processing matrix are repeated to obtain at least one intermediate processing result. Based on the intermediate processing results of at least one portion of the noise to be processed, the noise output result obtained after processing the noise by the current convolutional layer is obtained. The shape of the noise output result is (N, Cout, Dout, Hout, Wout), where N represents the number of samples, Cout represents the number of output channels, Dout represents the magnitude of the noise output result in the time dimension, Hout represents the magnitude of the noise output result in the width dimension, and Wout represents the magnitude of the noise output result in the height dimension. By dividing the matrix corresponding to the noise portion of the processing matrix and the processing matrix corresponding to the convolutional kernel into multiple sub-matrices and performing matrix calculations in blocks, the memory occupied by the intermediate processing results is reduced, avoiding memory overflow issues.
[0061] In this embodiment of the invention, when the current process type is a multi-process type, the target processing method is: according to the processing order of multiple convolutional layers, the column-segmented three-dimensional convolutional tensor and the row-segmented three-dimensional convolutional tensor are alternately processed, so as to process the noise to be processed based on the segmented convolutional layers and obtain the noise output result.
[0062] In multi-process scenarios, multiple GPUs may be used to process the convolutional operations of the video generation model in parallel. Sequential alternation can be understood as follows: if the current convolutional layer performs column-based 3D convolutional tensor processing, then the next convolutional layer performs row-based 3D convolutional tensor processing. Segmentation of 3D convolutional tensors is typically a data processing method used when performing convolution operations on noise. In 3D convolution, tensor segmentation usually does not refer to physically dividing the tensor into multiple parts, but rather to progressively processing different parts of the tensor corresponding to the noise during the convolution operation by sliding the convolution kernel. This segmentation method can be seen as an abstract description of the convolution operation.
[0063] The column-segmented 3D convolution tensor can be understood as using the number of output channels as the segmentation dimension in 3D convolution operations. The processing matrix corresponding to the 3D convolution kernel is then divided according to the number of output channels to obtain multiple partial matrices. Each partial matrix corresponds to an independent convolution task, which is processed by multiple GPUs.
[0064] Row-segmented 3D convolution tensors can be understood as using the number of input channels as a segmentation dimension in 3D convolution operations. The processing matrix corresponding to the 3D convolution kernel is divided according to the number of input channels to obtain multiple partial matrices. Each partial matrix corresponds to an independent convolution task, which is processed by multiple GPUs respectively.
[0065] Specifically, in the case of a multi-process scenario, where 3D convolution operations can be processed in parallel by multiple GPUs, the 3D convolution tensors are alternately split into columns and rows according to the sequence of multiple convolutional layers. That is, the current convolutional layer performs column splitting of the 3D convolution tensor, resulting in a column-splitting convolutional layer. The next convolutional layer then performs row splitting of the 3D convolution tensor, resulting in a row-splitting convolutional layer. This process continues until all convolutional layers have been split. The noise to be processed is then processed based on the split convolutional layers corresponding to each convolutional layer, yielding the noise output result. It should be noted that the split convolutional layers can be allocated to different GPUs for processing; that is, each GPU handles a portion of the convolutional layer's computational task. GPUs can exchange data and synchronize results via high-speed communication links to ensure the correctness of the noise output result.
[0066] Optionally, column-segmented 3D convolutional tensors and row-segmented 3D convolutional tensors are alternately performed according to the processing order of multiple convolutional layers, including: determining the column-segmented 3D convolutional tensor corresponding to the convolutional kernel in the current convolutional layer based on the number of multi-processes; and determining the row-segmented 3D convolutional tensor corresponding to the convolutional kernel in the next convolutional layer of the current convolutional layer, until the current convolutional layer is the last convolutional layer in the video generation model.
[0067] Here, the number of processes can be understood as the number of processes. For example, if the number of processes is 2, then the column-splitting 3D convolution tensor corresponding to the convolution kernel is also 2.
[0068] Specifically, based on the number of processes, the column-segmented 3D convolutional tensors corresponding to the convolutional kernels in the current convolutional layer are determined. This allows the noise to be processed to undergo 3D convolution with the column-segmented 3D convolutional tensors, yielding the noise output of the current convolutional layer. That is, the processing matrix corresponding to the convolutional kernels in the current convolutional layer is column-segmented according to the number of output channels, resulting in multiple partial matrices corresponding to the column segments. The number of partial matrices corresponding to the column segments is consistent with the number of processes. These partial matrices are allocated to the GPU corresponding to each process, enabling that process to perform 3D convolution on the noise based on the column-segmented partial matrices corresponding to the convolutional kernels, obtaining the output of that process. The outputs of multiple processes are used as the input features of the next convolutional layer, and 3D convolution is performed on the input features of the next convolutional layer and the row-segmented 3D convolutional tensors based on the process. That is, the processing matrix corresponding to the convolutional kernels in the next convolutional layer is row-segmented according to the number of input channels, resulting in multiple partial matrices corresponding to the row segments. The partial matrix corresponding to the row segmentation is assigned to the GPU corresponding to each process, so that the process performs 3D convolution operation on the corresponding input features based on the partial matrix of the row segmentation corresponding to the convolution kernel, and obtains the output result of the process. The output results of multiple processes are summed and communicated to obtain the noise output result corresponding to the next convolutional layer. The noise output result is used as the input feature of the next convolutional layer, and the above input feature and column segmentation 3D convolution tensor and row segmentation 3D convolution tensor are repeated to perform 3D convolution operation in an alternating order until the last convolutional layer of the video generation model completes the 3D convolution processing of the input features to obtain the target noise feature.
[0069] For example, since the processing order of multiple convolutional layers in the video generation model alternates between column-splitting and row-splitting of the 3D convolutional tensor, the explanation assumes the video generation model contains two convolutional layers. The column-splitting 3D convolutional tensor corresponding to the convolution kernel in the current convolutional layer, and the row-splitting 3D convolutional tensor corresponding to the next convolutional layer, are described below. In the case of a multi-process model with two processes, each process processes a portion of the convolutional operations of the current convolutional layer.
[0070] See the examples above. Figure 4 , Figure 4 Example diagram of 3D convolution processing for noise and column segmentation 3D convolution tensors. Figure 4The leftmost rectangle containing X represents the matrix X corresponding to the noise to be processed. The matrix of noise to be processed is an N×Cin matrix, where N represents the number of samples corresponding to the noise matrix and Cin represents the number of input channels. Figure 4 The two rectangles containing W0 and W1 represent the two partial matrices obtained by dividing the processing matrix corresponding to the convolution kernel of the current convolutional layer according to the number of output channels using the column-segmented 3D convolution tensor parallel algorithm: matrix W0 and matrix W1. The rows of matrix W0 represent the number of input channels, and the columns of matrix W1 represent the number of output channels. The sum of the number of columns in matrix W0 and matrix W1 represents the number of output channels. Figure 4 The rightmost matrix contains two matrices, Y0 and Y1, representing the matrices obtained after processing the noise from a partial matrix, i.e., the noise output of the current convolutional layer. The rows of matrix Y0 contain the number of samples N, and the rows of matrix Y1 contain the number of samples N. The sum of the number of columns in matrices Y0 and Y1 is the number of output channels.
[0071] In the dual-process scenario, the column-segmented 3D convolution tensor parallel algorithm partitions the processing matrix corresponding to the convolution kernel in the current convolutional layer according to the number of output channels, resulting in two partial matrices, W0 and W1. The current process performs matrix operations (3D convolution) on the X matrix corresponding to the input noise and the W0 matrix corresponding to the convolution kernel, yielding the output matrix Y0 for the current process. Similarly, the other process performs matrix operations on the X matrix corresponding to the input noise and the W1 matrix corresponding to the convolution kernel, yielding the output matrix Y1 for the other process. Based on this, during 3D convolution operations, the memory occupied by the intermediate partial matrices and the output of the current convolutional layer is reduced to half of the original amount, significantly decreasing memory consumption during convolution processing.
[0072] The next convolutional layer after the current convolutional layer is used as the current convolutional layer. See also Figure 5 , Figure 5 Example diagram of 3D convolution processing for output results and row-segmented 3D convolution tensors. Figure 5 The leftmost two rectangles, X0 and X1, correspond to the output matrices Y0 and Y1 of the previous convolutional layer, obtained from the two processes. That is, the X0 matrix corresponds to the output matrix Y0 of the previous convolutional layer, and the X1 matrix corresponds to the output matrix Y1 of the previous convolutional layer. The rows of the X0 and X1 matrices represent the number of samples N in the current convolutional layer, and the sum of the number of columns of the X0 and X1 matrices represents the number of input channels Cin of the current convolutional layer, which also corresponds to the number of output channels of the previous convolutional layer. Figure 5The two rectangles on the right, W0 and W1, represent the row-segmented 3D convolutional tensor parallel algorithm. This algorithm splits the processing matrix corresponding to the convolutional kernel in the current convolutional layer into two partial matrices, W0 and W1, based on the number of input channels. The columns of W0 and W1 represent the number of output channels (Cout) corresponding to the convolutional kernel in the current convolutional layer, and the sum of the number of rows in W0 and W1 represents the number of input channels (Cin) corresponding to the convolutional kernel in the current convolutional layer.
[0073] In the dual-process scenario, the row-segmented 3D convolution tensor parallel algorithm partitions the processing matrix corresponding to the convolution kernel in the current convolutional layer according to the number of input channels, resulting in two partial matrices, W0 and W1. A 3D convolution operation is then performed on the output matrix X0 of the previous convolutional layer and the W0 matrix corresponding to the convolution kernel of the current convolutional layer, yielding the output X0W0 for the current process. Similarly, a 3D convolution operation is performed on the output matrix X1 of the previous convolutional layer and the W1 matrix corresponding to the convolution kernel of the current convolutional layer, yielding the output X1W1 for the other process. AllReduce is then used to sum and communicate the outputs of the two processes, resulting in the noise output Y = AllReduce(X0W0, X1W1), which represents the target noise features based on the output of the last convolutional layer. It should be noted that AllReduce is a collective communication operation, primarily used in distributed computing environments, especially when handling large-scale data parallel tasks. By exchanging data between multiple GPUs and performing aggregation operations on the output results on all GPUs, such as summing, finding the maximum value, or finding the average value of features, each GPU ultimately obtains the same aggregation result, i.e., the target noise feature.
[0074] Optionally, after performing a three-dimensional convolution operation on the feature to be processed and a three-dimensional convolution tensor segmented by one column based on the process, the output is the concatenation result corresponding to the three-dimensional convolution tensor segmented by one column; based on the concatenation results of multiple processes, the noise output result is determined.
[0075] Here, the feature to be processed can be understood as the input feature of the current convolutional layer. Optionally, the feature to be processed can be the output feature of the previous convolutional layer, or it can be the noise to be processed from the input of the first convolutional layer. The concatenation result can be understood as the output result obtained after performing three-dimensional convolution processing on the feature to be processed based on the process.
[0076] Specifically, in a multi-process scenario, a 3D convolution operation is performed on the 3D convolution tensor corresponding to a column segment of the current process, based on the feature to be processed in the current process. This yields a concatenation result corresponding to the concatenated 3D convolution tensor of that column. The concatenation results from multiple processes are then concatenated to obtain the noise-reduced result.
[0077] For example, in conjunction with the above examples, see Figure 4 , Figure 4 The X matrix in the figure corresponds to the features to be processed mentioned above. Figure 4 The W0 or W1 matrix in the diagram corresponds to one of the columns of the 3D convolution tensor. Based on the current process, after performing 3D convolution processing using the feature matrix X to be processed and the 3D convolution tensor W0 to be split into one column, the concatenated result matrix Y0 is obtained. Correspondingly, based on another process, after performing 3D convolution processing using the feature matrix X to be processed and the 3D convolution tensor W1 to be split into one column, the concatenated result matrix Y1 is obtained. The two concatenated results are then concatenated to obtain the output of the current convolutional layer. Based on this, the computation and memory of the convolution processing are distributed across multiple processes, which not only improves the speed of convolution processing but also saves memory. For example, existing video generation models, limited to 80GB of memory per process, can only generate 720P videos. However, based on the target processing method controlled by the single-process type mentioned in this embodiment of the invention, 1080P videos can be generated. Under multi-process conditions, existing technologies can only process 720P videos, while the video generation model controlled by the target processing method corresponding to the multi-process type mentioned in the embodiments of this invention can generate 1080P videos, and the video generation speed is faster.
[0078] S240. Based on the target noise features and the target text features after decoding the text features to be processed, determine the target video corresponding to the video description text.
[0079] Among them, the target text features can be obtained by decoding the features to be processed by the decoding module of the video generation model, and are text features corresponding to the video description text.
[0080] Specifically, the text features to be processed are decoded to obtain the target text features. The noise to be processed is decoded by the decoding module to obtain the target noise features. Based on the target noise features and the target text features, the target video corresponding to the video description text is obtained.
[0081] The technical solution of this embodiment obtains video description text and random noise. The random noise is encoded using the encoding module in the video generation model to obtain noise to be processed. Furthermore, features of the video description text are extracted using the encoding module to obtain text features corresponding to the video description text. Based on the target processing method corresponding to the current process type, the current convolutional layer in the decoding module of the video generation model processes the noise to be processed, obtaining a noise output result. This noise output result is then input to the next convolutional layer, which is used as the current convolutional layer, until the last convolutional layer outputs the target noise features. Based on this, by using block convolution processing in a single process or distributing the computation and memory of convolution processing among various processes in a multi-process approach, not only can the speed of convolution processing be improved, but memory can also be saved. Based on the target noise features and the target text features after decoding the text features to be processed, the target video corresponding to the video description text is determined. This invention solves the problem in the prior art where, when performing video generation processing based on existing video generation models, the memory of the device deploying the existing video generation model is excessively occupied, resulting in limited model processing performance and low video generation efficiency and quality. By controlling the processing of the video generation model through the target processing method corresponding to the current process type, the system usage during model processing is reduced, thereby improving the efficiency and quality of video generation.
[0082] Example 3
[0083] Figure 6 This is a flowchart illustrating a video generation model training method provided in Embodiment 3 of the present invention. This embodiment, based on the above embodiments, allows for the prior training of a pre-trained video generation model before video generation processing. Specific implementation details can be found in the technical solution of this embodiment. Technical terms identical or corresponding to those in the above embodiments will not be repeated here. Figure 6 As shown, the method includes:
[0084] S310. Obtain multiple training samples, wherein the training samples include: sample videos and sample video description text corresponding to the sample videos.
[0085] The sample video can be a video used to train the video generation model. The sample video description text can be used to summarize or explain the sample video in detail. The content of the video to be generated can be determined through the sample video description text.
[0086] Specifically, before training the video generation model, multiple training samples can be obtained to train the model. To improve the accuracy of the video generation model, as many and varied training samples as possible can be obtained. This includes acquiring multiple sample videos and corresponding sample video description text.
[0087] S320. Input the training samples into the video generation model to be trained, and use the encoding module based on the video generation model to add noise to the sample video to obtain the noise of the sample to be processed. Also, use the encoding module to extract features from the description text of the sample video to obtain the sample text features corresponding to the description text of the sample video.
[0088] Here, the noise in the sample to be processed can be understood as the noise obtained by adding noise to the sample video. The sample text features can be the text features obtained by extracting features from the descriptive text of the sample video.
[0089] Specifically, training samples are input into the video generation model to be trained. The encoding module of the video generation model adds noise to the sample videos to obtain the noise of the processed samples. The encoding module then performs feature extraction on the descriptive text of the sample videos to obtain the sample text features corresponding to the descriptive text of the sample videos.
[0090] S330. Based on the target processing method corresponding to the current process type, control the decoding module of the video generation model to be trained to decode the noise and text features of the sample to be processed, respectively, to obtain the output noise features and output text features. Determine the output video based on the output text features and output noise features.
[0091] Specifically, the output noise features can be obtained by denoising the sample noise using the convolutional layers of the decoding module. The output text features can be obtained by decoding the sample text features. The output video can be understood as the generated video finally output by the video generation model to be trained.
[0092] Specifically, when the current process type is single-process, the target processing method corresponding to single-process is determined. Based on the target processing method, the decoding module of the video generation model to be trained is controlled to decode the noise and text features of the sample to be processed, respectively, to obtain the output noise features and output text features. Based on the output text features and output noise features, the output video is determined.
[0093] S340. Perform loss processing on each video frame in the output video and each video frame in the sample video to determine the loss value.
[0094] The loss value can be determined based on the degree of difference between the video frames of the sample video and the video frames of the output video.
[0095] Specifically, the loss value of the video generation model to be trained is determined based on the degree of difference between the video frames of the sample video and the video frames of the output video, and the model parameters of the video generation model to be trained are corrected based on the loss value.
[0096] S350. Based on the loss value, adjust the model parameters of the video generation model to be trained to obtain the trained video generation model.
[0097] Specifically, during the process of correcting the model parameters of the video generation model to be trained using the loss value, the convergence of the loss function can be used as a training objective. This includes checking if the training error is less than a preset error, if the error change tends to stabilize, or if the current number of iterations equals a preset number. If the convergence condition is met, such as the training error of the loss function being less than the preset error, or the error change trend tending to stabilize, it indicates that the video generation model to be trained has completed training, and iterative training can be stopped. If the convergence condition has not been met, other training samples can be obtained to continue training the video generation model until the training error of the loss function is within a preset range. When the training error of the loss function converges, the trained video generation model is obtained. That is, when the video description text and random noise are input into this video generation model, the target video can be accurately obtained.
[0098] The technical solution of this embodiment involves acquiring multiple training samples and inputting them into a video generation model to be trained. The encoding module of the video generation model adds noise to the sample videos to obtain noise samples. The encoding module also extracts features from the descriptive text of the sample videos to obtain sample text features corresponding to the descriptive text. Based on the target processing method corresponding to the current process type, the decoding module of the video generation model to be trained decodes the noise and text features of the sample videos to be processed, respectively, to obtain output noise features and output text features. The output video is determined based on the output text features and output noise features. Loss processing is performed on each video frame in the output video and each video frame in the sample video to determine the loss value. The model parameters of the video generation model to be trained are then corrected based on the loss value to obtain a trained video generation model. Through the above model training process, the accuracy and stability of the video generation model are improved.
[0099] Example 4
[0100] Figure 7 This is a schematic diagram of the structure of a video generation device provided in Embodiment 4 of the present invention. Figure 7 As shown, the device includes a text and noise acquisition module 410 and a target video determination module 420.
[0101] The text and noise acquisition module 410 is used to acquire video description text and random noise; the target video determination module 420 is used to control the video generation model to process the input video description text and random noise based on the target processing method corresponding to the current process type, so that the video generation model outputs a target video corresponding to the video description text; wherein, the video generation model includes an encoding module and a decoding module. The encoding module is used to encode the video description text and random noise and input the processing result to the decoding module. The decoding module includes at least one convolutional layer. The at least one convolutional layer includes a convolutional kernel and a processing matrix corresponding to the convolutional kernel. The processing matrix is constructed based on the processing parameters corresponding to the input and output channels of the convolutional kernel.
[0102] The technical solution of this embodiment acquires video description text and random noise, and controls the video generation model to process the input video description text and random noise according to the target processing method corresponding to the current process type, so that the video generation model outputs a target video corresponding to the video description text. The video generation model includes an encoding module and a decoding module. The encoding module encodes the video description text and random noise, and inputs the processing result to the decoding module, so that the decoding module obtains the target video based on the processing result. The decoding module includes at least one convolutional layer, and the at least one convolutional layer includes a convolutional kernel and a processing matrix corresponding to the convolutional kernel. The processing matrix is constructed based on the processing parameters corresponding to the input and output channels of the convolutional kernel. The processing matrix can accelerate the convolutional processing of corresponding features and reduce the memory usage of the device hosting the video generation model. This invention solves the problem in the prior art where, when performing video generation processing based on existing video generation models, the device's memory is excessively occupied, resulting in limited model processing performance and low video generation efficiency and quality. By controlling the processing of the video generation model by the target processing method corresponding to the current process type, the system usage during model processing is reduced, and the efficiency and quality of video generation are improved.
[0103] Based on the above embodiments, optionally, the target video determination module includes: an encoding module processing unit, used to encode random noise based on the encoding module in the video generation model to obtain noise to be processed, and to extract video description text features based on the encoding module to obtain text features to be processed corresponding to the video description text; a noise feature output unit, used to control the current convolutional layer in the decoding module of the video generation model to process the noise to be processed based on the target processing method corresponding to the current process type, to obtain a noise output result, and input the noise output result to the next convolutional layer, and use the next convolutional layer as the current convolutional layer, until the last convolutional layer outputs the target noise features; and a target video generation unit, used to determine the target video corresponding to the video description text based on the target noise features and the target text features after decoding the text features to be processed.
[0104] Optionally, the noise feature output unit includes: a single-process output result determination subunit, used to determine the processing matrix of the convolution kernel based on the number of input channels and the number of output channels of the convolution kernel in the current convolutional layer when the current process type is single-process type, wherein the processing matrix is an m×n matrix, m is the number of input channels and n is the number of output channels; divide the noise to be processed into sub-noise corresponding to the number of input channels, process the sub-noise based on the current convolutional layer, and determine the noise output result corresponding to the current convolutional layer based on the intermediate processing results and the processing matrix.
[0105] Optionally, the noise feature output unit includes a multi-process output result determination subunit, which is used to alternately perform column-segmented three-dimensional convolutional tensors and row-segmented three-dimensional convolutional tensors according to the processing order of multiple convolutional layers, so as to process the noise to be processed based on the segmented convolutional layers and obtain the noise output result.
[0106] Optionally, a multi-process output result determination subunit is used to determine the column-segmented 3D convolutional tensor corresponding to the convolutional kernel in the current convolutional layer based on the number of multi-processes; for the next convolutional layer of the current convolutional layer, the row-segmented 3D convolutional tensor corresponding to the convolutional kernel in the next convolutional layer is determined, until the current convolutional layer is the last convolutional layer in the video generation model.
[0107] Optionally, a multi-process output result determination subunit is used to perform three-dimensional convolution operations on the features to be processed and one column-segmented three-dimensional convolution tensor based on the process, and output the concatenation result corresponding to one column-segmented three-dimensional convolution tensor; and to determine the noise output result based on the concatenation results of multiple processes.
[0108] The video generation apparatus provided in this embodiment of the invention can execute the video generation method provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects of the method.
[0109] Example 5
[0110] Figure 8 This is a schematic diagram of the structure of an electronic device provided in Embodiment 5 of the present invention. The electronic device 10 is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (such as helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0111] like Figure 8 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 may also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0112] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0113] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as video generation methods.
[0114] In some embodiments, the video generation method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the video generation method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the video generation method by any other suitable means (e.g., by means of firmware).
[0115] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0116] Computer programs for implementing the video generation method of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The computer programs can be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0117] Example 6
[0118] Embodiment 6 of the present invention also provides a computer-readable storage medium storing computer instructions for causing a processor to execute a video generation method, the method comprising:
[0119] The process involves acquiring video description text and random noise; based on the target processing method corresponding to the current process type, controlling the video generation model to process the input video description text and random noise, so that the video generation model outputs a target video corresponding to the video description text; wherein, the video generation model includes an encoding module and a decoding module, the encoding module is used to encode the video description text and random noise, and input the processing result into the decoding module, the decoding module includes at least one convolutional layer, the at least one convolutional layer includes a convolutional kernel and a processing matrix corresponding to the convolutional kernel, the processing matrix is constructed based on the processing parameters corresponding to the input and output channels of the convolutional kernel.
[0120] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0121] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0122] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0123] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.
[0124] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.
[0125] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. A video generation method, characterized in that, The method includes: Obtain the video description text and random noise; Based on the target processing method corresponding to the current process type, the video generation model is controlled to process the input video description text and the random noise, so that the video generation model outputs a target video corresponding to the video description text; The video generation model includes an encoding module and a decoding module. The encoding module encodes and processes video description text and random noise, and inputs the processing result into the decoding module. The decoding module includes at least one convolutional layer, which includes a convolutional kernel and a processing matrix corresponding to the convolutional kernel. The processing matrix is constructed based on the processing parameters corresponding to the input and output channels of the convolutional kernel. The video generation model processes the input video description text and random noise to output a target video corresponding to the video description text, including: The random noise is encoded and processed by the encoding module in the video generation model to obtain noise to be processed; and the video description text features are extracted by the encoding module to obtain text features to be processed corresponding to the video description text. Based on the target processing method corresponding to the current process type, the current convolutional layer in the decoding module of the video generation model is controlled to process the noise to be processed, obtain the noise output result, and input the noise output result into the next convolutional layer, and use the next convolutional layer as the current convolutional layer, until the last convolutional layer outputs the target noise feature; Based on the target noise features and the target text features after decoding the text features to be processed, a target video corresponding to the video description text is determined; The step of controlling the current convolutional layer in the decoding module of the video generation model to process the noise to be processed, based on the target processing method corresponding to the current process type, to obtain a noise output result, includes: When the current process type is single-process, the target processing method is as follows: Based on the number of input channels and the number of output channels of the convolution kernel in the current convolutional layer, the processing matrix of the convolution kernel is determined, wherein the processing matrix is an m×n matrix, where m is the number of input channels and n is the number of output channels; The noise to be processed is divided into sub-noises corresponding to the number of input channels. The sub-noises are processed based on the current convolutional layer. Based on the intermediate processing results and the processing matrix, the noise output result corresponding to the current convolutional layer is determined. Wherein, if the current process type is a multi-process type, the target processing method is as follows: The three-dimensional convolutional tensor is alternately segmented into columns and rows according to the processing order of multiple convolutional layers. The noise to be processed is then processed based on the segmented convolutional layers to obtain the noise output result.
2. The method according to claim 1, characterized in that, The method of alternately splitting the 3D convolutional tensor into columns and rows according to the processing order of multiple convolutional layers includes: Based on the number of processes, determine the column-segmented 3D convolution tensor corresponding to the convolution kernel in the current convolutional layer; For the next convolutional layer of the current convolutional layer, determine the row-segmented 3D convolutional tensor corresponding to the convolutional kernel in the next convolutional layer, until the current convolutional layer is the last convolutional layer in the video generation model.
3. The method according to claim 2, characterized in that, The current convolutional layer in the decoding module of the video generation model processes the noise to be processed, and obtains the noise output result, including: After performing a three-dimensional convolution operation on the feature to be processed and one of the column-segmented three-dimensional convolution tensors, the output is the concatenation result corresponding to one of the column-segmented three-dimensional convolution tensors. The noise output result is determined based on the splicing results of multiple processes.
4. A video generation device, characterized in that, include: The text and noise acquisition module is used to acquire video description text and random noise; The target video determination module is used to control the video generation model to process the input video description text and the random noise based on the target processing method corresponding to the current process type, so that the video generation model outputs a target video corresponding to the video description text; The video generation model includes an encoding module and a decoding module. The encoding module encodes and processes video description text and random noise, and inputs the processing result into the decoding module. The decoding module includes at least one convolutional layer, which includes a convolutional kernel and a processing matrix corresponding to the convolutional kernel. The processing matrix is constructed based on the processing parameters corresponding to the input and output channels of the convolutional kernel. The target video determination module includes: an encoding module processing unit, used to encode random noise based on the encoding module in the video generation model to obtain noise to be processed, and to extract video description text features based on the encoding module to obtain text features to be processed corresponding to the video description text; The noise feature output unit is used to control the current convolutional layer in the decoding module of the video generation model to process the noise to be processed based on the target processing method corresponding to the current process type, to obtain the noise output result, and input the noise output result into the next convolutional layer, and use the next convolutional layer as the current convolutional layer, until the last convolutional layer outputs the target noise feature; The target video generation unit is used to determine the target video corresponding to the video description text based on the target noise features and the target text features after decoding the text features to be processed. The noise feature output unit includes a single-process output result determination subunit, which is used to determine the processing matrix of the convolution kernel based on the number of input channels and the number of output channels of the convolution kernel in the current convolutional layer when the current process type is single-process type. The processing matrix is an m×n matrix, where m is the number of input channels and n is the number of output channels. The noise to be processed is divided into sub-noises corresponding to the number of input channels. The sub-noises are processed separately based on the current convolutional layer. Based on the intermediate processing results and the processing matrix, the noise output result corresponding to the current convolutional layer is determined. The noise feature output unit further includes a multi-process output result determination subunit, which is used to alternately perform column-segmented three-dimensional convolutional tensors and row-segmented three-dimensional convolutional tensors according to the processing order of multiple convolutional layers, so as to process the noise to be processed based on the segmented convolutional layers and obtain the noise output result.
5. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the video generation method according to any one of claims 1-3.
6. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the video generation method according to any one of claims 1-3.
7. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the video generation method as described in any one of claims 1-3.
Citation Information
Patent Citations
Image generation method and device, electronic equipment, storage medium and program product
CN117437317A
Image generation method and electronic equipment
CN117611700A