A Video Spatial Expansion Method, Device, Equipment and Storage Medium
Through the methods of video boundary detection, generative model outscaling, video super-segment and fusion model fusion, the problems of poor dynamic video space expansion effect and large resource consumption in the prior art are solved, and the video space expansion effect with high quality and low resource consumption is achieved.
Patent Information
- Application Number
- CN202411701569.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-26
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2044-11-26
AI Technical Summary
The existing video space expansion method has poor results when processing dynamic videos and consumes too much resources, making it difficult to achieve high-quality and low-resource expansion of video space.
The original video is segmented through the video boundary detection model, and the video is out-scaled by downsampling using the generative model, extracted and super-scored by the video super-scored model. Finally, the video fusion model is used for gradient information matching and fusion to achieve high-quality video space expansion.
This method can effectively expand the video space, maintain high-quality pictures, reduce resource consumption, and achieve seamless integration of video content.
Smart Images

Figure CN119206422B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of video generation technology. Specifically, it relates to a video space expansion method, apparatus, device, and storage medium. Background Art
[0002] Video Outpainting is a technology based on deep learning and computer vision, aiming to generate additional regions from existing video content to expand it spatially, thereby obtaining a larger viewing angle or filling in missing parts in video frames. This technology is similar to Image Outpainting and is commonly used to enhance the richness of video content or adapt to different display devices. Video Outpainting is more complex than Image Outpainting because video has continuous frames and it is necessary to ensure the continuity and consistency of the expanded video over time.
[0003] Current video space expansion usually relies on manual editing or rule-based methods, such as interpolation techniques, which perform linear interpolation on missing or damaged video frames or temporal interpolation based on neighboring frames, or completion based on texture synthesis, manually or semi-automatically filling in missing regions by copying textures from existing images. Although these methods can handle static scenes, they often perform poorly for dynamic scenes, complex textures, and object movements and cannot process videos with rich details.
[0004] Video generation technologies based on denoising models have made significant progress in recent years and can generate high-quality videos using text or picture inputs. However, when these technologies are applied to video space expansion, they face some challenges, such as high computational resource requirements, high processing complexity, and difficulty in processing dynamic content. Therefore, current video space expansion methods based on video generation models perform poorly when processing dynamic videos and consume excessive resources. But there are still many technical challenges to achieve high-quality and low-resource-consuming video space expansion. Summary of the Invention
[0005] The purpose of this application is to provide a video space expansion method, apparatus, device, and storage medium to overcome the existing technical defects. By adopting a video generation model and an expansion strategy through video downsampling, it solves the problems that current video space expansion methods perform very poorly in dynamic videos and consume excessive resources.
[0006] The purpose of this application is achieved through the following technical solutions:
[0007] In a first aspect, this application proposes a video space expansion method, and the method includes:
[0008] The original video is segmented into multiple video segments based on the video boundary time points predicted by the video boundary detection model;
[0009] After downsampling the video segments, a generative model is used for video extrapolation to obtain a low-resolution extrapolated video;
[0010] The extrapolated part is extracted from the low-resolution extrapolated video and super-resolved by a video super-resolution model to obtain a high-resolution video of the extrapolated part;
[0011] The video fusion model is used to calculate and match the gradient information between the video segments and the high-resolution video of the extrapolated part, and the video segments are extended into the high-resolution video of the extrapolated part based on the gradient information to obtain a fused high-resolution extrapolated video.
[0012] In a possible implementation, the video boundary detection model includes a video feature extraction backbone, a feature similarity calculation module, and multiple classification heads;
[0013] The video feature extraction backbone includes a 3D convolutional module and a Transformers module for calculating the features of video frames;
[0014] The feature similarity calculation module calculates the feature similarity vector of adjacent frames and the RGB histogram similarity feature vector of adjacent frames;
[0015] Each classification head consists of multiple fully connected layers for predicting the probability that a video frame is a video boundary.
[0016] In a possible implementation, the generative model includes a video description model and a video extrapolation model. The step of using the generative model for video extrapolation after downsampling the video segments to obtain a low-resolution extrapolated video includes:
[0017] The video segments are downsampled to obtain low-resolution original video segments;
[0018] The video description model is used to obtain the corresponding text description for the low-resolution original video segments;
[0019] The video extrapolation model is used to perform video extrapolation in combination with the low-resolution original video segments and the corresponding text description to obtain a low-resolution extrapolated video.
[0020] In a possible implementation, the video extrapolation model includes a text feature extractor, a video feature extractor, a video encoder, a video decoder, and a video denoising model. The step of using the video extrapolation model to perform video extrapolation in combination with the text description and the low-resolution original video segments to obtain a low-resolution extrapolated video includes:
[0021] Use a text feature extractor to extract features from the text description to obtain text features;
[0022] Use a video feature extractor to process each frame of the low-resolution original video clip to obtain video features;
[0023] Initialize a video padding mask for the low-resolution original video clip with the region pixel value set to 1 and the outer expanded region pixel value set to 0;
[0024] Based on the video padding mask, pad the low-resolution original video clip according to the target size to be expanded, and then use a video encoder to process the padded video clip to obtain the first latent space feature;
[0025] Randomly initialize the target size to be expanded with Gaussian noise to obtain the second latent space feature;
[0026] Input the text features, video features, video padding mask, first latent space feature, and second latent space feature into a video denoising model for prediction to obtain noise;
[0027] Use a forward sampler to subtract the noise from the second latent space feature to obtain the third latent space feature, and input the third latent space feature, text features, video features, video padding mask, and first latent space feature into the video denoising model again for noise prediction and iterative loop to obtain the denoising result;
[0028] Use a video decoder to process the denoising result to obtain a low-resolution padded video.
[0029] In one possible implementation, the video denoising model adopts a U-shaped network structure, including a three-dimensional convolution module, a spatial attention module, a text attention module, and a layout attention module.
[0030] In one possible implementation, the step of using a video fusion model to calculate and match the gradient information between the video clip and the high-resolution video of the expanded part, and based on the gradient information, expand the video clip into the high-resolution video of the expanded part to obtain the fused high-resolution expanded video, includes:
[0031] Calculate the gradient field of the video clip frame by frame through the Sobel operator, and calculate the gradient divergence of the video clip using the Laplace operator;
[0032] Based on the gradient field and gradient divergence, use a discrete solution equation to calculate the edge pixel values of the high-resolution video of the expanded part;
[0033] Insert the edge pixel values into the high-resolution video of the expanded part to obtain the fused high-resolution expanded video.
[0034] In one possible implementation, the expression of the discrete solution equation is: + + , where is the abscissa of the pixel, is the ordinate of the pixel, is the coordinate after solution at the pixel value, is at the pixel value, is at the pixel value, is at the pixel value, is at the pixel value, is the gradient at in the video segment.
[0035] In a second aspect, the present application proposes a video spatial expansion device, and the device includes:
[0036] A video splitting module, configured to split the original video through the video boundary time points predicted by the video boundary detection model to obtain a plurality of video segments;
[0037] A video external expansion module, configured to perform video external expansion on the downsampled video segment by using a generative model to obtain a low-resolution externally expanded video;
[0038] A video super-resolution module, configured to extract the externally expanded part from the low-resolution externally expanded video and perform super-resolution processing on the externally expanded part through a video super-resolution model to obtain a high-resolution video of the externally expanded part;
[0039] A video fusion module, configured to use a video fusion model to calculate and match the gradient information between the video segment and the high-resolution video of the externally expanded part, and based on the gradient information, expand the video segment into the high-resolution video of the externally expanded part to obtain a fused high-resolution externally expanded video.
[0040] In a third aspect, the present application further proposes a computer device, and the computer device includes a processor and a memory. A computer program is stored in the memory, and the computer program is loaded and executed by the processor to implement the video spatial expansion method according to any one of the first aspects.
[0041] In a fourth aspect, the present application further proposes a computer-readable storage medium. A computer program is stored in the storage medium, and the computer program is loaded and executed by a processor to implement the video spatial expansion method according to any one of the first aspects.
[0042] The main solution of the present application and its various further alternative solutions can be freely combined to form multiple solutions, all of which are solutions that can be adopted and claimed in the present application; moreover, in the present application, (each non-conflicting alternative) alternatives can be freely combined with each other and with other alternatives. Those skilled in the art can understand that there are various combinations according to the prior art and common general knowledge after understanding the solutions of the present application, and all of them are the technical solutions to be protected by the present application, which will not be enumerated here.
[0043] The present application discloses a video spatial expansion method, device, equipment and storage medium, which relates to the technical field of video generation. First, the original video is segmented into multiple video segments through the video boundary time points predicted by the video boundary detection model. Secondly, after downsampling, the generative model is used for video expansion to obtain a low-resolution expanded video. The expanded part is extracted and the expanded part is super-resolved through the video super-resolution model to obtain a high-resolution video of the expanded part. The video fusion model is used to calculate and match the gradient information between the video segment and the high-resolution video of the expanded part, and the video segment is extended into the high-resolution video of the expanded part to obtain a fused high-resolution expanded video. The original video is seamlessly replaced into the expanded video through the video fusion method, ensuring high-quality spatial expansion, maintaining the content of the original video, utilizing the creativity of the video generation model and greatly reducing resource consumption. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required to be used in the embodiments. It should be understood that the following drawings only show some embodiments of the present application, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0045] Figure 1 FIG. shows a schematic flow chart of the video spatial expansion method proposed in the embodiment of the present application.
[0046] Figure 2 FIG. shows a schematic diagram of an example result of video super-resolution generation proposed in the embodiment of the present application.
[0047] Figure 3 FIG. shows a schematic diagram of an example result of first-frame image generation proposed in the embodiment of the present application.
[0048] Figure 4 FIG. shows a schematic diagram of an example result of video fusion generation proposed in the embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0049] The following specific examples illustrate the implementation manners of the present application. Those skilled in the art can easily understand other advantages and effects of the present application from the content disclosed in this specification. The present application can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present application. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other.
[0050] Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of the present application.
[0051] In the prior art, video generation technologies based on denoising models have made significant progress in recent years, and high-quality videos can be generated using text or picture inputs. However, when these technologies are applied to the spatial expansion of videos, they will face some challenges, such as high computational resource requirements, high processing complexity, and great difficulty in processing dynamic content. Therefore, the current video spatial expansion methods based on video generation models have poor effects when processing dynamic videos and consume excessive resources. However, there are still many technical challenges in achieving high-quality and low-resource-consumption video spatial expansion.
[0052] Therefore, to solve the above technical problems, the embodiments of the present application propose a video spatial expansion method, device, equipment, and storage medium. This solution uses a video shot segmentation method to divide a video into multiple video segments of different scenes, downsamples the video segments, and uses a generative video expansion method to achieve the spatial expansion of low-resolution videos. Then, a video super-resolution algorithm is used to improve the resolution of the expanded part of the video. Finally, a video fusion algorithm is used to replace the original video into the expanded video to obtain a high-quality video of the target size. High-quality expanded videos are obtained through low-resolution video expansion and video super-resolution, and the original video is seamlessly replaced into the expanded video through the video fusion method, which not only ensures high-quality spatial expansion but also maintains the content of the original video, fully utilizes the creativity of the video generation model, and greatly reduces resource consumption. The following will be a detailed description of it.
[0053] Please refer to Figure 1 , Figure 1 which shows a schematic flowchart of the video spatial expansion method proposed by the embodiments of the present application. The method includes:
[0054] Step S1: Segment the original video according to the video boundary time points predicted by the video boundary detection model to obtain multiple video segments.
[0055] The video boundary detection model includes a video feature extraction backbone, a feature similarity calculation module, and multiple classification heads;
[0056] The video feature extraction backbone includes a 3D convolutional module and a Transformers module for calculating the features of video frames;
[0057] The feature similarity calculation module calculates the feature similarity vector of adjacent frames and the RGB histogram similarity feature vector of adjacent frames;
[0058] Each classification head consists of multiple fully connected layers for predicting the probability that a video frame is a video boundary.
[0059] In step S1, first, a 3D convolutional module plus a Transformers module are used to extract the feature information of each frame from the original video. The 3D convolution can capture the changes in the time dimension (the relationship between consecutive frames), and the Transformers module is used to understand the global context information. Then, for two adjacent frames, the similarity between them based on the extracted feature vectors is calculated; at the same time, the difference in the RGB color histograms between adjacent frames is considered as supplementary features. Then, a classification head composed of multiple fully connected layers is used to analyze each pair of adjacent frames, and the possibility score of the existence of a boundary at this position is output. If the score at a certain position is high, it is considered that this is very likely to be a scene transition point or boundary. Finally, the true boundary time points are determined according to the probability scores obtained in the previous step, and the entire video is divided into several independent scene segments accordingly.
[0060] In step S2, after downsampling the video segment, a generative model is used to perform video upscaling to obtain a low-resolution upscaled video.
[0061] To reduce the data volume of the video and lower the computational complexity of subsequent processing, the original video frames are downsampled to reduce the size of each frame image, which is achieved by simply discarding some pixels or applying some form of filter (such as bilinear interpolation, Gaussian filtering, etc.) to achieve smooth transition.
[0062] To recover a high-resolution version from the low-resolution video and improve the spatial detail expressiveness of the video, a generative model (a deep learning model, such as a convolutional neural network CNN or a more advanced architecture such as GANs) is used to process the downsampled video segment. The model learns the mapping relationship between a large number of high-low resolution image pairs and predicts and generates the missing high-frequency information while maintaining the visual quality.
[0063] Step S2 includes:
[0064] Step S21: Downsample the video segment to obtain a low-resolution original video segment;
[0065] Step S22: Use a video description model to obtain the corresponding text description for the low-resolution original video segment;
[0066] Step S23: Use the video upscaling model to combine the low-resolution original video clip and the corresponding text description to perform video upscaling to obtain a low-resolution upscaled video.
[0067] The generative model includes a video description model and a video upscaling model.
[0068] To reduce the resolution of the video to reduce the data volume, video clip downsampling is adopted. Perform downsampling operations on the original video clip, and reduce the size of each frame image through methods such as average pooling and bilinear interpolation. For example, convert a high-definition video into a standard-definition or lower-resolution version.
[0069] After that, use the video description model to process the low-resolution original video clip to generate text descriptions corresponding to each video clip. These descriptions contain the main content or key information of the scene. Among them, a pre-trained video description model is adopted. Input the low-resolution video clip, and the output is a natural language description of the video clip.
[0070] Combine the text description and the low-resolution video clip to generate an upscaled video, and use additional context information (text description) to improve the quality of video upscaling. Take the low-resolution video clip and the corresponding text description as inputs, and use a deep learning model specifically designed for video upscaling as the video upscaling model. This model should not only be able to map images from low resolution to high resolution, but also be able to understand and integrate the given text information. The final output is a new version of the video with a higher resolution but still maintaining the main content and style of the original video.
[0071] Step S23 includes:
[0072] Use a text feature extractor to extract features from the text description to obtain text features;
[0073] Use a video feature extractor to process each frame of the low-resolution original video clip to obtain video features;
[0074] Initialize a video upscaling mask with the regional pixel values of the low-resolution original video clip being 1 and the upscaled regional pixel values being 0;
[0075] Based on the video upscaling mask, fill the low-resolution original video clip according to the target size to be expanded, and then use a video encoder to process the filled video clip to obtain the first latent space feature;
[0076] Randomly initialize the target size to be expanded with Gaussian noise to obtain the second latent space feature;
[0077] Input the text features, video features, video upscaling mask, first latent space feature, and second latent space feature into a video denoising model for prediction to obtain noise;
[0078] Use a forward sampler to subtract noise from the second latent space feature to obtain a third latent space feature, and input the third latent space feature, text feature, video feature, video padding mask, and first latent space feature into the video denoising model again for noise prediction cyclic iteration to obtain a denoising result;
[0079] Use a video decoder to process the denoising result to obtain a low-resolution padded video.
[0080] The video padding model includes a text feature extractor, a video feature extractor, a video encoder, a video decoder, and a video denoising model. The video denoising model adopts a U-shaped network structure, including a three-dimensional convolution module, a spatial attention module, a text attention module, and a layout attention module.
[0081] Figure 3 Shows a schematic diagram of an example result of the first-frame image generation proposed in the embodiment of the present application. Use a text feature extractor to analyze the provided text description and extract relevant feature vectors that can represent the meaning or context of the text. Then, for the input low-resolution original video segment, use a video feature extractor to process each frame to identify and extract important visual elements and structural features in each frame.
[0082] Create a mask (video padding mask) for the original video, where the pixel values in the original video area are set to 1, and the new area to be expanded is set to 0. According to the target size requirement, fill the original video to the corresponding size and convert it into a first latent space feature through a video encoder. Additionally, based on the same final size specification, randomly generate another latent representation as an initial guess (second latent space feature) using a Gaussian distribution.
[0083] Input all the prepared data (text feature, video feature, video padding mask, first latent space feature, and second latent space feature) into a model specifically designed for predicting noise. Using the forward sampling method, subtract the predicted noise part from the second latent space feature to obtain a new latent representation (third latent space feature). Then, input the updated third latent space feature together with other unchanged information back into the denoising model, and repeat the noise prediction and removal operations until an ideal denoising effect is achieved. Finally, use a video decoder to convert the data after multiple rounds of optimization processing back to the normal video format to obtain an enhanced video with a higher resolution than the original version.
[0084] Step S3: Extract the padded part from the low-resolution padded video and perform super-resolution processing on the padded part through a video super-resolution model to obtain a high-resolution video of the padded part.
[0085] Using the created video expansion mask, the expanded area is extracted from the low-resolution expanded video. This mask marks the original video area as 1 and the expanded area as 0. Therefore, the expanded part can be accurately located and extracted through the mask. Then, the extracted expanded part is input into a pre-trained video super-resolution model. The task of the video super-resolution model is to predict and generate high-resolution images by learning the mapping relationship between low-resolution images and high-resolution images. The model processes each frame of the expanded part, and after each frame is processed by the model, a high-resolution version is generated. Finally, these high-resolution frames are combined to form a high-resolution video of the expanded part.
[0086] Step S4: Use the video fusion model to calculate and match the gradient information between the video clip and the high-resolution video of the expanded part, and based on the gradient information, expand the video clip into the high-resolution video of the expanded part to obtain the fused high-resolution expanded video.
[0087] Figure 4 The figure shows a schematic diagram of the example result of video fusion generation proposed in the embodiment of the present application. The video fusion model is used to calculate the gradient information between the original video clip and the high-resolution video of the expanded part, which reflects the rate of change of the image or video frame in space and is particularly important for details such as edges and textures. Then, by matching the gradient information of these two parts, the visual consistency and continuity between the original video and the expanded part can be ensured. According to the calculated gradient information, the video fusion model will adjust and merge the original video clip and the high-resolution video of the expanded part. To achieve a natural fusion effect, other factors such as color consistency and lighting conditions are considered to further optimize the final result. The original video content is seamlessly embedded into the high-resolution video of the expanded part to form a complete and high-quality target-size video.
[0088] Step S4 includes:
[0089] Calculate the gradient field of the video clip frame by frame through the Sobel operator, and calculate the gradient divergence of the video clip using the Laplace operator;
[0090] Based on the gradient field and gradient divergence, use the discrete solution equation to calculate the edge pixel values of the high-resolution video of the expanded part;
[0091] Insert the edge pixel values into the high-resolution video of the expanded part to obtain the fused high-resolution expanded video.
[0092] The original low-resolution video clip is used as the source video, and the high-resolution video of the extended part after super-resolution processing is used as the target video. By applying the Sobel operator frame by frame, the gradient field of each frame can be obtained. By applying the Laplace operator to the gradient field, the gradient divergence can be calculated, which represents the rate of change of the gradient. The discrete solution equation is used to calculate the edge pixel values of the corresponding target video frame. The discrete solution equation maps the gradient information of the source video to the target video to ensure visual consistency between the two. These edge pixel values are inserted into the corresponding positions of the target video, thus achieving seamless fusion of the video clip and the extended part.
[0093] The expression of the discrete solution equation is: + + , where is the abscissa of the pixel, is the ordinate of the pixel, is the pixel value at the solved coordinate , is the pixel value at , is the pixel value at , is the pixel value at is the gradient at in the video clip.
[0094] In the discrete solution equation, the pixel values (up, down, left, right) and the gradient value at the current position in the source video are used to update the pixel value at the current position. In this way, it can be ensured that the pixel values in the target video are smooth in the local area and can match the gradient information in the source video.
[0095] Compared with the prior art, the embodiments of the present application have the following beneficial effects:
[0096] First, by splitting the video to perform shot boundary detection and segmentation on the original video, different scenes or shots can be effectively identified.
[0097] Second, through video extension, not only the spatial size of the video is increased, but also the key visual information is retained. At the same time, the demand for hardware resources can be reduced because the main operations are completed at a lower resolution.
[0098] Third, through video super-resolution, the previously obtained low-resolution extended part is further enhanced to a high-resolution state, thus ensuring that the entire output video has a unified and high picture quality level.
[0099] Fourth, a smooth transition from the original video to the extended version is achieved, and natural and coherent images can be created even in newly added areas that did not originally exist in the original video.
[0100] A possible implementation of a video spatial expansion device is given below. It is used to execute each execution step and corresponding technical effect of the video spatial expansion method shown in the above embodiments and possible implementations. The device includes:
[0101] A video splitting module, configured to split the original video into multiple video segments through the video boundary time points predicted by the video boundary detection model;
[0102] A video external expansion module, configured to perform video external expansion on the downsampled video segments using a generative model to obtain a low-resolution externally expanded video;
[0103] A video super-resolution module, configured to extract the externally expanded part from the low-resolution externally expanded video and perform super-resolution processing on the externally expanded part through a video super-resolution model to obtain a high-resolution video of the externally expanded part. Figure 2 The schematic diagram of the instance result of video super-resolution generation proposed in the embodiment of the present application is shown.
[0104] A video fusion module, configured to calculate and match the gradient information between the video segments and the high-resolution video of the externally expanded part using a video fusion model, and extend the video segments into the high-resolution video of the externally expanded part based on the gradient information to obtain a fused high-resolution externally expanded video.
[0105] This preferred embodiment provides a computer device. This computer device can implement the steps in any of the embodiments of the video spatial expansion method provided in the embodiments of the present application. Therefore, the beneficial effects of the video spatial expansion method provided in the embodiments of the present application can be achieved. For details, see the previous embodiments and will not be elaborated here.
[0106] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructions or by controlling relevant hardware through instructions. The instructions can be stored in a computer-readable storage medium and loaded and executed by a processor. For this reason, the embodiments of the present application provide a storage medium, in which multiple instructions are stored, and the instructions can be loaded by a processor to execute the steps in any of the embodiments of the video spatial expansion method provided in the embodiments of the present application.
[0107] Among them, the storage medium may include: read-only memory (ROM, Read Only Memory), random access memory (RAM, Random Access Memory), magnetic disk or optical disc, etc.
[0108] Since the instructions stored in the storage medium can execute the steps in any of the video spatial expansion method embodiments provided in the embodiments of the present application, the beneficial effects achievable by any of the video spatial expansion methods provided in the embodiments of the present application can be realized. For details, refer to the previous embodiments and will not be elaborated herein.
[0109] The foregoing are only preferred embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.
Claims
1. A video space expansion method, characterized in that: The method comprises: The original video is segmented into multiple video clips according to the video boundary time points predicted by the video boundary detection model; After downsampling the video clips, the generative model is used to perform video expansion to obtain a low-resolution expanded video; The generative model includes a video description model and a video expansion model. After downsampling the video clip, the generative model is used to expand the video to obtain a low-resolution expanded video, including: Downsampling the video clip to obtain a low-resolution original video clip; Use the video description model to obtain corresponding text descriptions for low-resolution original video clips; The video expansion model is used to combine the low-resolution original video clips and the corresponding text descriptions to expand the video to obtain a low-resolution expanded video; Extract the outward expansion part from the low-resolution outward expansion video and perform super-resolution processing on the outward expansion part through a video super-resolution model to obtain a high-resolution video of the outward expansion part; The video fusion model is used to calculate and match the gradient information between the video clip and the high-resolution video of the expanded part, and the video clip is expanded to the high-resolution video of the expanded part based on the gradient information to obtain a fused high-resolution expanded video.
2. The video space expansion method according to claim 1, characterized in that: The video boundary detection model includes a video feature extraction backbone, a feature similarity calculation module, and multiple classification heads; The video feature extraction backbone includes 3D convolution modules and Transformers modules, which are used to calculate the features of video frames; The feature similarity calculation module calculates the feature similarity vectors of adjacent frames and the RGB histogram similarity feature vectors of adjacent frames; Each classification head consists of multiple fully connected layers to predict the probability that a video frame is a video boundary.
3. The video space expansion method according to claim 1, characterized in that: The video expansion model includes a text feature extractor, a video feature extractor, a video encoder, a video decoder, and a video denoising model. The steps of expanding the video by combining the video expansion model with the text description and the low-resolution original video clip to obtain a low-resolution expanded video include: Use a text feature extractor to extract features from the text description to obtain text features; Use a video feature extractor to process low-resolution original video clips frame by frame to obtain video features; Initialize the video expansion mask with the area pixel value of the low-resolution original video clip as 1 and the pixel value of the expanded area as 0; Based on the video expansion mask, the low-resolution original video segment is padded according to the target size to be expanded, and then the padded video segment is processed by a video encoder to obtain the first latent space feature; The second latent space feature is obtained by randomly initializing the target size to be expanded through Gaussian noise; Inputting text features, video features, video expansion mask, first latent space features, and second latent space features into a video denoising model to predict noise; A forward sampler is used to subtract noise from the second latent space feature to obtain a third latent space feature, and the third latent space feature, text feature, video feature, video expansion mask, and first latent space feature are input into the video denoising model again to perform noise prediction loop iteration to obtain a denoising result; The denoising result is processed using a video decoder to obtain a low-resolution out-scaled video.
4. The video space expansion method according to claim 3, characterized in that: The video denoising model adopts a U-shaped network structure, including a three-dimensional convolution module, a spatial attention module, a text attention module, and a layout attention module.
5. The video space expansion method according to claim 1, characterized in that: The steps of using a video fusion model to calculate and match gradient information between the video clip and the high-resolution video of the expanded part, and expanding the video clip to the high-resolution video of the expanded part based on the gradient information to obtain a fused high-resolution expanded video include: The gradient field of the video clip is calculated frame by frame through the Sobel operator, and the gradient divergence of the video clip is calculated using the Laplace operator; Based on the gradient field and gradient divergence, the edge pixel values of the high-resolution video of the expanded part are calculated by using discrete solution equations; Inserting edge pixel values into the high-resolution video of the expanded part obtains a fused high-resolution expanded video.
6. The video space expansion method according to claim 5, characterized in that: The expression of the discrete solution equation is: + + ,in, is the pixel horizontal coordinate, is the pixel ordinate, To solve the coordinates The pixel value at for The pixel value at for The pixel value at for The pixel value at for The pixel value at For video clips The gradient at .
7. A video space expansion device, characterized in that: The device comprises: A video splitting module is used to split the original video into multiple video segments according to the video boundary time points predicted by the video boundary detection model; The video expansion module downsamples the video clip to obtain a low-resolution original video clip; Use the video description model to obtain corresponding text descriptions for low-resolution original video clips; The video expansion model is used to combine the low-resolution original video clips and the corresponding text descriptions to expand the video to obtain a low-resolution expanded video; The video super-resolution module is used to extract the outward expansion part from the low-resolution outward expansion video and perform super-resolution processing on the outward expansion part through the video super-resolution model to obtain a high-resolution video of the outward expansion part; The video fusion module is used to calculate and match the gradient information between the video clip and the high-resolution video of the expanded part using the video fusion model, and expand the video clip to the high-resolution video of the expanded part based on the gradient information to obtain a fused high-resolution expanded video.
8. A computer device, characterized in that: The computer device includes a processor and a memory, wherein a computer program is stored in the memory, and the computer program is loaded and executed by the processor to implement the video space expansion method according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, which is loaded and executed by a processor to implement the video space expansion method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Action detection classification method and system based on unified large model, equipment and medium
CN117671785A
Video generation method and device, electronic equipment and readable storage medium
CN118632088A