Video processing methods, apparatus, electronic devices and storage media

By applying a time-space variational autoencoder, the problem of poor video dimensionality reduction in existing technologies is solved, and a more efficient video dimensionality reduction effect is achieved.

CN119814957BActive Publication Date: 2025-12-02BEIJING LUCHEN TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411893902.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-20
Publication Date
2025-12-02
Estimated Expiration
2044-12-20

AI Technical Summary

Technical Problem

Existing variational autoencoders have shortcomings in video dimensionality reduction.

Method used

A temporal-spatial variational autoencoder is employed, which includes a temporal-spatial 3D convolution unit, a temporal-spatial self-attention mechanism unit, a temporal-spatial upsampling unit, and a temporal-spatial downsampling unit to perform convolution, temporal-spatial dependency capture, numerical interpolation, and downsampling on the video.

Benefits of technology

The temporal-spatial variational autoencoder achieves video dimensionality reduction in both temporal and spatial dimensions, significantly improving the video dimensionality reduction effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119814957B_ABST
    Figure CN119814957B_ABST
Patent Text Reader

Abstract

This invention discloses a video processing method, apparatus, electronic device, and storage medium. The method includes: acquiring a video to be processed; inputting the video to be processed into a pre-trained temporal-spatial variational autoencoder to obtain a video encoding result. The temporal-spatial variational autoencoder includes a temporal-spatial 3D convolutional unit, a temporal-spatial self-attention mechanism unit, a temporal-spatial upsampling unit, and a temporal-spatial downsampling unit. The temporal-spatial 3D convolutional unit is used to convolve the video to be processed in both temporal and spatial dimensions; the temporal-spatial self-attention mechanism unit is used to capture spatiotemporal dependencies in the video to be processed in both temporal and spatial dimensions; the temporal-spatial upsampling unit is used to perform numerical interpolation on the video to be processed in both temporal and spatial dimensions; and the temporal-spatial downsampling unit is used to downsample the video to be processed in both temporal and spatial dimensions through 3D convolution.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and in particular to a video processing method, apparatus, electronic device, and storage medium. Background Technology

[0002] Video generation refers to training artificial intelligence to automatically generate high-fidelity video content that matches the given description based on single-modal or multi-modal data such as text, images, and videos.

[0003] Video generation technology primarily relies on the architecture of diffusion models and variational autoencoders. Currently, Black Forest Labs has open-sourced Flux's spatial variational autoencoder for textual graphs, which only has the ability to reduce spatial dimensionality.

[0004] In the process of realizing this invention, it was found that at least the following technical problems exist in the prior art: existing variational autoencoders have the problem of poor video dimensionality reduction effect. Summary of the Invention

[0005] This invention provides a video processing method, apparatus, electronic device, and storage medium to improve the dimensionality reduction effect of video.

[0006] According to one aspect of the present invention, a video processing method is provided, comprising:

[0007] Get the video to be processed;

[0008] The video to be processed is input into a pre-trained temporal-spatial variational autoencoder to obtain a video encoding result, wherein the dimension of the video encoding result is smaller than the dimension of the video to be processed.

[0009] The temporal-space variational autoencoder includes a temporal-space 3D convolutional unit, a temporal-space self-attention mechanism unit, a temporal-space upsampling unit, and a temporal-space downsampling unit.

[0010] The time-space 3D convolutional unit is used to convolve the video to be processed from the time dimension and the spatial dimension.

[0011] The time-space self-attention mechanism unit is used to capture the spatiotemporal dependencies of the video to be processed from the time and space dimensions.

[0012] The time-space upsampling unit is used to perform numerical interpolation on the video to be processed from the time and space dimensions.

[0013] The time-space downsampling unit is used to downsample the video to be processed from the time and space dimensions through three-dimensional convolution.

[0014] According to another aspect of the present invention, a video processing apparatus is provided, comprising:

[0015] The video acquisition module is used to acquire videos to be processed.

[0016] The temporal-spatial variational autoencoder module is used to input the video to be processed into a pre-trained temporal-spatial variational autoencoder to obtain a video encoding result, wherein the dimension of the video encoding result is smaller than the dimension of the video to be processed.

[0017] The temporal-space variational autoencoder includes a temporal-space 3D convolutional unit, a temporal-space self-attention mechanism unit, a temporal-space upsampling unit, and a temporal-space downsampling unit.

[0018] The time-space 3D convolutional unit is used to convolve the video to be processed from the time dimension and the spatial dimension.

[0019] The time-space self-attention mechanism unit is used to capture the spatiotemporal dependencies of the video to be processed from the time and space dimensions.

[0020] The time-space upsampling unit is used to perform numerical interpolation on the video to be processed from the time and space dimensions.

[0021] The time-space downsampling unit is used to downsample the video to be processed from the time and space dimensions through three-dimensional convolution.

[0022] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:

[0023] At least one processor;

[0024] and a memory communicatively connected to the at least one processor;

[0025] The memory stores a computer program that can be executed by the at least one processor, which is then executed by the at least one processor to enable the at least one processor to perform the video processing method described in any embodiment of the present invention.

[0026] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the video processing method according to any embodiment of the present invention.

[0027] The technical solution of this invention acquires a video to be processed and then inputs it into a pre-trained temporal-spatial variational autoencoder to obtain a video encoding result. The temporal-spatial variational autoencoder includes a temporal-spatial 3D convolutional unit, a temporal-spatial self-attention mechanism unit, a temporal-spatial upsampling unit, and a temporal-spatial downsampling unit. The temporal-spatial 3D convolutional unit performs convolution on the video to be processed from both temporal and spatial dimensions. The temporal-spatial self-attention mechanism unit captures spatiotemporal dependencies in the video to be processed from both temporal and spatial dimensions. The temporal-spatial upsampling unit performs numerical interpolation on the video to be processed from both temporal and spatial dimensions. The temporal-spatial downsampling unit downsamples the video to be processed from both temporal and spatial dimensions through 3D convolution. This technical solution achieves video dimensionality reduction in both temporal and spatial dimensions through the temporal-spatial variational autoencoder, effectively improving the video dimensionality reduction effect.

[0028] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0029] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0030] Figure 1 This is a flowchart of a video processing method provided according to Embodiment 1 of the present invention;

[0031] Figure 2 This is a flowchart of a video processing method provided according to Embodiment 2 of the present invention;

[0032] Figure 3 This is a schematic diagram of a two-dimensional convolution kernel expansion according to an embodiment of the present invention;

[0033] Figure 4 This is a flowchart of a video processing method provided according to Embodiment 3 of the present invention;

[0034] Figure 5 This is a schematic diagram of the structure of a video processing device according to Embodiment 4 of the present invention;

[0035] Figure 6 This is a schematic diagram of the structure of an electronic device that implements the video processing method of the present invention. Detailed Implementation

[0036] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0037] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be used interchangeably where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or device that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or devices. The acquisition, storage, use, and processing of data in the technical solutions of this application all comply with the relevant provisions of national laws and regulations.

[0038] Example 1

[0039] Figure 1 This is a flowchart of a video processing method provided in Embodiment 1 of the present invention. This embodiment is applicable to the situation of video compression in video generation tasks. The method can be executed by a video processing device, which can be implemented in hardware and / or software, and can be configured in electronic devices such as terminals and / or servers. Figure 1 As shown, the method includes:

[0040] S110, Obtain the video to be processed.

[0041] In this embodiment of the invention, the video to be processed refers to the video to be subjected to dimensionality reduction processing, which may consist of multiple video frames.

[0042] For example, the video to be processed can be read from a preset storage path of an electronic device, or the video to be processed can be acquired in real time through a camera or other video acquisition device.

[0043] S120. Input the video to be processed into a pre-trained temporal-spatial variational autoencoder to obtain a video encoding result, wherein the dimension of the video encoding result is smaller than the dimension of the video to be processed.

[0044] The temporal-spatial variational autoencoder includes a temporal-spatial 3D convolutional unit, a temporal-spatial self-attention mechanism unit, a temporal-spatial upsampling unit, and a temporal-spatial downsampling unit. The temporal-spatial 3D convolutional unit is used to convolve the video to be processed from the temporal and spatial dimensions. The temporal-spatial self-attention mechanism unit is used to capture the spatiotemporal dependencies of the video to be processed from the temporal and spatial dimensions. The temporal-spatial upsampling unit is used to perform numerical interpolation on the video to be processed from the temporal and spatial dimensions. The temporal-spatial downsampling unit is used to downsample the video to be processed from the temporal and spatial dimensions through 3D convolution.

[0045] It should be noted that the temporal-spatial variational autoencoder in this embodiment of the invention is a three-dimensional variational autoencoder obtained by improving upon Flux's two-dimensional spatial variational autoencoder based on textual graphs. The difference between the temporal-spatial variational autoencoder and the spatial variational autoencoder lies in the convolutional units, self-attention mechanism units, upsampling units, and downsampling units. In other words, the convolutional units, self-attention mechanism units, upsampling units, and downsampling units in the temporal-spatial variational autoencoder all have an added temporal dimension.

[0046] Specifically, the temporal-spatial 3D convolution unit is a type of 3D convolution that can be used to convolve the video to be processed from both temporal and spatial dimensions. The temporal-spatial self-attention mechanism unit is a self-attention mechanism used to capture spatiotemporal dependencies in the video to be processed from both temporal and spatial dimensions. The temporal-spatial upsampling unit can perform numerical interpolation on the video to be processed from both temporal and spatial dimensions. The temporal-spatial downsampling unit can downsample the video to be processed from both temporal and spatial dimensions through 3D convolution.

[0047] The technical solution of this invention acquires a video to be processed and then inputs it into a pre-trained temporal-spatial variational autoencoder to obtain a video encoding result. The temporal-spatial variational autoencoder includes a temporal-spatial 3D convolutional unit, a temporal-spatial self-attention mechanism unit, a temporal-spatial upsampling unit, and a temporal-spatial downsampling unit. The temporal-spatial 3D convolutional unit performs convolution on the video to be processed from both temporal and spatial dimensions. The temporal-spatial self-attention mechanism unit captures spatiotemporal dependencies in the video to be processed from both temporal and spatial dimensions. The temporal-spatial upsampling unit performs numerical interpolation on the video to be processed from both temporal and spatial dimensions. The temporal-spatial downsampling unit downsamples the video to be processed from both temporal and spatial dimensions through 3D convolution. This technical solution achieves video dimensionality reduction in both temporal and spatial dimensions through the temporal-spatial variational autoencoder, effectively improving the video dimensionality reduction effect.

[0048] Example 2

[0049] Figure 2 This is a flowchart of a video processing method provided in Embodiment 2 of the present invention. The method of this embodiment can be combined with various optional schemes in the video processing methods provided in the above embodiments. The video processing method provided in this embodiment has been further optimized. Optionally, before inputting the video to be processed into a pre-trained temporal-spatial variational autoencoder to obtain the video encoding result, the method further includes: obtaining the weight information of the spatial variational autoencoder; copying the weight information of the spatial variational autoencoder into an initial temporal-spatial variational autoencoder to obtain an dilated temporal-spatial variational autoencoder; and training the dilated temporal-spatial variational autoencoder to obtain a trained temporal-spatial variational autoencoder.

[0050] like Figure 2 As shown, the method includes:

[0051] S210. Obtain the weight information of the spatial variational autoencoder.

[0052] Spatial variational autoencoders refer to two-dimensional variational autoencoders, which can be Flux's text-based two-dimensional spatial variational autoencoders or other two-dimensional spatial variational autoencoders, etc., without specific limitations here. Weight information refers to the model parameters of the spatial variational autoencoder, which may include, but is not limited to, two-dimensional convolutional kernel weights and linear projection weights.

[0053] S220. Copy the weight information of the spatial variational autoencoder to the initial time-spatial variational autoencoder to obtain the dilated time-spatial variational autoencoder.

[0054] The initial temporal-spatial variational autoencoder is a pre-built three-dimensional variational autoencoder architecture that is compatible with two-dimensional spatial variational autoencoders. The dilated temporal-spatial variational autoencoder is a variational autoencoder with the spatial compression capability of a spatial variational autoencoder.

[0055] It should be noted that by copying the weight information of the spatial variational autoencoder to the initial temporal-spatial variational autoencoder, an dilated temporal-spatial variational autoencoder can be obtained. This allows the dilated temporal-spatial variational autoencoder to have the same spatial compression capability as the spatial variational autoencoder from the beginning. After a small amount of training, a temporal-spatial variational autoencoder with excellent compression capability can be obtained, which greatly reduces the waste of computing resources.

[0056] S230. Train the dilated time-space variational autoencoder to obtain the trained time-space variational autoencoder.

[0057] S240, Obtain the video to be processed.

[0058] S250. Input the video to be processed into a pre-trained temporal-spatial variational autoencoder to obtain a video encoding result, wherein the dimension of the video encoding result is smaller than the dimension of the video to be processed.

[0059] Optionally, the weight information of the spatial variational autoencoder includes two-dimensional convolutional kernel weight information; the initial temporal-spatial variational autoencoder includes an initial three-dimensional convolutional kernel, which includes a target position and remaining positions; correspondingly, the weight information of the spatial variational autoencoder is copied into the initial temporal-spatial variational autoencoder to obtain a dilated temporal-spatial variational autoencoder, which includes: copying the two-dimensional convolutional kernel weight information to the target position of the initial three-dimensional convolutional kernel, and filling the remaining positions of the initial three-dimensional convolutional kernel with zeros to obtain a dilated three-dimensional convolutional kernel, wherein the dilated temporal-spatial variational autoencoder includes a dilated three-dimensional convolutional kernel.

[0060] The two-dimensional convolutional kernel weight information refers to the weights of the two-dimensional convolutional kernel in the spatial variational autoencoder. For example, the tensor dimension of the two-dimensional convolutional kernel can be 3×3, corresponding to H×W. The initial three-dimensional convolutional kernel refers to the convolutional kernel of the three-dimensional convolution in the initial temporal-spatial variational autoencoder. For example, the tensor dimension of the initial three-dimensional convolutional kernel can be 3×3×3, corresponding to T×H×W.

[0061] The initial temporal-spatial variational autoencoder can include the same number of 3D convolutions as the 2D convolutions in the spatial variational autoencoder, allowing for one-to-one replication. The target location refers to the position used to store the weights of the 2D convolutional kernel, while the remaining locations refer to the positions in the initial 3D convolutional kernel excluding the target location. The dilated 3D convolutional kernel refers to a unit that possesses the capabilities of the 2D convolutional kernel in the spatial variational autoencoder.

[0062] In an embodiment of the present invention, Figure 3 This is a schematic diagram of a two-dimensional convolution kernel expansion according to an embodiment of the present invention. Specifically, the target position can be the center position, the head position, or the tail position of the initial three-dimensional convolution kernel, respectively corresponding to... Figure 3 (A), (B), and (C) in the diagram. Furthermore, the average value of the two-dimensional convolution kernel can be copied to each position of the initial three-dimensional convolution kernel to obtain an dilated three-dimensional convolution kernel, corresponding to... Figure 3 (D) in the middle.

[0063] Optionally, the target location is the center of the initial 3D convolution kernel.

[0064] The mathematical expression for the center position weight replication scheme is as follows:

[0065] K3[0] = 0;

[0066] K3[1] = K2;

[0067] K3[2]=0;

[0068] Wherein, K2 represents the two-dimensional convolution kernel, K3 represents the initial three-dimensional convolution kernel, K3[0] represents the head position of the initial three-dimensional convolution kernel, K3[1] represents the center position of the initial three-dimensional convolution kernel, and K3[2] represents the tail position of the initial three-dimensional convolution kernel.

[0069] It should be noted that, compared with the head position weight replication scheme, the tail position weight replication scheme, and the average value weight replication scheme, the time-space variational autoencoder obtained by the center position weight replication scheme has the best compression capability.

[0070] Optionally, the weight information of the spatial variational autoencoder includes the query vector, key vector, and numerical vector of the two-dimensional self-attention mechanism; the initial temporal-spatial variational autoencoder includes an initial three-dimensional self-attention mechanism unit, which includes the query vector position, key vector position, and numerical vector position; correspondingly, copying the weight information of the spatial variational autoencoder to the initial temporal-spatial variational autoencoder to obtain the dilated temporal-spatial variational autoencoder includes: copying the query vector of the two-dimensional self-attention mechanism to the query vector position of the initial three-dimensional self-attention mechanism unit, copying the key vector of the two-dimensional self-attention mechanism to the key vector position of the initial three-dimensional self-attention mechanism unit, and copying the numerical vector of the two-dimensional self-attention mechanism to the numerical vector position of the initial three-dimensional self-attention mechanism unit to obtain the dilated three-dimensional self-attention mechanism unit, and the dilated temporal-spatial variational autoencoder includes the dilated three-dimensional self-attention mechanism unit.

[0071] In this context, the query vector (Query) represents the weight Q in the self-attention mechanism. The key vector (Key) represents the weight K in the self-attention mechanism. The value vector (Value) represents the weight V in the self-attention mechanism. An inflated 3D self-attention mechanism unit refers to a self-attention mechanism unit that completes weight replication.

[0072] For example, the mathematical expression for weight replication in a self-attention mechanism unit is as follows:

[0073] 3D(Q) = 2D(Q);

[0074] 3D(K) = 2D(K);

[0075] 3D(V) = 2D(V);

[0076] Wherein, 2D(Q) represents the query vector of the two-dimensional self-attention mechanism, 3D(Q) represents the query vector position of the initial three-dimensional self-attention mechanism unit, and 3D(Q) = 2D(Q) means copying the query vector of the two-dimensional self-attention mechanism to the query vector position of the initial three-dimensional self-attention mechanism unit; 2D(K) represents the key vector of the two-dimensional self-attention mechanism, 3D(K) represents the key vector position of the initial three-dimensional self-attention mechanism unit, and 3D(K) = 2D(K) means copying the key vector of the two-dimensional self-attention mechanism to the key vector position of the initial three-dimensional self-attention mechanism unit; 2D(V) represents the numerical vector of the two-dimensional self-attention mechanism, 3D(V) represents the numerical vector position of the initial three-dimensional self-attention mechanism unit, and 3D(V) = 2D(V) means copying the numerical vector of the two-dimensional self-attention mechanism to the numerical vector position of the initial three-dimensional self-attention mechanism unit.

[0077] Optionally, the self-attention sequence of the initial three-dimensional self-attention mechanism unit contains spatial and temporal dimensional information.

[0078] For example, the self-attention sequence of the two-dimensional self-attention mechanism is h×w, where h represents the video length, w represents the video width, and h×w represents the spatial dimension information. The self-attention sequence of the initial three-dimensional self-attention mechanism unit is t×h×w, where t represents the temporal dimension information.

[0079] The technical solution of this invention obtains an dilated time-space variational autoencoder by copying the weight information of the spatial variational autoencoder into the initial time-space variational autoencoder. This allows the dilated time-space variational autoencoder to have the same spatial compression capability as the spatial variational autoencoder from the beginning. After a small amount of training, a time-space variational autoencoder with excellent compression capability can be obtained, which greatly reduces the waste of computing resources.

[0080] Example 3

[0081] Figure 4 This is a flowchart of a video processing method provided in Embodiment 3 of the present invention. The method of this embodiment can be combined with various optional schemes in the video processing methods provided in the above embodiments. The video processing method provided in this embodiment has been further optimized. Optionally, after inputting the video to be processed into a pre-trained temporal-spatial variational autoencoder to obtain the video encoding result, the method further includes: training the diffusion model to be trained based on the video encoding result to obtain the trained diffusion model.

[0082] like Figure 4 As shown, the method includes:

[0083] S310, Obtain the video to be processed.

[0084] S320. Input the video to be processed into a pre-trained temporal-spatial variational autoencoder to obtain a video encoding result, wherein the dimension of the video encoding result is smaller than the dimension of the video to be processed.

[0085] S330. Based on the video encoding results, train the diffusion model to be trained to obtain the trained diffusion model.

[0086] Specifically, noise can be added to the video encoding result to obtain a noisy video encoding result. Then, the diffusion model outputs predicted noise based on the noisy video encoding result, calculates the training loss based on the predicted noise and the actual noise, updates the parameters of the diffusion model, until the training stopping condition is met, and obtains the trained diffusion model.

[0087] For the training process of the diffusion model, the embodiments of the present invention can significantly reduce the resources required for training the diffusion model and reduce the training difficulty of the diffusion model by using an efficient dimensionality-reducing time-space variational autoencoder.

[0088] Example 4

[0089] Figure 5 This is a schematic diagram of the structure of a video processing device provided in Embodiment 4 of the present invention. Figure 5 As shown, the device includes:

[0090] The video acquisition module 410 is used to acquire the video to be processed.

[0091] The temporal-spatial variational autoencoder module 420 is used to input the video to be processed into a pre-trained temporal-spatial variational autoencoder to obtain a video encoding result, wherein the dimension of the video encoding result is smaller than the dimension of the video to be processed.

[0092] The temporal-space variational autoencoder includes a temporal-space 3D convolutional unit, a temporal-space self-attention mechanism unit, a temporal-space upsampling unit, and a temporal-space downsampling unit.

[0093] The time-space 3D convolutional unit is used to convolve the video to be processed from the time dimension and the spatial dimension.

[0094] The time-space self-attention mechanism unit is used to capture the spatiotemporal dependencies of the video to be processed from the time and space dimensions.

[0095] The time-space upsampling unit is used to perform numerical interpolation on the video to be processed from the time and space dimensions.

[0096] The time-space downsampling unit is used to downsample the video to be processed from the time and space dimensions through three-dimensional convolution.

[0097] The technical solution of this invention acquires a video to be processed and then inputs it into a pre-trained temporal-spatial variational autoencoder to obtain a video encoding result. The temporal-spatial variational autoencoder includes a temporal-spatial 3D convolutional unit, a temporal-spatial self-attention mechanism unit, a temporal-spatial upsampling unit, and a temporal-spatial downsampling unit. The temporal-spatial 3D convolutional unit performs convolution on the video to be processed from both temporal and spatial dimensions. The temporal-spatial self-attention mechanism unit captures spatiotemporal dependencies in the video to be processed from both temporal and spatial dimensions. The temporal-spatial upsampling unit performs numerical interpolation on the video to be processed from both temporal and spatial dimensions. The temporal-spatial downsampling unit downsamples the video to be processed from both temporal and spatial dimensions through 3D convolution. This technical solution achieves video dimensionality reduction in both temporal and spatial dimensions through the temporal-spatial variational autoencoder, effectively improving the video dimensionality reduction effect.

[0098] In some alternative embodiments, the video processing apparatus further includes:

[0099] The weight information acquisition module is used to acquire the weight information of the spatial variational autoencoder.

[0100] The weight information copying module is used to copy the weight information of the spatial variational autoencoder to the initial time-spatial variational autoencoder to obtain the dilated time-spatial variational autoencoder.

[0101] The dilated time-space variational autoencoder training module is used to train the dilated time-space variational autoencoder to obtain the trained time-space variational autoencoder.

[0102] In some optional implementations, the weight information of the spatial variational autoencoder includes two-dimensional convolutional kernel weight information; the initial temporal-spatial variational autoencoder includes an initial three-dimensional convolutional kernel, which includes the target position and the remaining position;

[0103] Correspondingly, the weight information replication module includes:

[0104] A two-dimensional convolution kernel dilation unit is used to copy the weight information of the two-dimensional convolution kernel to the target position of the initial three-dimensional convolution kernel and fill the remaining positions of the initial three-dimensional convolution kernel with zeros to obtain a dilated three-dimensional convolution kernel. The dilated temporal-spatial variational autoencoder includes the dilated three-dimensional convolution kernel.

[0105] In some alternative implementations, the target location is the center location of the initial 3D convolutional kernel.

[0106] In some optional implementations, the weight information of the spatial variational autoencoder includes a query vector of the two-dimensional self-attention mechanism, a key vector of the two-dimensional self-attention mechanism, and a numerical vector of the two-dimensional self-attention mechanism; the initial temporal-spatial variational autoencoder includes an initial three-dimensional self-attention mechanism unit, which includes the query vector position, the key vector position, and the numerical vector position.

[0107] Correspondingly, the weight information replication module includes:

[0108] A two-dimensional self-attention mechanism expansion unit is used to copy the query vector of the two-dimensional self-attention mechanism to the query vector position of the initial three-dimensional self-attention mechanism unit, copy the key vector of the two-dimensional self-attention mechanism to the key vector position of the initial three-dimensional self-attention mechanism unit, and copy the numerical vector of the two-dimensional self-attention mechanism to the numerical vector position of the initial three-dimensional self-attention mechanism unit, thereby obtaining an expanded three-dimensional self-attention mechanism unit. The expanded time-space variational autoencoder includes the expanded three-dimensional self-attention mechanism unit.

[0109] In some optional implementations, the self-attention sequence of the initial three-dimensional self-attention mechanism unit includes spatial dimension information and temporal dimension information.

[0110] In some alternative embodiments, the video processing apparatus further includes:

[0111] The diffusion model training unit is used to train the diffusion model to be trained based on the video encoding results, so as to obtain the trained diffusion model.

[0112] The video processing apparatus provided in the embodiments of the present invention can execute the video processing method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of executing the method.

[0113] Example 5

[0114] Figure 6A schematic diagram of an electronic device 10 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0115] like Figure 6 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded into the RAM 13 from storage unit 18. The RAM 13 can also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An I / O interface 15 is also connected to the bus 14.

[0116] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0117] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as video processing methods, which include:

[0118] Get the video to be processed;

[0119] The video to be processed is input into a pre-trained temporal-spatial variational autoencoder to obtain a video encoding result, wherein the dimension of the video encoding result is smaller than the dimension of the video to be processed.

[0120] The temporal-space variational autoencoder includes a temporal-space 3D convolutional unit, a temporal-space self-attention mechanism unit, a temporal-space upsampling unit, and a temporal-space downsampling unit.

[0121] The time-space 3D convolutional unit is used to convolve the video to be processed from the time dimension and the spatial dimension.

[0122] The time-space self-attention mechanism unit is used to capture the spatiotemporal dependencies of the video to be processed from the time and space dimensions.

[0123] The time-space upsampling unit is used to perform numerical interpolation on the video to be processed from the time and space dimensions.

[0124] The time-space downsampling unit is used to downsample the video to be processed from the time and space dimensions through three-dimensional convolution.

[0125] In some embodiments, the video processing method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or mounted on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the video processing method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the video processing method by any other suitable means (e.g., by means of firmware).

[0126] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0127] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0128] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0129] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0130] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0131] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.

[0132] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0133] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A video processing method, characterized in that, include: Obtain the weight information of the spatial variational autoencoder; The weight information of the spatial variational autoencoder is copied into the initial time-spatial variational autoencoder to obtain the dilated time-spatial variational autoencoder; The dilated time-space variational autoencoder is trained to obtain the trained time-space variational autoencoder. Get the video to be processed; The video to be processed is input into a pre-trained temporal-spatial variational autoencoder to obtain a video encoding result, wherein the dimension of the video encoding result is smaller than the dimension of the video to be processed. The temporal-space variational autoencoder includes a temporal-space 3D convolutional unit, a temporal-space self-attention mechanism unit, a temporal-space upsampling unit, and a temporal-space downsampling unit. The time-space 3D convolutional unit is used to convolve the video to be processed from the time dimension and the spatial dimension. The time-space self-attention mechanism unit is used to capture the spatiotemporal dependencies of the video to be processed from the time and space dimensions. The time-space upsampling unit is used to perform numerical interpolation on the video to be processed from the time and space dimensions. The time-space downsampling unit is used to downsample the video to be processed from the time and space dimensions through three-dimensional convolution; The weight information of the spatial variational autoencoder includes two-dimensional convolutional kernel weight information; the initial temporal-spatial variational autoencoder includes an initial three-dimensional convolutional kernel, which includes the target position and the remaining position; Accordingly, the weight information of the spatial variational autoencoder is copied into the initial time-spatial variational autoencoder to obtain the dilated time-spatial variational autoencoder, including: The weight information of the two-dimensional convolutional kernel is copied to the target position of the initial three-dimensional convolutional kernel, and zeros are filled into the remaining positions of the initial three-dimensional convolutional kernel to obtain the dilated three-dimensional convolutional kernel. The dilated temporal-spatial variational autoencoder includes the dilated three-dimensional convolutional kernel. The target position is the center position of the initial three-dimensional convolution kernel.

2. The method according to claim 1, characterized in that, The weight information of the spatial variational autoencoder also includes the query vector, key vector, and value vector of the two-dimensional self-attention mechanism; the initial temporal-spatial variational autoencoder also includes an initial three-dimensional self-attention mechanism unit, which includes the query vector position, key vector position, and value vector position. Accordingly, the weight information of the spatial variational autoencoder is copied into the initial time-spatial variational autoencoder to obtain the dilated time-spatial variational autoencoder, which also includes: The query vector of the two-dimensional self-attention mechanism is copied to the query vector position of the initial three-dimensional self-attention mechanism unit, the key vector of the two-dimensional self-attention mechanism is copied to the key vector position of the initial three-dimensional self-attention mechanism unit, and the numerical vector of the two-dimensional self-attention mechanism is copied to the numerical vector position of the initial three-dimensional self-attention mechanism unit to obtain the dilated three-dimensional self-attention mechanism unit. The dilated time-space variational autoencoder also includes the dilated three-dimensional self-attention mechanism unit.

3. The method according to claim 2, characterized in that, The self-attention sequence of the initial three-dimensional self-attention mechanism unit contains spatial and temporal dimensional information.

4. The method according to any one of claims 1-3, characterized in that, After inputting the video to be processed into a pre-trained temporal-spatial variational autoencoder to obtain the video encoding result, the process further includes: The diffusion model to be trained is trained based on the video encoding results to obtain the trained diffusion model.

5. A video processing apparatus, characterized in that, include: The weight information acquisition module is used to acquire the weight information of the spatial variational autoencoder. The weight information copying module is used to copy the weight information of the spatial variational autoencoder to the initial time-spatial variational autoencoder to obtain the dilated time-spatial variational autoencoder. The dilated time-space variational autoencoder training module is used to train the dilated time-space variational autoencoder to obtain the trained time-space variational autoencoder. The video acquisition module is used to acquire videos to be processed. The temporal-spatial variational autoencoder module is used to input the video to be processed into a pre-trained temporal-spatial variational autoencoder to obtain a video encoding result, wherein the dimension of the video encoding result is smaller than the dimension of the video to be processed. The temporal-space variational autoencoder includes a temporal-space 3D convolutional unit, a temporal-space self-attention mechanism unit, a temporal-space upsampling unit, and a temporal-space downsampling unit. The time-space 3D convolutional unit is used to convolve the video to be processed from the time dimension and the spatial dimension. The time-space self-attention mechanism unit is used to capture the spatiotemporal dependencies of the video to be processed from the time and space dimensions. The time-space upsampling unit is used to perform numerical interpolation on the video to be processed from the time and space dimensions. The time-space downsampling unit is used to downsample the video to be processed from the time and space dimensions through three-dimensional convolution; The weight information of the spatial variational autoencoder includes two-dimensional convolutional kernel weight information; the initial temporal-spatial variational autoencoder includes an initial three-dimensional convolutional kernel, which includes the target position and the remaining position; Accordingly, the weight information copying module includes: A two-dimensional convolution kernel dilation unit is used to copy the weight information of the two-dimensional convolution kernel to the target position of the initial three-dimensional convolution kernel and fill the remaining position of the initial three-dimensional convolution kernel with zeros to obtain a dilated three-dimensional convolution kernel. The dilated temporal-spatial variational autoencoder includes the dilated three-dimensional convolution kernel. The target position is the center position of the initial three-dimensional convolution kernel.

6. An electronic device, characterized in that, The electronic device includes: At least one processor; and a memory communicatively connected to the at least one processor; The memory stores a computer program that can be executed by the at least one processor, which is then executed by the at least one processor to enable the at least one processor to perform the video processing method according to any one of claims 1-4.

7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the video processing method according to any one of claims 1-4.

Citation Information

Patent Citations

  • Method for determining anomaly based on multivariable time series data reconstruction

    CN116127391A

  • Video generation with latent diffusion models

    US20240169479A1