Multi-scale video prediction method, system, medium, product and device

By adopting a dual-branch optical flow module and a space-channel collaborative attention fusion strategy in the video prediction method, the problem that the existing technology is difficult to capture high-order interaction characteristics is solved, and higher video prediction accuracy and effect are achieved.

CN119418255BActive Publication Date: 2025-05-13SHANDONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510031102.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-09
Publication Date
2025-05-13
Estimated Expiration
2045-01-09

AI Technical Summary

Technical Problem

Existing video prediction methods are difficult to effectively capture the high-order interaction characteristics between spatial details and global dynamics, resulting in limited prediction accuracy in complex scenarios.

Method used

The multi-scale video prediction method is adopted, and the motion and spatial features are extracted separately through the dual-branch optical flow module, and the two branch features are deeply interacted with the space-channel coordinated attention fusion strategy to generate fusion features to improve prediction accuracy.

Benefits of technology

It significantly improves the accuracy and effect of video prediction, can better understand the changes of objects at different scales, and enhances the ability to understand complex scenes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119418255B_ABST
    Figure CN119418255B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of image processing technology. A multi-scale video prediction method, system, medium, product and device are proposed, wherein the previous frame image and the current frame image are respectively input into the two branches of a dual-branch optical flow module to obtain motion features and spatial features; based on the motion features and the spatial features, fusion features are obtained, wherein the fusion features include the reverse optical flow between the next frame image and the previous frame image, the reverse optical flow between the next frame image and the current frame image, and a weight map; based on the reverse optical flow between the next frame image and the previous frame image, the reverse optical flow between the next frame image and the current frame image, and the weight map, the prediction result of the next frame image is determined. The present invention captures the motion trend and spatial detail information of dynamic objects at different scales, and uses the space-channel collaborative attention fusion strategy to deeply interact the two-branch features, which significantly improves the accuracy and effect of video prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing technology, and in particular to a multi-scale video prediction method, system, medium, product and equipment. Background Art

[0002] The statements in this section merely provide background art related to the present invention and do not necessarily constitute prior art.

[0003] Video prediction technology is an important branch of computer vision and artificial intelligence, which aims to predict future video content based on existing video frame sequences. This task not only involves complex time series modeling problems, but also requires understanding the motion patterns of objects and the spatial structure of scenes in high-dimensional space. Video prediction is of great significance in many practical applications such as autonomous driving, robotic arm operation, and behavior prediction. Accurate future frame prediction can enhance the system's scene perception ability, thereby improving the accuracy and efficiency of intelligent decision-making.

[0004] At present, mainstream video prediction methods usually rely on deep learning models to solve these complex problems. Although the technology has made some progress in recent years, existing methods still have some shortcomings. For example, these methods usually require additional information to improve the prediction accuracy, have high computational resource overhead, and can often only capture features at a certain scale. Therefore, it is difficult for them to effectively obtain high-order interaction characteristics between spatial details and global dynamics, and the prediction accuracy in complex scenes is still limited. Summary of the invention

[0005] In order to address the shortcomings of the prior art, the present invention provides a multi-scale video prediction method, system, medium, product and equipment to capture the motion trend and spatial detail information of dynamic objects at different scales, and use the spatial-channel collaborative attention fusion strategy to deeply interact the two-branch features, significantly improving the accuracy and effect of video prediction.

[0006] In order to achieve the above object, the present invention adopts the following technical solution:

[0007] In a first aspect, the present invention provides a multi-scale video prediction method.

[0008] A multi-scale video prediction method includes the following steps:

[0009] Obtain the previous frame image and the current frame image of the video to be predicted;

[0010] Inputting the previous frame image and the current frame image into two branches of a dual-branch optical flow module respectively to obtain motion features and spatial features;

[0011] Obtaining fusion features according to the motion features and the spatial features, wherein the fusion features include a reverse optical flow between a next frame image and a previous frame image, a reverse optical flow between a next frame image and a current frame image, and a weight map;

[0012] The prediction result of the next frame image is determined according to the reverse optical flow between the next frame image and the previous frame image, the reverse optical flow between the next frame image and the current frame image, and the weight map.

[0013] As a further limitation of the first aspect of the present invention, the extraction of spatial features includes: extracting spatial detail information of the previous frame image and the current frame image, dividing the spatial detail information into two feature streams, performing feature extraction using depthwise separable convolution and dilated convolution respectively, and splicing the feature extraction results of the depthwise separable convolution and the dilated convolution to obtain the spatial features.

[0014] As a further limitation of the first aspect of the present invention, the extraction of motion features includes: performing residual feature extraction on the previous frame image and the current frame image, performing channel division on the residual features, pooling the features of each channel at different scales, extracting multi-scale features and restoring the original size, splicing the features of each channel after pooling and restoring the size, extracting convolution features on the splicing result and performing dot multiplication with the residual feature to obtain the motion feature.

[0015] As a further limitation of the first aspect of the present invention, obtaining a fusion feature according to the motion feature and the spatial feature includes:

[0016] Splicing the motion feature and the spatial feature to obtain a spliced ​​feature;

[0017] In combination with the splicing features, a query vector, a key vector and a value vector are obtained by using a depthwise separable convolution, a self-attention calculation is performed according to the query vector, the key vector and the value vector, and after maximum pooling is performed on the self-attention calculation result, a weighted calculation is performed with the splicing features to obtain a fused feature.

[0018] As a further limitation of the first aspect of the present invention, the query vector , key vector Sum value vector The calculation method includes: , , ;in, for Depthwise separable convolution, The splicing feature.

[0019] As a further limitation of the first aspect of the present invention, performing weighted calculation with the splicing feature includes: ,in, is the result of self-attention calculation, To use the self-attention calculation results The kernel size is used for maximum pooling. is the splicing feature, For fusion features.

[0020] In a second aspect, the present invention provides a multi-scale video prediction system.

[0021] A multi-scale video prediction system, comprising:

[0022] The image acquisition unit is configured to: acquire a previous frame image and a current frame image of the video to be predicted;

[0023] The feature extraction unit is configured to: input the previous frame image and the current frame image into two branches of a dual-branch optical flow module respectively to obtain motion features and spatial features;

[0024] A feature fusion unit is configured to obtain a fusion feature according to the motion feature and the spatial feature, wherein the fusion feature includes a reverse optical flow between a next frame image and a previous frame image, a reverse optical flow between a next frame image and a current frame image, and a weight map;

[0025] The image prediction unit is configured to determine the prediction result of the next frame image according to the reverse optical flow between the next frame image and the previous frame image, the reverse optical flow between the next frame image and the current frame image, and the weight map.

[0026] In a third aspect, the present invention provides a computer device, comprising: a processor and a computer-readable storage medium;

[0027] a processor adapted to execute a computer program;

[0028] A computer-readable storage medium, wherein a computer program is stored in the computer-readable storage medium, and when the computer program is executed by the processor, the multi-scale video prediction method as described in the first aspect of the present invention is implemented.

[0029] In a fourth aspect, the present invention provides a computer-readable storage medium storing a computer program, wherein the computer program is suitable for being loaded by a processor and executing the multi-scale video prediction method as described in the first aspect of the present invention.

[0030] In a fifth aspect, the present invention provides a computer program product, comprising a computer program, and when the computer program is executed by a processor, the multi-scale video prediction method as described in the first aspect of the present invention is implemented.

[0031] Compared with the prior art, the present invention has the following beneficial effects:

[0032] (1) Traditional methods can only capture features at a specific scale and cannot fully express the changes of objects at different scales, which leads to information loss. The present invention designs a dual-branch feature extraction strategy to capture the motion trend and spatial detail information of the object respectively, which significantly improves the model's ability to understand complex scenes.

[0033] (2) Existing methods lack sufficient flexibility in the multi-scale information fusion process and find it difficult to effectively capture the high-order interaction characteristics between spatial details and global dynamics. The present invention utilizes a spatial-channel collaborative attention fusion strategy to deeply interact the features of the two branches, which significantly improves the accuracy and effect of video prediction.

[0034] Advantages of additional aspects of the present invention will be given in part in the following description, and in part will become obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] The accompanying drawings in the specification, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.

[0036] Figure 1 A schematic diagram of a flow chart of a multi-scale video prediction method provided in Embodiment 1 of the present invention;

[0037] Figure 2 A schematic diagram of the principle of the multi-scale video prediction method provided in Embodiment 1 of the present invention;

[0038] Figure 3 A schematic diagram of a multi-scale video prediction system provided by Embodiment 2 of the present invention;

[0039] Figure 4 A schematic diagram of a computer device provided in Example 3 of the present invention. DETAILED DESCRIPTION

[0040] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0041] It should be noted that the following detailed descriptions are exemplary and are intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meanings as those commonly understood by those skilled in the art to which the present invention belongs.

[0042] In the absence of conflict, the embodiments of the present invention and the features of the embodiments may be combined with each other.

[0043] Embodiment 1:

[0044] Most pixels in a video are approximate copies of adjacent frames, and the mainstream solution is to use optical flow estimation to obtain the motion vector of the pixel. Some current methods use optical flow supervision to build models, but accurate optical flow images are usually difficult to obtain, which makes them impossible to apply in real scenarios. At the same time, due to the differences in size, motion amplitude, etc. of different moving objects, feature extraction at a single scale will lead to information loss, resulting in insufficient prediction accuracy in dynamically changing or detailed scenes. In view of this, this implementation proposes a multi-scale video prediction method, which captures the motion trend and spatial detail information of moving objects at different scales, and uses the spatial-channel collaborative attention fusion module to deeply interact the two-branch features, thereby enriching the dynamic feature information of the video frame, thereby significantly improving the accuracy and effect of video prediction. Specifically, such as Figure 1 and Figure 2 As shown, the following process is included:

[0045] S1: Obtain the previous frame image and the current frame image of the video to be predicted;

[0046] S2: Inputting the previous frame image and the current frame image into two branches of a dual-branch optical flow module respectively to obtain motion features and spatial features;

[0047] S3: obtaining a fusion feature according to the motion feature and the spatial feature, wherein the fusion feature includes a reverse optical flow between a next frame image and a previous frame image, a reverse optical flow between a next frame image and a current frame image, and a weight map;

[0048] S4: Determine the prediction result of the next frame image according to the reverse optical flow between the next frame image and the previous frame image, the reverse optical flow between the next frame image and the current frame image, and the weight map.

[0049] In step S1 of the present implementation, specifically, the previous frame image and the current frame image do not have to be two consecutive adjacent needle images, and may also be the previous frame and the current frame obtained by sampling each frame of the video at intervals (this method is selected in the present invention).

[0050] It can be understood that the processes S2-S4 in this implementation can be implemented by using a multi-scale video prediction model. The process of establishing the multi-scale video prediction model includes:

[0051] The recorded video is frame-decimated (the sampling rate is set to 10, and interval decimation is used to obtain a complete interval-sampled image sequence). The data is processed in a sliding window form to obtain multiple RGB image subsequences. Based on the RGB image subsequences, an end-to-end optical flow estimation method is used to obtain a multi-scale video prediction model.

[0052] In this implementation, the sliding window processing of the data is to move the fixed window size from the starting position to the right step by step to obtain an RGB image subsequence with temporal local dynamic information. In this implementation, the window length is selected as 7 and the step size is 1. After moving one frame to the right each time, the RGB frame in the window is saved as an RGB sub-video sequence. By iterating the original frame sequence, several sub-video frame sequences are obtained for training and testing.

[0053] In this implementation, preferably, an end-to-end optical flow estimation method is used to obtain a multi-scale video prediction model, and the previous frame and the current frame images are obtained and input into a dual-branch optical flow module to extract motion features and spatial features;

[0054] Through deep shared convolution and self-attention mechanism, multi-semantic spatial information fusion and inter-channel similarity calculation are performed respectively to obtain the optical flow field and weight map of the RGB image subsequence; more specifically, the optical flow of the previous frame and the current frame is estimated and the weight map is used to adjust the proportion to ensure the optimal prediction effect under different video content and motion modes;

[0055] The future video frame sequence is predicted through reverse optical flow and weight map, and multi-scale perceptual similarity evaluation is performed between it and the real video frame;

[0056] After the similarity evaluation meets the requirements, the model and model internal parameters corresponding to the best output result are saved and packaged into an executable program as a video prediction model.

[0057] In this implementation, preferably, the dual-branch optical flow module includes: enhancing the network's ability to capture long-distance dependencies through channel splitting and multiple pooling technology, and extracting motion features at multiple scales through a multi-scale feature extraction module. The calculation formula is as follows:

[0058] First, divide the residual features into channels:

[0059] (1);

[0060] Pool the features at different scales, extract multi-scale features and restore the original size:

[0061] (2);

[0062] Concatenate features and perform point multiplication weighting:

[0063] (3);

[0064] (4);

[0065] in, represents the residual features extracted by the residual feature extraction module, Indicates average division according to the channel dimension (where, is the division result of different channel dimensions), Representatives Perform 1x, 2x, 3x, and 4x maximum pooling and multi-scale modulation, where represents 1, 2, 3 and 4, Represents the change in pixel size. refers to Depthwise separable convolution, Indicates that the image features are upsampled to the original resolution by nearest neighbor interpolation. right (include ) to perform channel dimension splicing, and use convolution With residual features Perform point multiplication to finally obtain multi-scale motion features .

[0066] Use depthwise separable convolution and dilated convolution to enhance the network's ability to understand details and extract spatial features. Extract spatial detail information at high resolution and divide it into two feature streams:

[0067] (5);

[0068] Extract local features and regional feature information respectively:

[0069] (6);

[0070] in, and Respectively represent and Convolution, which is divided into and , respectively using depthwise separable convolution and dilated convolution Perform feature extraction and finally pass Splicing to get spatial features .

[0071] The calculation formula of the feature information interaction module based on spatial-channel collaborative attention is as follows:

[0072] Splicing motion features and spatial characteristics , get the splicing features :

[0073] (7);

[0074] Using depth-wise separable convolution to obtain , , And perform sub-attention calculation:

[0075] (8);

[0076] (9);

[0077] (10);

[0078] (11);

[0079] Exploit 7 7-size kernels are used for maximum pooling and concatenated features Perform weighted calculation to obtain fusion features :

[0080] (12).

[0081] Fusion Features The number of channels is 5, including the reverse optical flow and weight map output by the fusion module, which are arrive Optical Flow , arrive Optical Flow And the weight graph .

[0082] In this implementation, the backward warping method is used to obtain the corresponding frame and use the weight map Perform weighted fusion on the predicted frames:

[0083] (13).

[0084] In this implementation, after judging the multi-scale perceptual similarity evaluation, if the similarity does not meet the requirements, the result is fed back to each execution module of the dual-branch optical flow module, and the internal parameters of each execution module are adjusted to adjust the output result.

[0085] The multi-scale perceptual similarity is calculated as follows:

[0086] (14);

[0087] in, and Represent the predicted frame and the true value respectively, and Respectively represent the number of convolutional layers and the number of Laplacian pyramid layers, which are set to 5; is the modulation factor, which is 0.8; is the number of dual-branch optical flow modules, set to 9; Representative A dual-branch optical flow module, Represents the Laplace pyramid layer, Representative convolutional layers.

[0088] Embodiment 2:

[0089] like Figure 3 As shown, this implementation provides a multi-scale video prediction system, including:

[0090] The image acquisition unit is configured to: acquire a previous frame image and a current frame image of the video to be predicted;

[0091] The feature extraction unit is configured to: input the previous frame image and the current frame image into two branches of a dual-branch optical flow module respectively to obtain motion features and spatial features;

[0092] A feature fusion unit is configured to obtain a fusion feature according to the motion feature and the spatial feature, wherein the fusion feature includes a reverse optical flow between a next frame image and a previous frame image, a reverse optical flow between a next frame image and a current frame image, and a weight map;

[0093] The image prediction unit is configured to determine the prediction result of the next frame image according to the reverse optical flow between the next frame image and the previous frame image, the reverse optical flow between the next frame image and the current frame image, and the weight map.

[0094] The specific working process of each of the above units is described in Example 1 and will not be repeated here.

[0095] It is understandable that the above-mentioned units can be separately or completely combined into one or several other units to constitute, or one (some) of the units can be further divided into multiple smaller units in function to constitute, which can achieve the same operation without affecting the realization of the technical effects of the embodiments of the present application. The above-mentioned units are divided based on logical functions. In practical applications, the functions of one unit can also be implemented by multiple units, or the functions of multiple units can be implemented by one unit. In other embodiments of the present application, the system may also include other units. In practical applications, these functions can also be implemented with the assistance of other units, and can be implemented by the collaboration of multiple units.

[0096] According to another embodiment of the present application, the system described in this embodiment can be constructed, and the method of Example 1 of the present application can be implemented by running a computer program (including program code) capable of executing the steps involved in the corresponding method described in Example 1 on a general-purpose computing device such as a computer, which includes processing elements and storage elements such as a central processing unit (CPU), random access memory (RAM), and read-only memory (ROM). The computer program can be recorded on, for example, a computer-readable recording medium, and loaded into the above-mentioned computing device through the computer-readable recording medium and run therein.

[0097] Embodiment 3:

[0098] like Figure 4 As shown, this implementation provides an electronic device, which includes a processor 1001, a communication interface 1002, and a computer-readable storage medium 1003. The processor 1001, the communication interface 1002, and the computer-readable storage medium 1003 may be connected via a bus or other means.

[0099] Among them, the communication interface 1002 is used to receive and send data, the computer-readable storage medium 1003 can be stored in the memory of the electronic device, the computer-readable storage medium 1003 is used to store a computer program, the computer program includes program instructions, and the processor 1001 is used to execute the program instructions stored in the computer-readable storage medium 1003.

[0100] The processor 1001 (or CPU (Central Processing Unit)) is the computing core and control core of the electronic device, which is suitable for implementing one or more instructions, and specifically suitable for loading and executing one or more instructions to implement corresponding method processes or corresponding functions.

[0101] The processor 1001 is configured to execute the following process:

[0102] Obtain the previous frame image and the current frame image of the video to be predicted;

[0103] Inputting the previous frame image and the current frame image into two branches of a dual-branch optical flow module respectively to obtain motion features and spatial features;

[0104] Obtaining fusion features according to the motion features and the spatial features, wherein the fusion features include a reverse optical flow between a next frame image and a previous frame image, a reverse optical flow between a next frame image and a current frame image, and a weight map;

[0105] The prediction result of the next frame image is determined according to the reverse optical flow between the next frame image and the previous frame image, the reverse optical flow between the next frame image and the current frame image, and the weight map.

[0106] The specific process is described in Example 1 and will not be repeated here.

[0107] Embodiment 4:

[0108] This implementation provides a computer-readable storage medium (Memory), which is a memory device in an electronic device for storing programs and data. It is understandable that the computer-readable storage medium here can include both built-in storage media in the electronic device and, of course, extended storage media supported by the electronic device. The computer-readable storage medium provides a storage space that stores the processing system of the electronic device.

[0109] In addition, the storage space also stores one or more instructions suitable for being loaded and executed by the processor, and these instructions may be one or more computer programs (including program codes). It should be noted that the computer-readable storage medium here may be a high-speed RAM memory, or a non-volatile memory, such as at least one disk memory; optionally, it may also be at least one computer-readable storage medium located away from the aforementioned processor.

[0110] In one embodiment, the computer-readable storage medium stores one or more instructions; the processor loads and executes the one or more instructions stored in the computer-readable storage medium to implement the following process:

[0111] Obtain the previous frame image and the current frame image of the video to be predicted;

[0112] Inputting the previous frame image and the current frame image into two branches of a dual-branch optical flow module respectively to obtain motion features and spatial features;

[0113] Obtaining fusion features according to the motion features and the spatial features, wherein the fusion features include a reverse optical flow between a next frame image and a previous frame image, a reverse optical flow between a next frame image and a current frame image, and a weight map;

[0114] The prediction result of the next frame image is determined according to the reverse optical flow between the next frame image and the previous frame image, the reverse optical flow between the next frame image and the current frame image, and the weight map.

[0115] The specific process is described in Example 1 and will not be repeated here.

[0116] Embodiment 5:

[0117] The present implementation provides a computer program product or a computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the electronic device performs the following process:

[0118] Obtain the previous frame image and the current frame image of the video to be predicted;

[0119] Inputting the previous frame image and the current frame image into two branches of a dual-branch optical flow module respectively to obtain motion features and spatial features;

[0120] Obtaining fusion features according to the motion features and the spatial features, wherein the fusion features include a reverse optical flow between a next frame image and a previous frame image, a reverse optical flow between a next frame image and a current frame image, and a weight map;

[0121] The prediction result of the next frame image is determined according to the reverse optical flow between the next frame image and the previous frame image, the reverse optical flow between the next frame image and the current frame image, and the weight map.

[0122] The specific process is described in Example 1 and will not be repeated here.

[0123] A person skilled in the art can appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed in this application can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. A person skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0124] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on the computer, the process or function according to the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted through a computer-readable storage medium. The computer instructions can be transmitted from a website site, a computer, a server or a data center to another website site, a computer, a server or a data center by wired (e.g., coaxial cable, optical fiber, digital line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that a computer can access or a data processing device such as a server, a data center, etc. that includes one or more available media integrated. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive (SSD)), etc.

[0125] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A multi-scale video prediction method, characterized in that: The process includes: Obtain the previous frame image and the current frame image of the video to be predicted; Inputting the previous frame image and the current frame image into two branches of a dual-branch optical flow module respectively to obtain motion features and spatial features; Extraction of motion features, including: extracting residual features from the previous frame image and the current frame image, dividing the residual features into channels, pooling the features of each channel at different scales, extracting multi-scale features and restoring the original size, splicing the features of each channel after pooling and restoring the size, extracting convolution features from the spliced ​​results and performing dot multiplication with the residual features to obtain motion features; Extracting spatial features, including: extracting spatial detail information from the previous frame image and the current frame image, dividing the spatial detail information into two feature streams, respectively extracting features using depthwise separable convolution and dilated convolution, and splicing feature extraction results of the depthwise separable convolution and the dilated convolution to obtain the spatial features; Obtaining fusion features according to the motion features and the spatial features, wherein the fusion features include a reverse optical flow between a next frame image and a previous frame image, a reverse optical flow between a next frame image and a current frame image, and a weight map; According to the motion feature and the spatial feature, a fusion feature is obtained, including: Splicing the motion feature and the spatial feature to obtain a spliced ​​feature; In combination with the splicing feature, a query vector, a key vector and a value vector are obtained by using a depthwise separable convolution, self-attention calculation is performed according to the query vector, the key vector and the value vector, and after maximum pooling is performed on the self-attention calculation result, a weighted calculation is performed with the splicing feature to obtain a fusion feature; The prediction result of the next frame image is determined according to the reverse optical flow between the next frame image and the previous frame image, the reverse optical flow between the next frame image and the current frame image, and the weight map.

2. The multi-scale video prediction method according to claim 1, characterized in that: Query Vector , key vector Sum value vector The calculation method includes: , , ;in, for Depthwise separable convolution, The splicing feature.

3. The multi-scale video prediction method according to claim 1, characterized in that: The weighted calculation is performed with the splicing features, including: ,in, is the result of self-attention calculation, To use the self-attention calculation results The kernel size is used for maximum pooling. is the splicing feature, For fusion features.

4. A multi-scale video prediction system, characterized in that: include: The image acquisition unit is configured to: acquire a previous frame image and a current frame image of the video to be predicted; The feature extraction unit is configured to: input the previous frame image and the current frame image into two branches of a dual-branch optical flow module respectively to obtain motion features and spatial features; Extraction of motion features, including: extracting residual features from the previous frame image and the current frame image, dividing the residual features into channels, pooling the features of each channel at different scales, extracting multi-scale features and restoring the original size, splicing the features of each channel after pooling and restoring the size, extracting convolution features from the spliced ​​results and performing dot multiplication with the residual features to obtain motion features; Extracting spatial features, including: extracting spatial detail information from the previous frame image and the current frame image, dividing the spatial detail information into two feature streams, respectively extracting features using depthwise separable convolution and dilated convolution, and splicing feature extraction results of the depthwise separable convolution and the dilated convolution to obtain the spatial features; The feature fusion unit is configured to: obtain a fusion feature according to the motion feature and the spatial feature, wherein the fusion feature includes a reverse optical flow between a next frame image and a previous frame image, a reverse optical flow between a next frame image and a current frame image, and a weight map; obtain a fusion feature according to the motion feature and the spatial feature, including: Splicing the motion feature and the spatial feature to obtain a spliced ​​feature; In combination with the splicing feature, a query vector, a key vector and a value vector are obtained by using a depthwise separable convolution, self-attention calculation is performed according to the query vector, the key vector and the value vector, and after maximum pooling is performed on the self-attention calculation result, a weighted calculation is performed with the splicing feature to obtain a fusion feature; The image prediction unit is configured to determine the prediction result of the next frame image according to the reverse optical flow between the next frame image and the previous frame image, the reverse optical flow between the next frame image and the current frame image, and the weight map.

5. A computer device, characterized in that: include: a processor and a computer readable storage medium; a processor adapted to execute a computer program; A computer-readable storage medium, wherein a computer program is stored in the computer-readable storage medium, and when the computer program is executed by the processor, the multi-scale video prediction method according to any one of claims 1 to 3 is implemented.

6. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and the computer program is suitable for being loaded by a processor and executing the multi-scale video prediction method according to any one of claims 1 to 3.

7. A computer program product, characterized in that The computer program product comprises a computer program, and when the computer program is executed by a processor, the multi-scale video prediction method according to any one of claims 1 to 3 is implemented.