Video deblurring method based on 3D space-time grid perception, medium and equipment

The 3D spatiotemporal grid-based video deblurring method addresses sudden camera shakes in port operations by using a 3D spatiotemporal network to enhance image clarity and safety in real-time.

CN120318118AActive Publication Date: 2025-07-15BROAD VISION (XIAMEN) TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510802813.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-16
Publication Date
2025-07-15
Estimated Expiration
2045-06-16

AI Technical Summary

Technical Problem

The existing video anti-shake algorithm cannot effectively deal with the instantaneous blur problem caused by violent shaking in port container monitoring, and cannot meet the real-time requirements, affecting operational safety.

Method used

The video defuzzing method based on 3D space-time grid perception is adopted. By receiving a sequence of continuous N-frame video images, the shaking intensity is calculated and the grid blocks are adaptively divided. The 3D space-time network model is used for defuzzing. The fusion weight is calculated based on the pixel position and the center distance of the adjacent grid blocks, and the clear current frame image is output.

Benefits of technology

It realizes efficient real-time blurring of port container monitoring screens, improves operational safety and image clarity, and adapts to the special shaking characteristics of port scenes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318118A_ABST
    Figure CN120318118A_ABST
Patent Text Reader

Abstract

The invention discloses a video deblurring method based on 3D space-time grid perception, a medium and equipment. The method comprises the following steps: firstly, receiving a continuous N-frame video image sequence, and taking [N, H, W, 3] as an input dimension vector; the method comprises the following steps: calculating the shake intensity through the optical flow amplitude of adjacent frames, adaptively dividing each frame of image into S * S overlapped grid blocks, and extracting the grid blocks corresponding to pixel positions to form a [N, S, S, 3] sequence; inputting the grid block sequence into a 3D space-time network model for processing, and outputting deblurred grid blocks; and a fusion weight is calculated based on the distance between the pixel and the center of the adjacent grid block, the grid block is weighted and recombined, and a clear current frame image is output. According to the method, efficient and real-time deblurring is realized by utilizing a 3D space-time network and grid block processing in combination with scene shaking characteristics, the problem of blurring of port container monitoring pictures can be effectively solved, and the operation safety is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing, and particularly to a video deblurring method, medium and device based on 3D spatio-temporal grid perception. Background Art

[0002] In the scenarios of port container transportation and logistics operations, the cameras mounted on the container trolley frame play an important mission of real-time monitoring and recording the whole process of container loading, unloading, stacking, transfer, etc. Compared with the conventional monitoring scenarios, port container monitoring presents significant particularities, which are specifically manifested in the following aspects: Firstly, its scenario has the characteristics of high fixity and periodic repetition. The monitoring cameras in the port container yard usually remain fixedly installed, focusing on the same area for a long time, and the repeatability and predictability of the operation picture content are extremely strong. This characteristic provides the possibility for the system to carry out the memory, learning and prediction of historical scenarios.

[0003] Secondly, the special shaking mode becomes a monitoring problem. During the container operation process, the operations of heavy equipment such as cranes and gantry cranes will cause sudden and large-amplitude shaking. The duration of this kind of shaking is extremely short, generally less than 1 second, but the intensity far exceeds the common persistent slight jitter, and the monitoring picture can be seriously blurred instantly.

[0004] Furthermore, the real-time requirement is extremely strict. The high requirements for operation safety and efficiency in port container operations determine that the monitoring system must process the video stream immediately, and the traditional offline processing method or high-latency solution simply cannot meet the actual needs.

[0005] However, the current monitoring technical means are difficult to cope with these challenges. Most traditional anti-shake algorithms are designed based on persistent or slow shaking scenarios, and they are powerless when dealing with the instantaneous blurring problems caused by short-term severe shaking in the port container operation environment. The existing general deblurring technologies not only lack the targeted processing ability for the periodic short-term severe shaking in the port container yard, but also cannot make full use of the advantages of scene fixity and predictability; there are also shortcomings in terms of real-time performance, with a relatively high processing delay, which seriously affects the effectiveness of operation safety monitoring. Summary of the Invention

[0006] In view of the above problems, the present invention provides a video deblurring method, medium and device based on 3D spatio-temporal grid perception, so as to solve the technical problems that the existing video anti-shake algorithms cannot be applied to the anti-shake processing of video pictures with severe shaking, cannot meet the application in the port container monitoring scenario, and are prone to missing monitoring pictures and affecting safe operations.

[0007] To achieve the above object, in the first aspect, the present application provides a video deblurring method based on 3D spatio-temporal grid perception, and the method includes the following steps: Receive a sequence of N consecutive video images, and the input dimension vector is [N, H, W, 3], where N represents the number of input video image frames, H and W respectively represent the height and width, and 3 represents the RGB channels; Calculate the shaking intensity through the optical flow amplitude of adjacent frames, adaptively divide each video image into overlapping grid blocks of size S×S, and for the pixel position (i, j) in each video image, extract the grid block corresponding to this pixel position from N video images to form a grid block sequence with the shape of [N, S, S, 3], where N is a positive integer greater than 1, N - 1 frames are clear frame images, and the Nth frame image is the current frame image that needs to be deblurred; Input the grid block sequence into the trained 3D spatio-temporal network model for block processing, and output the deblurred grid blocks; Calculate the fusion weight based on the distance between the pixel position and the center of the adjacent grid block, perform weighted fusion recombination on the deblurred grid blocks, and output the clear current frame image.

[0008] Further, the 3D spatio-temporal network model includes a model generator, and the model generator includes: A 3D spatio-temporal feature extraction module that converts the input [N, S, S, 3] grid block sequence into a [1, 3, N, S, S] dimension to adapt to the 3D convolution input; A shaking perception encoder that extracts multi-scale features through three layers of downsampling convolution and outputs them to be saved in the skip connection layer; A residual block module that cascades 6 residual blocks, and each residual block contains two 3×3 convolutional layers and an identity skip connection; A spatio-temporal context decoder that performs upsampling through transposed convolution and uses skip connections to obtain the multi-scale features of the shaking perception encoder; A detail enhancement output layer that reduces the number of feature channels from 64 to 3 through a 3×3 convolution and outputs an RGB image with a range of [-1, 1] through the Tanh activation function.

[0009] Further, the shaking perception encoder includes a first encoding layer, a second encoding layer, and a third encoding layer. The input of the first encoding layer is a feature map with the original resolution, the input of the second encoding layer is a feature map with 2 times downsampling, and the third encoding layer is a feature map with 4 times downsampling; The output results of the first encoding layer and the second encoding layer are transmitted to the spatio-temporal context decoder through the skip connection layer, and the output result of the third encoding layer is transmitted to the residual block module for processing after adjusting the size through bilinear interpolation.

[0010] Further, the format of the RGB image output by the detail enhancement output layer is [batch, S, S, 3], where batch represents the output block batch, S represents the block size, and 3 represents the RGB channels.

[0011] Further, the 3D spatio-temporal network model includes a model discriminator, and the model discriminator includes: A 3D spatio-temporal feature perception module that combines the grid block sequence with the generator output image into a tensor of shape [1, 3, N + 1, S, S], applies 3D convolution with a convolution kernel size of (N + 1, 4, 4) and a stride of (1, 2, 2), and outputs a feature map of shape [1, 64, S / 2, S / 2], where S is the grid block size and is an even number; A 2D block GAN discrimination chain that performs multi-layer convolutional downsampling on the feature map and outputs a feature map of true and false degrees.

[0012] Further, the activation function of the 3D convolution is LeakyReLU.

[0013] Further, the 3D spatio-temporal network model is trained according to the following method: Obtain a sample image set, obtain a sample grid block sequence based on the sample image set, input the sample grid block sequence into the 3D spatio-temporal network model to be trained, and perform iterative training to obtain a trained 3D spatio-temporal network model.

[0014] Further, the sample image set includes a first sample image set, a second sample image set, and a third sample image set; The first sample image set is obtained according to the following method: Collect actual monitoring video streams from multiple port container yards, extract all video frame images of each video stream, calculate the clarity score of each video frame image, filter out the video frame images with an average clarity score higher than a preset clarity threshold, and extract a continuous video frame sequence from the filtered video frame images as the first sample image set; The second sample image set is obtained according to the following method: Set a linear shaking model, a jitter shaking model, and a combined shaking model. The linear shaking model includes a linear shaking blur kernel, the jitter shaking model includes a jitter shaking blur kernel, and the combined shaking model includes a linear shaking blur kernel and a jitter shaking blur kernel; Input the first sample image set into the linear shaking model, the jitter shaking model, and the combined shaking model respectively to obtain the second sample image set; The third sample image set is obtained according to the following method: Perform transformation processing on the images in the first sample image set and the second sample image set to obtain a third sample image set, where the transformation processing includes geometric transformation, illumination transformation, and adding noise.

[0015] In a second aspect, the present application provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the video deblurring method based on 3D spatio-temporal grid perception as described in the first aspect of the present application.

[0016] In a third aspect, the present application provides an electronic device, on which a computer program is stored, including a processor and a storage medium. The computer program stored on the storage medium, when executed by the processor, implements the video deblurring method based on 3D spatio-temporal grid perception as described in the first aspect of the present application.

[0017] Different from the prior art, the above technical solution provides a video deblurring method, medium, and device based on 3D spatio-temporal grid perception. The method first receives a continuous sequence of N video image frames with an input dimension vector of [N, H, W, 3]; calculates the shaking intensity through the optical flow amplitude of adjacent frames, adaptively divides each frame image into S×S overlapping grid blocks, and extracts the grid blocks corresponding to the pixel positions to form a sequence of [N, S, S, 3]. Input the grid block sequence into a 3D spatio-temporal network model for processing to output deblurred grid blocks; then calculate the fusion weights based on the distance between the pixel and the center of the adjacent grid blocks, and weighted recombine the grid blocks to output a clear current frame image. This method uses 3D spatio-temporal network and grid block processing, combined with the shaking characteristics of the scene, to achieve efficient and real-time deblurring, which can effectively solve the problem of blurred monitoring images of port containers and improve operation safety.

[0018] The above relevant descriptions of the invention content are only an overview of the technical solution of the present invention. In order to enable those of ordinary skill in the art to more clearly understand the technical solution of the present invention, and then implement it according to the content recorded in the description and the drawings, and in order to make the above objects, other objects, features, and advantages of the present invention more easily understood, the following is described in conjunction with the specific embodiments and drawings of the present invention. Description of the Drawings

[0019] The drawings are only used to illustrate the principles, implementation methods, applications, features, and effects of the specific embodiments of the present invention and other related contents, and should not be considered as a limitation of the present invention.

[0020] In the drawings of the description: Figure 1 It is the first flowchart of the video deblurring method based on 3D spatio-temporal grid perception involved in the specific embodiment; Figure 2 It is a schematic diagram of the spatio-temporal grid perception architecture involved in the specific embodiment; Figure 3 It is the second flowchart of the video deblurring method based on 3D spatio-temporal grid perception involved in the specific implementation manner; Figure 4 It is a schematic diagram of the distribution of grid blocks identified in the port monitoring screen involved in the specific implementation manner; Figure 5 It is the schematic principle diagram of 3D spatio-temporal grid perception for processing grid blocks involved in the specific implementation manner; Figure 6 It is the working flowchart of the model generator involved in the specific implementation manner; Figure 7 It is the working flowchart of the 3D spatio-temporal feature extraction module involved in the specific implementation manner; Figure 8 It is the working flowchart of the jitter perception encoder involved in the specific implementation manner; Figure 9 It is the working flowchart of the residual block module involved in the specific implementation manner; Figure 10 It is the working flowchart of a single residual block involved in the specific implementation manner; Figure 11 It is the working flowchart of the spatio-temporal context decoder involved in the specific implementation manner; Figure 12 It is the working flowchart of the detail enhancement output layer involved in the specific implementation manner; Figure 13 It is the working flowchart of the model discriminator involved in the specific implementation manner; Figure 14 It is the flowchart of self-supervised data generation involved in the specific implementation manner; Figure 15 It is the flowchart of constructing the port jitter pattern library involved in the specific implementation manner; Figure 16 It is the flowchart of the training process of the model generator involved in the specific implementation manner; Figure 17 It is the flowchart of the training process of the model discriminator involved in the specific implementation manner; Figure 18 It is the module schematic diagram of the electronic device described in the specific implementation manner; The descriptions of the reference numerals involved in the above-mentioned various drawings are as follows: 10. Electronic device; 101. Processor; 102. Storage medium. Specific implementation manner

[0021] To illustrate in detail the possible application scenarios, technical principles, specific implementable solutions, achievable objectives and effects of the present invention, the following will be described in detail with reference to the specific examples listed and in conjunction with the accompanying drawings. The embodiments described herein are only used to more clearly illustrate the technical solutions of the present invention, and thus are only examples and cannot be used to limit the protection scope of the present invention.

[0022] Reference to "embodiment" herein means that a particular feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of the present invention. The term "embodiment" appearing at various positions in the specification does not necessarily refer to the same embodiment, nor does it particularly limit its independence or relevance to other embodiments. In principle, in the present invention, as long as there is no technical contradiction or conflict, the various technical features mentioned in each embodiment can be combined in any way to form corresponding implementable technical solutions.

[0023] Unless otherwise defined, the meanings of the technical terms used herein are the same as those commonly understood by those skilled in the technical field to which the present invention belongs; the use of the relevant terms herein is only for describing specific embodiments and is not intended to limit the present invention.

[0024] In the description of the present invention, the term "and / or" is an expression used to describe the logical relationship between objects, indicating that three relationships can exist. For example, A and / or B means: there is A, there is B, and there is both A and B at the same time. In addition, the character " / " herein generally represents an "or" logical relationship between the associated objects before and after.

[0025] In the present invention, terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual quantity, primary-secondary or order relationship, etc. between these entities or operations.

[0026] Without further limitation, in the present invention, the use of the terms "include", "comprise", "have" or other similar expressions in a statement is intended to cover non-exclusive inclusion. These expressions do not exclude the possibility that there may be additional elements in the process, method or product including the said elements, so that a process, method or product including a series of elements may include not only those defined elements, but also other elements not explicitly listed, or elements inherent to such process, method or product.

[0027] In the present invention, expressions such as "greater than", "less than", "exceeding", etc. are understood not to include the number itself; expressions such as "above", "below", "within", etc. are understood to include the number itself. In addition, in the description of the embodiments of the present invention, the meaning of "a plurality of" is two or more (including two), and similar expressions related to "many" are also understood in this way, such as "multiple groups", "multiple times", etc., unless otherwise clearly and specifically defined.

[0028] In the description of the embodiments of the present invention, the spatially related expressions used, such as "center", "longitudinal", "lateral", "length", "width", "thickness", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "perpendicular", "top", "bottom", "inner", "outer", "clockwise", "counterclockwise", "axial", "radial", "circumferential", etc., the indicated orientation or positional relationship is based on the orientation or positional relationship shown in the specific embodiment or the drawing, and is only for the convenience of describing the specific embodiment of the present invention or facilitating the understanding of the reader, rather than indicating or implying that the device or component referred to must have a specific position, a specific orientation, or be constructed or operated in a specific orientation, and therefore cannot be understood as a limitation to the embodiments of the present invention.

[0029] Unless otherwise clearly specified or limited, in the description of the embodiments of the present invention, the terms "installed", "connected", "connected", "fixed", "set", etc. used shall be understood in a broad sense. For example, the "connection" may be a fixed connection, a detachable connection, or an integral setting; it may be a mechanical connection, an electrical connection, or a communication connection; it may be directly connected, or indirectly connected through an intermediate medium; it may be the communication inside two components or the interaction relationship between two components. For those skilled in the technical field to which the present invention belongs, the specific meanings of the above terms in the embodiments of the present invention can be understood according to specific circumstances.

[0030] As Figures 1-17 shown, in the first aspect, the present application provides a video deblurring method based on 3D spatio-temporal grid perception, and the method includes the following steps: Step S1: Receive a sequence of N consecutive video images, and the input dimension vector is [N, H, W, 3], where N represents the number of input video images, H and W respectively represent the height and width, and 3 represents the RGB channels; Step S2: Calculate the shaking intensity through the optical flow amplitude of adjacent frames, adaptively divide each frame of video image into overlapping grid blocks of size S×S, and for the pixel position (i, j) in each video image, extract the grid block corresponding to this pixel position from N frames of video images to form a grid block sequence with the shape of [N, S, S, 3], where N is a positive integer greater than 1, N - 1 frames are clear frame images, and the Nth frame image is the current frame image that needs to be deblurred. Step S3: Input the grid block sequence into the trained 3D spatio-temporal network model for block processing, and output the de-blurred grid blocks. Step S4: Calculate the fusion weights based on the pixel position and the distance from the center of adjacent grid blocks, and perform weighted fusion and recombination on the de-blurred grid blocks to output a clear current frame image.

[0031] The above method uses 3D spatio-temporal network and grid block processing, combined with the scene shaking characteristics, to achieve efficient real-time de-blurring, which can effectively solve the problem of blurred monitoring images of port containers and improve operation safety.

[0032] As Figure 2 shown, taking the monitoring video stream image of port container yard operation as an example, the spatio-temporal grid perception of the present application includes an input layer, a processing layer, an output layer and a training auxiliary module. The input information of the input layer is a multi-frame input monitoring video sequence of the port. The processing layer includes a preprocessing module and a generator network. The preprocessing module is used for frame alignment and grid segmentation. The generator network adopts a 3D-TSPatchNet architecture. When the network model is trained, the sample images can be obtained from the port scene video library. The discriminator network is a 3D-TSPatchGAN architecture, and the output of the model is a clear port image after de-blurring processing.

[0033] As Figure 3 shown, when the model receives a video image sequence of continuous N frames, first judge whether there is a video frame with a large shaking intensity in the current video image sequence through the shaking warning and detection module. If not, directly continue to output the original monitoring picture. If so, perform preprocessing and de-blurring processing through steps S1-S4, and then restore the clear current frame image. After that, replace the original current frame image with the clear current frame image and output the video monitoring picture.

[0034] In this embodiment, before performing the de-blurring processing of the current frame, it is necessary to complete the preprocessing of the input data and the grid division. Taking the port monitoring video as an example, the input data processing flow is as follows: (1) Original video input: First, receive a complete port monitoring video sequence, which contains continuous N frames (the value of N is preferably 3), then the original input dimension can be expressed as [N, H, W], where H and W respectively represent the height and width of the original frame, and the interval between each frame is about 33 ms (based on a 30 fps video).

[0035] (2) Adaptive grid division: According to the detected shaking intensity, divide each frame image into a series of overlapping grid blocks. The standard grid block size is S×S (such as 64×64 or 128×128 pixels), and the number of overlapping pixels of adjacent grid blocks is dynamically adjusted according to the shaking intensity (the standard is S / 4, such as 16 or 32 pixels).

[0036] (3) Grid block extraction and organization: For each position (i, j) in the video sequence, grid blocks corresponding to the position are extracted from N frames to form a grid block sequence: a tensor with the shape of [N, S, S, 3]. For example, assuming S = 64, the input for a single position is [3, 64, 64, 3], representing a sequence of 3 RGB images of 64×64 pixels.

[0037] (4) Grid block processing characteristics: In each grid block sequence, the first N - 1 frames (such as the first 2 frames) are usually relatively clear frames, and the last frame is the current frame to be deblurred. There are overlapping regions between adjacent grid blocks to ensure smooth fusion of the processing results. All grid blocks can be input into the 3D-TSPatchNet model in parallel for processing.

[0038] (5) Grid block preprocessing, which includes operations such as normalization, data format adjustment, and batch processing organization. Normalization means mapping the pixel values in the grid from [0, 255] to the range [-1, 1]. Data format adjustment is to adjust [N, S, S, 3] to the input format required by the model network. Batch processing organization is to organize grid blocks at multiple positions into batches and synchronously input them into the model network to improve processing efficiency.

[0039] Through this grid division and processing method, the 3D-TSPatchNet model can efficiently process blurred images with short-term severe shaking in port surveillance videos, restore clear surveillance images, and keep the computational resource requirements within a reasonable range. The divided grid images are as Figure 4 shown.

[0040] In some embodiments, the 3D spatio-temporal network model includes a model generator, and the model generator includes: A 3D spatio-temporal feature extraction module that converts the input [N, S, S, 3] grid block sequence into a [1, 3, N, S, S] dimension to adapt to 3D convolution input; A shaking perception encoder that extracts multi-scale features through three layers of downsampling convolution and saves the output to the skip connection layer; A residual block module that concatenates 6 residual blocks, and each residual block contains two 3×3 convolution layers and an identity skip connection; A spatio-temporal context decoder that performs upsampling through transposed convolution and uses skip connections to obtain the multi-scale features of the shaking perception encoder; A detail enhancement output layer that reduces the number of feature channels from 64 to 3 through 3×3 convolution and outputs an RGB image in the range [-1, 1] through the Tanh activation function.

[0041] In this embodiment, the model generator receives a multi-frame sequence of a single grid block Patch, processes spatio-temporal information, and outputs the corresponding de-blurred clear image. The shape of the multi-frame sequence of a single grid block is represented as [N, S, S, 3], and the value range is [-1, 1]. For example, [3, 64, 64, 3] represents a sequence of 3 RGB images of 64×64, where the first N - 1 frames are clear frame images, and the last frame is the current frame image to be de-blurred. The output of the model is the de-blurred image, with the shape represented as [S, S, 3], and the value range is the same as the input. The output is the clear version of the grid block corresponding to the last frame (the current frame image). Each grid block sequence input to the 3D-TSPatchNet comes from the front and back frame videos, in the form of Figure 5 as shown.

[0042] As Figure 6 shown, the present invention innovatively designs a 3D spatio-temporal perception network (3D-TSPatchNet), which is completely different from the traditional 2D U-Net architecture. A series of key optimizations and innovations have been carried out specifically for the de-blurring task of short-term severe shaking of port containers. This generator network captures information in both time and space dimensions through 3D convolution, and combines the grid block processing method to achieve efficient spatio-temporal feature extraction, multi-scale feature learning, residual learning, and spatio-temporal skip connection. The overall process is as follows: (1) 3D spatio-temporal feature fusion design: The traditional U-Net is mainly used for image segmentation tasks and only processes 2D information of a single image. The present invention has carried out a fundamental innovation by creating a new 3D-TSPatchNet architecture, enabling it to process features in three dimensions of time and space simultaneously, forming a real spatio-temporal three-dimensional fusion architecture. Specifically, it is manifested as: designing a special 3D-2D transition layer to retain temporal information while being compatible with the U-Net structure; introducing a transmission mechanism of temporal features in skip connections to ensure that temporal information is not lost in the deep network; adjusting the downsampling ratio in the encoder-decoder structure to adapt to the scale characteristics of shaking blur during the loading and unloading process of port containers.

[0043] (2) Shaking characteristic perception mechanism: The traditional U-Net lacks the ability to model motion characteristics. The present invention specifically designs a shaking characteristic perception component for port container operations. Specifically, it is manifested as: introducing a motion characteristic analysis branch in the encoder stage to extract the direction, intensity, and mode of the shaking of port container handling; designing a fusion mechanism of shaking characteristics and image features to enable the network to adjust the de-blurring strategy according to the unique shaking characteristics of the port; retaining shaking information during downsampling to ensure that deep features still contain motion clues caused by port equipment.

[0044] (3) Multi-scale Residual Learning: Innovatively combines residual learning and the U-Net architecture, enabling the network to more effectively learn the blur transformation caused by shaking in port container operations. Specifically: A special multi-scale residual block is designed to simultaneously handle the shaking blur of different scales of port containers; the residual connection structure is optimized to better suit the characteristics of short-term and intense shaking in the port environment; a channel attention mechanism is introduced to enhance the ability to extract key features of port containers.

[0045] (4) Spatio-temporal Context-aware Skip Connection: The skip connection of the traditional U-Net only transmits spatial information. The present invention designs an enhanced skip connection that can transmit spatio-temporal context information. Specifically: A special spatio-temporal feature fusion module is designed to simultaneously transmit the spatial and temporal information of the port scene in the skip connection; a learnable spatio-temporal feature selection mechanism is introduced to dynamically determine which features are more helpful for deblurring the port container monitoring video; the feature fusion algorithm is optimized to reduce the computational overhead while improving the information transmission efficiency.

[0046] (5) Lightweight and Efficient Design: To achieve real-time processing in the port environment, a series of lightweight innovations are made to the architecture. Specifically: Depthwise separable convolutions are used to replace some standard convolutions, significantly reducing the number of parameters and the computational amount; a balance mechanism between computational cost and effect is designed to allocate computational resources according to the importance of different layers; through the optimized design of the network depth and width, the computational complexity is minimized while ensuring the deblurring effect; all grid blocks share the same network weights, greatly reducing the number of parameters and memory occupancy.

[0047] The following elaborates on each module generated by the model in detail: First is the 3D spatio-temporal feature extraction module. As Figure 7 shown, this module is used to capture the spatio-temporal relationship and shaking pattern features of short-term and intense shaking sequences in the port container yard monitoring. In this application, the input multi-frame sequence adopts a special dimensional arrangement, converting the input with the shape of [N, S, S, 3] into the form of [1, 3, N, S, S], enabling the 3D convolution to simultaneously process time and space information. Here, N represents the number of frames (preferably 3 frames), S represents the grid block size, 3 represents the RGB channels, and 1 represents the batch size, indicating the current one sample (a multi-frame sequence). This arrangement enables the network to better capture the features of short-term and intense shaking in port container operations.

[0048] In addition, this application also designs a special 3D convolutional kernel with a shape of (n_frames, 3, 3). Different from traditional 2D convolution or sequential scanning methods, this design can perform feature extraction simultaneously in the time and space dimensions. The number of input channels is 3 (RGB), and the number of output channels is 64, achieving effective compression of information and enhancement of features. The relevant parameter settings are as follows: the stride is 1, the padding in the time dimension is 0 (to capture complete temporal information), and the padding in the space dimension is 1 (to retain edge features).

[0049] In this application, whether the video frame image is blurred can be determined by a clarity-blur information fusion module, which is a sub-module of the 3D spatio-temporal feature extraction module. This module can effectively distinguish and utilize the clarity and blur frame information in the input sequence, and through feature interaction in the time dimension, the port container structure information in the clarity frame can guide the restoration process of the blur frame. This clarity-blur information fusion is the key advantage of this invention over traditional single-frame deblurring methods.

[0050] Secondly, there is a jitter perception encoder, as Figure 8 shown. This encoder extracts deeper-level features layer by layer through a multi-layer downsampling convolutional network, while reducing the spatial resolution to provide multi-scale feature representations for subsequent processing.

[0051] Preferably, the jitter perception encoder includes a first coding layer, a second coding layer, and a third coding layer. The input of the first coding layer is a feature map with the original resolution, the input of the second coding layer is a feature map with 2-fold downsampling, and the third coding layer is a feature map with 4-fold downsampling.

[0052] The output results of the first coding layer and the second coding layer are passed to the spatio-temporal context decoder through a skip connection layer, and the output result of the third coding layer is passed to the residual block module for processing after being adjusted in size by bilinear interpolation.

[0053] Specifically, the input of the first coding layer (i.e., Figure 8 coding layer 1 in) is a feature map of [batch, 64, S, S]. The number of channels is expanded from 64 to 128 through a 3×3 convolutional kernel (stride 1, padding 1). After batch normalization to standardize the feature distribution and accelerate training, and then through the ReLU activation function to introduce non-linear transformation ability, a feature map of [batch, 128, S, S] with the same output resolution is output, and is saved to the decoder layer through a skip connection to achieve feature reuse.

[0054] The second coding layer (i.e., Figure 8The input of the encoding layer 2) is a feature map of [batch, 128, S, S]. The number of channels is expanded from 128 to 256 through a 3×3 convolutional kernel (stride 2, padding 1). After batch normalization to standardize the feature distribution and accelerate training, and then through the ReLU activation function to introduce non-linear transformation ability, a feature map of [batch, 256, S / 2, S / 2] with half the spatial size is output. This feature map is saved to the decoder layer through skip connections to achieve cross-level feature fusion and enhance the ability to restore details.

[0055] The third encoding layer (i.e., Figure 8 the encoding layer 3) in has an input feature map of [batch, 256, S / 2, S / 2]. A spatial downsampling operation is performed through a 3×3 convolutional kernel (stride 2, padding 1), and the number of channels is expanded from 256 to 512 to enhance the feature expression ability; after passing through the batch normalization layer to standardize the data distribution and accelerate model convergence, and then through the ReLU activation function to introduce non-linear transformation ability, a feature map with a spatial size halved to [batch, 512, S / 4, S / 4] is output. This feature map is passed to the subsequent residual block processing stage for deep feature extraction and fusion to improve the model's ability to model complex patterns.

[0056] As Figure 9 shown, the residual block module is set between the encoder and the decoder. Through multi-layer residual learning, the network's ability to express complex shaking patterns in port container operations is enhanced, and at the same time, the gradient vanishing problem of deep networks is effectively alleviated through residual connections.

[0057] Specifically, when configuring the residual block, in the deep feature extraction stage, the optimization balance between complexity and performance is achieved by serially configuring 6 residual blocks. Among them, as Figure 10 shown, each residual block contains a 3×3 convolutional layer, batch normalization, and the ReLU activation function. Through the above settings, the following beneficial effects can be achieved: (1) Complexity balance: Experimental verification shows that when the number of residual blocks is less than 4, the ability to model complex shaking patterns in port container operations is insufficient. When it exceeds 8, the computational cost increases significantly while the performance improvement is less than 5%. 6 residual blocks achieve the optimal feature expression ability while maintaining computational efficiency.

[0058] (2) Receptive field matching: The concatenation of 6 3×3 convolutional layers expands the theoretical receptive field of the network to 13×13 pixels, which precisely matches the typical blur kernel size (10 - 15 pixels) of short-term violent shaking caused by port container handling equipment, improving the pertinence of motion blur removal.

[0059] (3)Gradient stability: The residual connection structure allows gradients to propagate stably in deep networks. Tests show that for a 6-layer depth, without introducing additional regularization measures, the variance of the training loss convergence curve is reduced by 40%, effectively suppressing the vanishing gradient problem.

[0060] Residual blocks provide a shortcut for gradient backpropagation through identity mapping, enabling effective training even when the depth reaches 6 residual blocks. The network model learns the residual part rather than the complete mapping, making it easier for the model to approximate the ideal deblurring function. Special optimization for the jitter scenario in port container operations enables the model to more effectively handle different types of camera jitters, including random jitter, linear motion, and rotational motion, etc.

[0061] As Figure 11 shown, the decoder module gradually restores the spatial resolution through transposed convolution (deconvolution) operations, while using spatio-temporal context-aware skip connections to obtain multi-scale feature information from the encoder to achieve high-quality deblurring reconstruction. The skip connections in the traditional U-Net only transmit spatial information. The present invention designs enhanced skip connections that can transmit spatio-temporal context information. Information about the port container scenario can be effectively transmitted between different levels through these skip connections, ensuring that detailed information is not lost during deep network processing. Each skip connection contains a feature channel attention mechanism that can intelligently select and weight encoder features, enabling the information most useful for restoring port scene details to be preferentially utilized.

[0062] As Figure 12 shown, the detail enhancement output layer converts the feature map output by the decoder into the final RGB image, with particular attention to restoring the edge and texture details of port containers. The format of the RGB image output by the detail enhancement output layer is [batch, S, S, 3], where batch represents the output block batch, S represents the block size, and 3 represents the RGB channels.

[0063] The detail enhancement output layer reduces the number of feature channels from 64 to 3 (corresponding to the three RGB channels) through a 3×3 convolutional kernel, constrains the output value range to [-1, 1] through the Tanh activation function to achieve color dynamic balance, and then reconstructs the feature map into the standard RGB format [batch, S, S, 3] through a dimension adjustment operation. This design significantly improves the clarity and structural integrity of port container images by retaining high-frequency edge details and texture features.

[0064] In some embodiments, as Figure 13 shown, the 3D spatio-temporal network model includes a model discriminator, and the model discriminator includes: The 3D spatio-temporal feature perception module combines the grid block sequence with the generator output image into a tensor of shape [1, 3, N + 1, S, S], applies 3D convolution with a convolution kernel size of (N + 1, 4, 4) and a stride of (1, 2, 2), and outputs a feature map of shape [1, 64, S / 2, S / 2], where S is the grid block size and is an even number; The 2D block GAN discriminant chain performs multi-layer convolutional downsampling on the feature map and outputs a feature map of true and false degrees.

[0065] To solve the key defect that traditional 2D PatchGAN cannot model temporal dimension information, this application designs a 3D spatio-temporal feature perception module, which is used to capture the temporal correlation features between the port container monitoring input sequence (N frames) and the output image. The specific implementation process includes: (1) Combine the input sequence and the output image into a tensor of dimension [1, 3, N + 1, S, S]; (2) Apply 3D convolution operation with a convolution kernel size of (N + 1, 4, 4) and a stride of (1, 2, 2); (3) Through complete compression in the time dimension and spatial downsampling (stride 2), the output dimension is converted to [1, 64, S / 2, S / 2]; (4) Adopt the LeakyReLU activation function with a negative slope coefficient α = 0.2 to significantly enhance the network's sensitivity to the shaking features of the port scene and the stability of gradient propagation.

[0066] The 2D block GAN discriminant chain generates feature maps with decreasing spatial resolution through multi-layer convolutional downsampling to quantify the authenticity probability of the port scene image. The specific hierarchical structure is as follows: The first layer (spatial downsampling): The input dimension is [batch, 64, S / 2, S / 2]. A 4×4 convolution kernel and a stride of 2 are used to expand the number of channels from 64 to 128, and the output spatial resolution is reduced to 1 / 4 of the original size (i.e., [batch, 128, S / 4, S / 4]). This layer integrates batch normalization and the LeakyReLU activation function with a negative slope coefficient α = 0.2 to suppress the gradient disappearance problem in the low-light port scene.

[0067] The second layer (depth feature extraction): The input dimension is [batch, 128, S / 4, S / 4]. A 4×4 convolution kernel and a stride of 2 are used to increase the number of channels from 128 to 256, and the output resolution is further reduced to 1 / 8 ([batch, 256, S / 8, S / 8]). The batch normalization layer alleviates the feature distribution shift caused by port rain and fog interference during the training process, and LeakyReLU enhances the sensitivity to the container displacement features.

[0068] The third layer (channel expansion and feature retention): The input dimension is [batch, 256, S / 8, S / 8]. A 4×4 convolutional kernel with a stride of 1 and padding of 1 is used to expand the number of channels from 256 to 512, and the output maintains the resolution of [batch, 512, S / 8, S / 8] unchanged. This design avoids excessive loss of spatial information and retains high-frequency texture information for subsequent local detail discrimination.

[0069] The output layer (discriminant score generation): The input dimension is [batch, 512, S / 8, S / 8]. The number of channels is reduced from 512 to 1 through a 4×4 convolutional kernel with a stride of 1 and padding of 1, and the output is the original logits matrix with the dimension of [batch, 1, S / 16, S / 16]. The design without an activation function ensures that the output value range is not limited. Each spatial point corresponds to the authenticity probability of a local area (such as the corner fittings of a container, the edge of the container door), realizing pixel-level quality assessment.

[0070] The above solution proposes a 3D-TSPatchGAN discriminator dedicated to port monitoring, which realizes the joint discrimination of the input sequence and the output image through a 3D+2D hybrid architecture (3D convolution compresses the temporal dimension + 2D convolution processes spatial features). This design can effectively capture the dynamic temporal correlations in container operations (such as the continuity of the lifting trajectory), identify the temporal inconsistencies between the generated frames and the monitoring sequence, and significantly improve the temporal coherence; at the same time, a local grid discrimination mechanism (output resolution S / 16×S / 16) is adopted to focus on the edge texture of the container and the reflective details of the metal surface, enhancing the detection sensitivity to blurring residues and artifacts and improving the accuracy of container number recognition. In terms of lightweight, the number of parameters is much less than that of the global discriminator, meeting the latency requirements for real-time inspection of damaged containers at the port gate; its training stability is enhanced through batch normalization and local discrimination strategies, effectively suppressing the risk of mode collapse, and supporting dynamic adjustment of the input size (such as S = 512 / 1024), with small fluctuations in cross-resolution task performance.

[0071] In some embodiments, the 3D spatio-temporal network model is trained according to the following method: Obtain a sample image set, obtain a sample grid block sequence based on the sample image set, input the sample grid block sequence into the 3D spatio-temporal network model to be trained, and perform iterative training to obtain a trained 3D spatio-temporal network model.

[0072] Preferably, the sample image set includes a first sample image set, a second sample image set, and a third sample image set; The first sample image set is obtained according to the following method: Collect the actual monitoring video streams from multiple port container yards, extract all video frame images of each video stream, calculate the clarity score of each video frame image, filter out the video frame images with an average clarity score higher than the preset clarity threshold, and extract a continuous video frame sequence from the filtered video frame images as the first sample image set; The second sample image set is obtained according to the following method: Set a linear shaking model, a jitter shaking model, and a combined shaking model. The linear shaking model includes a linear shaking blur kernel, the jitter shaking model includes a jitter shaking blur kernel, and the combined shaking model includes a linear shaking blur kernel and a jitter shaking blur kernel; Input the first sample image set into the linear shaking model, the jitter shaking model, and the combined shaking model respectively to obtain the second sample image set; The third sample image set is obtained according to the following method: Perform transformation processing on the images in the first sample image set and the second sample image set to obtain the third sample image set. The transformation processing includes geometric transformation, illumination transformation, and adding noise.

[0073] This solution uses a dual-path dataset construction method to collect the samples required for model training: (1) Simulated data augmentation: Extract key frames from clear port monitoring videos, generate noisy samples by synthesizing various motion blur models (including typical interferences such as ship shaking and crane displacement), and construct fuzzy-clear frame pairs with time series matching; (2) Physical anti-shake comparison collection: Deploy a camera array equipped with mechanical optical image stabilization (OIS), synchronously capture the clear images after anti-shake processing and the original jittery images in the port operation scenario, and form accurate fuzzy-clear paired samples verified by physical mechanisms.

[0074] The collection process of the real scene data of the port container yard is as follows: Collect the actual monitoring videos from multiple port container yards, traverse the port monitoring video files in the dataset, extract all frame images of each video, calculate the clarity score of each frame (using methods such as gradient magnitude and frequency domain analysis), filter out the videos whose average clarity score of each included frame image is higher than the preset clarity threshold, and extract a continuous video sequence from the filtered videos as the basic training data.

[0075] The method for calculating the clarity score of video frame images is as follows: First, perform gradient magnitude evaluation , and the formula is as follows: ; Among them, is the height of the image (number of pixel rows), is the width of the image (number of pixel columns), is the row and the gray value (or brightness) of the pixel in the column, indicating the gradient of the image in the horizontal direction ( axis), indicating the gradient of the image in the vertical direction ( axis).

[0076] Secondly, the high-frequency energy ratio is calculated, and the formula is as follows: ; where is the frequency-domain representation (complex value) of the image after Fourier transform, is the coordinate in the frequency domain. The high-frequency region: usually refers to the region near the spectrum edge, corresponding to the detailed information of the image.

[0077] Then, the contrast evaluation is performed, and the formula is as follows: ; where is the mean value of the image gray value, is the standard deviation of the image gray value, is the grayscale image.

[0078] Finally, the comprehensive sharpness score is calculated, and the formula is as follows: , where are the weight coefficients of the three evaluation indicators respectively, satisfying (usually set by experience or optimized through training); , , are the gradient amplitude evaluation, high-frequency energy ratio, and contrast evaluation indicators calculated previously respectively; is the final comprehensive sharpness score of the image.

[0079] Of course, the model can also use the real-scene data of the port container yard to complete the self-supervised generation of data, and the generation method is as Figure 14 shown. And it can perform transformation processing based on the real-scene data and the data self-supervised generated by the model to obtain a new set of sample images, and the transformation processing includes geometric transformation, illumination transformation, and adding noise, etc.

[0080] This application specifically designs two special blur kernels to simulate the sudden shaking of the port container monitoring system: Linear motion blur kernel: Given the blur length parameter L and the angle parameter θ, a kernel matrix of size (2L+1)×(2L+1) is constructed. Taking the center of the matrix as the origin, L pixels are extended on both sides along the angle direction of θ, and the pixel value covered by the track is set to 1. The kernel matrix is generated after normalization processing to simulate the uniform linear motion smear caused by the translation operation of the port crane.

[0081] Jitter blur kernel: Given the jitter intensity I and complexity parameter C, C random connection points are generated from the center point of the kernel matrix to form a jitter trajectory, and the trajectory coverage pixel value is set to 1; a Gaussian filter with a standard deviation of I is applied to the trajectory matrix to simulate the irregular optical defocus blur caused by ship surge or mechanical vibration.

[0082] This design can cover the mainstream disturbance types of ports and adapt to different operating environments by adjusting the L / θ / I / C parameters to meet the real-time data enhancement needs. This fuzzy kernel generation method is used as a training data enhancement module to work with the 3D-TSPatchGAN discriminator to improve the model's anti-disturbance ability.

[0083] like Figure 15 As shown, when constructing a port-specific sway pattern library, the present application divides the core sway types into linear sway, jitter sway and compound sway, and generates a parameterized fuzzy template library based on these three core types. Through multi-dimensional physical modeling, the sudden disturbance characteristics in port loading and unloading operations (such as crane translation, ship bumps and compound mechanical vibrations) are accurately restored. Finally, the template library is applied to clear video frames to generate simulated sway degradation effects, providing customized training data covering the entire scene for the deblurring algorithm.

[0084] The training process of the model generator and the model discriminator involved in this application are respectively as follows: Figure 16 and Figure 17 shown.

[0085] This application innovatively proposes a 3D-TSPatchNet architecture, breaking through the limitations of traditional 2D U-Net, and enabling the network to have three-dimensional spatiotemporal information processing, efficient grid block calculation and multi-scale feature extraction capabilities; on this basis, this application designs a 3D spatiotemporal feature extraction front end specifically for the short-term sway characteristics of ports, effectively capturing and utilizing the port-specific sway pattern information; at the same time, it introduces spatiotemporal context-aware jump connections to improve the deblurring quality and port container detail restoration capabilities by transmitting spatiotemporal context information; develops a sway characteristic perception mechanism to enable the network to dynamically adjust the deblurring strategy according to the sway characteristics of port container operations; and innovatively designs a time-series-aware 3D-TSPatchGAN discriminator, using a 3D-2D hybrid architecture to achieve effective discrimination of the temporal continuity of port scenes, fundamentally solving the defect that traditional discriminators cannot process time dimension information, thereby comprehensively improving the robustness of port video deblurring and the accuracy of detail restoration.

[0086] In a second aspect, the present invention further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the video deblurring method based on 3D spatio-temporal grid perception as described in the first aspect of the present invention.

[0087] Wherein, the computer-readable storage medium may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories.

[0088] The non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a ferromagnetic random access memory (FRAM), a flash memory, a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD ROM); the magnetic surface memory may be a disk memory or a tape memory.

[0089] The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static random access memory (SRAM), synchronous static random access memory (SSRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), sync link dynamic random access memory (SLDRAM), direct rambus random access memory (DRRAM). The computer-readable storage medium described in the embodiments of the present invention is intended to include these and any other suitable types of memory.

[0090] As Figure 18 shown, in a third aspect, the present invention provides an electronic device 10, including a processor 101 and a storage medium 102, where a computer program is stored on the storage medium, and when the computer program is executed by the processor, it implements the video deblurring method based on 3D spatio-temporal grid perception as described in the first aspect of the present invention.

[0091] In some embodiments, the processor may be implemented by software, hardware, firmware, or a combination thereof, and at least one of a circuit, a single or multiple application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), central processing units (CPUs), controllers, microcontrollers, and microprocessors may be used, so that the processor can execute some steps, all steps, or any combination of the steps in the method for video deblurring based on 3D spatio-temporal grid perception described in various embodiments of the present application.

[0092] Finally, it should be noted that although the above embodiments have been described in the text and drawings of the specification of the present application, the patent protection scope of the present application cannot be limited thereby. Any equivalent structure or equivalent process substitution or modification made based on the substantial concept of the present application, using the content recorded in the text and drawings of the specification of the present application, as well as any technical solution directly or indirectly implementing the technical solutions of the above embodiments in other related technical fields, etc., are all included in the patent protection scope of the present application.

Claims

1. A video deblurring method based on 3D spatio-temporal grid perception, characterized in that, It includes the following steps: Receive a sequence of N consecutive video images, with the input dimension vector being [N, H, W, 3], where N represents the number of input video image frames, H and W represent the height and width respectively, and 3 represents the RGB channels; Calculate the shaking intensity through the optical flow amplitude of adjacent frames. Adaptive divide each video image into overlapping grid blocks of size S×S. For the pixel position (i, j) in each video image, extract the grid block corresponding to this pixel position from N video images to form a grid block sequence with the shape of [N, S, S, 3], where N is a positive integer greater than 1, N - 1 frames are clear frame images, and the Nth frame image is the current frame image to be deblurred; Input the grid block sequence into the trained 3D spatio-temporal network model for block processing, and output the deblurred grid blocks; Calculate the fusion weight based on the distance between the pixel position and the center of the adjacent grid blocks, perform weighted fusion recombination on the deblurred grid blocks, and output the clear current frame image.

2. The video deblurring method based on 3D spatio-temporal grid perception according to claim 1, wherein The 3D spatio-temporal network model includes a model generator, and the model generator includes: A 3D spatio-temporal feature extraction module that converts the input [N, S, S, 3] grid block sequence into a [1, 3, N, S, S] dimension to adapt to the 3D convolution input, where 1 represents the batch size, indicating inputting a multi-frame sequence; A shaking perception encoder that extracts multi-scale features through three layers of downsampling convolution, and saves the output to the skip connection layer; A residual block module that concatenates 6 residual blocks, and each residual block contains two 3×3 convolutional layers and an identity skip connection; A spatio-temporal context decoder that performs upsampling through transposed convolution and uses the skip connection to obtain the multi-scale features of the shaking perception encoder; A detail enhancement output layer that reduces the feature channels from 64 to 3 through a 3×3 convolution, and outputs an RGB image in the range [-1, 1] through the Tanh activation function.

3. The video deblurring method based on 3D spatio-temporal grid perception according to claim 2, characterized in that, The shaking perception encoder includes a first encoding layer, a second encoding layer, and a third encoding layer. The input of the first encoding layer is the feature map with the original resolution, the input of the second encoding layer is the feature map with 2 times downsampling, and the third encoding layer is the feature map with 4 times downsampling; The output results of the first encoding layer and the second encoding layer are passed to the spatio-temporal context decoder through the skip connection layer, and the output result of the third encoding layer is processed by the residual block module after the size is adjusted by bilinear interpolation.

4. The video deblurring method based on 3D spatio-temporal grid perception according to claim 2, wherein, The format of the RGB image output by the detail enhancement output layer is [batch, S, S, 3], where batch represents the output block batch, S represents the block size, and 3 represents the RGB channels.

5. The video deblurring method based on 3D spatio-temporal grid perception according to claim 2, wherein The 3D spatio-temporal network model includes a model discriminator, and the model discriminator includes: A 3D spatio-temporal feature perception module that combines the grid block sequence and the generator output image into a tensor with the shape of [1, 3, N + 1, S, S], applies a 3D convolution with a convolution kernel size of (N + 1, 4, 4) and a stride of (1, 2, 2), and outputs a feature map with the shape of [1, 64, S / 2, S / 2], where S is the grid block size and is an even number; The 2D block GAN discriminant chain performs multi-layer convolutional downsampling on the feature map and outputs a feature map of the true / false degree.

6. The video deblurring method based on 3D spatio-temporal grid perception according to claim 5, characterized in that The activation function of the 3D convolution is LeakyReLU.

7. The video deblurring method based on 3D spatio-temporal grid perception according to claim 1, wherein The 3D spatio-temporal network model is trained according to the following method: Obtain a sample image set, obtain a sample grid block sequence based on the sample image set, input the sample grid block sequence into the 3D spatio-temporal network model to be trained, and perform iterative training to obtain a trained 3D spatio-temporal network model.

8. The video deblurring method based on 3D spatio-temporal grid perception according to claim 7, wherein The sample image set includes a first sample image set, a second sample image set, and a third sample image set; The first sample image set is obtained according to the following method: Collect actual monitoring video streams from multiple port container yards, extract all video frame images of each video stream, calculate the clarity score of each video frame image, screen out the video frame images with an average clarity score higher than a preset clarity threshold, and extract a continuous video frame sequence from the screened video frame images as the first sample image set; The second sample image set is obtained according to the following method: Set a linear shaking model, a jitter shaking model, and a combined shaking model. The linear shaking model includes a linear shaking blur kernel, the jitter shaking model includes a jitter shaking blur kernel, and the combined shaking model includes a linear shaking blur kernel and a jitter shaking blur kernel; Input the first sample image set into the linear shaking model, the jitter shaking model, and the combined shaking model respectively to obtain the second sample image set; The third sample image set is obtained according to the following method: Perform transformation processing on the images in the first sample image set and the second sample image set to obtain the third sample image set. The transformation processing includes geometric transformation, illumination transformation, and adding noise.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, it implements the video deblurring method based on 3D spatio-temporal grid perception according to any one of claims 1 to 8.

10. An electronic device on which a computer program is stored, characterized in that, It includes a processor and a storage medium. A computer program is stored on the storage medium. When the computer program is executed by the processor, it implements the video deblurring method based on 3D spatio-temporal grid perception according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Image processing method and device, electronic equipment and storage medium

    CN113177889A

  • Video deblurring method and device and computing equipment

    CN113658062A

  • Method and device for generating video by using image, and storage medium

    CN114694074A

  • Image deblurring method and system based on cavity double-residual multi-scale deep network

    CN114723630A

  • Real-time full-frame video stabilization method, system and equipment

    CN117425073A

Cited By

  • Video deblurring processing method and device and electronic equipment

    CN120751274A

  • Method and system for correcting motion blur of endoscope

    CN121280275A