A multi-stage turbulent flow dynamic video restoration method based on physical models
Through the multi-stage turbulent dynamic video recovery method based on physical model, the decomposition of the recovery task is three stages: detillation, segmentation enhancement and defuzzing. Combined with the dual-stage registration strategy and the dynamic benefit index (DEI) quantization method, the problem of poor turbulent recovery effect in complex dynamic environments in the existing technology is solved, and efficient and stable video recovery effect is achieved.
Patent Information
- Application Number
- CN202510279320.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-03-11
AI Technical Summary
The prior art is difficult to effectively restore the impact of turbulence on images and videos in complex dynamic environments, especially when atmospheric turbulence and object motion distortion are superimposed, it is impossible to effectively distinguish turbulence interference from object motion, and physical model analysis is insufficient, so it is impossible to fully understand the mixed distortion for effective recovery.
A multi-stage turbulent dynamic video recovery method based on physical model is proposed. Through a multi-stage recovery framework, the recovery task is decomposed into three stages: detillation, segmentation enhancement and defuzzing. Combining the two-stage registration strategy and the dynamic benefit index (DEI) quantization method, effective recovery of turbulent dynamic video is achieved.
It significantly improves the quality of video recovery in complex distortions, enhances visual effects, makes the video clearer, natural, richer details, and can achieve more stable and accurate image alignment under conditions of high turbulence disturbance and lens jitter.
Smart Images

Figure CN119784648B_ABST
Abstract
Description
Technical Field
[0001] The invention relates to the technical field of image processing, in particular to a multi-stage turbulence dynamic video restoration method based on a physical model. Background Art
[0002] Atmospheric turbulence is an optical phenomenon caused by the spatiotemporal variation of the refractive index in the air. Its intensity is affected by many factors, including temperature, wind speed, air pressure, and the distance between the target and the imaging device. In complex dynamic environments such as extreme temperature gradients or rapid target motion, atmospheric turbulence is particularly disruptive to long-distance horizontal and oblique path imaging. This interference usually results in pixel shifts and local non-uniform blurring of geometric distortions in images and videos, significantly reducing the image quality. In recent years, although many image and video enhancement and restoration techniques have been proposed to deal with turbulence distortion, these methods are difficult to achieve effective and synchronous restoration in dynamic scenes because turbulent dynamic videos involve multiple complex distortion types.
[0003] Traditional multi-frame methods usually treat turbulence interference as a single type of distortion, and mitigate the impact of turbulence to a certain extent through strategies such as "lucky imaging" or geometric stabilization. However, these methods are difficult to completely eliminate the distortion caused by global pixel offset. End-to-end methods based on deep learning are trained through synthetic datasets to model turbulence distortion as a whole. Although the visual effect has been improved, the restoration of edge details is still not ideal in scenes with high turbulence intensity or complex motion.
[0004] Research based on physical analysis further attempts to decompose turbulence distortion into tilt distortion and blur distortion, and adopts a staged recovery strategy, such as combining optical flow algorithms with deep learning models. Such methods perform well in static scenes or slightly dynamic scenes, but their performance drops significantly in scenes with camera movement or violent target dynamics, and often cause time series artifacts at the edges of moving targets.
[0005] Existing research has the following limitations: (1) When atmospheric turbulence and object motion distortion are superimposed, it is difficult to effectively distinguish turbulent interference from object motion; (2) There is insufficient analysis of the physical model of turbulent dynamic video, which makes it impossible to fully understand the mixed distortion and effectively restore it; (3) The physical jitter of the lens and the turbulent optical disturbance affect the registration accuracy of multiple frame images, limiting the subsequent restoration effect.
[0006] Therefore, the present invention provides a physical model driven multi-stage restoration framework combined with a dual-stage registration strategy to solve the above technical problems. Summary of the invention
[0007] The purpose of the present invention is to overcome the shortcomings of the prior art and provide a multi-stage turbulence dynamic video restoration method based on a physical model, which solves the technical problems that traditional restoration methods cannot effectively restore larger complex distorted scenes, the physical model analysis of turbulent dynamic scenes is insufficient, and the image video cannot be effectively and comprehensively restored and effective dynamic intensity quantization cannot be implemented. In addition, due to the influence of lens jitter, turbulence disturbance and object motion, effective multi-frame image alignment cannot be performed.
[0008] To achieve the above object, the present invention proposes a multi-stage turbulence dynamic video restoration method based on a physical model, which performs multi-stage turbulence dynamic video restoration through the following steps:
[0009] Data collection: collect turbulence dynamic videos of various motion scenes and form a video set; including a synthetic video set and a real video set. The synthetic video set consists of a number of synthetic data and is used to train the multi-stage turbulence recovery model; the real video set consists of a number of real data and is used to verify the multi-stage turbulence recovery model;
[0010] Data preprocessing: Based on the physics-based deep learning model, the turbulence intensity of the turbulence dynamic video obtained during the data collection phase is analyzed. The computational quantification of the turbulence intensity , optical flow map DyOF and dynamic area ratio DPR to calculate the dynamic benefit index DEI to quantify the dynamic intensity of dynamic videos under the interference of atmospheric turbulence, and then construct a high-dynamic turbulence dataset through the dynamic benefit index DEI to enhance the pertinence of model training and evaluation;
[0011] Anti-shake mechanism: Based on the multi-stage restoration framework PMR, a two-stage registration strategy is used to process the video frames of turbulent dynamic videos. In the first stage of registration, template matching is performed on the video frames through global motion estimation before the multi-stage restoration task to achieve regional alignment and eliminate large-scale inter-frame displacement. In the second stage of registration, during the de-tilting and deblurring stages of the multi-stage restoration task, two-dimensional pixel alignment operations based on optical flow information are performed on the video frames to achieve synchronous optimization of restoration and alignment.
[0012] Multi-stage restoration task setting: By modeling the distortion of turbulent dynamic video, a forward physical model consisting of three factors superimposed: tilt, motion distortion, and blur is obtained; then, based on the multi-stage restoration framework PMR, the restoration task of turbulent dynamic video is decomposed into three restoration stages: de-tilting stage, motion segmentation enhancement stage, and deblurring stage;
[0013] De-tilting stage: A lightweight model is built based on the U-Net framework. The lightweight model outputs the pixel tilt field of the corresponding scale video frame at different decoding layers. Then, the tilt correction module is used to superimpose the pixel tilt field layer by layer to correct the position offset of the pixels in the image and generate a de-tilted video frame.
[0014] Segmentation enhancement stage: The foreground and background areas of the video frame are segmented by optical flow mask and processed independently; according to the turbulence intensity Calculate the Gaussian weight of the background area, stabilize the background area through Gaussian weighting, and use the segmentation mask to merge the foreground and background, and finally generate a segmentation-enhanced video frame;
[0015] Deblurring stage: A lightweight hybrid model is constructed by combining convolutional neural networks and Transformer. The global temporal attention mechanism and frequency domain spatial channel information interaction between multiple frames are established through the lightweight hybrid model. The local non-uniform blur is repaired by capturing the detail changes in the dynamic scene, and finally the dynamic video restoration frame with full restoration of turbulence effects is output.
[0016] As a further solution, during data preprocessing, a physics-based deep learning model is used to analyze the turbulence intensity of the turbulence dynamic video obtained during the data collection phase. The turbulence intensity is calculated by the following formula: :
[0017]
[0018] Where PFOV represents the pixel field of view, D represents the lens aperture diameter, L represents the distance to the target, P represents the turbulence constant, and V represents the image sequence. represents the variance of the image sequence, Represents the gradient of an image sequence.
[0019] As a further solution, during data preprocessing, the dynamic benefit index DEI is quantified through the following steps:
[0020] The pixel displacement between adjacent frames of the turbulent dynamic video is calculated using the pre-trained RAFT model to obtain preliminary optical flow estimation results.
[0021] According to turbulence intensity Scale the optical flow of each video frame;
[0022] Calculate the optical flow of N adjacent video frames, obtain N-1 optical flow maps and average them to obtain the average optical flow map of the current video frame;
[0023] After calculating the average optical flow map of each video frame, the optical flow values in the average optical flow map are normalized and the average distance between the optical flow value and the predefined threshold of 0.5 is maximized. ;
[0024] The pixels with average distance close to 1 are set as dynamic pixels, and after processing all pixels, a dynamic area composed of dynamic pixels is obtained;
[0025] Calculate the optical flow intensity DyOF of the dynamic area and the spatial proportion of the dynamic area in the video frame, and combine them to calculate the dynamic benefit index DEI;
[0026] The turbulence dynamic video data is classified based on the dynamic benefit index DEI and the empirical threshold T, including high dynamic benefit index video and low dynamic benefit index video.
[0027] As a further solution, the dynamic benefit index DEI is calculated by the following formula:
[0028]
[0029]
[0030]
[0031]
[0032]
[0033] In the formula, is a constant coefficient, C is a quantitative constant for the influence of dynamic proportion on dynamic intensity, is the total number of video frames, N is the total number of video frames, is the optical flow intensity of the dynamic region of the current video frame i, DPR is the spatial proportion of the dynamic region in the video frame, is the average optical flow map of the current video frame i, , represents adjacent video frames, RAFT() is a pixel displacement calculation function based on the RAFT model, represents the maximum value function, Indicates the calculation of the average distance. Represents a range function.
[0034] As a further solution, the forward physical model of the three factors of tilt, motion distortion and blur is expressed as:
[0035]
[0036] in, For clear images, is the two-dimensional spatial coordinate corresponding to the number i, t is the time domain of the continuous video frame; M is the motion distortion, It is a composite of blur and tilt distortion; The turbulence dynamic image is formed by superimposing the distortion in the order of tilt distortion T, motion distortion M, and blur distortion B. .
[0037] As a further solution, the following steps are used to perform template matching on the video frames to achieve region alignment:
[0038] Calculate the video consisting of N consecutive video frames The video frame mean is taken, and the video frame mean is subtracted from each video frame to obtain the corresponding cropped frame ;
[0039] Video The first frame is used as the reference frame , and transform the reference frame into With crop frame Perform template matching and multiply the pixel values at the same position one by one to get the similarity , and the similarity corresponding to each pixel at each position Composition similarity graph HW;
[0040] Find the maximum index value of the similarity graph HW. The maximum index value corresponds to the pixel position (x, y) of the pixel, which is the reference frame. and crop frame The best matching position (x, y);
[0041] Apply the transformation matrix M to the current video frame, mapping each pixel position (x, y) to a new position frame by frame. , and finally obtain the video frame after global motion alignment.
[0042] As a further solution, the following steps are performed to perform 2D pixel alignment on the video frames based on optical flow information:
[0043] Use depth-separable convolution blocks to extract image feature information from the input video frame;
[0044] Use 3D convolution to map feature information to two-dimensional space, represented as the optical flow field of the current video frame , and use the optical flow field Transform the current video frame to obtain :
[0045] Will The new coordinates are normalized to range, and according to the changes The coordinates are used to sample and transform the image of the current video frame in an interpolation manner;
[0046] Detailed features and aligned transformations are again enhanced through depthwise separable convolutional blocks.
[0047] As a further solution, the following steps are performed to correct the positional offset of pixels in the image and generate a de-tilted high-quality video frame:
[0048] The optical flow correction 3D convolution block is used to extract multi-frame features in the spatiotemporal dimension for the input N consecutive video frames, and the pixel optical flow position is also corrected;
[0049] The feature information extracted by the encoder is downsampled in space and time by wavelet, and the high-frequency edge and detail information of each frame and the low-frequency global structure are obtained by performing independent wavelet decomposition frame by frame.
[0050] The processed video frames are stacked along the time dimension. The stacked tensor retains the wavelet features of each frame. The time dimension and the spatial dimension will be jointly learned during the subsequent feature extraction.
[0051] For N video frames processed by spatiotemporal wavelet downsampling, wavelet 3D convolution block is used to extract the features of spatiotemporal and frequency domain channels;
[0052] Multi-level wavelet decomposition is performed through wavelet convolution to obtain multi-scale frequency components; each multi-scale frequency component is independently convolved to capture spatial features and frequency domain information at different scales;
[0053] Through inverse wavelet transform, each frequency component is fused, multi-scale features are integrated and reconstructed into the input space;
[0054] The multi-scale features are further convolved point by point to fuse channel information, and layer normalization and activation functions are combined to improve the stability of multi-scale features and enhance the nonlinear expression ability of the model;
[0055] After the activation function is activated, point-by-point convolution is used to optimize the information interaction between channels;
[0056] Through skip connection, the feature information of the encoding layer E3 is fused with the feature information of the decoding layer D4 after 3D deconvolution upsampling;
[0057] Use partial depth 3D convolution blocks to perform depth convolution on some channels, then concatenate with the original unprocessed features along the channel dimension, and finally use depth separable convolution to process the overall features;
[0058] In the decoding layers D1, D2, and D3 that have passed through the partial depth 3D convolutional blocks, pixel tilt fields of continuous video frames at the current scale are generated respectively, and the tilt fields are averaged in the time domain to obtain a stable tilt field representation;
[0059] The tilt field mean value is superimposed on the input turbulence image, and its pixel values in the horizontal and vertical directions are normalized to the interval [-1, 1] respectively;
[0060] The normalized result is processed by interpolation mapping to generate a corrected image after tilt correction at the current scale;
[0061] In decoding layer D1, decoding layer D2, and decoding layer D3, three pixel tilt fields of different scales are generated layer by layer, and the turbulence image is restored by multi-scale stacking through the tilt correction module;
[0062] By integrating the rectified images at each scale, a de-skewed video frame is generated.
[0063] As a further solution, the foreground and background regions are segmented and processed independently through the following steps, and then fused to output the enhanced video frame:
[0064] The RAFT model is used to calculate the optical flow map of consecutive frames, and the minimum optical flow deviation between consecutive frames is calculated to determine the optimal optical flow mask;
[0065] Calculate the grayscale distribution of the optimal mask and accumulate it to obtain the grayscale histogram H;
[0066] Based on the grayscale histogram H, the maximum value of the inter-class variance is calculated and the optimal segmentation threshold is determined , through the optimal segmentation threshold Consistently segment all video frames to obtain a segmentation mask containing foreground and background regions;
[0067] Perform 3D convolution on the segmentation mask of the video frame to expand the effective area of the foreground and background in the segmentation mask. The size of the convolution kernel is adaptively set according to the turbulence intensity, and the kernel dimension is an odd number to ensure center alignment during the convolution operation.
[0068] Background area according to turbulence intensity Calculate the Gaussian weight and perform pixel-level multiplication with the current video frame to achieve Gaussian weighting to stabilize the background and reduce the impact of turbulence;
[0069] The segmentation mask is used to fuse the background area with the foreground area, and the segmentation-enhanced video frame is output.
[0070] As a further solution, the following steps are performed to obtain a dynamic video recovery frame that fully recovers the turbulence effect:
[0071] Divide the video frames into three levels according to different resolutions to form a multi-level image input ;
[0072] Input images of different resolutions Input the encoding layer in sequence ,Through the fusion of multi-scale information features, the context interaction between different resolutions is realized;
[0073] Encoding layer The optical flow correction 3D convolution block is used to extract multi-frame features in the spatiotemporal dimension, and the pixel optical flow position is also corrected;
[0074] For the coding layer The extracted feature information is subjected to spatiotemporal wavelet downsampling; among them,
[0075] By performing wavelet decomposition independently on each frame, high-frequency edge and detail information and low-frequency global structure of each frame are extracted;
[0076] The processed frames are stacked along the time dimension, and the stacked tensor retains the wavelet features of each frame;
[0077] In subsequent feature extraction, the model will jointly learn the features of the time dimension and the space dimension, and transform the image Channel fusion is performed with the feature map after wavelet downsampling and input into the encoding layer ;
[0078] The encoding layer E2 uses the spatiotemporal channel Transformer block to encode the image Perform feature extraction; among them,
[0079] Perform a single-channel convolution operation on the input image through 3D convolution, combine temporal features with channel features, and use linear layers for modeling to capture the complex spatiotemporal interaction between frames;
[0080] The extracted feature map is divided into query Q, key K and value V along the channel dimension. The query Q and key K are fused with channel and time information, and divided into multiple subspaces, which are calculated in parallel through the multi-head attention mechanism.
[0081] Performing a dot product operation using the attention weight Attn and the value V to generate a weighted feature; wherein the weighted feature is added to the input feature through a residual connection;
[0082] In the feedforward neural network, the number of feature channels is first expanded to generate high-dimensional hidden features; then, the local spatiotemporal features are extracted using depthwise separable convolution; finally, the original information is retained again through residual connections and the results are output;
[0083] At the decoding layer , use the spatiotemporal channel Transformer block to reconstruct the spatiotemporal features of the current scale for the image that fuses the features of the coding layer E3 and the feature map of the coding layer E4 after convolution permutation and upsampling;
[0084] The image output by the decoding layer is further enhanced with the spatiotemporal channel Transformer block to enhance the overall and local details. Subsequently, the enhanced result is combined with the original image in the optical flow-corrected 3D convolution block to generate a dynamic video recovery frame that fully restores the turbulence effect by aligning the pixel positions again.
[0085] Compared with related technologies, the multi-stage turbulence dynamic video restoration method based on physical model provided by the present invention has the following advantages:
[0086] 1. The present invention decomposes the restoration task of turbulent dynamic video into three stages: de-tilting, segmentation enhancement and deblurring through a multi-stage restoration framework (PMR), and optimizes different types of distortions respectively. In the de-tilting stage, the lightweight model based on the U-Net framework can effectively correct pixel offset and eliminate geometric distortion. In the segmentation enhancement stage, the foreground and background are processed independently through motion segmentation to avoid miscorrection of moving targets. At the same time, Gaussian weighting is used to stabilize the background area, significantly reducing the phenomenon of smear and artifacts. In the deblurring stage, a hybrid model is constructed by combining CNN and Transformer to capture the detailed changes in dynamic scenes and repair local non-uniform blur. This staged processing method can deal with complex distortions more accurately. Compared with the traditional single-stage restoration method, it can significantly improve the quality of the restored video, enhance the visual effect, and make the video clearer, more natural, and richer in details.
[0087] 2. The present invention proposes a dynamic benefit index (DEI) quantification method based on turbulence intensity, optical flow map and dynamic area ratio; by calculating the optical flow estimation through the pre-trained RAFT model and scaling the optical flow in combination with the turbulence intensity, the interference of turbulence on dynamic estimation can be effectively reduced, thereby achieving accurate quantification of dynamic features; this method can accurately distinguish whether the pixel offset is caused by turbulent disturbance or object motion, and provide an objective and accurate quantitative index for the dynamic intensity of turbulent dynamic videos; based on the quantification results of DEI, videos can be classified into two categories: high dynamic benefit index and low dynamic benefit index, which provides a basis for dynamic analysis under different turbulence conditions, helps to optimize the recovery strategy, and improves the recovery efficiency and effect.
[0088] 3. The present invention designs a two-stage registration strategy, from global motion alignment to pixel-level registration, which effectively solves the problem of multi-frame image misalignment caused by lens jitter, turbulence disturbance and object motion in real scenes; first, regional alignment is achieved through template matching, and then two-dimensional pixel alignment operation is performed using optical flow information, which can accurately correct the optical flow pixel position between video frames; this dual registration mechanism not only improves the accuracy of multi-frame image alignment, but also enhances the robustness of the restoration method in complex dynamic scenes; compared with traditional methods, the present invention can achieve more stable and accurate image alignment under conditions of high turbulence disturbance and lens jitter, providing more reliable input for subsequent restoration processing, thereby significantly improving the restoration effect, making it more advantageous in practical applications. BRIEF DESCRIPTION OF THE DRAWINGS
[0089] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0090] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the related technologies, the drawings required for use in the embodiments or the related technical descriptions are briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0091] Figure 1 A schematic diagram of the PMR structure of a multi-stage recovery framework provided by the present invention;
[0092] Figure 2 A schematic diagram of a lightweight model structure based on the U-Net framework provided by the present invention;
[0093] Figure 3 A schematic diagram of the steps of a motion segmentation enhancement method based on optical flow provided by the present invention;
[0094] Figure 4 A schematic diagram of the structure of a lightweight hybrid model combining CNN and Transformer provided by the present invention;
[0095] Figure 5 A schematic diagram of the structure of a 3D convolution block for optical flow correction provided by the present invention;
[0096] Figure 6 A schematic diagram of a wavelet 3D convolution block structure provided by the present invention;
[0097] Figure 7 A schematic diagram of a partial depth 3D convolution block structure provided by the present invention;
[0098] Figure 8A schematic diagram of a spatiotemporal channel Transformer block structure provided by the present invention;
[0099] Fig. 9 This is a schematic diagram of the test results provided by the present invention.
[0100] The purpose, features and advantages of this application will be further described in conjunction with the embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0101] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Generally, the components of the embodiments of the present invention described and shown in the drawings here can be arranged and designed in various different configurations.
[0102] See also Figure 1 The embodiment of the present application provides a multi-stage turbulence dynamic video restoration method based on a physical model, and performs multi-stage turbulence dynamic video restoration through the following steps:
[0103] Data collection: collect turbulence dynamic videos of various motion scenes and form a video set; including a synthetic video set and a real video set. The synthetic video set consists of a number of synthetic data and is used to train the multi-stage turbulence recovery model; the real video set consists of a number of real data and is used to verify the multi-stage turbulence recovery model;
[0104] Data preprocessing: Based on the physics-based deep learning model, the turbulence intensity of the turbulence dynamic video obtained during the data collection phase is analyzed. The computational quantification of the turbulence intensity , optical flow map DyOF and dynamic area ratio DPR to calculate the dynamic benefit index DEI to quantify the dynamic intensity of dynamic videos under the interference of atmospheric turbulence, and then construct a high-dynamic turbulence dataset through the dynamic benefit index DEI to enhance the pertinence of model training and evaluation;
[0105] Multi-stage task setting: First, the anti-shake stage is set to deal with the physical shake of the lens and the influence of turbulent optical disturbance; then, by modeling the distortion of the turbulent dynamic video, a forward physical model consisting of three factors superimposed: tilt, motion distortion and blur is obtained; then, based on the multi-stage restoration framework PMR, the restoration task of the turbulent dynamic video is decomposed into three restoration stages: de-tilting stage, segmentation enhancement stage and deblurring stage;
[0106] Anti-shake stage: Based on the multi-stage restoration framework PMR, a two-stage registration strategy is used to process the video frames of turbulent dynamic videos. First, template matching is performed on the video frames through global motion estimation to achieve regional alignment. Then, a two-dimensional pixel alignment operation based on optical flow information is performed on the video frames to correct the optical flow pixel positions between video frames.
[0107] De-tilting stage: A lightweight model is built based on the U-Net framework. The lightweight model outputs the pixel tilt field of the corresponding scale video frame at different decoding layers. Then, the tilt correction module is used to superimpose the pixel tilt field layer by layer to correct the position offset of the pixels in the image and generate a de-tilted video frame.
[0108] Segmentation enhancement stage: the foreground and background areas of the video frame are segmented by optical flow mask and processed independently; according to the turbulence intensity Calculate the Gaussian weight of the background area, stabilize the background area through Gaussian weighting, and use the segmentation mask to merge the foreground and background, and finally generate a segmentation-enhanced video frame;
[0109] Deblurring stage: A lightweight hybrid model is constructed by combining convolutional neural networks and Transformer. The global temporal attention mechanism and frequency domain spatial channel information interaction between multiple frames are established through the lightweight hybrid model. The local non-uniform blur is repaired by capturing the detail changes in the dynamic scene, and finally the dynamic video restoration frame with full restoration of turbulence effects is output.
[0110] It should be noted that in recent years, the application of deep learning in the field of turbulence restoration has attracted much attention. Most methods regard turbulent disturbances as a single type of distortion and complete restoration in an end-to-end manner. For example, CNN-based methods can effectively learn turbulence features, and the multi-frame averaging effect is better than a single frame. However, due to the static filter weights and limited receptive field of the CNN model, its performance is limited when dealing with spatial dynamic changes caused by turbulence. To solve the above problems, Mao et al. proposed TurbNet, which combines a physics-inspired degradation and reconstruction module, and significantly improved the turbulence modeling capability by introducing a Transformer layer to replace the convolutional layer in the classic encoder-decoder. Wu et al. further proposed a semi-supervised self-attention model that fully exploits the potential of unlabeled data using the "average teacher" method. They found that the trained network basis functions have a high correlation similar to Zernike polynomials in the spatial dimension. In addition, the TSR-WGAN model proposed by Jin et al. combines time and space information, inputs turbulent videos as three-dimensional tensors into the network, and learns the residual representation between observed data and ideal data. Ettedgui and Yitzhaky integrated a recurrent neural network (RNN) into the generator of GAN to predict the optical flow field of dynamic objects, and the generation effect was better than AT-Net. Jaiswal et al. achieved the latest performance in the field of turbulence restoration by combining Swin-Transformer and diffusion model. These studies laid the foundation for the single-stage turbulence restoration method and showed good turbulence interference repair ability in general scenes.
[0111] Compared with the single-stage method, the multi-stage turbulence restoration strategy can more accurately restore the edge details of the video frame and significantly improve the visual effect by analyzing the turbulence physical model and decomposing the mixed distortion. For example, the TMT proposed by Zhang et al. combines the advantages of CNN and Transformer. Specifically, TMT first uses the UNet-based CNN module to correct the spatial geometric distortion, and then optimizes the blur correction and detail reconstruction through the Transformer network. The AT-Net proposed by Yasarla and Patel adopts a dual UNet architecture. The first network estimates the inter-frame distortion map, and the second network uses the distortion map to complete the geometric correction and blur removal. Shimizu et al. proposed a three-stage image restoration method, combining the time-averaged reference image, B-spline registration and multi-frame high-resolution restoration to complete the final restoration, but this method has a high computational complexity, and time averaging is prone to blur diffusion.
[0112] In summary, the shortcomings of the existing technical solutions mainly include:
[0113] 1. The existing technology is mainly aimed at restoring turbulent videos of static scenes and scenes with small motion. There is no effective method for complex distorted scenes with large motion.
[0114] 2. Currently, for the restoration of turbulent scenes, most deep learning methods treat turbulent disturbances as a single type of distortion for restoration, rarely considering its forward physical model, poorly restoring edges and details, and average overall visual effects. In particular, there is insufficient analysis of the physical model of turbulent dynamic scenes, and it is impossible to fully understand mixed distortions for effective restoration.
[0115] 3. Under the interference of atmospheric turbulence, image pixels will randomly shift in any direction over time. Therefore, when quantifying the dynamic intensity, it is difficult to distinguish whether the pixels belong to turbulent disturbance or object movement, and it is impossible to implement effective dynamic intensity quantization of image videos.
[0116] 4. In real-world turbulent dynamic scenes, images and videos are affected by lens jitter, turbulent disturbances and object motion. Existing methods cannot effectively align multiple frames of images, which affects the subsequent video image restoration effect.
[0117] To this end, this embodiment models the turbulent dynamic video from the perspective of video restoration, decomposes the turbulent disturbance into the superposition of tilt distortion T and blur distortion B, and models the distortion of the dynamic scene as the motion distortion of the dynamic foreground on the static background, which is formalized as:
[0118]
[0119] in, For clear images, is the two-dimensional spatial coordinate corresponding to the number i, t is the time domain of continuous video frames; M is the motion distortion, It is a composite of blur and tilt distortion; The turbulence dynamic image is formed by superimposing the distortion in the order of tilt distortion T, motion distortion M, and blur distortion B. .
[0120] Based on this model, we propose a multi-stage restoration framework PMR, and restore the mixed distortion in turbulent dynamic videos in stages, removing tilt, motion distortion and blur in turn, and finally restoring clear dynamic video frames, providing an effective restoration method for complex distortion scenes with large motion.
[0121] 1) De-tilting stage: Fusing multi-scale optical flow information to correct geometric distortion and pixel shift caused by turbulence;
[0122] 2) Segmentation and enhancement stage: segment the foreground and background, enhance the target area, and highlight its salient features;
[0123] 3) Deblurring stage: Combine the spatiotemporal dynamic characteristics to capture local details, correct the blurred areas, and generate high-quality restored images.
[0124] Among them, the DET model of a lightweight U-Net network is designed in the de-tilting stage, and the DEB hybrid lightweight model of CNN and Transformer is designed in the deblurring stage.
[0125] Furthermore, in order to solve the problem that it is difficult to effectively quantify the dynamic intensity of dynamic videos under the interference of atmospheric turbulence, the present invention proposes a dynamic benefit index DEI, which combines the turbulence intensity , optical flow map and dynamic area ratio, accurately distinguish whether the pixel offset is caused by turbulence disturbance or object motion, and provide objective and accurate quantitative indicators for the dynamic intensity of videos under different turbulence conditions.
[0126] In addition, we have designed a two-stage registration strategy from global motion alignment to pixel-level registration. We achieve overall alignment of multiple frames by performing template matching on cropped edges, and use inter-frame optical flow information to perform two-dimensional spatial correction on pixels to solve the problem of image misalignment caused by lens jitter in real scenes.
[0127] During data preprocessing, a physics-based deep learning model is used to analyze the turbulence intensity of the turbulence dynamic video obtained during the data collection phase. The turbulence intensity is calculated by the following formula: :
[0128]
[0129] Where PFOV represents the pixel field of view, D represents the lens aperture diameter, L represents the distance to the target, P represents the turbulence constant, and V represents the image sequence. represents the variance of the image sequence, Represents the gradient of the image sequence under different convolution n.
[0130] Atmospheric turbulence interference can cause the position of pixels in the video to shift randomly over time, forming a "pixel dancing" phenomenon. When quantifying the dynamic intensity of turbulent dynamic videos, existing methods have difficulty accurately distinguishing whether pixel shifts are caused by turbulent disturbances or object motion. To this end, we quantify the dynamic efficiency index (DEI) combined with the turbulence intensity during data preprocessing. , dynamic area optical flow map DyOF and dynamic area proportion DPR, by correcting pixel offset, effectively reduce the interference of turbulence on dynamic estimation, thereby achieving accurate quantification of dynamic features.
[0131] The specific steps are as follows:
[0132] First, the pixel displacement between adjacent frames of the turbulent dynamic video is calculated using the pre-trained RAFT model to obtain the preliminary optical flow estimation result.
[0133] Secondly, due to the disturbance caused by atmospheric turbulence, each pixel may randomly shift in different directions, resulting in unstable optical flow estimation.
[0134] In order to reduce the interference of turbulence on optical flow estimation, according to the turbulence intensity Scale the optical flow of each video frame;
[0135] Finally, the optical flow of N adjacent video frames is calculated to obtain N-1 optical flow maps and average them to obtain the average optical flow map of the current video frame to smooth the changes between frames and reduce the influence of noise;
[0136] The following formula can represent this step:
[0137]
[0138] in, is the average optical flow map of video frame i, N is the number of adjacent video frames, , represents adjacent video frames, and RAFT() is a pixel displacement calculation function based on the RAFT model.
[0139] After calculating the average optical flow map of each video frame, the optical flow values in the average optical flow map are normalized, and the average distance between the optical flow value and the predefined threshold of 0.5 is maximized to achieve dynamic segmentation of the image; pixels with an average distance close to 1 are set as dynamic pixels, and after processing all pixels, a dynamic area composed of dynamic pixels is obtained; after calculating the average optical flow map of each video frame, the optical flow values in the average optical flow map are normalized, and then the average distance between the optical flow value and the predefined threshold of 0.5 is calculated. , and a binarization process is applied, where values close to 1 represent dynamic pixels for dynamic region segmentation.
[0140] Calculate the optical flow map of a single frame dynamic area The specific calculation formula of the dynamic area ratio DPR of the video is as follows;
[0141]
[0142]
[0143] in, is the total number of video frames, N is the total number of video frames, The larger it is, the more dynamic the video is; represents the maximum value function, davg represents the calculation of the average distance, Representing range functions
[0144] However, in actual calculation In the process of dynamic content detection, relying solely on the numerical value of the dynamic ratio cannot fully reflect the dynamic benefit index DEI of the video, because the sensitivity of visual perception to dynamic content is nonlinear.
[0145] In order to objectively evaluate the dynamic strength of the video, this embodiment also proposes a method that combines physical indicators with human visual perception. In the experiment, we further studied The impact on the dynamic benefit index DEI is quantified in the form of a piecewise function. The range of the dynamic area is defined by the following piecewise function:
[0146]
[0147] Through experimental verification, it is found that when DPR The dynamic efficiency index DEI of the video is usually most obvious when the dynamic area is between 100 and 100, and within this range, the dynamic area can show significant motion characteristics without causing redundancy or interference in visual information due to excessive dynamic proportion.
[0148] Based on the above calculation, DyOF and DPR are used as the main indicators to quantify the dynamic intensity of turbulent dynamic video. However, due to the interference of turbulence on pixel offset, the directly calculated DEI value may be higher than the actual motion intensity. , and introduce the turbulence constant Dynamic normalization is performed to achieve objective quantification of turbulent dynamic video. The calculation formula is:
[0149]
[0150] In the formula, is a constant coefficient, C is the quantitative constant of the influence of dynamic proportion on dynamic intensity, N is the total number of video frames, the larger the DEI is, the higher the dynamic intensity of the turbulent dynamic video is, and vice versa, the lower the dynamic intensity is.
[0151] Based on the DEI calculation results and experimental statistical analysis, the present invention selects an empirical threshold T to divide the collected multi-motion scene videos under atmospheric turbulence into two categories: high dynamic benefit index and low dynamic benefit index. The high dynamic benefit index video has a significant dynamic area and high motion intensity, which is suitable for studying the impact of atmospheric turbulence on significant dynamic scenes.
[0152] We deeply analyze the atmospheric turbulence physical model and motion scenes, decompose the turbulence dynamic video restoration into multi-stage tasks, start from the perspective of image processing, deeply explore the forward physical model of atmospheric turbulence dynamic video, and decompose the distortion in the turbulent environment into the superposition of tilt and blur, where tilt is caused by phase distortion causing pixel offset, and blur is caused by high-order aberrations causing image smoothing; and the distortion in the motion scene is regarded as the motion distortion of the dynamic foreground on the background; based on this, this embodiment models the distortion of turbulence dynamic video as the superposition of three factors: tilt, motion distortion and blur.
[0153]
[0154] Then, for this forward physical model, considering the complexity of superposition of multiple distortion factors, the present invention proposes a multi-stage recovery framework PMR, such as Figure 1 As shown in the figure, the main restoration stages of the framework are: de-tilting, segmentation enhancement, and deblurring. In view of camera shake and high turbulence disturbances in real-world scenarios, a two-stage registration strategy is designed to achieve image registration, further improving the application of the PMR method in the real world.
[0155] Two-stage registration strategy
[0156] In real-world video data, factors such as wind speed changes and human interference can cause lens jitter, resulting in video frame confusion. In the task of restoring turbulent dynamic videos, this jitter will further destroy the consistency of optical flow and feature information expression between adjacent frames, increasing the difficulty of restoration. To solve this problem, we designed a two-stage registration strategy based on the PMR framework, from global motion alignment to local optical flow correction, aiming to achieve regional alignment and pixel-level registration of consecutive video frames, thereby eliminating the impact of factors such as lens jitter on turbulent dynamic video restoration.
[0157] The specific reasoning steps are as follows:
[0158] Sub-step 1: Perform template matching on the video frames to achieve region alignment through the following steps:
[0159] Calculate the video consisting of N consecutive video frames The video frame mean is taken, and the video frame mean is subtracted from each video frame to obtain the corresponding cropped frame ; Its formula is:
[0160]
[0161] Video The first frame is used as the reference frame , and transform the reference frame into With crop frame Perform template matching and multiply the pixel values at the same position one by one to get the similarity , and the similarity corresponding to each pixel at each position The similarity graph HW is composed of:
[0162]
[0163] in, Indicates the matching similarity at that position, (u,v) indicates traversing all pixel positions of the reference frame
[0164] Find the maximum index value of the similarity graph HW. The maximum index value corresponds to the pixel position (x, y) of the pixel, which is the reference frame. and crop frame The best matching position (x, y);
[0165] Apply the transformation matrix M to the current video frame, mapping each pixel position (x, y) to a new position frame by frame. , and finally obtain the video frame after global motion alignment; its formula is expressed as: .
[0166] Sub-step 2: If Figure 5 As shown in the figure, in the restoration model of the de-tilting stage and the deblurring stage, the optical flow correction 3D convolution block performs a two-dimensional pixel alignment operation based on the optical flow information on the input and output videos to correct the optical flow pixel position between frames and realize multi-frame image pixel calibration:
[0167] Use depth-separable convolution blocks to extract image feature information from the input video frame;
[0168] Use 3D convolution to map feature information to two-dimensional space, represented as the optical flow field of the current video frame , and use the optical flow field Transform the current video frame to obtain ; Its formula is:
[0169]
[0170] Will The new coordinates are normalized to range, and according to the changes The coordinates are used to sample and transform the image of the current video frame in an interpolation manner;
[0171] Detailed features and aligned transformations are again enhanced through depthwise separable convolutional blocks.
[0172] Phase One of Multi-Phase Recovery: De-Tilt
[0173] Turbulence disturbances can cause the pixel positions of each area in the image to shift randomly in all directions over time, causing the same pixel to appear in different positions at different time points, which visually appears as wave-like distortion and jitter, namely tilt distortion.
[0174] The goal of the de-tilting stage is to correct the random fluctuations of image pixels over time so that they are aligned as closely as possible in consecutive image frames, thereby eliminating the visual tilt and distortion effects and achieving a more temporally consistent image presentation.
[0175] Based on the above visual characteristics of the tilted distorted image and the characteristic that the pixel position offset obeys the zero-mean Gaussian process, the present invention proposes a lightweight model DET based on the U-Net framework, such as Figure 2 The basic framework of the model is a U-Net encoder-decoder structure with a depth of 4, in which the encoding layers E1 and E4 extract features and correct pixel displacement through optical flow correction 3D convolution blocks; other encoding layers use wavelet 3D convolution blocks combined with spatiotemporal wavelet downsampling to extract feature information in the frequency domain and spatiotemporal channels.
[0176] The decoder uses 3D deconvolution to achieve upsampling, and performs feature fusion through partial depth 3D convolution blocks in the D1, D2, and D3 layers, and integrates the feature information of the corresponding coding layer through jump connections during the fusion process. The pixel tilt field of the corresponding scaled video frame is output at different decoding layers, and the tilt field records the pixel offset information; then the tilt correction module superimposes the tilt field layer by layer to correct the position offset of the pixels in the image, and finally generates a de-tilted high-quality image.
[0177] The specific reasoning steps of the DET model in the de-tilting stage are as follows:
[0178] like Figure 5 As shown, the optical flow correction 3D convolution block is used to extract multi-frame features in the spatiotemporal dimension for the input N consecutive video frames, and the pixel optical flow position is also corrected;
[0179] Step 1: Perform spatiotemporal wavelet downsampling on the feature information extracted by the encoder, and obtain the high-frequency edge and detail information and low-frequency global structure of each frame by performing independent wavelet decomposition frame by frame;
[0180] Step 2: Stack the processed video frames along the time dimension. The stacked tensor retains the wavelet features of each frame. The time dimension and the space dimension will be jointly learned during the subsequent feature extraction. The formula is:
[0181]
[0182] Among them, Harr represents wavelet decomposition, t represents the time dimension, and the image tensor is composed of I -> .
[0183] Step 3: If Figure 6 As shown, for N video frames that have been processed by spatiotemporal wavelet downsampling, a wavelet 3D convolution block is used to extract the features of spatiotemporal and frequency domain channels;
[0184] Multi-level wavelet decomposition is performed through wavelet convolution to obtain multi-scale frequency components; each multi-scale frequency component is independently convolved to capture spatial features and frequency domain information at different scales;
[0185] Through inverse wavelet transform, each frequency component is fused, multi-scale features are integrated and reconstructed into the input space;
[0186] On this basis, point-by-point convolution further integrates channel information, and combines layer normalization and activation functions to improve feature stability and enhance the nonlinear expression ability of the model; and after activation, point-by-point convolution is used to optimize the information interaction between channels.
[0187] Step 4: After being processed in the second and third steps in the encoder, image features of multiple resolutions are gradually extracted and obtained;
[0188] Step 5: In the decoding layer D3, firstly, the features of the encoding layer E3 are fused with the feature information of the decoding layer D4 after 3D deconvolution upsampling through skip connection.
[0189] Then use some of the deep 3D convolution blocks to perform deep convolution on some channels, then concatenate them with the original unprocessed features along the channel dimension, and finally use deep separable convolution to process the overall features, such as Figure 7 shown.
[0190] Step 6: After the decoding layers D1, D2, and D3 of the partial depth 3D convolution block, the pixel tilt fields of the continuous frame images at the current scale are generated respectively, and the tilt fields are averaged in the time domain to obtain a stable tilt field representation.
[0191] Subsequently, the tilt field mean is superimposed on the input turbulence image, and its pixel values in the horizontal and vertical directions are normalized to the [-1, 1] interval, respectively.
[0192] Finally, the normalized result is processed by interpolation mapping to generate a tilt-corrected image at this scale.
[0193] Step 7: In the decoding layers D1 to D3, three pixel tilt fields of different scales are generated layer by layer according to step 5, and the turbulence image is restored by multi-scale stacking through the tilt correction module. Finally, the de-tilted video frame is generated by integrating the corrected images of each scale.
[0194] The second stage of multi-stage recovery: segmentation enhancement
[0195] In the process of turbulent dynamic video restoration, the de-tilting stage corrects the pixel position by superimposing the average tilt fields of different scales, which can repair the random pixel offset problem caused by turbulence to a certain extent.
[0196] However, for motion blur, especially the smear and artifact problems that occur at the junction of the moving foreground and background, when the overall image is uniformly corrected through the tilt field, unnecessary adjustments are made to the moving edges of the object, thereby blurring the boundary between the foreground and the background, causing the smear phenomenon to become more significant.
[0197] Therefore, the present invention proposes a motion segmentation enhancement method DEM based on optical flow; Figure 3 As shown, this method processes the foreground and background independently through optical flow mask segmentation, where the dynamic area retains the original details to ensure motion authenticity, while the static area applies Gaussian weighting to stabilize the background and eliminate turbulence interference, effectively reducing the impact of smearing and artifacts.
[0198] The specific reasoning steps in the segmentation enhancement stage are as follows:
[0199] The RAFT model is used to calculate the optical flow map of consecutive frames, as shown in formula (1), and the minimum optical flow deviation (BD) between consecutive frames is calculated to determine the optimal optical flow mask; the formula is as follows:
[0200]
[0201] in, is the average optical flow map of the current video frame i, is the global maximum value of the pixel, and mean is the average value of the deviation; the smaller BD is, the closer the optical flow amplitude distribution is to the uniform distribution, and the corresponding optical flow map is the optimal mask.
[0202] Calculate the grayscale distribution of the optimal mask and accumulate it to obtain the grayscale histogram H; calculate the maximum value of the inter-class variance based on the grayscale histogram H and determine the optimal segmentation threshold , through the optimal segmentation threshold Consistently segment all video frames to obtain a segmentation mask containing the foreground area and the background area; the specific formula is as follows:
[0203]
[0204] in, () is the maximum value selection function, and are the background and foreground weights, respectively. is the overall mean, is the background mean, and t is the possible segmentation threshold. is the frequency of pixel number i, the grayscale histogram H contains 256 gray levels, each gray level corresponds to the frequency of a pixel.
[0205] A 3D convolution operation is performed on the segmentation mask of the video frame to expand the effective area of the foreground and background in the segmentation mask. The size of the convolution kernel is adaptively set according to the turbulence intensity, and the dimension of the kernel is an odd number to ensure center alignment during the convolution operation. The specific formula is: ;
[0206] Background area according to turbulence intensity Calculate Gaussian weights , and multiply it with the current video frame at the pixel level to achieve Gaussian weighting, thereby stabilizing the background and reducing the impact of turbulence; Gaussian weight The calculation formula is:
[0207]
[0208] The closer the current frame i is to the continuous frame array N, The smaller it is, The bigger. The bigger, the The smaller it is.
[0209] The segmentation mask is used to fuse the background area with the foreground area, and the segmentation-enhanced video frame is output.
[0210] The third stage of multi-stage restoration: deblurring
[0211] After completing the de-tilting and segmentation enhancement stages, the pixel offset and boundary overlap artifacts in the image have been basically eliminated. However, there is still a non-uniform blur problem caused by high-order aberrations in the local area. In order to solve this problem, the present invention designs a lightweight hybrid model DEB that combines convolutional neural network (CNN) and Transformer at this stage, such as Figure 4 As shown in the figure. This model can effectively capture the detail changes in dynamic scenes by constructing a global temporal attention mechanism between multiple frames and interacting with the frequency domain spatial channel information, thereby further repairing the local non-uniform blur and ultimately achieving full restoration of the turbulence effect.
[0212] The specific reasoning steps in the defuzzification stage are as follows:
[0213] Divide the video frames into three levels according to different resolutions to form a multi-level image input ;
[0214] Input images of different resolutions Input the encoding layer in sequence ,Through the fusion of multi-scale information features, the context interaction between different resolutions is realized, thereby extracting rich spatiotemporal features;
[0215] Encoding layer The optical flow correction 3D convolution block is used to extract multi-frame features in the spatiotemporal dimension, and the pixel optical flow position is also corrected, such as Figure 5 As shown;
[0216] For the coding layer The extracted feature information is downsampled by spatiotemporal wavelet. First, wavelet decomposition is performed independently frame by frame to extract high-frequency edge and detail information and low-frequency global structure of each frame. Then, the processed frames are stacked along the time dimension, and the stacked tensor retains the wavelet features of each frame. In subsequent feature extraction, the model will jointly learn the features of the time dimension and the space dimension. The formula is shown in (2).
[0217] Image Channel fusion is performed with the feature map after wavelet downsampling, and at the encoding layer The spatiotemporal channel Transformer block is used for feature extraction. The structure of this module is as follows Figure 8 As shown. Specifically, first, a single-channel convolution operation is performed on the input image through 3D convolution, combining temporal features with channel features, and using linear layers for modeling to capture the complex spatiotemporal interaction relationship between frames. Subsequently, the extracted feature map is split into query Q, key K, and value V along the channel dimension. By fusing Q and K with channel and temporal information, it is divided into multiple subspaces and calculated in parallel through a multi-head attention mechanism. The weight calculation formula for each attention head is as follows:
[0218]
[0219] in: is the dot product between the query and the key, is the channel dimension of each attention head.
[0220] Next, the dot product operation is performed using the attention weight Attn and the value V to generate weighted features. The weighted features are added to the input features through residual connections, thereby retaining the original information of the input features and further improving the feature representation capability. The calculation formula is: ;in, For image input The input features of .
[0221] In the feedforward neural network, the number of channels of the feature is first expanded to generate high-dimensional hidden features. Then, the local spatiotemporal features are extracted using depthwise separable convolution. The extracted feature map is split along the channel dimension into ,right Apply the activation function and then Multiply element by element to enhance the nonlinear expression and feature interaction capabilities. Finally, the residual connection is used again to retain the original information and output the result. The formula is as follows:
[0222]
[0223] in, for Apply an activation function.
[0224] The same operations are used in other coding layers to extract multi-scale image time-frequency channel features layer by layer.
[0225] At the decoding layer , use the spatiotemporal channel Transformer block to reconstruct the spatiotemporal features of the current scale for the image that fuses the features of the coding layer E3 and the feature map of the coding layer E4 after convolution permutation and upsampling;
[0226] The image output by the decoding layer is further enhanced with the spatiotemporal channel Transformer block to enhance the overall and local details. Subsequently, the enhanced result is combined with the original image in the optical flow-corrected 3D convolution block to generate a dynamic video recovery frame that fully restores the turbulence effect by aligning the pixel positions again.
[0227] In summary, the present invention starts from three improvement perspectives and organically combines them to obtain a multi-stage turbulent dynamic video restoration method for dynamic video, so that the present invention has the following technical advantages:
[0228] 1. Advantages of segmentation enhancement and multi-stage recovery sequence based on physical models
[0229] Most existing turbulent dynamic video restoration methods adopt an overall modeling strategy, taking the mixed distortion of turbulent dynamic videos as a unified restoration target and restoring them through an end-to-end deep learning model; some methods are based on physical models, decomposing turbulent distortion into tilt distortion and blur distortion, and adopt a staged restoration strategy. Although these methods can achieve good restoration results in static or slightly dynamic scenes, the restoration performance is significantly reduced in scenes with large movements, and time series artifacts are easily generated at the edges of moving targets, affecting the continuity and visual quality of the video.
[0230] In view of the above problems, in order to more effectively restore dynamic video under turbulence interference, the present invention models turbulent dynamic video from the perspective of video restoration while fully considering motion distortion, as shown in the following formula.
[0231]
[0232] In the above formula, we decompose the turbulence distortion into tilt distortion T and blur distortion B; meanwhile, the motion distortion M is regarded as the distortion of the moving foreground relative to the background. According to the characteristics of tilt distortion and blur distortion, the present invention uses lightweight models DET and DEB for restoration respectively. The motion distortion B is processed by the motion segmentation enhancement method based on optical flow; specifically, the optical flow mask is used to achieve accurate segmentation of dynamic areas and background; the dynamic area retains the original details to ensure the authenticity of the motion, while the static area is enhanced according to the turbulence intensity. Gaussian weighting is adaptively applied to smooth the background and reduce motion ghosting and artifacts.
[0233] In addition, existing studies have verified that the restoration of turbulent distortion should follow the order of "de-tilting first, then de-blurring". On this basis, the present invention further demonstrates the optimal restoration order of turbulent dynamic video mixed distortion, that is, de-tilting first, then de-motion distortion, and finally de-blurring. On the public dataset TMT, the present invention tested 1,184 turbulent dynamic videos with 32 different motion types. The experimental results are as follows: Fig. 9 As shown in the figure, by following the restoration order of first removing the tilt, then removing the motion distortion, and finally removing the blur, the optimal restoration index can be obtained, which can significantly improve the visual quality and restoration effect of the turbulent dynamic video.
[0234] 2. Advantages of the dual-stage registration strategy
[0235] Existing image registration methods mainly rely on key feature matching or deep learning end-to-end registration mapping, which is usually performed in the data preprocessing stage and can show good alignment effects under ideal conditions. However, in real-world turbulent dynamic scenes, video images are simultaneously affected by lens jitter, turbulent disturbances, and object motion, making it difficult for existing methods to accurately identify stable key features, and thus unable to achieve high-precision multi-frame alignment, affecting the quality of subsequent video restoration.
[0236] To this end, the present invention proposes a two-stage registration strategy from global motion alignment to pixel-level registration, which improves the registration accuracy and stability of video restoration through a staged alignment mechanism. In the first stage, global motion alignment is performed before video restoration to eliminate large-scale inter-frame displacements. The present invention achieves overall alignment between multiple frames by cropping image edges and combining template matching technology. In the second stage, in the DET model in the de-tilting stage and the DEB model in the deblurring stage, the present invention introduces an optical flow corrected 3D convolution block, and uses optical flow information to perform two-dimensional pixel-level alignment to achieve simultaneous optimization of restoration and alignment. This strategy can effectively solve the problem of image misalignment caused by factors such as lens jitter in real turbulent dynamic scenes, ensure the consistency of multi-frame fusion, thereby improving the stability and accuracy of video restoration and enhancing the visual quality after restoration.
[0237] 3. Advantages of DEI indicators
[0238] Existing methods for calculating dynamic intensity mainly rely on optical flow to estimate pixel motion between frames or use deep learning for end-to-end calculation. However, under the interference of atmospheric turbulence, image pixels will randomly shift over time, making it difficult for existing methods to distinguish whether pixel displacement is caused by turbulent disturbances or object motion when quantifying the dynamic intensity of turbulent dynamic videos, thus failing to effectively quantify the true dynamic intensity of the video.
[0239] To this end, the present invention proposes a method based on turbulence intensity Dynamic Benefit Index (DEI). This index corrects the pixel offset caused by turbulent interference and combines the dynamic area optical flow map and the dynamic area ratio as the core calculation basis to accurately quantify the dynamic intensity of turbulent dynamic videos. The DEI index can not only accurately evaluate the impact of turbulence on video quality, but also serve as an objective and reliable dynamic intensity quantification indicator, which is suitable for videos under different turbulent conditions. In addition, the present invention constructs a high-dynamic turbulence dataset based on DEI, which further improves the training and evaluation targeting of the model in complex turbulent environments, thereby enhancing its generalization ability and practical application effect.
[0240] 4. Overall process advantages
[0241] Existing methods mainly focus on restoring static or low-dynamic turbulence videos, and there are relatively few studies on the restoration of high-dynamic turbulence videos. The few restoration methods for turbulent dynamic videos often have difficulty maintaining good generalization capabilities when training from synthetic datasets to real-world turbulent dynamic videos. This limitation is mainly due to the lack of datasets and the lack of effective processing mechanisms for complex situations such as the superposition of motion distortion and turbulence distortion and lens jitter in real scenes, which limits the restoration effect.
[0242] In response to the above problems, the present invention proposes a multi-stage restoration framework based on the physical model of turbulent dynamic video, which gradually removes mixed distortion and improves the accuracy and stability of restoration. At the same time, the dynamic efficiency index (DEI) is introduced to classify the data set, and intensive training is performed based on the high DEI data set, so that the model can more effectively learn the mixed distortion characteristics in high dynamic scenes, thereby showing stronger adaptability in real-world turbulent environments and restoring clearer edge contours and detail information. In addition, in response to the problem of video frame misalignment caused by lens jitter and other force majeure factors in real scenes, the present invention further combines a two-stage registration strategy, from global motion alignment to pixel-level correction, to achieve more accurate inter-frame alignment, and effectively improve the quality and visual effect of real-world turbulent dynamic video restoration. Through the above innovative methods, the present invention exhibits excellent restoration effects in complex turbulent dynamic scenes in the real world, ensuring the clear details, stability and realism of the restored video.
[0243] The above are only some embodiments of the present application, and are not intended to limit the patent scope of the present application. All equivalent structural changes made using the contents of the present application specification and drawings under the technical concept of the present application, or direct / indirect application in other related technical fields are included in the patent protection scope of the present application.
Claims
1. A multi-stage turbulence dynamic video restoration method based on physical model, characterized in that: Multi-stage turbulence dynamic video restoration is performed through the following steps: Data collection: collect turbulence dynamic videos of various motion scenes and form a video set; including a synthetic video set and a real video set. The synthetic video set consists of a number of synthetic data and is used to train the multi-stage turbulence recovery model; the real video set consists of a number of real data and is used to verify the multi-stage turbulence recovery model; Data preprocessing: Based on the physics-based deep learning model, the turbulence intensity of the turbulence dynamic video obtained during the data collection phase is analyzed. The computational quantification of the turbulence intensity , optical flow map DyOF and dynamic area ratio DPR to calculate the dynamic benefit index DEI to quantify the dynamic intensity of dynamic videos under the interference of atmospheric turbulence, and then construct a high-dynamic turbulence dataset through the dynamic benefit index DEI to enhance the pertinence of model training and evaluation; Anti-shake mechanism: Based on the multi-stage restoration framework PMR, a two-stage registration strategy is used to process the video frames of turbulent dynamic videos. In the first stage of registration, template matching is performed on the video frames through global motion estimation before the multi-stage restoration task to achieve regional alignment and eliminate large-scale inter-frame displacement. In the second stage of registration, during the de-tilting and deblurring stages of the multi-stage restoration task, two-dimensional pixel alignment operations based on optical flow information are performed on the video frames to achieve synchronous optimization of restoration and alignment. Multi-stage restoration task setting: By modeling the distortion of turbulent dynamic video, a forward physical model consisting of three factors superimposed: tilt, motion distortion, and blur is obtained; then, based on the multi-stage restoration framework PMR, the restoration task of turbulent dynamic video is decomposed into three restoration stages: de-tilting stage, motion segmentation enhancement stage, and deblurring stage; De-tilting stage: A lightweight model is built based on the U-Net framework. The lightweight model outputs the pixel tilt field of the corresponding scale video frame at different decoding layers. Then, the tilt correction module is used to superimpose the pixel tilt field layer by layer to correct the position offset of the pixels in the image and generate a de-tilted video frame. Segmentation enhancement stage: The foreground and background areas of the video frame are segmented by optical flow mask and processed independently; according to the turbulence intensity Calculate the Gaussian weight of the background area, stabilize the background area through Gaussian weighting, and use the segmentation mask to merge the foreground and background, and finally generate a segmentation-enhanced video frame; Deblurring stage: A lightweight hybrid model is constructed by combining convolutional neural networks and Transformer. The global temporal attention mechanism and frequency domain spatial channel information interaction between multiple frames are established through the lightweight hybrid model. The local non-uniform blur is repaired by capturing the detail changes in the dynamic scene, and finally the dynamic video restoration frame with full restoration of turbulence effects is output.
2. According to the physical model-based multi-stage turbulence dynamic video restoration method of claim 1, it is characterized in that: During data preprocessing, the turbulence intensity of the turbulence dynamic video obtained during the data collection phase is analyzed based on the physical deep learning model. The turbulence intensity is calculated by the following formula: : Where PFOV represents the pixel field of view, D represents the lens aperture diameter, L represents the distance to the target, P represents the turbulence constant, V represents the image sequence, represents the variance of the image sequence, Represents the gradient of an image sequence.
3. According to the physical model-based multi-stage turbulence dynamic video restoration method of claim 1, it is characterized in that: During data preprocessing, the dynamic benefit index DEI is quantified through the following steps: The pixel displacement between adjacent frames of the turbulent dynamic video is calculated using the pre-trained RAFT model to obtain preliminary optical flow estimation results. According to the turbulence intensity Scale the optical flow of each video frame; Calculate the optical flow of N adjacent video frames, obtain N-1 optical flow maps and average them to obtain the average optical flow map of the current video frame; After calculating the average optical flow map of each video frame, the optical flow values in the average optical flow map are normalized and the average distance between the optical flow value and the predefined threshold of 0.5 is maximized. ; The pixels with average distance close to 1 are set as dynamic pixels, and after processing all pixels, a dynamic area composed of dynamic pixels is obtained; Calculate the optical flow intensity DyOF of the dynamic area and the spatial proportion of the dynamic area in the video frame, and combine them to calculate the dynamic benefit index DEI; The turbulence dynamic video data is classified based on the dynamic benefit index DEI and the empirical threshold T, including high dynamic benefit index video and low dynamic benefit index video.
4. According to the physical model-based multi-stage turbulence dynamic video restoration method of claim 1, it is characterized in that: The dynamic benefit index DEI is calculated by the following formula: In the formula, is a constant coefficient, C is a quantitative constant for the influence of dynamic proportion on dynamic intensity, is the total number of video frames, N is the total number of video frames, is the optical flow intensity of the dynamic region of the current video frame i, DPR is the spatial proportion of the dynamic region in the video frame, is the average optical flow map of the current video frame i, , represents adjacent video frames, RAFT() is a pixel displacement calculation function based on the RAFT model, represents the maximum value function, Indicates the calculation of the average distance. Represents the normalization function.
5. According to the physical model-based multi-stage turbulence dynamic video restoration method of claim 1, it is characterized in that: The forward physical model composed of the three factors of tilt, motion distortion and blur is expressed as: in, For clear images, is the two-dimensional spatial coordinate corresponding to the number i, t is the time domain of continuous video frames; M is the motion distortion, It is a composite of blur and tilt distortion; The turbulence dynamic image is formed by superimposing the distortion in the order of tilt distortion T, motion distortion M, and blur distortion B. .
6. The multi-stage turbulence dynamic video restoration method based on a physical model according to claim 1, characterized in that: The following steps are used to perform template matching on video frames to achieve region alignment: Calculate the video consisting of N consecutive video frames The video frame mean is taken, and the video frame mean is subtracted from each video frame to obtain the corresponding cropped frame ; Video The first frame is used as the reference frame , and transform the reference frame into With crop frame Perform template matching and multiply the pixel values at the same position one by one to get the similarity , and the similarity corresponding to each pixel at each position Composition similarity graph HW; Find the maximum index value of the similarity graph HW. The maximum index value corresponds to the pixel position (x, y) of the pixel, which is the reference frame. and crop frame The best matching position (x, y); Apply the transformation matrix M to the current video frame, mapping each pixel position (x, y) to a new position frame by frame. , and finally obtain the video frame after global motion alignment.
7. The multi-stage turbulence dynamic video restoration method based on physical model according to claim 1 is characterized in that: Perform the following steps to perform two-dimensional pixel alignment on the video frame based on optical flow information: Use depth-separable convolution blocks to extract image feature information from the input video frame; Use 3D convolution to map feature information to two-dimensional space, represented as the optical flow field of the current video frame , and use the optical flow field Transform the current video frame to obtain : Will The new coordinates are normalized to range, and according to the changes The coordinates are used to sample and transform the image of the current video frame in an interpolation manner; Detailed features and aligned transformations are again enhanced through depthwise separable convolutional blocks.
8. The multi-stage turbulence dynamic video restoration method based on physical model according to claim 1 is characterized in that: The following steps are used to correct the positional shift of pixels in the image and generate a de-tilted high-quality video frame: The optical flow correction 3D convolution block is used to extract multi-frame features in the spatiotemporal dimension for the input N consecutive video frames, and the pixel optical flow position is also corrected; The feature information extracted by the encoder is downsampled in space and time by wavelet, and the high-frequency edge and detail information of each frame and the low-frequency global structure are obtained by performing independent wavelet decomposition frame by frame. The processed video frames are stacked along the time dimension. The stacked tensor retains the wavelet features of each frame. The time dimension and the spatial dimension will be jointly learned during the subsequent feature extraction. For N video frames processed by spatiotemporal wavelet downsampling, wavelet 3D convolution block is used to extract the features of spatiotemporal and frequency domain channels; Multi-level wavelet decomposition is performed through wavelet convolution to obtain multi-scale frequency components; each multi-scale frequency component is independently convolved to capture spatial features and frequency domain information at different scales; Through inverse wavelet transform, each frequency component is fused, multi-scale features are integrated and reconstructed into the input space; The multi-scale features are further convolved point by point to fuse channel information, and layer normalization and activation functions are combined to improve the stability of multi-scale features and enhance the nonlinear expression ability of the model; After the activation function is activated, point-by-point convolution is used to optimize the information interaction between channels; Through skip connection, the feature information of the encoding layer E3 is fused with the feature information of the decoding layer D4 after 3D deconvolution upsampling; Use partial depth 3D convolution blocks to perform depth convolution on some channels, then concatenate with the original unprocessed features along the channel dimension, and finally use depth separable convolution to process the overall features; In the decoding layers D1, D2, and D3 that have passed through the partial depth 3D convolutional blocks, pixel tilt fields of continuous video frames at the current scale are generated respectively, and the tilt fields are averaged in the time domain to obtain a stable tilt field representation; The tilt field mean value is superimposed on the input turbulence image, and its pixel values in the horizontal and vertical directions are normalized to the interval [-1, 1] respectively; The normalized result is processed by interpolation mapping to generate a corrected image after tilt correction at the current scale; In decoding layer D1, decoding layer D2, and decoding layer D3, three pixel tilt fields of different scales are generated layer by layer, and the turbulence image is restored by multi-scale stacking through the tilt correction module; By integrating the rectified images at each scale, a de-skewed video frame is generated.
9. The multi-stage turbulence dynamic video restoration method based on physical model according to claim 1, characterized in that: The foreground and background regions are segmented and processed independently through the following steps, and then fused to output the enhanced video frame: The RAFT model is used to calculate the optical flow map of consecutive frames, and the minimum optical flow deviation between consecutive frames is calculated to determine the optimal optical flow mask; Calculate the grayscale distribution of the optimal mask and accumulate it to obtain the grayscale histogram H; Based on the grayscale histogram H, the maximum value of the inter-class variance is calculated and the optimal segmentation threshold is determined , through the optimal segmentation threshold Consistently segment all video frames to obtain a segmentation mask containing foreground and background regions; A 3D convolution operation is performed on the segmentation mask of the video frame to expand the effective area of the foreground and background in the segmentation mask; wherein the size of the convolution kernel is based on the turbulence intensity Adaptive setting, and the kernel dimension is an odd number to ensure center alignment during convolution operation; Background area according to turbulence intensity Calculate the Gaussian weight and perform pixel-level multiplication with the current video frame to achieve Gaussian weighting to stabilize the background and reduce the impact of turbulence; The segmentation mask is used to fuse the background area with the foreground area, and the segmentation-enhanced video frame is output.
10. The multi-stage turbulence dynamic video restoration method based on physical model according to claim 1, characterized in that: The dynamic video restoration frame with full restoration of turbulence effects is obtained through the following steps: Divide the video frames into three levels according to different resolutions to form a multi-level image input ; Input images of different resolutions Input the encoding layer in sequence ,Through the fusion of multi-scale information features, contextual interaction between different resolutions is achieved; Encoding layer The optical flow correction 3D convolution block is used to extract multi-frame features in the spatiotemporal dimension, and the pixel optical flow position is also corrected; For the coding layer The extracted feature information is downsampled by spatiotemporal wavelet; By performing wavelet decomposition independently on each frame, high-frequency edge and detail information and low-frequency global structure of each frame are extracted; The processed frames are stacked along the time dimension, and the stacked tensor retains the wavelet features of each frame; In subsequent feature extraction, the model will jointly learn the features of the time dimension and the space dimension, and transform the image Channel fusion is performed with the feature map after wavelet downsampling and input into the encoding layer ; The encoding layer E2 uses the spatiotemporal channel Transformer block to encode the image Perform feature extraction; among them, Perform a single-channel convolution operation on the input image through 3D convolution, combine temporal features with channel features, and use linear layers for modeling to capture the complex spatiotemporal interaction between frames; The extracted feature map is divided into query Q, key K and value V along the channel dimension. The query Q and key K are fused with channel and time information, and divided into multiple subspaces, which are calculated in parallel through the multi-head attention mechanism. Performing a dot product operation using the attention weight Attn and the value V to generate a weighted feature; wherein the weighted feature is added to the input feature through a residual connection; In the feedforward neural network, the number of feature channels is first expanded to generate high-dimensional hidden features; then, the local spatiotemporal features are extracted using depthwise separable convolution; finally, the original information is retained again through residual connections and the results are output; At the decoding layer , use the spatiotemporal channel Transformer block to reconstruct the spatiotemporal features of the current scale for the image that fuses the features of the coding layer E3 and the feature map of the coding layer E4 after convolution permutation and upsampling; The image output by the decoding layer is further enhanced with the spatiotemporal channel Transformer block to enhance the overall and local details. Subsequently, the enhanced result is combined with the original image in the optical flow-corrected 3D convolution block to generate a dynamic video recovery frame that fully restores the turbulence effect by aligning the pixel positions again.
Citation Information
Patent Citations
Turbulence degraded image restoration method and related device
CN116777777A
Target and method of detecting, identifying, and determining 3-d pose of the target
US20130259304A1
Cited By
Event-image dual-mode fusion video turbulence correction method, medium and system
CN121639532A
Event-image bimodal fusion video turbulence correction method, medium and system
CN121639532B