Multi-spectral guided atmospheric turbulence correction system and method
By using a multispectral-guided atmospheric turbulence correction system, and leveraging cross-modal feature alignment and fusion mechanisms and spatiotemporal joint modeling, the image quality degradation problem of photoelectric imaging under atmospheric turbulence conditions was solved, achieving efficient image correction results.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-03-03
- Publication Date
- 2026-03-31
AI Technical Summary
Existing technologies suffer from image quality degradation in photoelectric imaging under atmospheric turbulence conditions, including spatial blurring, nonlinear geometric distortion, and temporal instability. Furthermore, existing correction methods are characterized by hardware complexity, high cost, poor robustness, and limited generalization ability.
A multispectral-guided atmospheric turbulence correction system is adopted. By using stable structural information in multispectral images to provide structural priors for RGB image correction, and by utilizing cross-modal feature alignment and fusion mechanisms, combined with spatiotemporal joint modeling methods, the system achieves synergy between multispectral structural information and RGB image features, thereby alleviating the spatial distortion and temporal instability caused by turbulence.
Without adding complex optical hardware, the system's engineering deployability and application adaptability are improved, the robustness of structural representation is enhanced, image distortion caused by turbulence is suppressed, and the visual realism and consistency of the image are improved.
Smart Images

Figure CN121767210A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of photoelectric imaging technology, and more specifically to a multispectral guided atmospheric turbulence correction system and method. Background Technology
[0002] In long-distance photoelectric imaging, light waves inevitably traverse large areas of non-uniform atmospheric media along their propagation path. When light waves propagate in these non-uniform media, their wavefronts are distorted by turbulent structures at different scales. Small-scale turbulence mainly affects the high-frequency components of the light waves, causing blurring of image details; while large-scale turbulence causes overall deflection of the light path, resulting in geometric distortion of the image.
[0003] The aforementioned wavefront distortion manifests as a significant degradation in image quality in imaging systems, specifically in the following aspects: 1. Spatial blur degradation: High-frequency texture and detail information in the image are suppressed, and edges become blurred; 2. Nonlinear geometric distortion: Irregular deformation occurs in the whole or part of the image, destroying the original geometric structure relationship; 3. Enhanced temporal instability: Continuous frame images exhibit obvious jitter, flickering, and structural inconsistencies in the temporal dimension.
[0004] The aforementioned degradation phenomena are usually not isolated, but rather superimposed and mutually influential during actual imaging, making image correction under atmospheric turbulence conditions highly complex.
[0005] To address the imaging degradation caused by atmospheric turbulence, scholars and engineers both domestically and internationally have proposed various correction techniques. These existing techniques can be broadly categorized as follows: (1) Correction methods based on optical hardware are represented by adaptive optics systems. These systems measure wavefront distortion caused by atmospheric turbulence in real time using wavefront sensors and compensate for it using optical devices such as deformable mirrors. Although adaptive optics systems have achieved certain results in specific application areas such as astronomical imaging, they still have the following shortcomings: complex system structure, high hardware cost, and difficult maintenance; strict requirements for optical path stability and environmental conditions; and difficulty in large-scale deployment in general optoelectronic imaging platforms. Therefore, the applicability of this type of method in practical engineering applications is greatly limited.
[0006] (2) Atmospheric turbulence correction methods based on traditional image processing usually rely on redundant information between multiple frames of images to achieve turbulence mitigation through image registration, frame selection, statistical fusion, etc. Typical methods include Lucky Imaging, multi-frame averaging, non-rigid registration and reconstruction, etc. However, these methods generally have the following problems: they are sensitive to changes in turbulence intensity and are prone to performance degradation under strong turbulence conditions; they rely on manually designed features and rules, resulting in limited robustness and generalization ability; and they are difficult to ensure both spatial structure restoration and temporal consistency.
[0007] (3) Deep learning-based image deturbulence methods. These methods learn the mapping relationship between turbulence degradation and clear images to a certain extent through large-scale data training. However, most existing deep learning methods only use RGB images as a single input mode, and they still have the following shortcomings under complex turbulence conditions: they are prone to geometric distortion; their generalization ability to different turbulence intensities and imaging conditions is limited; and they are prone to introducing inter-frame inconsistency problems when processing time series images. Summary of the Invention
[0008] To address the aforementioned problems, the present invention aims to propose a multispectral-guided atmospheric turbulence correction system and method. Utilizing the stable structural information contained in multispectral images, a reliable structural prior is provided for the RGB image correction process. Through cross-modal feature alignment and fusion mechanisms, effective synergy between multispectral structural information and RGB image features is achieved. Combined with a spatiotemporal joint modeling method, the spatial distortion and temporal instability caused by turbulence are simultaneously mitigated. This improves the system's engineering deployability and application adaptability without introducing complex optical hardware.
[0009] The system includes: The system includes a data acquisition module, a data preprocessing module, a multispectral structural feature extraction module, an RGB image feature extraction module, a cross-modal feature alignment and fusion module, a spatiotemporal feature modeling module, and an image decoding and correction module. The data acquisition module is used to acquire RGB image datasets. and multispectral image datasets ; The data preprocessing module receives RGB image datasets and multispectral image datasets, performs preprocessing, and obtains preprocessed RGB image datasets. and preprocessed multispectral image dataset ; The multispectral structural feature extraction module is used to receive... and will Divided into G channel groups; This is used to sequentially extract and fuse the structural features of each channel group to obtain multi-scale, multispectral structural features. ; The RGB image feature extraction module is used to receive... and according to Extracting multi-scale RGB features ; The cross-modal feature alignment and fusion module is used to receive and and will and After sequentially performing spatial layer alignment, channel layer alignment, and structural layer alignment, the fused features are output. ; The spatiotemporal feature modeling module is used to obtain And, with Adjacent Frame to Structural correlations between frame fusion features ; by As a constraint factor, construct cross-modal space attention weights. , and Weighted fusion yields spatiotemporal consistency fusion features. ; The image decoding and correction module is used to receive... and according to Correct the RGB image to obtain the corrected RGB image.
[0010] Furthermore, the data acquisition module includes a visible light imaging unit and a multispectral imaging unit; the multispectral imaging unit adopts a single-chip multispectral sensor structure.
[0011] Furthermore, the preprocessing includes: radiometric correction and normalization, pixel-level registration, and time series processing.
[0012] Furthermore, the channel group includes several image channels in adjacent bands; Extracting the structural features of each channel group sequentially involves: constructing a structural feature encoding branch for each channel group. Extract the structural features of each channel group in sequence. , , The index value representing the time. Indicates the index value of the channel. Indicates the first At that moment, the Multispectral data of the channel, Includes several layers of residual convolutional blocks; The multi-scale multispectral structural features The formula for calculation is: ,in , Indicates the first Multispectral data fusion weights for channels.
[0013] Furthermore, the multi-scale RGB features The formula for calculation is: ,in, Indicates multi-layer feature coding branches, It includes several residual feature extraction units, and any two residual feature extraction units are connected by a downsampling and cross-layer connection process.
[0014] Furthermore, the spatial layer alignment specifically refers to: and All samples are mapped to a uniform spatial resolution via an upsampling module. Then, through the spatial alignment function of deformable convolution, it is mapped to the mapping of the upsampling module. The corresponding spatial feature domain; The spatial alignment function of the deformable convolution includes: a parameter localization network and a bilinear interpolation sampler; Spatial alignment and , recorded as and ; The channel layer alignment specifically involves: using a learnable channel mapping matrix. and , respectively and Mapped to a unified channel dimension D; and Each includes 1 convolution; After channel layer alignment and , recorded as and ; The structural layer alignment specifically involves: using a structural guidance function. according to Generate structural constraint weights and apply them to... , obtain fusion features , ,in, This indicates a feature fusion operation. It includes: the first convolutional module, the second convolutional module, and the Sigmoid module.
[0015] Furthermore, the spatiotemporal feature modeling module includes N spatiotemporal modeling units; ; Each spatiotemporal modeling unit includes a time dimension branch and a spatial dimension branch; The time dimension branch is obtained through the cosine similarity function. And, with Adjacent Frame to Structural correlation between fused features of any frame in a frame ; Spatial dimension branch will Divided into several spatial feature units, based on multispectral structural features As a constraint factor, through cross-modal attention mechanism Calculate the cross-modal spatial attention weights for each spatial feature unit. ; Each spatiotemporal modeling unit will and By combining and fusing, the spatiotemporal consistency fusion characteristics of the current spatiotemporal modeling unit are obtained. ; The formula for calculation is: ; express The corresponding learnable weights.
[0016] Furthermore, the atmospheric turbulence correction system also includes a model training and optimization module; The model training and optimization module is used to construct the generative network, the discriminative network, and the joint loss function; This is used to perform end-to-end adversarial training on the generator network and the discriminator network based on a joint loss function, and to iteratively update the system.
[0017] Furthermore, the generating network is the atmospheric turbulence correction system; The discriminant network performs a comprehensive evaluation based on the corrected RGB image output by the generator network; The joint loss function is iteratively optimized based on the comprehensive evaluation results. , , and ; The iteration ends when the comprehensive evaluation results stabilize.
[0018] A multispectral-guided method for correcting atmospheric turbulence, the method being implemented using the aforementioned system; The method includes the following steps: S1. Acquire RGB image dataset and multispectral image datasets ; S2. Preprocess the data collected in step S1 to obtain the preprocessed RGB image dataset. and preprocessed multispectral image dataset ; S3, will The system is divided into G channel groups. The structural features of each channel group are extracted and fused sequentially to obtain multi-scale multispectral structural features. ; S4. Extract through multi-layer feature coding branches Multiscale RGB features ; S5, will and After sequentially performing spatial layer alignment, channel layer alignment, and structural layer alignment, the fused features are output. ; S6, Obtain And, with Adjacent Frame to Structural correlations between frame fusion features ; S7, with As a constraint factor, construct cross-modal space attention weights. , and Weighted fusion yields spatiotemporal consistency fusion features. ; S8, according to Correct the RGB image to obtain the corrected RGB image.
[0019] The beneficial effects of the method described in this invention are as follows: (1) This invention adopts a single-chip multispectral sensor structure to simultaneously acquire RGB images and multispectral images, so that each spectral channel completes imaging under the same optical path conditions, thereby ensuring the consistency of different modal images in space and time at the physical level, overcoming the defects of spatiotemporal registration error in multimodal data acquisition in the prior art, and providing a reliable data foundation for subsequent cross-modal fusion.
[0020] (2) By introducing a spectral grouping processing mechanism, the present invention groups the multispectral channels according to the principle of similar bands and extracts structural features independently. Then, adaptive weighted fusion is performed through learnable fusion weights, which effectively solves the problem of unstable structural information caused by the difference in the degree of influence of turbulence on different spectral channels and enhances the robustness of structural representation. (3) This invention constructs a three-layer hierarchical alignment mechanism of spatial layer, channel layer and structural layer, uses deformable convolution to achieve pixel-level spatial registration, eliminates the dimensional differences between modes through learnable channel mapping matrix, and finally uses structural guidance function to generate constraint weight map to perform pixel-level modulation of RGB features. This overcomes the structural constraint failure problem caused by direct splicing or simple weighted fusion in the prior art, and realizes that multispectral structural information participates in RGB feature modeling in the form of constraint rather than simple additional information.
[0021] (4) This invention sets up a time dimension branch and a space dimension branch in the spatiotemporal joint modeling module. The time dimension branch uses cosine similarity to capture the structural association between adjacent frames to suppress random flicker. The space dimension branch uses multispectral structural features as a constraint factor to limit geometric deformation through a cross-modal spatial attention mechanism. The structural anchor points are injected layer by layer into each spatiotemporal modeling unit for real-time calibration. This overcomes the problem that it is difficult to handle turbulent coupling effects and deep network feature drift when using space or time modeling alone in the prior art. At the same time, it alleviates the inter-frame inconsistency phenomenon that is easy to occur under strong turbulent conditions in traditional methods.
[0022] (5) By introducing a generative adversarial constraint mechanism during the training phase, the present invention enables the discriminant network to comprehensively evaluate the correction results from the perspectives of overall statistical characteristics, structural consistency and texture distribution, thereby further improving the visual realism of the corrected image. The adversarial constraint is only used during the training phase and does not increase inference overhead, thus ensuring the real-time performance and engineering deployability of the system. Attached Figure Description
[0023] Figure 1 This is a schematic diagram of the system structure described in this invention. Detailed Implementation
[0024] The technical solution of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0025] Example 1 This embodiment provides a multispectral guided atmospheric turbulence correction system, the structural schematic of which is shown below. Figure 1 As shown, the system is based on a multispectral structure-guided RGB image atmospheric turbulence correction system, and adopts a modular and hierarchical structural design. The system forms a complete, closed-loop data processing flow from data acquisition, preprocessing, feature extraction, cross-modal fusion, spatiotemporal modeling to result output.
[0026] Through the above modular design, the input and output relationships between each functional module are clear, which not only facilitates the overall implementation of the system, but also allows for the independent replacement or optimization of individual modules in different application scenarios, thereby improving the system's engineering adaptability and scalability.
[0027] The system includes: a data acquisition module, a data preprocessing module, a multispectral structural feature extraction module, an RGB image feature extraction module, a cross-modal feature alignment and fusion module, a spatiotemporal feature modeling module, and an image decoding and correction module; The data acquisition module is used to synchronously acquire RGB image datasets of the target scene at the same time. With multispectral image datasets It is the data source of the entire system, among which, , Represents the real number field. Indicates the height of the RGB image. Indicates the width of the RGB image. , This indicates the number of channels in a multispectral image.
[0028] The data acquisition module includes a visible light imaging unit and a multispectral imaging unit. The multispectral imaging unit preferably adopts a single-chip multispectral sensor structure based on the same optical path design, which is used to simultaneously acquire image data of multiple spectral channels under the same optical path conditions, so as to achieve spatiotemporal alignment of RGB images and multispectral images at the physical level.
[0029] The data preprocessing module receives RGB image datasets and multispectral image datasets, performs preprocessing, and obtains preprocessed RGB image datasets. and preprocessed multispectral image dataset ; The preprocessing includes: radiometric correction and normalization, pixel-level registration, and time series processing.
[0030] The radiometric correction and normalization process involves performing radiometric correction on images from different channels and normalizing pixel values to a uniform range to eliminate the influence of different sensor response characteristics.
[0031] The pixel-level registration process specifically involves performing pixel-level spatial registration between RGB images and multispectral images to ensure that pixels in different modalities have a one-to-one correspondence at the same spatial location.
[0032] Time series processing: Organizes continuous images into a fixed-length input sequence according to a preset number of frames, providing structured input for subsequent spatiotemporal modeling modules.
[0033] After being processed by the data preprocessing module, the output data is transmitted in a standardized format to the multispectral structure feature extraction module and the RGB image feature extraction module.
[0034] The multispectral structure feature extraction module is used to extract relatively stable scene structure information under atmospheric turbulence conditions from multispectral images, providing reliable structural priors for RGB image correction.
[0035] To accommodate the unique characteristics of multispectral data, the multispectral structural feature extraction module incorporates a spectral grouping processing mechanism at its input end. The multispectral structural feature extraction module receives... and will The C spectral channels are grouped into G channel groups according to the principle of similar bands. In this embodiment, the value of G is 9. Each channel group contains several image channels of adjacent bands to characterize the structural information within the band range. Features are extracted independently from each group and then fused.
[0036] Construct a structural feature encoding branch for each channel group. Extract the structural features of each channel group in sequence. , , The index value representing the time. Indicates the index value of the channel. Indicates the first At that moment, the Multispectral data of the channel, This embodiment represents a deep coding network composed of several layers of residual convolutional blocks (ResNet Blocks) cascaded together. It contains 4 layers of residual convolutional blocks. Since each structural feature encoding branch processes the grouped spectral channels, the input data dimension is relatively small. Using a 4-layer residual convolutional block structure can ensure extraction accuracy while avoiding an overly bloated model, thus improving the system's real-time correction efficiency.
[0037] The features extracted by the multispectral structure feature extraction module are not simple grayscale value mappings, but rather through convolutional layers ( This process filters out radiative differences between different spectral bands, preserving shared geometric boundaries, object contours, and spatial topological information across modalities. This information is more invariant than RGB color information under turbulent interference. The output is a multi-scale feature map sequence. The multi-scale approach aims to capture comprehensive structural priors, from small-scale textures (slightly affected by turbulence) to large-scale contours (strongly affected by turbulence), achieving consistent extraction across the spectrum.
[0038] Due to differences in spatial resolution and response intensity among different spectral channel groups, the structural characteristics output by each channel group vary. The scale and numerical distribution are not entirely consistent. Due to differences in spatial resolution and response intensity among different spectral channel groups, the structural characteristics output by each channel group vary. Since they are not entirely consistent in scale and numerical distribution, learnable fusion weights are introduced. The structural features of each channel group are weighted and fused to form a unified multispectral structural feature representation: : ,in , Indicates the first The multispectral data fusion weights for each channel are generated through a channel attention mechanism. By weighing the contributions of different spectral channels to the structure, a final structural feature map is generated. The fused multispectral structural features serve as an important input to the subsequent cross-modal feature alignment module. This fusion mechanism enables the system to automatically adjust the contribution of different spectral channel groups in structure guidance based on training data, thereby enhancing the stability and robustness of the structural representation.
[0039] Under strong turbulent conditions, high-frequency textures and local geometric structures in RGB images are highly susceptible to random distortion, leading to unstable feature extraction results.
[0040] To accurately characterize the degradation properties of RGB images under turbulent conditions, this invention uses an RGB image feature extraction module to receive... and according to Perform feature encoding and extract multi-scale RGB features. (Texture, edge, and geometric information); The multi-scale RGB features The formula for calculation is: ,in, Indicates multi-layer feature coding branches, It includes several residual feature extraction units (Residual Blocks). Any two residual feature extraction units are connected by a downsampling and skip connection. In this embodiment... It includes 10 residual feature extraction units, and the downsampling process is strided convolution or max pooling to achieve "multi-scale". Designed as a pyramid structure to output feature maps of different resolutions at different depths.
[0041] A multi-level cascaded deep convolutional neural network (CNN) encoder structure is used for the input (damaged) data. Then, spatial features are extracted step by step through continuous convolution operations. As the network depth increases, the receptive field gradually expands, forming a multi-scale RGB feature representation. Low-level features are obtained by extracting fine-grained preliminary features such as texture and edges that are severely damaged, and mainly represent local texture information, while high-level features contain richer semantic and structural information.
[0042] High-level features are extracted by drawing global features with semantic information. Cross-layer connections (SkipConnection) ensure the integrity of multi-scale information, containing richer semantic and structural information. Specifically: During the downsampling (resolution reduction) process, in order to prevent the loss of fine-grained edge and texture information, this invention directly transmits shallow features and fuses them with deep features through cross-layer connections, thereby suppressing turbulence noise while preserving the realism of the original image.
[0043] Pixel-level RGB data degraded by turbulence ( Mapped to a high-dimensional feature space ( This lays the foundation for subsequent cross-modal fusion; During the extraction process, some high-frequency noise is initially smoothed out, while preserving as much color and content detail as possible from the original image.
[0044] Adjusting the number of channels and spatial resolution of the RGB features to a dimension that matches the multispectral structural features facilitates guided modulation.
[0045] Under strong turbulence conditions, since RGB images contain only three wide bands with similar wavelengths, the statistical characteristics of each channel affected by turbulence are highly similar, and there is a lack of information sources that can provide independent and stable references. Therefore, relying solely on RGB features for deturbulence modeling makes it difficult to maintain structural consistency in strong turbulence scenarios.
[0046] To address the problem that relying solely on RGB features makes it difficult to maintain structural consistency in highly turbulent scenarios, this invention uses multispectral structural features as a priori structural constraint on RGB features.
[0047] Because RGB features and multispectral structural features differ significantly in statistical distribution, feature dimension, and response scale, directly splicing features or simple weighted fusion can easily lead to the failure of structural constraints.
[0048] To address the aforementioned issues, this invention employs a cross-modal feature alignment and fusion module to eliminate the differences between RGB features and multispectral structural features in the feature space. From three levels—spatial layer, channel layer, and structural constraint layer—RGB features and multispectral structural features are aligned and collaboratively modeled at each level, achieving effective fusion between different modal features.
[0049] The cross-modal feature alignment and fusion module receives... and and will and After sequentially performing spatial layer alignment, channel layer alignment, and structural layer alignment, the fused features are output. ; The spatial layer alignment specifically refers to: and All samples are mapped to a uniform spatial resolution via an upsampling module. Then, using the spatial alignment function of deformable convolution (DCN), it is mapped to the mapping of the upsampling module. The corresponding spatial feature domain refers to the set of physical spatial locations corresponding to the activation values of each neuron in the feature map tensor. The purpose of alignment is to ensure that each "point" in the multispectral feature map numerically corresponds to a "physical point" in the RGB feature map, enabling the fusion of the two types of features in the same spatial coordinate system. In this embodiment, the upsampling module includes: a bilinear interpolation layer and... Convolution module.
[0050] The upsampling module is located at the front end of the spatial alignment process, and its function is to first achieve the desired result through a bilinear interpolation layer. and Align the resolution, and then through The convolution module performs downsampling to unify the resolution.
[0051] The bilinear interpolation sampler is the core internal component of the DCN spatial alignment function, used to achieve pixel-level position alignment after resolution unification.
[0052] The spatial alignment function of the deformable convolutional network (DCN) includes a localization network and a bilinear interpolation sampler; in this embodiment, the bilinear interpolation sampler is specifically a grid sampling operator. It is the core component of the DCN spatial alignment function, responsible for performing sub-pixel-level precise sampling on the feature map based on the offsets output by the localization network. Based on these sub-pixel-level offsets, the feature values at the corresponding positions are "grabbed" from the original feature map using a bilinear interpolation algorithm. This achieves "spatial alignment".
[0053] In the spatial alignment function of deformable convolution (DCN), the parameter localization network learns... and The relative displacement between the two sensors generates an offset field. Then, a bilinear interpolation sampler is used to perform nonlinear deformation sampling of the multispectral features based on the offset (offset field). This compensates for the nonlinear geometric distortions between the multispectral sensor and the visible light sensor caused by physical position (parallax), focal length differences, or turbulence, achieving precise point-to-point overlap of the two modal features in the pixel coordinate system.
[0054] Spatial alignment and , recorded as and ; Spatial layer alignment ensures that the geometric constraints provided by multispectral structural features can accurately apply to the spatial location corresponding to the RGB image.
[0055] After completing spatial layer alignment, it is still necessary to address the issue of differences in the expression of different modal features in the channel dimension.
[0056] To address this, the present invention introduces a channel layer alignment mechanism in the cross-modal fusion process, which aligns RGB features and multispectral structural features in the channel dimension through a learnable channel mapping matrix.
[0057] The channel layer alignment specifically involves: using a learnable channel mapping matrix. and , respectively and Mapped to a unified channel dimension D; and Each includes 1 convolution( ConvolutionalLayer); and The input features are linearly weighted and combined using matrix multiplication. This involves projecting the original channels of varying numbers (e.g., 3 channels for RGB and multiple channels for multispectral structural features) onto a vector space D of the same dimension. This eliminates the channel number differences between different sensor modes, allowing them to be used in subsequent arithmetic operations (such as addition and splicing). During the mapping process, the contribution weights of different spectral channels to atmospheric turbulence restoration are learned, preserving key information and suppressing noise.
[0058] After channel layer alignment and , recorded as and ; After completing the alignment of the spatial layer and the channel layer, this invention further introduces a structural constraint layer alignment mechanism, injecting multispectral structural features as explicit constraints into the RGB feature modeling process.
[0059] The structural layer alignment specifically involves: using a structural guidance function. according to Generate structural constraint weights and apply them to... , obtain fusion features , ,in, This indicates a feature fusion operation. It includes: the first convolutional module, the second convolutional module, and the Sigmoid module.
[0060] A spatial attention weight generator, consisting of two convolutional layers and a sigmoid activation function, is employed. The input is aligned multispectral features. The geometric saliency is extracted through convolution, and finally a weight map (Attention Map) with values between [0, 1] is output through the Sigmoid function. This serves as a "geometric guide," modulating the RGB features at the pixel level. Weights are increased at locations identified as "sharp edges" in the multispectral analysis, and decreased at locations with "blurred / distorted" features.
[0061] By aligning the structural layers, multispectral structural features are no longer simply additional information, but participate in the RGB feature update process as constraints, thereby suppressing structural distortion caused by turbulence at the feature level.
[0062] After completing cross-modal alignment, the features are fused. It also includes texture information from RGB images and structural constraint information provided by multispectral images.
[0063] After completing multispectral structural feature extraction and cross-modal collaborative modeling, the fused features may still be affected by residual atmospheric turbulence disturbances in both spatial and temporal dimensions. These disturbances not only manifest as local geometric distortions but may also cause structural jumps and stability degradation between consecutive frames in the temporal dimension.
[0064] The effects of atmospheric turbulence on images are not simply spatial distortions or temporal jitters, but rather the result of their coupling and superposition. Therefore, it is difficult to achieve stable results under complex turbulent conditions by using spatial or temporal modeling methods alone.
[0065] To this end, the present invention introduces a spatiotemporal joint modeling mechanism guided by multispectral structure into the overall system. By simultaneously modeling spatial and temporal relationships within a unified feature space, it achieves further suppression of turbulence degradation.
[0066] This invention constructs a spatiotemporal joint modeling module, which jointly models the changes of features in the spatial and temporal dimensions within a unified network framework. This enables the model to comprehensively utilize the spatial structural information within the same frame and the relatively stable structural information in previous and subsequent frames when processing the current frame.
[0067] The spatiotemporal feature modeling module is composed of N spatiotemporal modeling units stacked together in depth; Each spatiotemporal modeling unit includes a time dimension branch and a spatial dimension branch; The time dimension branch is obtained through the cosine similarity function. And, with Adjacent Frame to Structural correlation between fused features of any frame in a frame ; In this embodiment ; The cosine similarity function represents a time-dimensional attention mechanism based on a sliding window. Similarity calculation is equivalent to obtaining the attention weights between the current frame and its neighboring frames. Cosine similarity measures the feature relevance of two frames at the same spatial location; higher relevance (greater similarity) indicates more reliable supplementary information provided by neighboring frames at that location, thus assigning them higher weights during weighted aggregation. This is essentially a non-local mean processing method constrained by a time window.
[0068] , ; Indicates the first In each spatiotemporal modeling unit, and The index value of the adjacent frame to which the operation is performed. In each spatiotemporal modeling unit, the time dimension branch only applies to... and The adjacent frames are processed.
[0069] N time-dimensional branches search for clearer and more accurate feature blocks at the pixel location in adjacent frames along the time axis. Through weighted aggregation, information from multiple frames is fused into the current frame. Finally, the temporal consistency enhancement feature is obtained. This feature effectively suppresses random flickering and abrupt changes caused by atmospheric turbulence, significantly improving the visual smoothness and clarity of videos or image sequences.
[0070] Spatial dimension branch will Divided into several spatial feature units, based on multispectral structural features As a constraint factor, through cross-modal attention mechanism Calculate the cross-modal spatial attention weights for each spatial feature unit. ; ,in, express The square root of the dimension Indicates the query transformation function. Represents the key transformation function. Represents the value transformation function, This indicates transpose processing. The spatial dimension branch utilizes stable edges and contours provided by multispectral features as "keys" and "values" to guide the spatial reconstruction of damaged RGB fusion features. Finally, a geometrically constrained spatial enhancement feature is obtained. This feature eliminates local pixel shifts caused by turbulence, restoring the image's geometric contours to a near-stable state close to that of the multispectral image.
[0071] Spatial dimension branch will Defined as a "resident structural anchor," this anchor is not only input to the first layer of the module, but is synchronously input to each cascaded spatiotemporal modeling unit through a dedicated feature injection path. This path refers to a skip connection path in a deeply cascaded network architecture, independent of the backbone feature flow, specifically for the transmission of multispectral structural priors. It ensures that the original multispectral structural features can cross layers, bypassing deep nonlinear degradation, and be directly "broadcast" and input to the modulation port of each spatiotemporal modeling unit.
[0072] exist As the system flows through each spatiotemporal modeling unit, it uses the structural anchor points corresponding to that layer to perform pixel-level spatial modulation on the currently generated intermediate features. This modulation calculates the spatial correlation between the current feature and the structural anchor points, and uses dynamic weighted residual injection to perform real-time calibration and correction of feature points that have experienced geometric shifts.
[0073] The dynamic weighted residual injection specifically involves an additive operation following "pixel-level spatial modulation of intermediate features using structural anchors." First, the calculated "structural association weights" are multiplied pixel-by-pixel with the "structural anchor features" to modulate the spatial dimension of the multispectral structural prior, thereby generating a structural enhancement correction signal. Then, this structural enhancement correction signal is used as a residual term and superimposed onto the RGB feature map of the current layer via residual concatenation. This design ensures that the system does not simply replace RGB information, but rather corrects and compensates for RGB features using the geometric priors provided by multispectral analysis while preserving the original texture.
[0074] The core purpose of the aforementioned spatial dimension branch "layer-by-layer modulation" design is: (1) Ensure that the stable geometric topology information provided by multispectral analysis is not attenuated or lost throughout the nonlinear evolution of the deep neural network.
[0075] (2) It forces each layer of feature updates to be performed under the "monitoring" of the multispectral structure, so that the final output correction features can achieve spatiotemporal continuity while having extremely high geometric fidelity, effectively solving the feature drift problem common in deep networks.
[0076] Each spatiotemporal modeling unit will and By combining and fusing, the spatiotemporal consistency fusion characteristics of the current spatiotemporal modeling unit are obtained. ; The formula for calculation is: ; express The corresponding learnable weights. It is a set of adaptive weight parameters that change with network depth. During the training phase of the system described in this invention, it is updated synchronously with the convolutional layer weights using gradient descent by minimizing the total joint loss function. The system automatically learns the weights at the 1st... Layer-time spatial enhancement features ( ) and temporal enhancement features ( The optimal fusion ratio is determined to achieve the best performance balance.
[0077] The image decoding and correction module is used to receive... and according to Correct the RGB image to obtain the corrected RGB image.
[0078] The image decoding and correction module uses several deconvolution layers to... The resolution is gradually restored to its original size, and the number of channels is compressed to 3 (RGB), ultimately outputting a corrected image with accurate geometry and clear details. In this embodiment, the number of deconvolution layers is 3, used to match the downsampling rate of the front end to achieve resolution restoration.
[0079] Example 2 This embodiment further defines Embodiment 1. To further improve the realism of the correction results at the visual level, this embodiment introduces a generative adversarial constraint mechanism during the system training phase described in Embodiment 1, and performs additional optimization on the model output through adversarial learning.
[0080] This embodiment includes a model training and optimization module; The model training and optimization module is used to construct the generative network, the discriminative network, and the joint loss function; The generating network is the atmospheric turbulence correction system; the generating network is responsible for outputting the corrected RGB image, and the discriminant network is responsible for judging the overall authenticity of the input image (the corrected RGB image).
[0081] The discriminant network performs downsampling operations sequentially through four 3x3 convolutional modules from input to output. The network consists of four cascaded convolutional downsampling modules, each employing a classic structure and using stride convolution for downsampling. The discriminant network receives the image to be evaluated (the corrected RGB image and the ground truth image). After four downsampling operations, the network outputs a probability feature map, where each value represents the confidence level that the input image is a true, sharp image.
[0082] The ground truth image was obtained manually based on experience.
[0083] Instead of compressing the entire corrected RGB image into a single value, the discriminative network evaluates the corrected RGB image and outputs a probabilistic feature map. Each pixel value represents the probability that the corresponding region in the generated image is a "real, sharp image." The cross-entropy loss between this probabilistic feature map and the all-ones matrix (representing the true label) serves as input to the joint loss function for adversarial loss.
[0084] The discriminant network does not only focus on pixel-level differences, but also comprehensively evaluates the generated results from multiple perspectives, such as overall statistical characteristics, structural consistency, and texture distribution.
[0085] The joint loss function is used to perform end-to-end adversarial training on the generator network and the discriminator network, and the system is iteratively updated.
[0086] During the training phase of the system, the discrimination network evaluates the realism of the generated images and outputs a probability. The joint loss function weights and sums the adversarial loss, pixel reconstruction loss, structural consistency loss, and time-series consistency loss to calculate the total gradient. A backpropagation algorithm is then used to propagate this total gradient back through each network layer. Using Adam, the weights, biases, and learnable parameters of each convolutional kernel in the generation network are dynamically adjusted based on the gradient direction. , , and The goal is to continuously minimize the total loss until the generated image can "fool" the discriminator and maintain geometric accuracy.
[0087] The pixel reconstruction loss is obtained by calculating the L1 distance between the corrected RGB image and the ground truth image, aiming to ensure the quality of the underlying image.
[0088] Structural consistency loss is calculated using the Structural Similarity Index Measure (SSIM) to compare the corrected RGB image with multispectral structural anchor points. The similarity measure between the two is obtained to ensure the accuracy of the geometric contours.
[0089] Temporal consistency loss is obtained by calculating the feature deviation between the correction results of adjacent frames of the corrected RGB image using optical flow estimation methods, aiming to eliminate video flicker and maintain smoothness in the temporal dimension.
[0090] The above three losses, together with the adversarial loss, constitute the joint loss function, which guides the end-to-end optimization of the entire system after being weighted and summed.
[0091] The iteration ends when the evaluation results of the discriminant network stabilize.
[0092] Example 3 This embodiment further defines Embodiment 1. This embodiment provides a multispectral guided atmospheric turbulence correction method, which is implemented using the system described in Embodiment 1. The system includes the following steps: S1. Acquire RGB image dataset and multispectral image datasets ; S2. Preprocess the data collected in step S1 to obtain the preprocessed RGB image dataset. and preprocessed multispectral image dataset ; S3, will The system is divided into G channel groups. The structural features of each channel group are extracted and fused sequentially to obtain multi-scale multispectral structural features. ; S4. Extract through multi-layer feature coding branches Multiscale RGB features ; S5, will and After sequentially performing spatial layer alignment, channel layer alignment, and structural layer alignment, the fused features are output. ; S6, Obtain And, with Adjacent Frame to Structural correlations between frame fusion features ; S7, with As a constraint factor, construct cross-modal space attention weights. , and Weighted fusion yields spatiotemporal consistency fusion features. ; S8, according to Correct the RGB image to obtain the corrected RGB image.
[0093] Example 4 This embodiment further defines Embodiment 1. In this embodiment, the same low-turbulence image and medium-turbulence image are simultaneously corrected using the method described in this invention and the pyramid attention method. The correction results are shown in Table 1, where PSNR represents the peak signal-to-noise ratio and SSIM represents the structural similarity index. Table 1
[0094] As shown in Table 1, the correction effect of the method described in this invention has achieved significant improvement in both PSNR and SSIM.
Claims
1. A multispectral guided atmospheric turbulence correction system, characterized in that, The system includes: a data acquisition module, a data preprocessing module, a multispectral structural feature extraction module, an RGB image feature extraction module, a cross-modal feature alignment and fusion module, a spatiotemporal feature modeling module, and an image decoding and correction module; The data acquisition module is used to acquire RGB image datasets. and multispectral image datasets ; The data preprocessing module receives RGB image datasets and multispectral image datasets, performs preprocessing, and obtains preprocessed RGB image datasets. and preprocessed multispectral image dataset ; The multispectral structural feature extraction module is used to receive... and will Divided into G channel groups; This is used to sequentially extract and fuse the structural features of each channel group to obtain multi-scale, multispectral structural features. ; The RGB image feature extraction module is used to receive... and according to Extracting multi-scale RGB features ; The cross-modal feature alignment and fusion module is used to receive and and will and After sequentially performing spatial layer alignment, channel layer alignment, and structural layer alignment, the fused features are output. ; The spatiotemporal feature modeling module is used to obtain And, with Adjacent Frame to Structural correlations between frame fusion features , The index value representing the time. by As a constraint factor, construct cross-modal space attention weights. , and Weighted fusion yields spatiotemporal consistency fusion features. ; The image decoding and correction module is used to receive... and according to Correct the RGB image to obtain the corrected RGB image.
2. The multispectral guided atmospheric turbulence correction system according to claim 1, characterized in that, The data acquisition module includes a visible light imaging unit and a multispectral imaging unit; the multispectral imaging unit adopts a single-chip multispectral sensor structure.
3. The multispectral guided atmospheric turbulence correction system according to claim 2, characterized in that, The preprocessing includes: radiometric correction and normalization, pixel-level registration, and time series processing.
4. The multispectral guided atmospheric turbulence correction system according to claim 3, characterized in that, The channel group contains several image channels in adjacent bands; Extracting the structural features of each channel group sequentially involves: constructing a structural feature encoding branch for each channel group. Extract the structural features of each channel group in sequence. , , Indicates the index value of the channel. Indicates the first At that moment, the Multispectral data of the channel, Includes several layers of residual convolutional blocks; The multi-scale multispectral structural features The formula for calculation is: ,in , Indicates the first Multispectral data fusion weights for channels.
5. A multispectral guided atmospheric turbulence correction system according to claim 4, characterized in that, The multi-scale RGB features The formula for calculation is: ,in, Indicates multi-layer feature coding branches, It includes several residual feature extraction units, and any two residual feature extraction units are connected by a downsampling and cross-layer connection process.
6. The multispectral guided atmospheric turbulence correction system according to claim 5, characterized in that, The spatial layer alignment specifically refers to: and All samples are mapped to a uniform spatial resolution via an upsampling module. Then, through the spatial alignment function of deformable convolution, it is mapped to the mapping of the upsampling module. The corresponding spatial feature domain; The spatial alignment function of the deformable convolution includes: a parameter localization network and a bilinear interpolation sampler; Spatial alignment and , recorded as and ; The channel layer alignment specifically involves: using a learnable channel mapping matrix. and , respectively and Mapped to a unified channel dimension D; and Each includes 1 convolution; After channel layer alignment and , recorded as and ; The structural layer alignment specifically involves: using a structural guidance function. according to Generate structural constraint weights and apply them to... , obtain fusion features , ,in, This indicates a feature fusion operation. It includes: the first convolutional module, the second convolutional module, and the Sigmoid module.
7. A multispectral guided atmospheric turbulence correction system according to claim 6, characterized in that, The spatiotemporal feature modeling module includes N spatiotemporal modeling units; ; Each spatiotemporal modeling unit includes a time dimension branch and a spatial dimension branch; The time dimension branch is obtained through the cosine similarity function. And, with Adjacent Frame to Structural correlation between fused features of any frame in a frame ; Spatial dimension branch will Divided into several spatial feature units, based on multispectral structural features As a constraint factor, through cross-modal attention mechanism Calculate the cross-modal spatial attention weights for each spatial feature unit. ; Each spatiotemporal modeling unit will and By combining and fusing, the spatiotemporal consistency fusion characteristics of the current spatiotemporal modeling unit are obtained. ; The formula for calculation is: ; express The corresponding learnable weights.
8. A multispectral guided atmospheric turbulence correction system according to claim 7, characterized in that, The atmospheric turbulence correction system also includes a model training and optimization module; The model training and optimization module is used to construct the generative network, the discriminative network, and the joint loss function; This is used to perform end-to-end adversarial training on the generator network and the discriminator network based on a joint loss function, and to iteratively update the system.
9. A multispectral guided atmospheric turbulence correction system according to claim 8, characterized in that, The generating network is the atmospheric turbulence correction system; The discriminant network performs a comprehensive evaluation based on the corrected RGB image output by the generator network; The joint loss function is iteratively optimized based on the comprehensive evaluation results. , , and ; The iteration ends when the comprehensive evaluation results stabilize.
10. A multispectral-guided method for correcting atmospheric turbulence, characterized in that, The method is implemented using the system described in any one of claims 1-9; The method includes the following steps: S1. Acquire RGB image dataset and multispectral image datasets ; S2. Preprocess the data collected in step S1 to obtain the preprocessed RGB image dataset. and preprocessed multispectral image dataset ; S3, will The system is divided into G channel groups. The structural features of each channel group are extracted and fused sequentially to obtain multi-scale multispectral structural features. ; S4. Extract through multi-layer feature coding branches Multiscale RGB features ; S5, will and After sequentially performing spatial layer alignment, channel layer alignment, and structural layer alignment, the fused features are output. ; S6, Obtain And, with Adjacent Frame to Structural correlations between frame fusion features ; S7, with As a constraint factor, construct cross-modal space attention weights. , and Weighted fusion yields spatiotemporal consistency fusion features. ; S8, according to Correct the RGB image to obtain the corrected RGB image.