A method for correcting flow-driven multimodal drone image enhancement

CN122415357BActive Publication Date: 2026-09-15UNIV OF JINAN
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610874670.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-17
Publication Date
2026-09-15
Estimated Expiration
2046-06-17

AI Technical Summary

Technical Problem

[0002]随着无人机在灾害监测、应急救援和低空遥感等领域的广泛应用,无人机图像质量直接影响目标识别、场景感知和后续智能分析任务的准确性,然而,在夜间弱光、烟雾遮挡、恶劣天气以及复杂背景等场景下,无人机获取的图像容易出现亮度不足、对比度下降、边缘模糊以及细节缺失等问题,导致目标可辨识度降低,难以满足复杂环境下的实际应用需求,因此,如何提升复杂场景下无人机图像的质量和环境适应能力,成为当前图像增强领域的重要研究方向

Benefits of technology

[0021]Compared with existing technologies, the beneficial effects of this invention are as follows: By comprehensively utilizing the texture structure information of visible light images and the thermal target response information of infrared images, this invention addresses the problems of low illumination, low contrast, local misalignment, structural blurring, and insufficient target saliency in UAV images under complex flight environments. It establishes cross-modal structural correspondences through a parallax-tolerant structural anchor co-encoder, simultaneously achieving spatial calibration and confidence coding of dual-modal features. Furthermore, it introduces a corrected flow latent space evolution mechanism to continuously and dynamically enhance and optimize the joint latent features, and combines it with a conditional velocity field predictor to achieve time-related multi-scale feature modeling and stable latent space updates. Simultaneously, it uses a hierarchical selective multi-flow fusion enhancement head to collaboratively enhance and adaptively fuse multi-level intermediate features and enhanced latent features. This method possesses advantages such as strong cross-modal structural alignment capability, good latent space dynamic modeling capability, full utilization of multi-scale information, and natural and stable enhancement results. It can enhance the target region response while ensuring structural consistency and detail integrity, achieving stable improvement in UAV image quality under complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122415357B_ABST
    Figure CN122415357B_ABST
Patent Text Reader

Abstract

The application provides a flow-driven multi-modal unmanned aerial vehicle image enhancement method for correcting, and belongs to the field of image enhancement. First, multi-modal information is introduced, and specifically, images of the same region are collected by an unmanned aerial vehicle on-board visible light imaging unit and an infrared imaging unit and are preprocessed. Then, a parallax tolerant structure anchor point cooperative encoder is constructed, double-modal latent features are extracted, and structure alignment and main and auxiliary confidence encoding are completed. Further, the double-modal latent features are subjected to noise addition, speed prediction and iterative updating through a correction flow latent space evolution module. Finally, an enhanced image is generated through a hierarchical selective multi-flow fusion enhancement head, and an end-to-end training is performed by using a multi-target joint optimization function, so that the definition, contrast and target saliency of the unmanned aerial vehicle image are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of image enhancement, and specifically relates to a corrected flow-driven multimodal UAV image enhancement method. Background Technology

[0002] With the widespread application of drones in disaster monitoring, emergency rescue, and low-altitude remote sensing, the quality of drone images directly affects the accuracy of target recognition, scene perception, and subsequent intelligent analysis tasks. However, in scenarios such as low light at night, smoke obscuration, severe weather, and complex backgrounds, images acquired by drones are prone to problems such as insufficient brightness, decreased contrast, blurred edges, and missing details, resulting in reduced target recognizability and difficulty in meeting the practical application needs in complex environments. Therefore, how to improve the quality and environmental adaptability of drone images in complex scenarios has become an important research direction in the field of image enhancement.

[0003] In recent years, visible light and infrared multimodal image enhancement technology has provided a new solution for UAV visual perception in complex environments. Existing methods usually utilize the rich texture details of visible light images and the stable thermal target response information of infrared images for joint enhancement. However, due to the differences in the imaging mechanisms of the two modalities, UAVs are easily affected by factors such as platform vibration, viewpoint changes, and sensor installation errors during flight, resulting in local spatial misalignment and structural shift between modalities. At the same time, most existing enhancement methods lack the ability to dynamically model the latent space enhancement process, making it difficult to adaptively adjust the enhancement strategy according to different degradation levels. As a result, the enhancement results are still insufficient in terms of detail preservation, structural consistency, and target prominence. Therefore, it is urgent to propose a multimodal UAV image enhancement method that can achieve cross-modal structural alignment, dynamic latent space enhancement, and multi-level feature collaborative optimization. Summary of the Invention

[0004] This invention provides a corrected flow-driven multimodal UAV image enhancement method. Based on the visible light master mode and infrared auxiliary mode, this method constructs a parallax-tolerant structured anchor co-encoder, a corrected flow latent space evolution module, and a hierarchical selective multi-flow fusion enhancement head, and constructs a multi-objective joint optimization function to train the network end-to-end, thereby improving the detail preservation, structural consistency, and target prominence of UAV images. The method includes the following steps.

[0005] S1. Control the UAV to fly over the area to be observed to obtain the original visible light image and the original infrared image; then perform preprocessing to obtain the preprocessed visible light and infrared UAV image pairs.

[0006] S2. Construct a disparity-tolerant structural anchor co-encoder FDSAE, including multi-level latent feature encoding, cross-modal structural anchor extraction, local disparity offset estimation, anchor constraint space calibration, and master-slave confidence coding output.

[0007] S3. Construct the corrected flow latent space evolution module RFLE, including latent space noise addition, conditional velocity prediction, and velocity-guided latent feature update.

[0008] S4. Construct a conditional velocity field predictor CVFP, which has a U-shaped hierarchical structure and a core feature processing unit based on time conditions.

[0009] S5. Construct a hierarchical selective multi-stream fusion enhancement head (HSMF), including intermediate feature fusion, latent feature fusion, and fusion enhancement feature generation.

[0010] S6. Construct a multi-objective joint optimization function for end-to-end training of FDSAE, RFLE, and HSMF. The multi-objective joint optimization function includes aggregated enhancement loss and corrected flow velocity field regression term. The aggregated enhancement loss includes image content preservation term, cross-modal edge enhancement term, and structural consistency term.

[0011] Preferably, in step S2, constructing the disparity-tolerant structural anchor co-encoder FDSAE ​​specifically includes: S21 and FDSAE ​​include multi-level latent feature coding, cross-modal structure anchor point extraction, local disparity offset estimation, anchor point constraint space calibration, and master-slave confidence coding output. The multi-level latent feature coding part is used to extract visible light and infrared features at different levels, and the input is a preprocessed visible light image. and infrared images The specific process is as follows: First, the multi-level latent feature encoding part has a multi-level structure and includes visible light and infrared dual-mode branches. Let the total number of levels be... The multi-level latent feature encoding part will be executed iteratively in the visible light and infrared modes respectively. Next, regarding the hierarchy Modality The iterative process is as follows: ,in, Representative level Modal The output characteristics, Representative level Modal The output characteristics, Representative level Modal The feature extraction network; secondly, the hierarchy Modal The feature extraction network has the following structure: First, it extracts... First, local texture features are obtained through a local convolutional extraction layer. Second, structural enhancements are performed on these local texture features in the horizontal, vertical, and neighborhood directions. Third, the results of the structural enhancements in the horizontal, vertical, and neighborhood directions are fused to obtain directional enhancement features. Fourth, channel weights are calculated using these directional enhancement features, and channel modulation is applied to them to obtain adaptive enhancement features. Fifth, the adaptive enhancement features and... Residual fusion yields ; S22. For the cross-modal structural anchor point extraction part, edge responses and structural abrupt response of visible light and infrared features are calculated to generate a structural anchor point map. The cross-modal structural anchor point extraction part is only for the hierarchical level. Modality The specific process is as follows: First, extract the edges, contours, or texture boundaries from visible light and infrared features. The specific formula is: ,in For the absolute value operation, and These represent the first-order gradient operators in the horizontal and vertical directions, respectively. The first step is to create an edge response map; the second step is to capture locations in the image where brightness, texture, or thermal radiation intensity changes rapidly, using the following formula: ,in It is a second-order gradient operator. The first step is to generate the structural catastrophe response map. Then, the edge response map and the structural catastrophe response map are concatenated and input into the structural anchor point generation network to obtain the structural anchor point map. The specific formula is as follows: ,in Representative level The structural anchor point generation network, It is the Sigmoid activation function. This indicates a channel splicing operation; S23. For the local disparity shift estimation part, it is used to predict the local displacement of infrared features relative to visible light features at each spatial location, so that subsequent infrared auxiliary features can be aligned with the visible light main mode. The local disparity shift estimation part is only for the hierarchical level. The specific process is as follows: First, calculate the cross-modal structure correlation diagram. ,in, and These are structural anchor point diagrams showing visible light and infrared characteristics, respectively. For convolution operations, This is element-wise multiplication; subsequently, based on the aforementioned cross-modal structure correlation diagram... Predicting local disparity offset fields using a convolutional offset prediction network ; S24. The anchor point constraint space calibration part is used to align the structural information in the infrared auxiliary mode to the spatial coordinates of the visible light main mode, reducing local misalignment. The anchor point constraint space calibration part is only for the hierarchy. The specific process is as follows: First, the sampling coordinates of the infrared features are determined based on the local parallax offset field. The specific formula is: ,in, Represents the spatial position in the visible light characteristic coordinate system. Indicates the location Local disparity offset at that location This represents the infrared feature sampling coordinates obtained through mapping; secondly, spatial sampling calibration is performed on the infrared features and the structural anchor point map of the infrared features, using the following formula: , ,in, This indicates the calibrated infrared signature. Indicates position The calibrated infrared signature This is the calibrated infrared structure anchor point diagram. Indicates position The calibrated infrared structural anchor point value, Infrared mode output characteristics, For bilinear interpolation sampling, finally, the structural consistency weight graph is calculated. ,in Indicates exponentiation; S25. For the primary and secondary confidence coding output section, only for the hierarchy... The specific process is as follows: First, calculate the confidence weights for the visible light and infrared modes, using the following formula: , , ,in, To predict the response with confidence, The output characteristics are those of the visible light mode. and They are respectively In the response components of visible and infrared modes, and These represent the confidence weights for the visible light and infrared modes, respectively. Represent the convolutional mapping unit; subsequently, compute the latent features at the visible light ends. and infrared terminal latent features The specific formula is as follows: , .

[0012] Preferably, in step S2, a parallax-tolerant structural anchor co-encoder is proposed to establish the structural correspondence between the visible light mode and the infrared mode while encoding features. The encoder first extracts multi-level latent features from the dual-modal images, then constructs cross-modal structural anchors through edge response and structural abrupt response, and combines local parallax offset estimation to achieve spatial calibration from the infrared auxiliary mode to the visible light master mode. On this basis, a dual-modal latent feature representation is generated through structural consistency constraints and master-slave confidence coding mechanism. This encoder can effectively alleviate the local misalignment problem caused by sensor perspective differences, platform vibration and different imaging mechanisms during UAV flight. While maintaining the structural details of the visible light image, it makes full use of the target response information of the infrared mode, providing a reliable feature basis for subsequent latent space enhancement.

[0013] Preferably, in step S3, the corrected flow latent space evolution module RFLE is constructed, specifically including: RFLE includes latent space noise addition, conditional velocity prediction, and velocity-guided latent feature update, the specific process being: First, latent space noise addition is performed, using standard Gaussian noise latent variables... Combined latent features with dual modes A corrected linearity-adding noisy state is constructed between these states, and the time steps obtained through random sampling are obtained. Noisy Joint Latent Features The specific formula is as follows: ,in express to Uniform distribution over the interval; second, and Input the conditional velocity field predictor to obtain the initial velocity field prediction result. and according to Calculate the regression term of the corrected flow velocity field The specific formula is as follows: , ,in, This represents a conditional velocity field predictor and shares parameters within RFLE. Represents the L2 norm; third, for go through The joint iteration consists of several iterations, each including conditional velocity prediction and velocity-guided latent feature update, with the third iteration being the specific iteration. Next iteration: First, determine the time step corresponding to the current iteration. and discrete evolution step size The specific formula is as follows: Secondly, Noisy Joint Latent Features Corresponding to Time Steps and Input conditional velocity field predictor, obtain Corresponding predicted latent space velocity field and intermediate feature set The specific formula is as follows: , and when hour, Finally, regarding Update and get Joint latent features of the next iteration The specific formula is as follows: Fourth, after After several joint iterations, the enhanced joint latent features are obtained. and the intermediate feature set In the Kth joint iteration, the first This iteration is only used for advancing the latent feature state and does not participate in the gradient backpropagation of the fusion loss. The next iteration participates in backpropagation, where This represents the number of iterations involved in gradient backpropagation.

[0014] Preferably, in step S3, a corrected flow latent space evolution module (RFLE) is constructed to dynamically enhance and optimize the state of the dual-modal joint latent features. This module maps the dual-modal latent features to a continuous evolution trajectory by constructing a corrected flow noisy state in the joint latent space, and gradually updates the latent features using the conditional velocity field prediction results. During the iterative evolution process, the latent features can be continuously corrected and optimized according to the current time state, achieving degradation information suppression and effective information enhancement. This module can break through the limitations of traditional static mapping enhancement methods, enabling the latent feature enhancement process to have continuous dynamic modeling capabilities, thereby improving the ability to restore image details, enhance target saliency, and preserve structure in complex environments.

[0015] Preferably, in step S4, the conditional velocity field predictor CVFP is constructed, specifically including: S41 and CVFP have a U-shaped hierarchical structure, including overlapping convolutional embeddings, encoding paths, bottleneck paths, decoding paths, cross-layer skip connections, and velocity field output mappings. The feature processing units in the encoding, bottleneck, and decoding paths are all Temporal Conditional Feature Processing Units (TCTBs). For the overall U-shaped hierarchical structure, the specific process for the k-th iteration is as follows: First, for the k-th iteration... The noisy joint latent features corresponding to the next iteration are embedded by overlapping convolution to obtain the initial embedded features. Second, in the encoding path, the initial embedded features are encoded at multiple scales, with a total encoding scale of [missing information]. Regarding the first Each coding scale, and The specific formula is as follows: , ,in, Indicates the first Output features at each scale, and when hour , Indicates the first Multiple TCTBs are concatenated in one coding scale. Indicates the first The encoded features output at each encoding scale Indicates the first Output features at each scale Indicates the first Downsampling operations at each scale, after After several encoding scale iterations, the final output features of the encoding scale are obtained. Third, Input a bottleneck path consisting of multiple TCTBs connected in series to obtain bottleneck features. Fourth, in the decoding path, bottleneck features are decoded at multiple scales, with a total decoding scale of [missing information]. Regarding the first Each decoding scale, and The specific formula is as follows: , ,in, For the first Features after upsampling and skip connections at each decoding scale For the first Upsampling operation at each decoding scale For the first Output features of each decoding scale, and hour, , In order to be with the first The encoded features output by the encoding scale corresponding to each decoding scale. For the first Output features of each decoding scale For the first Multiple TCTBs are concatenated in each decoding scale, after... Each decoding scale yields the final output features of that decoding scale. Fifth, Mapped to latent space velocity field Simultaneously, the bottleneck features and the output features of each decoding scale are combined to form an intermediate feature set. ; S42. For the TCTB part, it is used to perform temporal conditional modulation, self-attention modeling, and feedforward update on the input features according to the current time step. TCTB is a two-branch structure, including a self-attention branch and a feedforward branch. For the k-th iteration, the specific process is as follows: First, through... The specific formula for generating modulation parameters is as follows: ,in, For time embedding functions, For linear mapping layer, and These are the offset modulation parameters and scaling modulation parameters for the self-attention branch, respectively. and The first part describes the offset modulation parameters and scaling modulation parameters of the feedforward branch; the second part describes the input features of the TCTB. The updated features are obtained by performing temporal conditional modulation in the self-attention branch. The specific formula is as follows: ,in, For self-attention operations, For layer normalization operation; third, for The output characteristics of TCTB are obtained by performing time-conditional modulation in the feedforward branch. The specific formula is as follows: ,in, This represents a feedforward network.

[0016] Preferably, in step S4, a conditional velocity field predictor (CVFP) is constructed to provide time-dependent velocity field guidance for the latent space evolution process of the correction flow. This predictor adopts a U-shaped hierarchical structure, constructs multi-scale feature representations through encoding paths, bottleneck paths, and decoding paths, and introduces time-conditional feature processing units at each level to achieve joint modeling of temporal information and spatial features. At the same time, it fully preserves shallow detail information and deep semantic information through cross-layer jump connections, and outputs the latent space velocity field and multi-scale intermediate feature sets. This predictor can accurately characterize the changing trends of latent features at different evolutionary stages, improve the stability and accuracy of latent space state updates, and provide rich feature support for subsequent multi-level enhancement information utilization.

[0017] Preferably, in step S5, constructing the Hierarchical Selective Multistream Fusion Enhancement Head (HSMF) specifically includes: S51 and HSMF include intermediate feature fusion, latent feature fusion, and fusion-enhanced feature generation. For the intermediate feature fusion part, the input is... The specific process is as follows: First, traverse the layers in order from deepest to shallowest. Each intermediate feature in, where Indicates intermediate features; second, for hierarchy The intermediate features are used for cross-level feature residual fusion, and the specific formula is as follows: ,in Representative level The fusion of input features hierarchical The third step involves fusing and enhancing the output features; and finally, performing dual-round selective scanning and feedforward updating on the fused input features. , ,in, and Representing levels The output features are obtained through dual-wheel selective scanning and feedforward updating. Fourth, the output features of the feedforward update are split into visible light responses according to the channel dimension, using a dual-wheel selective scanning method. and infrared response Then calculate the hierarchy. Fusion Enhancement Output Features ,in This indicates selective kernel feature fusion, which is a multi-branch fusion based on selective weights; fifth, the second, third, and fourth processes... The next iteration yields the final fused and enhanced output features. ; S52. For dual-wheel selective scanning, the specific process is as follows: First, the input features are split according to the channel dimension and local position enhancement is performed to obtain the position-enhanced visible light two-dimensional features. and infrared two-dimensional features Second, the two-dimensional features of visible light and infrared light are respectively unfolded into one-dimensional sequences. and The specific formula is as follows: ,in The third step involves performing a flattening operation on the visible light and infrared one-dimensional sequences, respectively, to obtain the first-stage visible light output sequence after the scan. and infrared output sequence The specific formula is as follows: ,in The fourth step involves performing a second-stage Mamba selective state-space scan on the interleaved visible and infrared output sequences of the first stage to obtain a second-stage mixed output sequence. The specific formula is as follows: ,in Represents interleaving operations. Fifth, the mixed output sequence represents the embedding of learnable modal sources. By deinterlacing and restoring the two-dimensional sequence from the one-dimensional sequence, the output features of the dual-wheel selective scanning are obtained by splicing along the channel dimension. S53. For the latent feature fusion and fusion-enhanced feature generation part, the specific process is as follows: First, the enhanced joint latent features are... Decomposed along the channel dimension into enhanced visible light latent features and infrared latent features And perform fusion enhancement to obtain latent feature fusion enhancement features. Second, the final fusion enhances the output features. Enhanced results of latent feature fusion Channel stitching and normalization are performed to obtain the enhanced UAV image. .

[0018] Preferably, in step S5, a hierarchical selective multi-stream fusion enhancement head (HSMF) is constructed to collaboratively enhance the multi-scale intermediate features generated during the correction stream evolution process and the enhanced dual-modal latent features. This enhancement head gradually aggregates feature information from different levels through a cross-level residual fusion mechanism and establishes a long-distance dependency between visible light and infrared responses using dual-wheel selective scanning. On this basis, adaptive enhancement of multi-source information is achieved through selective kernel feature fusion, which can fully explore the complementary relationship between features at different levels while taking into account both local detail preservation and global structural expression, thereby further improving the target prominence capability of the enhanced image.

[0019] Preferably, in step S6, a multi-objective joint optimization function is constructed, specifically including: First, the multi-objective joint optimization function includes an aggregated enhancement loss and a corrected flow velocity field regression term. The aggregated enhancement loss includes an image content preservation term, a cross-modal edge enhancement term, and a structural consistency term, which respectively constrain the enhancement results in terms of preserving intensity information, edge details, and structural similarity. The corrected flow velocity field regression term is used to constrain the velocity prediction process of the conditional velocity field predictor. Second, in one training round, the corrected flow velocity field regression term only acts on the first conditional velocity prediction of the corrected flow latent space evolution module, providing overall directional constraints.

[0020] Preferably, in step S6, a multi-objective joint optimization function is constructed to constrain the end-to-end training process. By aggregating the enhanced loss and the corrected flow velocity field regression term, the network can improve its latent space modeling ability and enhance the stability of the results while maintaining image structural details and target saliency.

[0021] Compared with existing technologies, the beneficial effects of this invention are as follows: By comprehensively utilizing the texture structure information of visible light images and the thermal target response information of infrared images, this invention addresses the problems of low illumination, low contrast, local misalignment, structural blurring, and insufficient target saliency in UAV images under complex flight environments. It establishes cross-modal structural correspondences through a parallax-tolerant structural anchor co-encoder, simultaneously achieving spatial calibration and confidence coding of dual-modal features. Furthermore, it introduces a corrected flow latent space evolution mechanism to continuously and dynamically enhance and optimize the joint latent features, and combines it with a conditional velocity field predictor to achieve time-related multi-scale feature modeling and stable latent space updates. Simultaneously, it uses a hierarchical selective multi-flow fusion enhancement head to collaboratively enhance and adaptively fuse multi-level intermediate features and enhanced latent features. This method possesses advantages such as strong cross-modal structural alignment capability, good latent space dynamic modeling capability, full utilization of multi-scale information, and natural and stable enhancement results. It can enhance the target region response while ensuring structural consistency and detail integrity, achieving stable improvement in UAV image quality under complex environments. Attached Figure Description

[0022] Figure 1 This is a flowchart of a correction flow-driven multimodal UAV image enhancement method provided by the present invention.

[0023] Figure 2 This is a structural diagram of the parallax-tolerant structure anchor point co-encoder FDSAE ​​provided by the present invention.

[0024] Figure 3 This is a structural diagram of the corrected flow latent space evolution module RFLE provided by the present invention.

[0025] Figure 4 This is a structural diagram of the conditional velocity field predictor CVFP provided by the present invention.

[0026] Figure 5 This is a structural diagram of the intermediate feature fusion portion of the Hierarchical Selective Multistream Fusion Enhancement Head (HSMF) provided by the present invention.

[0027] Figure 6 This is a structural diagram of the dual-wheel selective scanning of the Hierarchical Selective Multistream Fusion Enhancement Head (HSMF) provided by the present invention.

[0028] Figure 7 This is a structural diagram of the latent feature fusion and fusion enhancement feature generation part of the Hierarchical Selective Multistream Fusion Enhancement Head (HSMF) provided by the present invention.

[0029] Figure 8 This is a structural diagram of the multi-objective joint optimization function provided by the present invention. Detailed Implementation

[0030] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.

[0031] Please see Figures 1 to 8 This invention provides a corrected flow-driven multimodal UAV image enhancement method. Based on visible light and infrared multimodal information, it achieves UAV image enhancement through a designed parallax-tolerant structure anchor point co-encoder, a corrected flow latent space evolution module, a conditional velocity field predictor, and a hierarchical selective multi-flow fusion enhancement head.

[0032] S1. Control the UAV to fly over the area to be observed to obtain the original visible light image and the original infrared image; then perform preprocessing to obtain the preprocessed visible light and infrared UAV image pairs.

[0033] Furthermore, in S1, the UAV cruises and scans the area to be observed according to a preset flight path, acquiring scene texture information through a visible light imaging device and scene thermal radiation information through an infrared imaging device. Before formal data acquisition, the visible light and infrared imaging devices are calibrated separately, and a dual-modal imaging correspondence is established based on the installation position relationship of the two sensors on the UAV platform to reduce the impact of differences in imaging from different sensors. For the acquired dual-modal image data, quality screening is first performed according to preset rules to remove invalid images caused by drastic changes in aircraft posture, motion blur, or environmental interference. Subsequently, preliminary spatial alignment of the dual-modal images is performed according to the sensor correspondence, and uniform size transformation and data format conversion are completed. Further, the visible light and infrared images are numerically normalized and standardized to ensure that the data from different modalities have similar data distributions. Finally, the processed visible light and infrared images are used to construct dual-modal UAV image pairs according to the correspondence relationship for input into the subsequent augmentation network.

[0034] S2. Construct a disparity-tolerant structural anchor co-encoder FDSAE, including multi-level latent feature encoding, cross-modal structural anchor extraction, local disparity offset estimation, anchor constraint space calibration, and master-slave confidence coding output.

[0035] Furthermore, in S2, a parallax-tolerant structural anchor point co-encoder FDSAE ​​is constructed, the specific process of which includes...

[0036] S21 and FDSAE ​​include multi-level latent feature coding, cross-modal structure anchor point extraction, local disparity offset estimation, anchor point constraint space calibration, and master-slave confidence coding output. The multi-level latent feature coding part is used to extract visible light and infrared features at different levels, and the input is a preprocessed visible light image. and infrared images The specific process is as follows: First, the multi-level latent feature encoding part has a multi-level structure and includes visible light and infrared dual-mode branches. Let the total number of levels be... In this embodiment, The multi-level latent feature encoding part will be executed iteratively in the visible light and infrared modes respectively. Next, regarding the hierarchy Modality The iterative process is as follows: ,in, Representative level Modal The output characteristics, Representative level Modal The output characteristics, Representative level Modal The feature extraction network; secondly, the hierarchy Modal The feature extraction network has the following structure: First, it extracts... Local texture features are obtained through a local convolutional extraction layer. In this embodiment, the local convolutional extraction layer consists of a convolutional kernel with a size of [missing information]. The convolution, batch normalization, and GELU activation function are sequentially concatenated. Secondly, the local texture features are structurally enhanced in the horizontal, vertical, and neighborhood directions. In this embodiment, the horizontal structural enhancement uses a convolution kernel with a size of [missing information]. Lateral depthwise convolution, vertical structure enhancement using convolution kernel size of Vertical depthwise convolution, neighborhood orientation structure enhancement is employed Third, the results of the structural enhancements in the horizontal, vertical, and neighborhood directions are fused to obtain directional enhancement features. In this embodiment, the fusion is achieved by concatenating the channel dimensions and using a convolution kernel size of [missing information]. The convolutions are sequentially concatenated. Fourth, channel weights are calculated using the directional enhancement features, and channel modulation is applied to the directional enhancement features to obtain adaptive enhancement features. In this embodiment, the channel weight calculation consists of global average pooling, a first fully connected layer, a GELU activation function, a second fully connected layer, and a Sigmoid activation function concatenated. Channel modulation is obtained by element-wise multiplication of the channel weights and the directional enhancement features. Fifth, the adaptive enhancement features and... Residual fusion yields .

[0037] S22. For the cross-modal structural anchor point extraction part, edge responses and structural abrupt response of visible light and infrared features are calculated to generate a structural anchor point map. The cross-modal structural anchor point extraction part is only for the hierarchical level. Modality The specific process is as follows: First, extract the edges, contours, or texture boundaries from visible light and infrared features. The specific formula is: ,in For the absolute value operation, and These represent the first-order gradient operators in the horizontal and vertical directions, respectively. For the edge response map, in this embodiment, the Sobel operator is used as the first-order gradient operator; secondly, the locations in the image where brightness, texture, or thermal radiation intensity changes rapidly are captured, specifically by the following formula: ,in It is a second-order gradient operator. To generate the structural abrupt response map, in this embodiment, the second-order gradient operator is the Laplacian operator. Subsequently, the edge response map and the structural abrupt response map are concatenated and input into the structural anchor point generation network to obtain the structural anchor point map. The specific formula is as follows: ,in Representative level The structural anchor point generation network, It is the Sigmoid activation function. This indicates a channel splicing operation. In this embodiment, the structural anchor point generation network consists of a convolutional kernel size of... The convolution and convolution kernel size are The convolutions are connected in series.

[0038] S23. For the local disparity shift estimation part, it is used to predict the local displacement of infrared features relative to visible light features at each spatial location, so that subsequent infrared auxiliary features can be aligned with the visible light main mode. The local disparity shift estimation part is only for the hierarchical level. The specific process is as follows: First, calculate the cross-modal structure correlation diagram. ,in, and These are structural anchor point diagrams showing visible light and infrared characteristics, respectively. For convolution operations, For element-wise multiplication, in this embodiment, the convolution operation uses a kernel size of [size missing]. The convolution; subsequently, based on the cross-modal structure correlation graph. Predicting local disparity offset fields using a convolutional offset prediction network In this embodiment, the convolutional offset prediction network is a local displacement regression network with a three-level convolutional structure.

[0039] S24. The anchor point constraint space calibration part is used to align the structural information in the infrared auxiliary mode to the spatial coordinates of the visible light main mode, reducing local misalignment. The anchor point constraint space calibration part is only for the hierarchy. The specific process is as follows: First, the sampling coordinates of the infrared features are determined based on the local parallax offset field. The specific formula is: ,in, Represents the spatial position in the visible light characteristic coordinate system. Indicates the location Local disparity offset at that location This represents the infrared feature sampling coordinates obtained through mapping; secondly, spatial sampling calibration is performed on the infrared features and the structural anchor point map of the infrared features, using the following formula: , ,in, This indicates the calibrated infrared signature. Indicates position The calibrated infrared signature This is the calibrated infrared structure anchor point diagram. Indicates position The calibrated infrared structural anchor point value, Infrared mode output characteristics, For bilinear interpolation sampling, finally, the structural consistency weight graph is calculated. ,in This indicates exponentiation.

[0040] S25. For the primary and secondary confidence coding output section, only for the hierarchy... The specific process is as follows: First, calculate the confidence weights for the visible light and infrared modes, using the following formula: , , ,in, To predict the response with confidence, The output characteristics are those of the visible light mode. and They are respectively In the response components of visible and infrared modes, and These represent the confidence weights for the visible light and infrared modes, respectively. This represents a convolutional mapping unit. In this embodiment, the convolutional mapping unit consists of a convolutional kernel with a size of [missing information]. The convolutional layer, batch normalized layer, GELU activation function, and convolutional kernel size are... The convolutions are sequentially concatenated; subsequently, the latent features at the visible light end are calculated. and infrared terminal latent features The specific formula is as follows: , .

[0041] S3. Construct the corrected flow latent space evolution module RFLE, including latent space noise addition, conditional velocity prediction, and velocity-guided latent feature update.

[0042] Furthermore, in S3, a corrected flow latent space evolution module RFLE is constructed. The specific process includes: RFLE includes latent space noise addition, conditional velocity prediction, and velocity-guided latent feature update. The specific process is as follows: First, latent space noise addition is performed, using standard Gaussian noise latent variables. Combined latent features with dual modes A corrected linearity-adding noisy state is constructed between these states, and the time steps obtained through random sampling are obtained. Noisy Joint Latent Features The specific formula is as follows: ,in express to Uniform distribution over the interval; second, and Input the conditional velocity field predictor to obtain the initial velocity field prediction result. and according to Calculate the regression term of the corrected flow velocity field The specific formula is as follows: , ,in, This represents a conditional velocity field predictor and shares parameters within RFLE. Represents the L2 norm; third, for go through In this embodiment, the joint iteration... Each iteration includes conditional velocity prediction and velocity-guided latent feature update, targeting the first... Next iteration: First, determine the time step corresponding to the current iteration. and discrete evolution step size The specific formula is as follows: Secondly, Noisy Joint Latent Features Corresponding to Time Steps and Input conditional velocity field predictor, obtain Corresponding predicted latent space velocity field and intermediate feature set The specific formula is as follows: , and when hour, Finally, regarding Update and get Joint latent features of the next iteration The specific formula is as follows: Fourth, after After several joint iterations, the enhanced joint latent features are obtained. and the intermediate feature set ,exist In the first joint iteration This iteration is only used for advancing the latent feature state and does not participate in the gradient backpropagation of the fusion loss. The next iteration participates in backpropagation, where In this embodiment, the number of iterations for gradient backpropagation is... .

[0043] S4. Construct a conditional velocity field predictor CVFP, which has a U-shaped hierarchical structure and a core feature processing unit based on time conditions.

[0044] Furthermore, in S4, a conditional velocity field predictor (CVFP) is constructed, the specific process of which includes...

[0045] S41 and CVFP have a U-shaped hierarchical structure, including overlapping convolutional embeddings, encoding paths, bottleneck paths, decoding paths, cross-layer skip connections, and velocity field output mappings. The feature processing units in the encoding, bottleneck, and decoding paths are all Temporal Conditional Feature Processing Units (TCTBs). For the overall U-shaped hierarchical structure, the specific process for the k-th iteration is as follows: First, for the k-th iteration... The noisy joint latent features corresponding to the next iteration are embedded by overlapping convolution to obtain the initial embedded features. Second, in the encoding path, the initial embedded features are encoded at multiple scales, with a total encoding scale of [missing information]. In this embodiment, Regarding the first Each coding scale, and The specific formula is as follows: , ,in, Indicates the first Output features at each scale, and when hour , Indicates the first Multiple TCTBs are concatenated in one coding scale. Indicates the first The encoded features output at each encoding scale Indicates the first Output features at each scale Indicates the first Downsampling operations at each scale, after After several encoding scale iterations, the final output features of the encoding scale are obtained. In this embodiment, At that time, the number of TCTBs connected in series is , At that time, the number of TCTBs connected in series is , At that time, the number of TCTBs connected in series is Third, Input a bottleneck path consisting of multiple TCTBs connected in series to obtain bottleneck features. In this embodiment, the number of TCTBs connected in series is: Fourth, in the decoding path, bottleneck features are decoded at multiple scales, with a total decoding scale of [missing information]. Regarding the first Each decoding scale, and The specific formula is as follows: , ,in, For the first Features after upsampling and skip connections at each decoding scale For the first Upsampling operation at each decoding scale For the first Output features of each decoding scale, and hour, , In order to be with the first The encoded features output by the encoding scale corresponding to each decoding scale. For the first Output features of each decoding scale For the first In this embodiment, multiple TCTBs are concatenated within a single decoding scale. At that time, the number of TCTBs connected in series is , At that time, the number of TCTBs connected in series is , At that time, the number of TCTBs connected in series is ,go through Each decoding scale yields the final output features of that decoding scale. Fifth, Mapped to latent space velocity field In this embodiment, the kernel size of the convolution operation is . Simultaneously, the bottleneck features and the output features of each decoding scale are combined to form an intermediate feature set. .

[0046] S42. For the TCTB part, it is used to perform temporal conditional modulation, self-attention modeling, and feedforward update on the input features according to the current time step. TCTB is a two-branch structure, including a self-attention branch and a feedforward branch. For the k-th iteration, the specific process is as follows: First, through... The specific formula for generating modulation parameters is as follows: ,in, For time embedding functions, For linear mapping layer, and These are the offset modulation parameters and scaling modulation parameters for the self-attention branch, respectively. and The first part describes the offset modulation parameters and scaling modulation parameters of the feedforward branch; the second part describes the input features of the TCTB. The updated features are obtained by performing temporal conditional modulation in the self-attention branch. The specific formula is as follows: ,in, For self-attention operations, For layer normalization operation; third, for The output characteristics of TCTB are obtained by performing time-conditional modulation in the feedforward branch. The specific formula is as follows: ,in, This represents a feedforward network.

[0047] S5. Construct a hierarchical selective multi-stream fusion enhancement head (HSMF), including intermediate feature fusion, latent feature fusion, and fusion enhancement feature generation.

[0048] Furthermore, in S5, a hierarchical selective multi-stream fusion enhancement head (HSMF) is constructed, the specific process of which includes...

[0049] S51 and HSMF include intermediate feature fusion, latent feature fusion, and fusion-enhanced feature generation. For the intermediate feature fusion part, the input is... The specific process is as follows: First, traverse the layers in order from deepest to shallowest. Each intermediate feature in, where Indicates intermediate features; second, for hierarchy The intermediate features are used for cross-level feature residual fusion, and the specific formula is as follows: ,in Representative level The fusion of input features hierarchical The third step involves fusing and enhancing the output features; and finally, performing dual-round selective scanning and feedforward updating on the fused input features. , ,in, and Representing levels The output features are obtained through dual-wheel selective scanning and feedforward updating. Fourth, the output features of the feedforward update are split into visible light responses according to the channel dimension, using a dual-wheel selective scanning method. and infrared response Then calculate the hierarchy. Fusion Enhancement Output Features ,in This indicates selective kernel feature fusion, which is a multi-branch fusion based on selective weights; fifth, the second, third, and fourth processes... The next iteration yields the final fused and enhanced output features. .

[0050] S52. For dual-wheel selective scanning, the specific process is as follows: First, the input features are split according to the channel dimension and local position enhancement is performed to obtain the position-enhanced visible light two-dimensional features. and infrared two-dimensional features In this embodiment, local location enhancement is achieved by a convolution kernel size of... The first part consists of depthwise convolutions and residual connections; the second part unfolds the visible light and infrared two-dimensional features into one-dimensional sequences respectively. and The specific formula is as follows: ,in The third step involves performing a flattening operation on the visible light and infrared one-dimensional sequences, respectively, to obtain the first-stage visible light output sequence after the scan. and infrared output sequence The specific formula is as follows: ,in The fourth step involves performing a second-stage Mamba selective state-space scan on the interleaved visible and infrared output sequences of the first stage to obtain a second-stage mixed output sequence. The specific formula is as follows: ,in Represents interleaving operations. Fifth, the mixed output sequence represents the embedding of learnable modal sources. By deinterlacing and restoring the one-dimensional sequence to a two-dimensional sequence, the output features of the dual-wheel selective scanning are obtained after splicing along the channel dimension.

[0051] S53. For the latent feature fusion and fusion-enhanced feature generation part, the specific process is as follows: First, the enhanced joint latent features are... Decomposed along the channel dimension into enhanced visible light latent features and infrared latent features And perform fusion enhancement to obtain latent feature fusion enhancement features. Second, the final fusion enhances the output features. Enhanced results of latent feature fusion Channel stitching and normalization are performed to obtain the enhanced UAV image. In this embodiment, the normalized output is determined by the convolution kernel size of The convolution and tanh activation function are connected in series.

[0052] S6. Construct a multi-objective joint optimization function for end-to-end training of FDSAE, RFLE, and HSMF. The multi-objective joint optimization function includes aggregated enhancement loss and corrected flow velocity field regression term. The aggregated enhancement loss includes image content preservation term, cross-modal edge enhancement term, and structural consistency term.

[0053] Furthermore, in S6, a multi-objective joint optimization function is constructed for end-to-end training of FDSAE, RFLE, and HSMF. The specific process includes: First, the multi-objective joint optimization function includes an aggregated enhancement loss and a corrected flow velocity field regression term. The aggregated enhancement loss includes an image content preservation term, a cross-modal edge enhancement term, and a structural consistency term, which respectively constrain the enhancement results' ability to preserve intensity information, edge details, and structural similarity. The corrected flow velocity field regression term is used to constrain the velocity prediction process of the conditional velocity field predictor. Second, in one training round, the corrected flow velocity field regression term only applies to the initial conditional velocity prediction of the corrected flow latent space evolution module, providing overall directional constraints. In this embodiment... ,in For a multi-objective joint optimization function, To aggregate and enhance loss, and ,in, The coefficients of the image content preservation term and , For image content preservation items, The coefficients of the cross-modal edge enhancement term and , For cross-modal edge enhancement terms, The coefficient of the structural consistency term and , As a structural consistency term, in this embodiment, the image content preservation term is calculated as follows: ,in To take the maximum value for each pixel, For L1 norm, the cross-modal edge enhancement term is calculated as follows: The structural consistency term is calculated as follows: ,in and This represents the structural weights adaptively assigned based on the visible light and infrared gradient responses. This is a function for calculating the structural similarity index.

[0054] Furthermore, a corrected flow-driven multimodal UAV image enhancement method is developed using the PyCharm application and Python code, employing the PyTorch framework. The model input is a pair of visible light and bimodal images with a resolution of 512×512.

[0055] The above are merely preferred embodiments of the present invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the inventive concept of the present invention, and these modifications and improvements all fall within the protection scope of the present invention.

Claims

1. A corrected flow-driven multimodal UAV image enhancement method, characterized in that, Includes the following steps: S1. Control the UAV to fly over the area to be observed to obtain the original visible light image and the original infrared image; then perform preprocessing to obtain the preprocessed visible light and infrared UAV image pairs. S2. Construct a disparity-tolerant structural anchor co-encoder (FDSAE), including multi-level latent feature encoding, cross-modal structural anchor extraction, local disparity offset estimation, anchor constraint space calibration, and master-slave confidence coding output. For the cross-modal structural anchor extraction part, it is used to calculate the edge response and structural abrupt response of visible light and infrared features to generate a structural anchor map. The cross-modal structural anchor extraction part is only for the hierarchical level. Modality The specific process is as follows: First, extract the edges, contours, or texture boundaries from visible light and infrared features. The specific formula is: ,in, Representative level Modal The output characteristics, For the absolute value operation, and These represent the first-order gradient operators in the horizontal and vertical directions, respectively. The first step is to create an edge response map; the second step is to capture locations in the image where brightness, texture, or thermal radiation intensity changes rapidly, using the following formula: ,in It is a second-order gradient operator. The first step is to generate the structural catastrophe response map. Then, the edge response map and the structural catastrophe response map are concatenated and input into the structural anchor point generation network to obtain the structural anchor point map. The specific formula is as follows: ,in Representative level The structural anchor point generation network, It is the Sigmoid activation function. This indicates a channel stitching operation. For the local disparity shift estimation part, it is used to predict the local displacement of infrared features relative to visible light features at each spatial location, enabling subsequent infrared auxiliary features to align with the dominant visible light mode. The local disparity shift estimation part only applies to the hierarchical level. The specific process is as follows: First, calculate the cross-modal structure correlation diagram. ,in, and These are structural anchor point diagrams showing visible light and infrared characteristics, respectively. For convolution operations, This is element-wise multiplication; subsequently, based on the aforementioned cross-modal structure correlation diagram... Predicting local disparity offset fields using a convolutional offset prediction network The anchor point constraint space calibration section is used to align the structural information in the infrared auxiliary mode to the spatial coordinates of the visible light main mode, reducing local misalignment. This anchor point constraint space calibration section only applies to the hierarchical level. The specific process is as follows: First, the sampling coordinates of the infrared features are determined based on the local parallax offset field. The specific formula is: ,in, Represents the spatial position in the visible light characteristic coordinate system. Indicates the location Local disparity offset at that location This represents the infrared feature sampling coordinates obtained through mapping; secondly, spatial sampling calibration is performed on the infrared features and the structural anchor point map of the infrared features, using the following formula: , ,in, This indicates the calibrated infrared signature. Indicates position The calibrated infrared signature This is the calibrated infrared structure anchor point diagram. Indicates position The calibrated infrared structural anchor point value, Infrared mode output characteristics, For bilinear interpolation sampling, finally, the structural consistency weight graph is calculated. ,in Indicates exponentiation; S3. Construct the corrected flow latent space evolution module RFLE, including latent space noise addition, conditional velocity prediction, and velocity-guided latent feature update; S4. Construct a conditional velocity field predictor (CVFP) to provide time-related velocity field guidance for correcting the latent space evolution process. The predictor adopts a U-shaped hierarchical structure, constructs multi-scale feature representations through encoding paths, bottleneck paths, and decoding paths, and introduces time-conditional feature processing units at each level to achieve joint modeling of time information and spatial features. At the same time, it fully preserves shallow detail information and deep semantic information through cross-layer jump connections, and outputs the latent space velocity field and multi-scale intermediate feature set. S5. Construct a hierarchical selective multi-stream fusion enhancement head (HSMF), including intermediate feature fusion, latent feature fusion, and fusion enhancement feature generation. The hierarchical selective multi-stream fusion enhancement head (HSMF) is used to synergistically enhance the multi-scale intermediate features generated during the correction flow evolution process and the enhanced dual-modal latent features. This enhancement head gradually aggregates feature information from different levels through a cross-level residual fusion mechanism, and establishes a long-distance dependency between visible light and infrared response using dual-wheel selective scanning. On this basis, adaptive enhancement of multi-source information is achieved through selective kernel feature fusion. S6. Construct a multi-objective joint optimization function for end-to-end training of FDSAE, RFLE, and HSMF. The multi-objective joint optimization function includes aggregated enhancement loss and corrected flow velocity field regression term. The aggregated enhancement loss includes image content preservation term, cross-modal edge enhancement term, and structural consistency term.

2. The corrected flow-driven multimodal UAV image enhancement method according to claim 1, characterized in that, In step S2, a disparity-tolerant structured anchor co-encoder FDSAE ​​is constructed: for the multi-level latent feature coding part, it is used to extract visible light and infrared features at different levels, with the input being the preprocessed visible light image. and infrared images The specific process is as follows: First, the multi-level latent feature encoding part has a multi-level structure and includes visible light and infrared dual-mode branches. Let the total number of levels be... The multi-level latent feature encoding part will be executed iteratively in the visible light and infrared modes respectively. Next, regarding the hierarchy Modality The iterative process is as follows: ,in, Representative level Modal The output characteristics, Representative level Modal The output characteristics, Representative level Modal The feature extraction network; secondly, the hierarchy Modal The feature extraction network has the following structure: First, it extracts... First, local texture features are obtained through a local convolutional extraction layer. Second, structural enhancements are performed on the local texture features in the horizontal, vertical, and neighborhood directions. Third, the results of the structural enhancements in the horizontal, vertical, and neighborhood directions are fused to obtain directional enhancement features. Fourth, channel weights are calculated using the directional enhancement features, and channel modulation is applied to the directional enhancement features to obtain adaptive enhancement features. Fifth, the adaptive enhancement features and... Residual fusion yields For the primary and secondary confidence coding output section, only the hierarchy is considered. The specific process is as follows: First, calculate the confidence weights for the visible light and infrared modes, using the following formula: , , ,in, To predict the response with confidence, The output characteristics are those of the visible light mode. and They are respectively In the response components of visible and infrared modes, and These represent the confidence weights for the visible light and infrared modes, respectively. Represent the convolutional mapping unit; subsequently, compute the latent features at the visible light ends. and infrared terminal latent features The specific formula is as follows: , .

3. The corrected flow-driven multimodal UAV image enhancement method according to claim 1, characterized in that, In step S3, the corrected flow latent space evolution module RFLE is constructed. RFLE includes latent space denoising, conditional velocity prediction, and velocity-guided latent feature update. The specific process is as follows: First, latent space denoising is performed using standard Gaussian noise latent variables. Combined latent features with dual modes A corrected linearity-adding noisy state is constructed between these states, and the time steps obtained through random sampling are obtained. Noisy Joint Latent Features The specific formula is as follows: ,in express to Uniform distribution over the interval; second, and Input the conditional velocity field predictor to obtain the initial velocity field prediction result. and according to Calculate the regression term of the corrected flow velocity field The specific formula is as follows: , ,in, This represents a conditional velocity field predictor and shares parameters within RFLE. Represents the L2 norm; third, for go through The joint iteration consists of several iterations, each including conditional velocity prediction and velocity-guided latent feature update, with the third iteration being the specific iteration. Next iteration: First, determine the time step corresponding to the current iteration. and discrete evolution step size The specific formula is as follows: Secondly, Noisy Joint Latent Features Corresponding to Time Steps and Input conditional velocity field predictor, obtain Corresponding predicted latent space velocity field and intermediate feature set The specific formula is as follows: , and when hour, Finally, regarding Update and get Joint latent features of the next iteration The specific formula is as follows: Fourth, after After several joint iterations, the enhanced joint latent features are obtained. and the intermediate feature set In the Kth joint iteration, the first This iteration is only used for advancing latent feature states and does not participate in the gradient backpropagation of the aggregation enhancement loss. The next iteration participates in backpropagation, where This represents the number of iterations involved in gradient backpropagation.

4. The corrected flow-driven multimodal UAV image enhancement method according to claim 1, characterized in that, In step S4, the conditional velocity field predictor CVFP is constructed: S41 and CVFP have a U-shaped hierarchical structure, including overlapping convolutional embeddings, encoding paths, bottleneck paths, decoding paths, cross-layer skip connections, and velocity field output mappings. The feature processing units in the encoding, bottleneck, and decoding paths are all Temporal Conditional Feature Processing Units (TCTBs). For the overall U-shaped hierarchical structure, the specific process for the k-th iteration is as follows: First, for the k-th iteration... The noisy joint latent features corresponding to the next iteration are embedded using overlapping convolutions to obtain the initial embedded features. Second, in the encoding path, the initial embedded features are encoded at multiple scales, with a total encoding scale of [missing information]. Regarding the first Each coding scale, and The specific formula is as follows: , ,in, Indicates the first Output features at each scale, and when hour , Indicates the first Multiple TCTBs are concatenated in one coding scale. Indicates the first The encoded features output at each encoding scale Indicates the first Output features at each scale Indicates the first Downsampling operations at each scale, after After several encoding scale iterations, the final output features of the encoding scale are obtained. Third, Input a bottleneck path consisting of multiple TCTBs connected in series to obtain bottleneck features. Fourth, in the decoding path, bottleneck features are decoded at multiple scales, with a total decoding scale of [missing information]. Regarding the first Each decoding scale, and The specific formula is as follows: , ,in, For the first Features after upsampling and skip connections at each decoding scale For the first Upsampling operation at each decoding scale For the first The output features of each decoding scale, and hour, , In order to be with the first The encoded features output by the encoding scale corresponding to each decoding scale. For the first Output features of each decoding scale For the first Multiple TCTBs are concatenated in each decoding scale, after... Each decoding scale yields the final output features of that decoding scale. Fifth, Mapped to latent space velocity field Simultaneously, the bottleneck features and the output features of each decoding scale are combined to form an intermediate feature set. ; S42. For the TCTB part, it is used to perform temporal conditional modulation, self-attention modeling, and feedforward update on the input features according to the current time step. TCTB is a two-branch structure, including a self-attention branch and a feedforward branch. For the k-th iteration, the specific process is as follows: First, through... The specific formula for generating modulation parameters is as follows: ,in, For time embedding functions, For linear mapping layer, and These are the offset modulation parameters and scaling modulation parameters for the self-attention branch, respectively. and The first part describes the offset modulation parameters and scaling modulation parameters of the feedforward branch; the second part describes the input features of the TCTB. The updated features are obtained by performing temporal conditional modulation in the self-attention branch. The specific formula is as follows: ,in, For self-attention operations, For layer normalization operation; third, for The output characteristics of TCTB are obtained by performing time-conditional modulation in the feedforward branch. The specific formula is as follows: ,in, This represents a feedforward network.

5. The corrected flow-driven multimodal UAV image enhancement method according to claim 1, characterized in that, In step S5, the Hierarchical Selective Multistream Fusion Enhancement Header (HSMF) is constructed: S51 and HSMF include intermediate feature fusion, latent feature fusion, and fusion-enhanced feature generation. For the intermediate feature fusion part, the input is... The specific process is as follows: First, traverse the layers in order from deepest to shallowest. Each intermediate feature in, where Indicates intermediate features; second, for hierarchy The intermediate features are used for cross-level feature residual fusion, and the specific formula is as follows: ,in Representative level The fusion of input features hierarchical The third step involves fusing and enhancing the output features; and finally, performing dual-round selective scanning and feedforward updating on the fused input features. , ,in, and Representing levels The output features are obtained through dual-wheel selective scanning and feedforward updating. Fourth, the output features of the feedforward update are split into visible light responses according to the channel dimension, using a dual-wheel selective scanning method. and infrared response Then calculate the hierarchy. Fusion Enhancement Output Features ,in This indicates selective kernel feature fusion, which is a multi-branch fusion based on selective weights; fifth, the second, third, and fourth processes... The next iteration yields the final fused and enhanced output features. ; S52. For dual-wheel selective scanning, the specific process is as follows: First, the input features are split according to the channel dimension and local position enhancement is performed to obtain the position-enhanced visible light two-dimensional features. and infrared two-dimensional features Second, the two-dimensional features of visible light and infrared light are respectively unfolded into one-dimensional sequences. and The specific formula is as follows: ,in The third step involves performing a flattening operation on the visible light and infrared one-dimensional sequences, respectively, to obtain the first-stage visible light output sequence after the scan. and infrared output sequence The specific formula is as follows: ,in The fourth step involves performing a second-stage Mamba selective state-space scan on the interleaved visible and infrared output sequences of the first stage to obtain a second-stage mixed output sequence. The specific formula is as follows: ,in Represents interleaving operations. Fifth, the mixed output sequence represents the embedding of learnable modal sources. By deinterlacing and restoring the two-dimensional sequence from the one-dimensional sequence, the output features of the dual-wheel selective scanning are obtained by splicing along the channel dimension. S53. For the latent feature fusion and fusion-enhanced feature generation part, the specific process is as follows: First, the enhanced joint latent features are... Decomposed along the channel dimension into enhanced visible light latent features and infrared latent features And perform fusion enhancement to obtain latent feature fusion enhancement features. Second, the final fusion enhances the output features. Enhanced results of latent feature fusion Channel stitching and normalization are performed to obtain the enhanced UAV image. .

6. The corrected flow-driven multimodal UAV image enhancement method according to claim 1, characterized in that, In step S6, a multi-objective joint optimization function is constructed: First, the multi-objective joint optimization function includes an aggregated enhancement loss and a corrected flow velocity field regression term. The aggregated enhancement loss includes an image content preservation term, a cross-modal edge enhancement term, and a structural consistency term, which respectively constrain the enhancement results in terms of preserving intensity information, edge details, and structural similarity. The corrected flow velocity field regression term is used to constrain the velocity prediction process of the conditional velocity field predictor. Second, in one training round, the corrected flow velocity field regression term only applies to the first conditional velocity prediction of the corrected flow latent space evolution module, providing an overall directional constraint.

Citation Information

Patent Citations

  • Multi-view driven three-dimensional generation method

    CN120259547A

  • Infrared and visible light image fusion method for composite degraded scene

    CN122222832A