Complete polarization synthetic aperture radar and visible light image text description method

Through deep learning technology and cross-modal self-attention mechanism, the resolution difference and noise interference problems in the fusion of fully polarized synthetic aperture radar and visible light images are solved, and efficient text description of multimodal remote sensing images is achieved, and the intelligent semantic analysis capability of remote sensing images is improved.

CN120339781AActive Publication Date: 2025-07-18CHINA UNIV OF MINING & TECH +1

Patent Information

Application Number
CN202510822446.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-07-18
Estimated Expiration
2045-06-19

AI Technical Summary

Technical Problem

The prior art is difficult to fully integrate the multimodal features of fully polarized synthetic aperture radar and visible light images, resulting in resolution differences, noise interference and information loss in remote sensing image processing, making it difficult to achieve efficient image text description.

Method used

Using deep learning technology, multi-layer feature extraction and natural language description of fully polarized synthetic aperture radar and visible light images are achieved through dual encoder-decoder structure and cross-modal self-attention mechanism, combined with Li's filtered denoising, adaptive weighted feature fusion and text generation model.

Benefits of technology

It significantly improves the fusion feature expression ability and the accuracy of text description of multimodal remote sensing images, adapts to complex remote sensing environments, and is applied to environmental monitoring, disaster assessment and target recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339781A_ABST
    Figure CN120339781A_ABST
Patent Text Reader

Abstract

The invention discloses a complete polarization synthetic aperture radar and visible light image text description method. The method comprises the following steps: acquiring a complete polarization synthetic aperture radar image and a visible light image; carrying out Lee filtering denoising processing on the image, fusing the features of a plurality of polarization channels, inputting the fused features into a visual converter encoder for feature extraction, and carrying out multi-layer converter feature extraction on the visible light image; and performing weighted fusion on the multilayer features of the visible light image and the synthetic aperture radar image, and inputting the fused multi-modal features into a text to a text migration converter model text encoder to generate natural language description. According to the method, the complementary advantages of the fully polarized synthetic aperture radar and the visible light image are fully exerted through the deep learning technology, dynamic fusion and semantic interaction of multi-modal features are achieved, the accuracy and robustness of automatic text description of the remote sensing image are remarkably improved, and the method is suitable for multi-modal remote sensing image analysis and understanding in a complex environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of remote sensing image processing and natural language generation, and particularly relates to a method for text description of fully polarimetric synthetic aperture radar and visible light images. Background Art

[0002] With the rapid development of remote sensing technology and artificial intelligence, fully polarimetric synthetic aperture radar (SAR) and visible light images, as two important remote sensing data sources, play a crucial role in fields such as surface monitoring, environmental monitoring, and military reconnaissance. The data of fully polarimetric synthetic aperture radar has the advantages of penetrating clouds and operating all-weather, and can provide the electromagnetic scattering characteristics of targets, while visible light images have rich texture and color information, which is convenient for intuitively identifying ground object features. The combination of the two can achieve multimodal information complementarity and improve the understanding and analysis ability of remote sensing images.

[0003] Traditional multimodal remote sensing image fusion methods mostly rely on manually designed feature extraction and fusion strategies, which are difficult to fully capture the complex semantic associations between different modalities, and have problems of low efficiency, limited fusion accuracy, and insufficient robustness when dealing with high-dimensional large-scale data. In addition, there are significant differences in resolution, observation angle, noise nature, etc. between synthetic aperture radar and visible light images, resulting in widespread information loss, registration errors, and noise interference in the data fusion process, seriously restricting the comprehensive application effect of multimodal remote sensing images.

[0004] In recent years, the development of deep learning technology has provided new solutions for the fusion and understanding of multimodal remote sensing data. Based on the encoder-decoder architecture of deep neural networks, by automatically learning multi-level and multi-scale feature representations, the internal associations and complementary information of different modality data can be effectively mined. At the same time, by using advanced technologies such as adaptive weighting mechanisms, attention mechanisms, and generative adversarial networks, the modality weights can be dynamically adjusted, information loss can be compensated, and registration errors can be corrected, significantly improving the accuracy and robustness of fusion.

[0005] However, existing methods still have deficiencies in the fine-grained level of multimodal feature fusion and cross-modal information interaction, and are difficult to fully utilize the advantages of fully polarimetric synthetic aperture radar and visible light images, especially limited in achieving accurate and natural image text description. To address this challenge, designing a multimodal fusion text description method based on a dual encoder-decoder structure, fusing the deep semantic features of fully polarimetric synthetic aperture radar and visible light images, and leveraging advanced cross-modal self-attention mechanisms and text generation architectures will greatly promote the development and application of remote sensing image intelligent interpretation technology. Summary of the Invention

[0006] The object of the present invention is to provide a method for text description of fully polarimetric synthetic aperture radar and visible light images. This method realizes the efficient fusion and semantic understanding of multi-modal data through deep learning technology, solves the problems existing in the resolution difference, registration error, noise interference, etc. between fully polarimetric synthetic aperture radar and visible light images, and improves the expression ability of fusion features and the accuracy and naturalness of text description. This method can adapt to complex remote sensing environments, realize intelligent semantic analysis and high-quality text generation of multi-modal remote sensing images, and is widely used in fields such as environmental monitoring, disaster assessment, and target recognition.

[0007] To achieve the above object, the present invention provides a method for text description of fully polarimetric synthetic aperture radar and visible light images, including the following steps: S1. Collect visible light images and fully polarimetric synthetic aperture radar images of the same scene, where the fully polarimetric synthetic aperture radar images include multiple polarization channels (such as HH, HV, VH, VV); S2. Perform Lee filtering denoising processing on each polarization channel of the fully polarimetric synthetic aperture radar image to remove noise and speckles; S3. Fuse the denoised multi-channel polarization features to form multi-channel fusion features, and input them into a vision transformer encoder for multi-layer transformer feature extraction; S4. Use a contrastive language-image pre-training model encoder to perform multi-layer transformer feature extraction on the visible light image; S5. Through a hierarchical feature fusion method with adaptive weighting coefficients, perform weighted fusion on the multi-layer transformer features of the visible light image and the synthetic aperture radar image respectively, and dynamically adjust the contribution weights of each layer of features; S6. Design a cross-modal self-attention mechanism to achieve deep bidirectional information interaction between the fusion features of the visible light image and the fusion features of the synthetic aperture radar image, and further improve the expression ability of the multi-modal fusion features; S7. Input the fused multi-modal features into a text encoder based on the text-to-text transfer transformer model architecture to generate corresponding natural language text descriptions; S8. Combine a temperature adjustment mechanism and a reinforcement learning strategy based on the reinforcement learning policy gradient algorithm to optimize the diversity and accuracy of the generated text and improve the final description quality.

[0008] Preferably, the step of performing Lee filtering denoising processing on each polarization channel of the fully polarimetric synthetic aperture radar image in S2 is: S2.1-1. Collect the full-polarization synthetic aperture radar image data under the same ground object scene, which includes four polarization channels, namely horizontal transmit - horizontal receive (HH), horizontal transmit - vertical receive (HV), vertical transmit - horizontal receive (VH), and vertical transmit - vertical receive (VV). The original images of these four polarization channels are respectively expressed as: , ; Each polarization channel reflects different scattering characteristics of the ground object and is the basis for subsequent fusion analysis; S2.1-2. For each polarization channel image, use the Lee filter algorithm for noise suppression and speckle removal. The Lee filter is based on local statistical characteristics. By calculating the local mean and variance, it dynamically adjusts the filtering weights to distinguish the signal and noise parts of the image. The specific filtering calculation formula is as follows: ; Among them, represents the pixel value of the original polarization channel to be processed; represents the pixel mean within a small window centered on pixel and reflects the average intensity of the local image; represents the variance of the pixels within this small window and reflects the degree of change in local intensity; K is the estimated value of the noise variance, usually obtained through statistical methods or prior knowledge.

[0009] The core idea of this filter is to use local statistical information to judge the ratio of signal to noise in the pixel value, calculate a weight factor to balance the relationship between maintaining signal edges and suppressing noise, thereby effectively reducing the influence of speckle noise in the synthetic aperture radar image.

[0010] S2.1-3. Perform Lee filtering on the four polarization channels respectively to obtain the denoised polarization channel images: ; Each denoised image has a high signal-to-noise ratio, retains the unique scattering information and image details of polarization, and provides quality assurance for the subsequent fusion of polarization channels and high-level feature extraction.

[0011] Preferably, the steps of fusing the denoised multiple polarization channel features in S3 specifically include: S3.1-1. After completing the Lee filtering denoising process for each polarization channel (HH, HV, VH, VV) of the full-polarization synthetic aperture radar image, we obtain four high-quality denoised polarization features, which are respectively expressed as , these four polarization channels respectively reflect the electromagnetic wave scattering characteristics of ground objects from different polarization angles and have obvious complementary information. It is difficult to fully explore the internal correlation between polarization information by separately processing the characteristics of these channels, which limits the semantic understanding and target detection performance of subsequent remote sensing images. By introducing learnable weight parameters different importance weights are assigned to the four polarization channels, and these weight parameters are automatically optimized through model training to adapt to different scenarios and task requirements. The weights satisfy the normalization constraint: ; S3.1-2. Concatenate the weighted four polarization channel features in the channel dimension to form a complete fused feature representation: ; The dimension of the concatenated feature vector is four times that of the single-channel feature dimension, retaining the fine-grained information of each polarization channel, enhancing the richness and expressiveness of the features. This weighted concatenation fusion method overcomes the problems of information redundancy and feature conflict that may be caused by simple concatenation or summation, fully retains the diversity and complementarity of the polarization channels, and at the same time enhances the discriminability and robustness of the features through adaptive weight adjustment; S3.1-3. The vision transformer encoder is mainly composed of multiple transformer modules. Each layer of the transformer abstracts and enhances the input features through the multi-head self-attention mechanism and the feed-forward network. Assume that the vision transformer encoder contains L layers of transformer modules, and the output feature of each layer is represented as: ; Among them, represents the polarized features after denoising and fusion of synthetic aperture radar data.

[0012] The feature dimension of each layer of the transformer module remains the same, which is convenient for the subsequent transfer and fusion of information between layers. Multiple layers of transformer modules sequentially capture the local and global context information of the image, gradually refining more advanced and more semantic visual features. The finally obtained multi-layer feature sequence serves as the multi-scale and multi-level feature representation of the synthetic aperture radar image.

[0013] Preferably, the steps of performing multi-layer transformer feature extraction on the visible light image in S4 specifically include: S4.1-1. Collect a sequence of visible light images in the same scene. Assume that its original input data is represented as . Since the pixel value range and resolution of visible light images may vary, before performing feature extraction, preprocessing operations are performed on the input images, including normalization processing. The specific normalization formula is: ; Among them, and are the mean and standard deviation of the visible light image pixels respectively, which are used to unify the numerical scales between different images; S4.1-2. Input the normalized visible light image into the pre-trained contrastive language-image pre-training model encoder for feature extraction. The contrastive language-image pre-training model encoder is mainly composed of multiple Transformer modules. Each layer of the Transformer abstracts and enhances the input features through the multi-head self-attention mechanism and the feed-forward network. Assume that the contrastive language-image pre-training model encoder contains L layers of Transformer modules, and the output feature of each layer is expressed as: ; Among them, represents the input image features after preliminary linear transformation and position encoding; S4.1-3. The feature dimensions of each layer of the Transformer module are kept consistent to facilitate the subsequent transfer and fusion of information between layers. The multiple Transformer modules sequentially capture the local and global context information of the image, gradually refining more advanced and semantic visual features. The finally obtained multi-layer feature sequence serves as the multi-scale and multi-level feature representation of the visible light image.

[0014] Preferably, the steps of the hierarchical feature fusion method using adaptive weighting coefficients in S5 specifically include: S5.1-1. Both the visible light image encoder (using the contrastive language-image pre-training model architecture) and the full-polarization synthetic aperture radar image encoder (using the vision Transformer architecture) contain L layers of Transformer modules. The output features of each layer of the module carry the expressions of different scales and semantic levels of the input image at that layer, forming serialized feature representations: ; Among them, N and M are the lengths of the visible light and synthetic aperture radar feature sequences respectively, and d is the feature dimension; S5.1-2. To overcome the drawback that the contributions of the feature expressions of each layer to the final task are unbalanced, a dynamic weight is assigned to each layer of features. The weight is obtained through training, specifically by designing a multi-layer perceptron (MLP) network for each layer of features to calculate the weight scores: ; Among them, and represent the weight scores of the k-th layer in the visible light and synthetic aperture radar modalities respectively. To ensure the rationality and comparability of the weights, the softmax function is used to normalize the weight scores of all layers: ; In this way, the sum of the weights of all layers is 1, ensuring stability during fusion; S5.1-3. Use the normalized weights to perform weighted summation on the features of each layer to obtain the single-modal fusion features: ; This fusion method comprehensively considers the contributions of different layers, taking into account both the detailed information of the shallow layer and the semantic information of the deep layer, enhancing the richness and discriminability of the fusion features.

[0015] S5.1-4. To avoid inconsistent feature scales after weighted fusion, further perform layer normalization on the fusion result: ; Normalization enhances the numerical stability of the features, facilitating subsequent cross-modal fusion calculations.

[0016] Preferably, the steps of implementing the deep fusion of the visible light image and the full-polarization synthetic aperture radar image features by using the cross-modal self-attention mechanism in S6 are specifically as follows: S6.1-1. Assume that after the hierarchical feature fusion with the adaptive weighting coefficient, the visible light modal fusion feature and the synthetic aperture radar modal fusion feature are respectively obtained, where N and M are the lengths of the visible light and synthetic aperture radar feature sequences, d is the feature dimension, and the two-modal fusion features contain the high-level semantic information extracted by the multi-layer transformer modules of their respective encoders, which are the basis for subsequent cross-modal fusion; S6.1-2. To achieve the deep semantic interaction of the visible light modality with respect to the synthetic aperture radar modality, design a cross-modal self-attention mechanism. Specifically, map the visible light modal fusion feature to a query matrix, and map the synthetic aperture radar modal fusion feature to a key and a value matrix. The calculation formula is as follows: ; Among them, is the linear projection matrix learned by the model, which projects the input feature into the attention space; S6.1-3. Based on the above , calculate the cross-modal self-attention weights, and weight the query of the visible light modality against the key of the synthetic aperture radar modality to generate an attention response: ; S6.1-4. Perform a residual connection between the attention response and the fusion feature of the visible light modality to avoid information loss and promote gradient propagation during model training, forming an improved fusion feature: ; S6.1-5. For the final fusion feature It is passed as input to the subsequent text encoder to generate a natural language description that combines multimodal information, improving the accuracy and richness of the description.

[0017] Preferably, the step of inputting the fused multimodal features into the text encoder based on the text-to-text transfer Transformer model architecture in S7 specifically includes: S7.1-1. Assume that the final multimodal fusion feature representation obtained through deep fusion by the cross-modal self-attention mechanism is: ; where L represents the length of the feature sequence, d is the feature dimension, and this fusion feature comprehensively reflects the multi-layer semantic information extracted by the visible light and full-polarization synthetic aperture radar image encoders, containing rich multimodal context semantics, which is the basic input for the subsequent text encoder to generate descriptive text; S7.1-2. Input the fusion feature into the text encoder based on the text-to-text transfer Transformer model architecture. This text encoder consists of multiple Transformer encoding layers. Each layer deeply processes and semantically abstracts the input features through the self-attention mechanism and the feed-forward network. The output of the i-th layer is denoted as: ; where is the total number of encoder layers, and the initial input is ; S7.1-3. The output of the last layer of the encoder is passed as a context vector to the text decoder. The text decoder generates a natural language description based on the autoregressive mechanism, predicting the probability of each word in turn: ; where represents the sequence of previously generated words.

[0018] Preferably, the step of optimizing text generation by combining the temperature adjustment mechanism and the reinforcement learning strategy based on the reinforcement learning policy gradient algorithm in S8 specifically includes: S8.1-1. During the process of the text decoder based on the text-to-text transfer Transformer model architecture generating text, in order to improve the diversity of the generated text and control the randomness in the sampling process, a temperature adjustment mechanism is introduced. Specifically, for the unnormalized prediction value of the i-th word in the vocabulary generated by the decoder at time step t, its probability distribution is adjusted through the temperature parameter T>0, and the calculation formula is: ; Among them, V represents the vocabulary size, and the temperature parameter T adjusts the smoothness of the probability distribution. When T > 1, the probability distribution becomes flatter, and the sampling diversity increases; when T < 1, the probability distribution becomes sharper, and it is more inclined to select the word with the highest probability, and the generated text is more certain; S8.1-2. To directly optimize the quality and semantic relevance of the generated text, in combination with the reinforcement learning strategy based on the reinforcement learning policy gradient algorithm, a reward function is defined for evaluating the generated sequence The quality of the reward can be based on bilingual evaluation alternative metrics, recall-oriented summary evaluation metrics, or task-related metrics. The optimization goal of the model is to maximize the expected reward of the generated text: ; Among them, represents the model parameters, is the generation probability distribution based on the current parameters; S8.1-3. According to the reinforcement learning policy gradient algorithm, calculate the gradient of the model parameters: ; Use the sampled text sequence and its corresponding reward for gradient estimation, and optimize the model parameters through gradient ascent; S8.1-4. To reduce the variance of gradient estimation, introduce a baseline function b and improve the gradient estimation to: ; Among them, is the number of sampling times, is the m-th sampled text sequence, and the baseline usually selects the mean value of the rewards, which helps to stabilize the training process.

[0019] Through the combination of the temperature adjustment mechanism and the reinforcement learning strategy of the reinforcement learning policy gradient algorithm, the text generation process is made to have both rich diversity and be optimized in terms of quality and semantics, significantly improving the intelligent text description effect of multi-modal remote sensing images.

[0020] Beneficial effects: The present invention collects full-polarization synthetic aperture radar images and visible light image data in the same scene; performs denoising processing on each polarization channel of the collected full-polarization synthetic aperture radar images respectively using Lee filtering to reduce speckle noise interference, and at the same time performs normalization preprocessing and multi-layer transformer feature extraction on the visible light images; adopts a hierarchical feature fusion method with adaptive weighting coefficients to fuse the features of each polarization channel of the full-polarization synthetic aperture radar and the multi-layer features of the visible light respectively; designs a cross-modal self-attention mechanism to achieve deep interaction of the fusion features between the visible light and the full-polarization synthetic aperture radar, enhancing the semantic complementarity between modalities; inputs the fused multi-modal features into a text encoder based on the text-to-text transfer transformer model architecture for high-level semantic encoding and natural language text generation; introduces a temperature adjustment mechanism to control the sampling diversity of text generation, and combines a reinforcement learning strategy based on the reinforcement learning policy gradient algorithm to optimize the quality and diversity of the generated text. Through the method of the present invention, it is possible to effectively fuse remote sensing image data of different modalities, significantly improve the expression ability of the fusion features and the accuracy of text description, adapt to multi-modal information processing and intelligent semantic understanding in complex remote sensing environments, and be widely applied in fields such as environmental monitoring, disaster assessment, and target recognition. Brief Description of the Drawings

[0021] Figure 1 is a schematic flowchart of the present invention; Figure 2 is a structural diagram of a dual-encoder network model; Figure 3 is a schematic structural diagram of a cross-modal self-attention mechanism; Figure 4 is a flowchart of text generation based on the text-to-text transfer transformer model. Detailed Embodiments

[0022] The present invention will be further described below with reference to the accompanying drawings.

[0023] As Figure 1 shown, a method for text description of full-polarization synthetic aperture radar and visible light images includes the following steps: S1. Collect visible light images and full-polarization synthetic aperture radar images in the same scene, where the full-polarization synthetic aperture radar images include multiple polarization channels (such as HH, HV, VH, VV); S2. Perform Lee filtering denoising processing on each polarization channel of the full-polarization synthetic aperture radar images respectively to remove noise and speckles; S3. Fuse the features of the multiple denoised polarization channels to form a multi-channel fusion feature, and input it into a vision transformer encoder for multi-layer transformer feature extraction; S4. Use the contrastive language-image pre-trained model encoder to perform multi-layer Transformer feature extraction on visible light images; S5. Through the hierarchical feature fusion method with adaptive weighting coefficients, perform weighted fusion on the multi-layer Transformer features of visible light images and synthetic aperture radar images respectively, and dynamically adjust the contribution weights of each layer of features; S6. Design a cross-modal self-attention mechanism to achieve deep bidirectional information interaction between the fusion features of visible light images and the fusion features of synthetic aperture radar images, and further improve the expression ability of multi-modal fusion features; S7. Input the fused multi-modal features into the text encoder based on the text-to-text transfer Transformer model architecture to generate corresponding natural language text descriptions; S8. Combine the temperature adjustment mechanism and the reinforcement learning strategy based on the reinforcement learning policy gradient algorithm to optimize the diversity and accuracy of the generated text and improve the final description quality.

[0024] Further, the steps of performing Lee filtering denoising processing on each polarization channel of the full-polarization synthetic aperture radar image in S2 are as follows: S2.1-1. Collect full-polarization synthetic aperture radar image data of the same ground object scene, including four polarization channels, namely horizontal transmit-horizontal receive HH, horizontal transmit-vertical receive HV, vertical transmit-horizontal receive VH, and vertical transmit-vertical receive VV. The original images of these four polarization channels are respectively represented as: ; Each polarization channel reflects different scattering characteristics of the ground object and is the basis for subsequent fusion analysis; S2.1-2. For each polarization channel image, use the Lee filtering algorithm for noise suppression and speckle removal. The Lee filter is based on local statistical characteristics. By calculating the local mean and variance, the filtering weights are dynamically adjusted to distinguish the signal and noise parts of the image. The specific filtering calculation formula is as follows: ; Among them, represents the pixel value of the original polarization channel to be processed; represents the pixel mean within a small window centered on pixel , reflecting the average intensity of the local image; represents the variance of the pixels within this small window, reflecting the degree of change of the local intensity; K is the noise variance estimate value, usually obtained through statistical methods or prior knowledge; The core idea of this filter is to use local statistical information to judge the ratio of signal to noise in the pixel value, calculate a weight factor to balance the relationship between maintaining the signal edge and suppressing noise, so as to effectively reduce the influence of speckle noise in the synthetic aperture radar image; S2.1-3. Perform Lee filtering on the four polarization channels respectively to obtain the denoised polarization channel images: ; Each denoised image has a high signal-to-noise ratio, retains the scattering information and image details unique to polarization, and provides quality assurance for the subsequent fusion of polarization channels and high-level feature extraction; Furthermore, the steps of fusing the denoised multiple polarization channel features in S3 specifically include: S3.1-1. After completing the Lee filtering denoising process for each polarization channel (HH, HV, VH, VV) of the full-polarization synthetic aperture radar image, we obtain four high-quality denoised polarization features, which are respectively denoted as , and these four polarization channels respectively reflect the electromagnetic wave scattering characteristics of ground objects from different polarization angles, with obvious complementary information. It is difficult to fully explore the internal relationship between each polarization information by processing these channel features alone, which limits the semantic understanding and target detection performance of subsequent remote sensing images. By introducing learnable weight parameters to assign different importance weights to the four polarization channels, and this weight parameter is automatically optimized through model training to adapt to different scenarios and task requirements, and the weights satisfy the normalization constraint: ; S3.1-2. Concatenate the weighted four polarization channel features in the channel dimension to form a complete fused feature representation: ; The dimension of the concatenated feature vector is four times that of the single-channel feature dimension, retaining the fine-grained information of each polarization channel, enhancing the richness and expressiveness of the features. This weighted concatenation fusion method overcomes the problems of information redundancy and feature conflict that may be caused by simple concatenation or summation, fully retains the diversity and complementarity of the polarization channels, and at the same time enhances the discriminability and robustness of the features through adaptive weight adjustment; S3.1-3. The vision transformer encoder is mainly composed of multiple transformer modules. Each layer of the transformer abstracts and enhances the input features through the multi-head self-attention mechanism and the feed-forward network. Assuming that the vision transformer encoder contains L layers of transformer modules, the output feature of each layer is expressed as: ; Among them, represents the polarization feature after denoising and fusion of synthetic aperture radar data; The feature dimensions of each layer of the transformer module are kept consistent, facilitating the subsequent transfer and fusion of information between layers. The multi-layer transformer modules sequentially capture the local and global context information of the image, gradually refining higher-level and more semantic visual features, and finally obtaining a multi-layer feature sequence as a multi-scale and multi-level feature representation of the synthetic aperture radar image.

[0025] Furthermore, the steps of performing multi-layer transformer feature extraction on the visible light image in S4 specifically include: S4.1-1. Collect a sequence of visible light images in the same scene. Assume that its original input data is represented as . Since the pixel value ranges and resolutions of visible light images may vary, before performing feature extraction, preprocessing operations are performed on the input images, including normalization. The specific normalization formula is: ; where and are the mean and standard deviation of the pixels of the visible light image respectively, used to unify the numerical scales between different images; S4.1-2. Input the normalized visible light image into the encoder of the pre-trained contrastive language-image pre-training model for feature extraction. The encoder of the contrastive language-image pre-training model mainly consists of multi-layer transformer modules. Each layer of the transformer abstracts and enhances the input features through the multi-head self-attention mechanism and the feed-forward network. Assume that the encoder of the contrastive language-image pre-training model contains L layers of transformer modules, and the output feature of each layer is represented as: ; where represents the input image features after preliminary linear transformation and position encoding; S4.1-3. The feature dimensions of each layer of the transformer module are kept consistent, facilitating the subsequent transfer and fusion of information between layers. The multi-layer transformer modules sequentially capture the local and global context information of the image, gradually refining higher-level and more semantic visual features, and finally obtaining a multi-layer feature sequence as a multi-scale and multi-level feature representation of the visible light image.

[0026] Furthermore, the steps of adopting the hierarchical feature fusion method with adaptive weighting coefficients in S5 specifically include: S5.1-1. Both the visible light image encoder (adopting the contrastive language-image pre-training model architecture) and the fully polarimetric synthetic aperture radar image encoder (adopting the vision transformer architecture) contain L layers of transformer modules. The output features of each layer of the module carry different-scale and different-semantic-level expressions of the input image at that layer, forming a serialized feature representation: ; where N and M are the lengths of the visible light and synthetic aperture radar feature sequences respectively, and d is the feature dimension; S5.1-2. To overcome the drawback that the feature expressions of each layer contribute unevenly to the final task, assign a dynamic weight to each layer of features , and the weight is obtained through training and learning. Specifically, calculate the weight score by designing a multi-layer perceptron (MLP) network for each layer of features: ; where and respectively represent the weight scores of the k-th layer in the visible light and synthetic aperture radar modalities. To ensure the rationality and comparability of the weights, use the softmax function to normalize the weight scores of all layers: ; The sum of the weights of all layers is 1 to ensure stability during fusion; S5.1-3. Use the normalized weights to perform weighted summation on each layer of features to obtain the single-modal fusion feature: ; This fusion method comprehensively considers the contributions of different layers, takes into account the detailed information of the shallow layer and the semantic information of the deep layer, and improves the richness and discriminability of the fusion feature; S5.1-4. To avoid inconsistent feature scales after weighted fusion, further perform layer normalization on the fusion result: ; Normalization enhances the numerical stability of the features and facilitates subsequent cross-modal fusion calculations.

[0027] Furthermore, the steps of using the cross-modal self-attention mechanism to achieve deep fusion of the features of visible light images and fully polarized synthetic aperture radar images in S6 are specifically as follows: S6.1-1. Assume that after hierarchical feature fusion with adaptive weighting coefficients, the visible light modality fusion feature and the synthetic aperture radar modality fusion feature are respectively obtained, where N and M are the lengths of the visible light and synthetic aperture radar feature sequences respectively, and d is the feature dimension. The fusion features of the two modalities contain the high-level semantic information extracted by the multi-layer transformer modules of their respective encoders and are the basis for subsequent cross-modal fusion; S6.1-2. To achieve in-depth semantic interaction between the visible light modality and the synthetic aperture radar modality, a cross-modal self-attention mechanism is designed. Specifically, the fused feature of the visible light modality is mapped to the query Q matrix, and the fused feature of the synthetic aperture radar modality is mapped to the key K and value V matrices. The calculation formulas are as follows: ; Among them, is the linear projection matrix learned by the model, which projects the input features into the attention space; S6.1-3. Based on the above , calculate the cross-modal self-attention weights, and the query of the visible light modality weights the key of the synthetic aperture radar modality to generate an attention response: ; S6.1-4. Perform a residual connection between this attention response and the fused feature of the visible light modality to avoid information loss and promote gradient propagation in model training, forming an improved fused feature: ; S6.1-5. Use the final fused feature as the input and pass it to the subsequent text encoder to generate a natural language description that combines multi-modal information, improving the accuracy and richness of the description.

[0028] Furthermore, the step of inputting the fused multi-modal features into the text encoder based on the text-to-text transfer Transformer model architecture in S7 specifically includes: S7.1-1. Assume that the final multi-modal fused feature obtained by in-depth fusion through the cross-modal self-attention mechanism is represented as: ; Among them, L represents the length of the feature sequence, d is the feature dimension, and this fused feature comprehensively reflects the multi-layer semantic information extracted by the visible light and full-polarization synthetic aperture radar image encoders, containing rich multi-modal context semantics, which is the basic input for the subsequent text encoder to generate descriptive text. S7.1-2. Input the fused feature into the text encoder based on the text-to-text transfer Transformer model architecture. This text encoder consists of multiple Transformer encoding layers, and each layer deeply processes and semantically abstracts the input features through the self-attention mechanism and the feed-forward network. The output of the i-th layer is denoted as: ; Among them, is the total number of encoder layers, and the initial input is ; S7.1-3. The output of the last layer of the encoder It is passed to the text decoder as a context vector, and the text decoder generates a natural language description based on the autoregressive mechanism, predicting each word in turn probability: ; wherein represents the sequence of previously generated words.

[0029] Furthermore, the steps of optimizing text generation by combining the temperature adjustment mechanism and the reinforcement learning strategy based on the reinforcement learning policy gradient algorithm in S8 are specifically as follows: S8.1-1. During the process of the text decoder generating text based on the text-to-text transfer transformer model architecture, in order to improve the diversity of the generated text and control the randomness in the sampling process, a temperature adjustment mechanism is introduced. Specifically, for the unnormalized prediction value of the i-th word in the vocabulary generated by the decoder at time step t, its probability distribution is adjusted by the temperature parameter T>0, and the calculation formula is: ; where V represents the size of the vocabulary, and the temperature parameter T adjusts the smoothness of the probability distribution. When T>1, the probability distribution becomes flatter and the sampling diversity increases; when T<1, the probability distribution becomes sharper and is more inclined to select the word with the highest probability, and the generated text is more certain; S8.1-2. In order to directly optimize the quality and semantic relevance of the generated text, a reinforcement learning strategy based on the reinforcement learning policy gradient algorithm is combined, and a reward function is defined to evaluate the quality of the generated sequence . The reward can be based on bilingual evaluation alternative metrics, recall-oriented summary evaluation metrics, or task-related metrics. The optimization goal of the model is to maximize the expected reward of the generated text: ; wherein represents the model parameters, is the generation probability distribution based on the current parameters; S8.1-3. According to the reinforcement learning policy gradient algorithm, calculate the gradient of the model parameters: ; Estimate the gradient using the sampled text sequence and its corresponding reward, and optimize the model parameters through gradient ascent; S8.1-4. In order to reduce the variance of gradient estimation, a baseline function b is introduced to improve the gradient estimation to: ; wherein is the number of sampling times, is the m-th sampled text sequence, baseline Generally, the mean value of the rewards is selected, which helps to stabilize the training process. By combining the temperature adjustment mechanism with the reinforcement learning policy gradient algorithm, the text generation process is made to be both rich and diverse, and optimized in terms of quality and semantics, significantly improving the intelligent text description effect of multi-modal remote sensing images.

Claims

1. A method for text description of fully polarized synthetic aperture radar and visible light images, characterized in that, Including the following steps: S1. Collect visible light images and fully polarized synthetic aperture radar images in the same scene, where the fully polarized synthetic aperture radar images contain multiple polarization channels; S2. Perform Lee filtering denoising processing on each polarization channel of the fully polarized synthetic aperture radar images respectively to remove noise and speckles; S3. Fuse the denoised features of multiple polarization channels to form multi-channel fusion features, and input them into a vision transformer encoder for multi-layer transformer feature extraction; S4. Use a contrastive language-image pre-training model encoder to perform multi-layer transformer feature extraction on the visible light images; S5. Through a hierarchical feature fusion method with adaptive weighting coefficients, perform weighted fusion on the multi-layer transformer features of the visible light images and synthetic aperture radar images respectively, and dynamically adjust the contribution weights of each layer of features; S6. Design a cross-modal self-attention mechanism to achieve deep bidirectional information interaction between the fusion features of the visible light images and the fusion features of the synthetic aperture radar images; S7. Input the fused multi-modal features into a text encoder based on the text-to-text transfer transformer model architecture to generate corresponding natural language text descriptions; S8. Combine the temperature adjustment mechanism and the reinforcement learning strategy based on the reinforcement learning policy gradient algorithm to optimize the diversity and accuracy of the generated text and improve the final description quality.

2. A method for text description of a full-polarization synthetic aperture radar and visible light image according to claim 1, characterized in that, The steps of performing Lee filtering denoising processing on each polarization channel of the fully polarized synthetic aperture radar images in S2 are as follows: S2.1-1. Collect the full-polarization synthetic aperture radar image data under the same ground object scene, which includes four polarization channels, namely horizontal transmit-horizontal receive (HH), horizontal transmit-vertical receive (HV), vertical transmit-horizontal receive (VH), and vertical transmit-vertical receive (VV). The original images of these four polarization channels are respectively represented as: , , , ; S2.1-2. For each polarization channel image, use the Lee filtering algorithm for noise suppression and speckle removal. The Lee filtering is based on local statistical characteristics. By calculating the local mean and variance, the filtering weights are dynamically adjusted to distinguish the signal and noise parts of the image. The specific filtering calculation formula is as follows: ; Among them, represents the original polarized channel pixel value to be processed; represents the pixel mean within a small window centered on pixel and reflects the average intensity of the local image; represents the variance of the pixels within this small window and reflects the degree of variation of the local intensity; K is the estimated value of the noise variance, usually obtained through statistical methods or prior knowledge; S2.1-3. Perform Lee filtering processing on the four polarization channels respectively to obtain the denoised polarization channel images: 。 3. A method for text description of a full-polarization synthetic aperture radar and visible light image according to claim 1, characterized in that, The steps of fusing the denoised features of multiple polarization channels in S3 specifically include: S3.1-1. By introducing learnable weight parameters , different importance weights are assigned to the four polarization channels. The weight parameters are automatically optimized through model training, and the weights satisfy the following normalization constraints: ; S3.1-2. Concatenate the weighted features of the four polarization channels in the channel dimension to form a complete fusion feature representation: ; S3.1-3. The vision transformer encoder is mainly composed of multiple transformer modules. Each layer of the transformer abstracts and enhances the input features through the multi-head self-attention mechanism and the feed-forward network. Assume that the vision transformer encoder contains L layers of transformer modules, and the output feature representation of each layer is: ; Among them, represents the polarization feature after denoising and fusion of synthetic aperture radar data.

4. A method for text description of a fully polarimetric synthetic aperture radar and visible light image according to claim 1, characterized in that, The steps of performing multi-layer transformer feature extraction on the visible light images in S4 specifically include: S4.1-1. Collect the visible light image sequence in the same scene. Assume that its original input data is represented as , and before feature extraction, perform preprocessing operations on the input image, including normalization, and the formula is: ; wherein, and are the mean and standard deviation of the visible light image pixels respectively, and are used to unify the numerical scales between different images; S4.1-2. Input the normalized visible light images into a pre-trained contrastive language-image pre-training model encoder for feature extraction. The contrastive language-image pre-training model encoder is mainly composed of multiple transformer modules. Each layer of the transformer abstracts and enhances the input features through the multi-head self-attention mechanism and the feed-forward network. Assume that the contrastive language-image pre-training model encoder contains L layers of transformer modules, and the output feature representation of each layer is: ; Among them, represents the input image features after preliminary linear transformation and positional encoding; S4.1-3. The feature dimensions of each layer of the transformer module are kept consistent. Multiple layers of transformer modules sequentially capture the local and global context information of the image, gradually refining more advanced and more semantic visual features, and finally obtaining a multi-layer feature sequence as the multi-scale and multi-level feature representation of the visible light image.

5. A method for text description of a full-polarization synthetic aperture radar and visible light image according to claim 1, characterized in that The steps of using the hierarchical feature fusion method with adaptive weighting coefficients in S5 specifically include: S5.1-1. Both the visible light image encoder and the full-polarization synthetic aperture radar image encoder contain L-layer transformer modules. The output features of each layer carry the expressions of different scales and semantic levels of the input image at that layer, forming a serialized feature representation: ; where N and M are the lengths of the visible light and synthetic aperture radar feature sequences respectively, and d is the feature dimension; S5.1-2. Assign a dynamic weight to each layer of features , and the weights are obtained through training. Specifically, a multi-layer perceptron (MLP) network for each layer of features is designed to calculate the weight scores: ; Among them, and respectively represent the weight scores of the k-th layer in the visible light and synthetic aperture radar modalities, and the softmax function is used to normalize the weight scores of all layers: ; The sum of the weights of all layers is 1 to ensure stability during fusion; S5.1-3. Use the normalized weights to perform weighted summation on the features of each layer to obtain the single-modal fusion feature: ; S5.1-4. To avoid inconsistent feature scales after weighted fusion, further perform layer normalization on the fusion result: 。 6. A method for text description of a full-polarization synthetic aperture radar and visible light image according to claim 1, characterized in that, The steps of using the cross-modal self-attention mechanism to achieve deep fusion of the features of the visible light image and the full-polarization synthetic aperture radar image in S6 are specifically as follows: S6.1-1. Assume that after the hierarchical feature fusion with adaptive weighting coefficients, the fused features of the visible light modality and the synthetic aperture radar modality are obtained respectively. and the fused features of the synthetic aperture radar modality , where N and M are the lengths of the visible light and synthetic aperture radar feature sequences respectively, and d is the feature dimension. S6.1-2. To achieve the deep semantic interaction of the visible light modality with respect to the synthetic aperture radar modality, design a cross-modal self-attention mechanism. Map the fused feature of the visible light modality to the query Q matrix, and map the fused feature of the synthetic aperture radar modality to the key K and value V matrices. The calculation formulas are as follows: ; Among them, is a linear projection matrix learned by the model, which projects the input features into the attention space; S6.1-3. Based on the above , calculate the cross-modal self-attention weights, where the queries in the visible light modality weight the keys in the synthetic aperture radar modality to generate an attention response: ; S6.1-4. Perform a residual connection on the attention response and the fused feature of the visible light modality to form an improved fused feature: ; S6.1-5. Pass the final fused feature as input to the subsequent text encoder to generate a natural language description that combines multi-modal information, improving the accuracy and richness of the description.

7. A method for text description of a full-polarization synthetic aperture radar and visible light image according to claim 1, characterized in that The steps of inputting the fused multi-modal features into the text encoder based on the text-to-text transfer transformer model architecture in S7 are specifically as follows: S7.1-1. Assume that the final multi-modal fusion feature representation obtained through deep fusion by the cross-modal self-attention mechanism is: ; where L represents the length of the feature sequence, and d is the feature dimension; S7.1-2. Input the fusion feature into a text encoder based on the text-to-text transfer Transformer model architecture, which consists of multiple layers of Transformer encoding layers. Each layer deeply processes and semantically abstracts the input features through a self-attention mechanism and a feed-forward network. The output of the i-th layer is denoted as: ; Among them, is the total number of encoder layers, and the initial input is ; S7.1-3. Output of the last layer of the encoder It is passed to the text decoder as a context vector, and the text decoder generates a natural language description based on the autoregressive mechanism, predicting the probability of each word in sequence as follows: ; Among them, represents the previously generated word sequence.

8. A method for text description of a full-polarization synthetic aperture radar and visible light image according to claim 1, characterized in that, The steps of optimizing text generation by combining the temperature adjustment mechanism and the reinforcement learning strategy based on the reinforcement learning policy gradient algorithm in S8 are specifically as follows: S8.1-1. During the process of generating text by the text decoder based on the text-to-text transfer transformer model architecture, a temperature adjustment mechanism is introduced. For the unnormalized prediction value of the i-th word in the vocabulary generated by the decoder at time step t , its probability distribution is adjusted by the temperature parameter T > 0, and the calculation formula is: ; where V represents the size of the vocabulary, and the temperature parameter T adjusts the smoothness of the probability distribution; S8.1-2. Define the reward function in combination with the reinforcement learning policy based on the reinforcement learning policy gradient algorithm for evaluating the quality of the generated sequence The reward can be based on bilingual evaluation alternative metrics, recall-oriented summary evaluation metrics or task-related metrics. The optimization objective of the model is to maximize the expected reward of the generated text: ; Among them, represents model parameters, is the generation probability distribution based on the current parameters; S8.1-3. According to the reinforcement learning policy gradient algorithm, calculate the gradient of the model parameters: ; Use the sampled text sequence and its corresponding reward for gradient estimation, and optimize the model parameters through gradient ascent; S8.1-4. Introduce the baseline function b to improve the gradient estimation to: ; Among them, is the number of sampling times, is the m-th sampled text sequence, and the baseline usually selects the mean value of the rewards.

Citation Information

Patent Citations

  • Multi-band image description generation method based on self-supervised learning and feature decoupling

    CN118736577A

  • Address information feature extraction method based on deep neural network model

    US20210012199A1

Cited By

  • Small target detection method, system and equipment based on visible light and SAR fusion

    CN121236370A