Image translation method, system and electronic device
By combining the global-local cascade extractor and the cross-scale collaborative compensation module with perception-related learning, the problems of insufficient pixel-level registration and feature mining in SAR image translation are solved, high-quality SAR image to optical image translation is achieved, and the visual consistency and semantic expression of the image are improved.
Patent Information
- Application Number
- CN202511020711.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-24
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2045-07-24
AI Technical Summary
Existing SAR image translation methods rely on a large number of paired image training, making it difficult to achieve high-quality pixel-level alignment. They also fail to fully exploit the unique signal characteristics of SAR images, resulting in distortion in the translated images in terms of structural restoration and regional identification, and failing to improve readability and semantic consistency.
A global-local cascade extractor and a cross-scale collaborative compensation module are adopted, combined with perception-related learning. Multi-scale contextual information is captured through a global feature extractor, a local feature extractor, and a self-gated feedforward network. Features are reconstructed through a cross-scale conversion collaborative module and a self-selection compensation mechanism. Combined with a composite loss function optimization model, the translation of SAR images to optical images is achieved.
It improves the visual consistency and semantic expression ability of SAR images, enhances the model's ability to model complex terrain structures, improves the structure restoration and detail retention capabilities of translated images, overcomes the problems of information dilution and feature conflict in traditional methods, and generates more realistic optical images.
Smart Images

Figure CN120526233B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image interpretation technology, and in particular to an image translation method, system and electronic equipment. Background Art
[0002] Optical imagery, relying on its rich spectral and color information, has a wide range of applications in surveillance, transportation, agriculture, disaster assessment, and environmental monitoring. However, optical imagery is highly dependent on lighting and climatic conditions, and its usability decreases significantly in low-visibility scenarios such as cloudy skies or at night. In contrast, Synthetic Aperture Radar (SAR) offers the advantage of all-weather, all-day imaging, capable of stably acquiring surface information in complex environments. However, SAR images are typically grayscale, affected by electromagnetic wave backscattering, often accompanied by speckle noise, and lack intuitive semantic expression, limiting interpretation by non-professionals.
[0003] To improve the readability and application value of SAR imagery, SAR-to-optical image translation has become a promising research direction in remote sensing. However, most current methods employ supervised learning frameworks and rely on a large number of paired images for training. However, due to differences in imaging principles, high-quality pixel-level registration of SAR and optical images is difficult to achieve, limiting the widespread application of this technology. Unsupervised image translation methods alleviate the paired data dependency issue, but most are designed based on natural images and fail to fully consider the specificity and complex scene characteristics of SAR images. Furthermore, existing methods have significant limitations in feature extraction and modeling, resulting in different landform types being incorrectly mapped to the same type, causing distortion in the translated images in terms of structural restoration and regional identification. Furthermore, most existing methods lack the integration of prior knowledge in SAR images during the pre-training phase, failing to fully exploit their unique signal characteristics, impacting overall visual quality and semantic consistency. Summary of the Invention
[0004] The present invention solves the above-mentioned technical problem by providing an image translation method, comprising the following steps:
[0005] S1. Data preprocessing: preprocessing and data enhancement operations on the original data;
[0006] S2. Feature information encoding: A global-local cascade extractor is used, including a global feature extractor, a local feature extractor, and a self-gated feedforward network, to model global channel dependencies and local fine-grained textures, respectively, to capture multi-scale contextual information;
[0007] S3. Feature Information Decoding: Reconstructing multi-scale features from the frequency characteristics and spatial perception dimensions through a cross-scale collaborative compensation module, including extracting frequency components through a cross-scale conversion collaborative module and strengthening key areas through a self-selected compensation mechanism;
[0008] S4, Feature Information Reconstruction: Adopting the adaptive information fusion mechanism, the multi-stage output of the decoding part is adaptively fused and spliced with the final output of the decoding part to integrate the feature representations of different levels;
[0009] S5. Loss function correction: A composite loss function is used, including the generator's adversarial loss, perception-related learning loss, and perceptual loss, as well as the discriminator's adversarial loss, to optimize model convergence.
[0010] S6. Evaluate network performance: Evaluate translation quality through peak signal-to-noise ratio, structural similarity, and perceptual image similarity metrics, and iteratively optimize the model.
[0011] Furthermore, in step S1, the preprocessing operation includes data cleaning, normalization, and noise reduction; the data enhancement operation includes geometric transformation and illumination change enhancement.
[0012] Furthermore, in step S2, the global feature extractor implicitly captures global context information in the image by calculating the cross-covariance relationship between feature channels, modeling the dependencies and mutual influences between channels;
[0013] The local feature extractor constructs feature association relationships at different scales by capturing fine-grained texture and edge information in a local area;
[0014] The self-gated feedforward network guides and optimizes the information flow of features between layers, enhances the interaction between spatial neighborhood pixels, and enables the model to focus more on detail restoration at different stages.
[0015] Furthermore, in step S3, the cross-scale collaborative compensation module includes a cross-scale conversion collaboration module and a self-selection compensation mechanism;
[0016] The cross-scale conversion collaborative module captures the multi-level structural contours and local details of the image by extracting the characteristic components of each frequency.
[0017] The self-selection compensation mechanism adaptively strengthens key areas to ensure that the characteristics of each area are globally contrasting.
[0018] Furthermore, in step S4, the information adaptive fusion mechanism fully integrates the feature information extracted at different scales and different receptive fields to compensate for the local loss or weakening of the features in a single stage.
[0019] Furthermore, in step S5, the loss functions adopted by the generator include adversarial loss, perceptual-related learning loss, and perceptual loss; the discriminator adopts adversarial loss;
[0020] The perception-related learning loss uses a pre-trained perception network to extract deep semantic features and models the perceptual similarity between the input image and the translated image at this feature level;
[0021] The query sample, positive sample and Negative samples are mapped to dimensional vectors, denoted as 、 and ; established a For the classification problem of the class, the probability of selecting positive samples by excluding negative samples using cross entropy is as follows:
[0022] ;
[0023] in, To query the sample, is a positive sample, is a negative sample, is the temperature coefficient.
[0024] The Kolmogorov-Arnold Network is designed to replace the multi-layer perceptron. The input is mapped into a combination of basis functions and spline functions, and nonlinear transformation is applied to enhance the feature expression capability.
[0025] Perception Network Output passes through the comparison network Generate a more discriminative feature stack from the perception network Select The output of the layer is used as input and fed into the contrast network To get the feature set ,in, Representation-aware network No. The output of the layer; each spatial position in each layer is recorded as ,in, is the number of spatial locations per layer; the positive sample of each query sample is , the negative sample is ,in, is the number of channels; the translated image is represented as , the formula for perceptual related learning loss is as follows:
[0026] ;
[0027] in, To query sample features, is the positive sample feature, is the negative sample feature, is the number of feature layers, is the number of spatial positions.
[0028] The perceptual loss optimizes the visual consistency between the translated image and the ground truth through high-level feature constraints; the formula of the perceptual loss is as follows:
[0029]
[0030] in, The first The features extracted by the layer, and are the height and width of the feature map, The generated translation image.
[0031] The adversarial loss formula is as follows:
[0032] .
[0033] in, is the input SAR image, The generated translation image.
[0034] To solve the above technical problems, the present invention further proposes an image translation system for executing the above image translation method, comprising:
[0035] Data acquisition and preprocessing module, used to collect SAR and optical images and perform data cleaning and enhancement operations;
[0036] Feature extraction and data construction module, extracting spatial texture features of SAR images and constructing cross-modal mapping relationships;
[0037] A deep learning model building module that integrates a global-local cascade extractor and a cross-scale collaborative compensation module to achieve SAR to optical image translation;
[0038] System evaluation and feedback module, which evaluates translation results through subjective and objective indicators and provides optimization suggestions;
[0039] Model iterative optimization module, adjusting the network structure and training strategy based on feedback;
[0040] The result visualization module displays the translation results through a graphical interface and verifies the applicability of the task.
[0041] To solve the above technical problems, the present invention also proposes an image translation electronic device, which mainly includes a memory, a processor, a communication interface and a bus; wherein the memory, the processor and the communication interface are connected to each other through the bus; and the processor is used to execute the image translation method as described above.
[0042] Compared with the existing technology, the present invention provides an image translation method and system based on perception-related learning and global-local feature collaboration, which has the following beneficial effects:
[0043] 1. The perception-related learning proposed in this paper uses a pre-trained perception network to extract deep semantic features of the input and translated images and model their perceptual similarities in the feature space, optimizing the model's discriminative ability and the visual consistency of the translated images, and improving the model's nonlinear expression and computational efficiency.
[0044] 2. The global-local cascade extractor proposed in this invention realizes the collaborative learning of the global structure and detailed texture of the image, enhances the modeling ability of complex ground structures, and improves the model's structure restoration and detail retention capabilities in high-density scenes.
[0045] 3. The cross-scale collaborative compensation module proposed in this invention realizes deep fusion and semantic compensation of multi-scale features under the dual modeling of frequency features and spatial attention, overcomes the problems of information dilution and feature conflict in traditional feature stacking, and enhances the model's perception of key areas of the image.
[0046] 4. The information adaptive fusion mechanism proposed in this invention fully integrates the feature information extracted at different scales and different receptive fields, compensates for the local loss or weakening of features in a single stage, and improves the integrity and discriminability of the overall feature representation. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the structures shown in these drawings without paying any creative work.
[0048] Figure 1 is a flowchart of the steps of the image translation method of the present invention;
[0049] Figure 2 A network structure diagram of the image translation method of the present invention;
[0050] Figure 3 Schematic diagram of the structure of the global-local cascade extractor of the present invention;
[0051] Figure 4 Schematic diagram of the structure of the global feature extractor of the present invention;
[0052] Figure 5 Schematic diagram of the structure of the local feature extractor of the present invention;
[0053] Figure 6 Schematic diagram of the structure of the self-gated feedforward network of the present invention;
[0054] Figure 7 Schematic diagram of the structure of the cross-scale collaborative compensation module of the present invention;
[0055] Figure 8 This is a diagram showing the working principle of perception-related learning in the present invention;
[0056] Figure 9 Schematic diagram for comparing relevant indicators of the image translation method of the present invention;
[0057] Figure 10 Schematic diagram of the image translation system of the present invention;
[0058] Figure 11 Schematic diagram of the internal structure of the image translation electronic device of the present invention. DETAILED DESCRIPTION
[0059] The present invention proposes an image translation method, system and electronic device, aiming to design an image translation method based on perception-related learning and global-local feature collaboration.
[0060] The image translation method proposed by the present invention will be described in detail below in a specific embodiment:
[0061] Example 1:
[0062] An image translation method, such as Figure 1 As shown, the following steps are included:
[0063] S1. Data preprocessing: preprocessing and data enhancement operations on the original data;
[0064] S2. Feature information encoding: A global-local cascade extractor is used, including a global feature extractor, a local feature extractor, and a self-gated feedforward network, to model global channel dependencies and local fine-grained textures, respectively, to capture multi-scale contextual information;
[0065] S3. Feature Information Decoding: Reconstructing multi-scale features from the frequency characteristics and spatial perception dimensions through a cross-scale collaborative compensation module, including extracting frequency components through a cross-scale conversion collaborative module and strengthening key areas through a self-selected compensation mechanism;
[0066] S4, Feature Information Reconstruction: Adopting the adaptive information fusion mechanism, the multi-stage output of the decoding part is adaptively fused and spliced with the final output of the decoding part to integrate the feature representations of different levels;
[0067] S5. Loss function correction: A composite loss function is used, including the generator's adversarial loss, perception-related learning loss, and perceptual loss, as well as the discriminator's adversarial loss, to optimize model convergence.
[0068] S6. Evaluate network performance: Evaluate translation quality through peak signal-to-noise ratio, structural similarity, and perceptual image similarity metrics, and iteratively optimize the model.
[0069] Furthermore, in step S1, the preprocessing operation includes data cleaning, normalization, and noise reduction; the data enhancement operation includes geometric transformation and illumination change enhancement.
[0070] Furthermore, in step S2, the global feature extractor implicitly captures global context information in the image by calculating the cross-covariance relationship between feature channels, modeling the dependencies and mutual influences between channels;
[0071] The local feature extractor constructs feature association relationships at different scales by capturing fine-grained texture and edge information in a local area;
[0072] The self-gated feedforward network guides and optimizes the information flow of features between layers, enhances the interaction between spatial neighborhood pixels, and enables the model to focus more on detail restoration at different stages.
[0073] Furthermore, in step S3, the cross-scale collaborative compensation module includes a cross-scale conversion collaboration module and a self-selection compensation mechanism;
[0074] The cross-scale conversion collaborative module captures the multi-level structural contours and local details of the image by extracting the characteristic components of each frequency.
[0075] The self-selection compensation mechanism adaptively strengthens key areas to ensure that the characteristics of each area are globally contrasting.
[0076] Furthermore, in step S4, the information adaptive fusion mechanism fully integrates the feature information extracted at different scales and different receptive fields to compensate for the local loss or weakening of the features in a single stage.
[0077] Furthermore, in step S5, the loss functions adopted by the generator include adversarial loss, perceptual-related learning loss, and perceptual loss; the discriminator adopts adversarial loss;
[0078] The perception-related learning loss uses a pre-trained perception network (such as VGG-16) to extract deep semantic features and models the perceptual similarity between the input image and the translated image at this feature level;
[0079] The query sample, positive sample and Negative samples are mapped to dimensional vectors, denoted as 、 and ; established a For the classification problem of the class, the probability of selecting positive samples by excluding negative samples using cross entropy is as follows:
[0080] ;
[0081] in, To query the sample, is a positive sample, is a negative sample, is the temperature coefficient.
[0082] The Kolmogorov-Arnold Network is designed to replace the multi-layer perceptron. The input is mapped into a combination of basis functions and spline functions, and nonlinear transformation is applied to enhance the feature expression capability.
[0083] Perception Network Output passes through the comparison network Generate a more discriminative feature stack from the perception network Select The output of the layer is used as input and fed into the contrast network To get the feature set ,in, Representation-aware network No. The output of the layer; each spatial position in each layer is recorded as ,in, is the number of spatial locations per layer; the positive sample of each query sample is , the negative sample is ,in, is the number of channels; the translated image is represented as , the formula for perceptual related learning loss is as follows:
[0084] ;
[0085] in, To query sample features, is the positive sample feature, is the negative sample feature, is the number of feature layers, is the number of spatial positions.
[0086] The perceptual loss optimizes the visual consistency between the translated image and the ground truth through high-level feature constraints; the formula of the perceptual loss is as follows:
[0087]
[0088] in, The first The features extracted by the layer, and are the height and width of the feature map, The generated translation image.
[0089] The adversarial loss formula is as follows:
[0090] .
[0091] in, is the input SAR image, The generated translation image.
[0092] Example 2:
[0093] An image translation method comprises the following steps:
[0094] Step 1, data preprocessing: By preprocessing the collected raw data, the quality and consistency of the training data are ensured. Various data augmentation techniques are used to enhance the robustness of the model and prevent overfitting.
[0095] Specifically, the preprocessing operations include data cleaning, normalization, standardization, noise reduction and data balancing; the data enhancement operations include geometric transformation, illumination change and hybrid enhancement;
[0096] Step 2: Feature Information Encoding: A global-local cascade extractor is used to model global and local features separately to capture the multi-scale contextual information of the image. This allows for collaborative modeling of local texture and global semantics, improving the model's ability to restore structure and retain detail in complex scenes.
[0097] Specifically, the global-local cascade extractor includes a global feature extractor, a local feature extractor and a self-gated feedforward network;
[0098] The global feature extractor implicitly captures the global context information in the image by calculating the cross-covariance relationship between feature channels, modeling the dependencies and mutual influences between channels;
[0099] The local feature extractor captures fine-grained texture and edge information in a local area, constructs feature associations at different scales, and improves the richness and accuracy of feature expression.
[0100] The self-gated feedforward network guides and optimizes the information flow of features between layers, enhances the interaction between spatial neighborhood pixels, and enables the model to focus more on detail restoration at different stages, thereby improving image reconstruction quality.
[0101] Step 3: Feature Information Decoding: A cross-scale collaborative compensation module is used to receive multi-level features from the encoding part and deeply reconstruct the multi-scale features from two dimensions: frequency characteristic modeling and spatial area perception. This strengthens the transmission of key features at different scales and improves the effectiveness and stability of multi-scale feature fusion.
[0102] Specifically, the cross-scale collaborative compensation module includes a cross-scale conversion collaboration module and a self-selection compensation mechanism;
[0103] The cross-scale conversion collaboration module captures the multi-level structural contours and local details of the image by extracting the characteristic components of each frequency, thereby improving the consistency between cross-scale semantics and the efficiency of information transfer.
[0104] The self-selection compensation mechanism adaptively strengthens key areas, ensuring that the characteristics of each area have global contrast, enhancing the reconstruction capability of local areas and the fine-grainedness of spatial modeling.
[0105] Step 4: Feature Information Reconstruction: Adopting an adaptive information fusion mechanism, the multi-stage outputs of the decoding part are adaptively fused and concatenated with the final output of the decoding part, fully integrating feature representations at different levels to improve the accuracy and visual quality of image translation.
[0106] Specifically, the information adaptive fusion mechanism fully integrates the feature information extracted at different scales and different receptive fields, compensates for the local loss or weakening of the features in a single stage, and improves the integrity and discriminability of the overall feature representation.
[0107] Step 5, loss function correction: By designing a reasonable loss function, minimize the loss function of the network output image and label, and optimize the convergence effect of the model;
[0108] The loss function is a composite loss function. The loss functions used by the generator include adversarial loss, perceptual-related learning loss, and perceptual loss; the discriminator uses adversarial loss.
[0109] The perception-related learning loss uses a pre-trained perception network (such as VGG-16) to extract deep semantic features and models the perceptual similarity between the input image and the translated image at this feature level. In addition, the information-noise contrast loss is used to optimize the model's discrimination ability in the perceptual feature space. By constructing a self-supervised contrastive learning framework, the model can automatically identify and match semantically similar feature pairs while distinguishing dissimilar features, thereby improving the visual perception consistency between the translated image and the original image. Negative samples are mapped to dimensional vectors, denoted as 、 and At this point, a The probability of selecting positive samples by excluding negative samples is calculated by cross entropy. The formula is as follows:
[0110] ;
[0111] in, To query the sample, is a positive sample, is a negative sample, is the temperature coefficient.
[0112] In addition, the Kolmogorov-Arnold Network is designed to replace the multi-layer perceptron. By mapping the input into a combination of basis functions and spline functions and applying nonlinear transformations, the feature expression capability is enhanced, and it has stronger nonlinear modeling capabilities and higher interpretability. Output passes through the comparison network Generating a more discriminative feature stack helps maintain image structure, achieve style consistency translation, and generate more realistic optical images. Select The output of the layer is used as input and fed into the contrast network To get the feature set ,in, Representation-aware network No. The output of the layer. Each spatial position in each layer is recorded as ,in, is the number of spatial locations in each layer. The positive sample of each query sample is , the negative sample is ,in, is the number of channels. Similarly, the translated image is represented as Therefore, the formula for the perceptually relevant learning loss is as follows:
[0113] .
[0114] in, To query sample features, is the positive sample feature, is the negative sample feature, is the number of feature layers, is the number of spatial positions.
[0115] Step 6: Evaluate network performance: Select appropriate evaluation metrics to measure the quality of the algorithm, image similarity, and image distortion, comprehensively evaluate model performance, and optimize and improve the model's performance in actual scenarios.
[0116] Example 3:
[0117] An image translation method, the method specifically comprising the following steps:
[0118] Step 1: Data preprocessing: By preprocessing the collected raw data, the quality and consistency of the training data are ensured. A variety of data enhancement techniques are used to enhance the robustness of the model and prevent overfitting.
[0119] Preprocessing operations include data cleaning, normalization, standardization, noise reduction, and data balancing. Data cleaning requires checking and screening the original collected SAR images to remove samples with ambiguity, interference, or invalid information to ensure the validity and representativeness of the input data. Normalization and standardization map pixel values to a fixed range (such as [0, 1] or [-1, 1]) to improve the model's sensitivity and stability to the data. Noise reduction removes random noise from the image through methods such as Gaussian filtering, mean filtering, or median filtering to ensure the integrity of image detail information. Data balancing improves the balance of training data by adjusting the data distribution.
[0120] Data augmentation techniques include geometric transformation, illumination change, and hybrid enhancement. Geometric transformation can simulate different perspectives and layout changes of objects in real scenes through random rotation and random region cropping, thereby improving the adaptability of the model. Illumination change simulates the response under different ambient lighting conditions by adjusting brightness and contrast, thereby enhancing the robustness of the model under complex lighting conditions. Hybrid enhancement generates new samples by mixing two images at the pixel or region level, which can improve the generalization ability of the model.
[0121] Step 2: Feature Information Encoding: A global-local cascade extractor is used to model global and local features separately to capture the multi-scale contextual information of the image. This allows for collaborative modeling of local texture and global semantics, improving the model's ability to restore structure and retain detail in complex scenes.
[0122] like Figure 2 、 Figure 8 As shown in the figure, the network structure of the image translation method based on perception-related learning and global-local feature collaboration, where convolutional layers 1 and 2 increase the model's expressive power by increasing the dimension, enabling the network to better learn the complex features of the input data. The convolutional kernel size is 3×3; convolutional layer 3 reconstructs the details and texture of the image. The convolution kernel size is 3×3; the T-type function prevents image pixel value overflow;
[0123] like Figure 3 As shown, the global-local cascade extractor includes a global feature extractor, a local feature extractor, and a self-gated feedforward network;
[0124] like Figure 4 As shown in the figure, the global feature extractor models the dependency and mutual influence between channels by calculating the cross-covariance relationship between feature channels, implicitly capturing the global context information in the image. It consists of a normalization layer, convolution layer 1, convolution layer 2, convolution layer 3, convolution layer 4, convolution layer 5, convolution layer 6, convolution layer 7, pixel-level addition operation, matrix multiplication operation and reshaping operation, and the convolution kernel size is 3×3 and 1×1;
[0125] like Figure 5 As shown in the figure, the local feature extractor captures fine-grained texture and edge information in the local area, constructs feature correlation relationships at different scales, and improves the richness and accuracy of feature expression. It consists of a layer normalization layer, convolution layer 1, convolution layer 2, convolution layer 3, convolution layer 4, convolution layer 5, convolution layer 6, convolution layer 7, pixel-level addition operation, matrix multiplication operation and reshaping operation, and the convolution kernel size is 3×3 and 1×1;
[0126] like Figure 6 As shown in the figure, the self-gated feedforward network enhances the interaction between spatial neighborhood pixels by guiding and optimizing the information flow of features between each layer, so that the model can focus more on detail restoration at different stages and improve the image reconstruction quality. It consists of a layer normalization layer, convolution layer 1, convolution layer 2, convolution layer 3, convolution layer 4, convolution layer 5, pixel-level addition operation, pixel-level multiplication operation, G-type function and expansion factor. The convolution kernel size is 3×3 and 1×1; the expansion factor can control the information of the jump connection;
[0127] Step 3: Feature Information Decoding: A cross-scale collaborative compensation module is used to receive multi-level features from the encoding part and deeply reconstruct the multi-scale features from two dimensions: frequency characteristic modeling and spatial area perception. This strengthens the transmission of key features at different scales and improves the effectiveness and stability of multi-scale feature fusion.
[0128] like Figure 7 As shown in the figure, the cross-scale collaborative compensation module includes a cross-scale conversion collaboration module and a self-selection compensation mechanism. The cross-scale conversion collaboration module captures the multi-level structural contours and local details of the image by extracting the characteristic components of each frequency, thereby improving the consistency between cross-scale semantics and the efficiency of information transmission. It consists of discrete cosine transform (DCT), convolution layer 1, global average pooling operation, global maximum pooling operation and channel splicing operation, and the convolution kernel size is 3×3. The self-selection compensation mechanism adaptively strengthens the key areas to ensure that the features of each area have global contrast, enhances the reconstruction ability of the local area and the fine-grainedness of spatial modeling, and consists of convolution layer 1, convolution layer 2, convolution layer 3, S-type function 1, matrix multiplication operation and pixel-level addition operation, and the convolution kernel size is 3×3.
[0129] Step 4: Feature Information Reconstruction: Adopting an adaptive information fusion mechanism, the multi-stage outputs of the decoding part are adaptively fused and concatenated with the final output of the decoding part, fully integrating feature representations at different levels to improve the accuracy and visual quality of image translation.
[0130] like Figure 2 As shown in the figure, the self-information adaptive fusion mechanism fully integrates the feature information extracted at different scales and different receptive fields, compensates for the local loss or weakening of the features in a single stage, and improves the integrity and discriminability of the overall feature representation. It consists of upsampling operations and channel splicing operations.
[0131] In order to ensure the robustness of the network, retain more structural information, and fully extract image features, the present invention uses three activation functions, namely G-type function, S-type function and T-type function. The definitions of G-type function, S-type function and T-type function are as follows:
[0132]
[0133] .
[0134] Step 5, loss function correction: By designing a reasonable loss function, minimize the loss function of the network output image and label, and optimize the convergence effect of the model; the loss function is a composite loss function. The loss function used by the generator includes adversarial loss, perceptual-related learning loss, and perceptual loss; the discriminator uses adversarial loss;
[0135] Perceptual-related learning loss uses a pre-trained perceptual network (such as VGG-16) to extract deep semantic features and models the perceptual similarity between the input image and the translated image at this feature level. In addition, the information-noise contrast loss is used to optimize the model's discrimination ability in the perceptual feature space. By building a self-supervised contrastive learning framework, the model can automatically identify and match semantically similar feature pairs while distinguishing dissimilar features, thereby improving the visual perception consistency between the translated image and the original image. Negative samples are mapped to dimensional vectors, denoted as 、 and At this point, a The probability of selecting positive samples by excluding negative samples is calculated by cross entropy. The formula is as follows:
[0136] ;
[0137] in, To query the sample, is a positive sample, is a negative sample, is the temperature coefficient.
[0138] In addition, the Kolmogorov-Arnold Network is designed to replace the multi-layer perceptron. By mapping the input into a combination of basis functions and spline functions and applying nonlinear transformations, the feature expression capability is enhanced, and it has stronger nonlinear modeling capabilities and higher interpretability. Output passes through the comparison network Generating a more discriminative feature stack helps maintain image structure, achieve style consistency translation, and generate more realistic optical images. Select The output of the layer is used as input and fed into the contrast network To get the feature set ,in, Representation-aware network No. The output of the layer. Each spatial position in each layer is recorded as ,in, is the number of spatial locations in each layer. The positive sample of each query sample is , the negative sample is ,in, is the number of channels. Similarly, the translated image is represented as Therefore, the formula for the perceptually relevant learning loss is as follows:
[0139] ;
[0140] in, To query sample features, is the positive sample feature, is the negative sample feature, is the number of feature layers, is the number of spatial positions.
[0141] To improve the similarity between the generated translation image and the ground truth, this paper employs an adversarial loss to make the generated image more realistic and detailed, enhancing its detail. Furthermore, to reduce differences in subjective perception, a perceptual loss optimizes the visual consistency between the translated image and the ground truth through high-level feature constraints. The combination of these three loss functions helps generate optical images with higher visual fidelity and accuracy. The mathematical formulas for the adversarial and perceptual losses are shown below:
[0142]
[0143]
[0144] in, is the input SAR image, The first The features extracted by the layer, are the height and width of the feature map, The generated translation image.
[0145] Step 6: Evaluate network performance: Select appropriate evaluation metrics to measure the quality of the algorithm, image similarity, and image distortion, comprehensively evaluate model performance, and optimize and improve the model's performance in real-world scenarios. During the training of the network model, use appropriate evaluation metrics to evaluate the quality, image similarity, and image distortion of the algorithm-translated images.
[0146] Suitable evaluation indicators include peak signal-to-noise ratio (PSNR), structural similarity, perceptual image similarity (PSI), Fréchet distance, and natural image quality assessment. The PSNR is used to measure the difference between the translation result and the real image. The higher the PSNR, the better the translation quality. Structural similarity quantifies the contour preservation during the translation process by calculating the structural similarity between the translation result and the real image. The higher the structural similarity, the smaller the difference. Perceptual image similarity focuses on the perceptual similarity between the translation result and the real image. The lower the PSI, the more similar the translation result and the real image. Fréchet distance evaluates the perceptual quality of the generated image by comparing the distance between the translation result and the real image distribution in the deep feature embedding. The lower the PSI, the higher the similarity, and the better the translation quality. Natural image quality assessment expresses the natural image quality assessment indicator of a given test image as the distance between the multivariate Gaussian model of the natural statistical features extracted from the test image and the multivariate Gaussian model of the quality perception features extracted from the natural image corpus. The definitions of PSNR, structural similarity, perceptual image similarity, Fréchet distance, and natural image quality assessment are as follows:
[0147]
[0148] in, , Represents images respectively and The mean and variance of and Represents images respectively and The standard deviation of Representing an image and The covariance of and is a constant; for and The distance between is a trainable weight parameter; and represent the generated image and the real image respectively, and represents the mean of each eigenvector, and represents the covariance matrix of the respective eigenvectors, represents the trace of the matrix; and Represent the mean vector and covariance matrix of the natural multivariate Gaussian model and the distorted image multivariate Gaussian model respectively.
[0149] All experiments were conducted on an Ubuntu 22.04.6 LTS system, and the algorithm was accelerated using an NVIDIA GeForce RTX 4090. The training cycle was set to 200 rounds, and the learning rates of the generator and discriminator were set to 1e−4. The upper limit of the number of images input to the network at each time is mainly determined by the performance of the computer's graphics processor. Generally, the number of images input to the network at each time is within the range of 4-8, which can make network training more stable and achieve better training results, and can ensure rapid network fitting. The Adam optimizer was selected as the network parameter optimizer. Its advantages mainly lie in its simple implementation, efficient computation, low memory requirements, and parameter updates that are not affected by gradient scaling, resulting in relatively stable parameters. When the discriminator's ability to detect fake images is balanced with the generator's ability to generate images that deceive the discriminator, the network is considered to have been basically trained. After network training is completed, all network parameters need to be saved, and then the translated optical image can be obtained by inputting the SAR image to be translated into the network. The network has no requirements for the input image size; any size can be used.
[0150] Among them, the implementation of convolution, activation function, splicing operation and batch normalization are algorithms well known to those skilled in the art. The specific processes and methods can be found in corresponding textbooks or technical literature;
[0151] The present invention constructs an image translation method based on perceptual correlation learning and global-local feature collaboration, which can directly translate SAR images into optical images without going through other intermediate steps, thus avoiding the need for manual design of relevant translation rules. Under the same conditions, the feasibility and superiority of the method are further verified by calculating the relevant indicators of the image obtained by the existing method. The relevant indicators of the existing technology and the method proposed in the present invention are compared. Figure 9 As shown;
[0152] from Figure 9It can be seen that the method proposed in the present invention has a higher peak signal-to-noise ratio, higher structural similarity, lower perceived image similarity, lower Fréchet distance, lower natural image quality assessment, fewer parameters and fewer floating-point operations than existing methods. These indicators further demonstrate that the method proposed in the present invention has better translation quality and lower computational complexity.
[0153] Example 4:
[0154] An image translation system, such as Figure 10 As shown, the method for performing the image translation method described in any one of embodiments 1 to 3 includes:
[0155] Data acquisition and preprocessing module, used to collect SAR and optical images and perform data cleaning and enhancement operations;
[0156] Specifically, multi-scene and multi-type SAR images and optical images are collected through satellites; the acquired raw data are preprocessed to improve data consistency and quality; data diversity is increased through data enhancement to prevent model overfitting; preprocessing mainly includes image cropping, image flipping and image translation, etc., and the ratio of the training set and test set is 5:1.
[0157] Feature extraction and data construction module, extracts the spatial texture features of SAR images and constructs the mapping relationship between SAR images and optical images;
[0158] A deep learning model building module integrates a global-local cascade extractor and a cross-scale collaborative compensation module to achieve SAR to optical image translation. It guides model convergence through a loss function and improves model performance through hyperparameter tuning.
[0159] System evaluation and feedback module, which evaluates translation results through subjective and objective indicators and provides optimization suggestions;
[0160] The model iterative optimization module adjusts the network structure and training strategy based on feedback; each image is resized from any size in the dataset to a fixed size of 256×256; a total of 200 rounds of training are performed with a batch size of 4; the number of filters in the first convolutional layer in the generator and discriminator is set to 32; the Adam optimizer is used; the discriminator and generator are trained alternately until the composite loss function converges.
[0161] The result visualization module displays the translation results through a graphical interface and verifies the applicability of the task.
[0162] Example 5:
[0163] An image translation electronic device, such as Figure 11As shown, it mainly includes a memory, a processor, a communication interface and a bus; wherein the memory, the processor and the communication interface realize communication connection with each other through the bus;
[0164] The memory may be a ROM, a static storage device, a dynamic storage device, or a RAM; the memory may store a program. When the program stored in the memory is executed by the processor, the processor and the communication interface are used to perform the various steps of the training method for the SAR to optical image translation network according to an embodiment of the present invention.
[0165] The processor may be a CPU, a microprocessor, an ASIC, a GPU, or one or more integrated circuits, and is used to execute relevant programs to implement the functions required to be performed by the units in the SAR to optical image translation training system of the present invention, or to perform the SAR to optical image translation training method of the present invention;
[0166] The processor may also be an integrated circuit chip with signal processing capabilities. During implementation, each step of the SAR to optical image translation training method of the present invention may be completed by hardware integrated logic circuits or software instructions in the processor. The aforementioned processor may also be a general-purpose processor, DSP, ASIC, FPGA, or other programmable logic device, discrete gate or transistor logic device, or discrete hardware component. The SAR to optical image translation method, steps, and logic block diagram of the present invention may be implemented or executed. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the SAR to optical image translation method of the present invention may be directly embodied as being executed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software module may be located in a storage medium mature in the art, such as a random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or register. The storage medium is located in a memory, and the processor reads information in the memory and, in combination with its hardware, completes the functions required to be performed by the units included in the SAR to optical image translation training system of the present invention, or executes the SAR to optical image translation training method of the present invention.
[0167] The communication interface uses a transceiver system such as, but not limited to, a transceiver to achieve communication between the system and other devices or communication networks; for example, the image to be processed or the initial feature map of the image to be processed can be obtained through the communication interface;
[0168] A bus may include a pathway that transfers information between various components of a system (e.g., memory, processor, communication interface);
[0169] The present invention also provides a computer-readable storage medium for image translation based on perception-related learning and global-local feature collaboration. The computer-readable storage medium may be the computer-readable storage medium included in the system described in the above embodiment; or it may be a separate computer-readable storage medium that is not assembled into a device. The computer-readable storage medium stores one or more programs, and the programs are used by one or more processors to execute the method described in the present invention.
[0170] It should be noted that although Figure 11 The electronic device shown only shows the memory, processor, and communication interface. However, in the specific implementation process, those skilled in the art should understand that the system also includes other devices necessary for normal operation. At the same time, according to specific needs, those skilled in the art should understand that the system may also include hardware devices that implement other additional functions. In addition, those skilled in the art should understand that the system may also include only the devices necessary to implement the embodiments of the present invention, and does not necessarily include Figure 11 All devices shown in .
[0171] The above description is merely a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present invention should be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.
Claims
1. An image translation method, characterized in that: The following steps are involved: S1. Data preprocessing: preprocessing and data enhancement operations on the original data; S2. Feature information encoding: A global-local cascade extractor is used, including a global feature extractor, a local feature extractor, and a self-gated feedforward network, to model global channel dependencies and local fine-grained textures, respectively, to capture multi-scale contextual information; S3. Feature Information Decoding: Reconstructing multi-scale features from the frequency characteristics and spatial perception dimensions through a cross-scale collaborative compensation module, including extracting frequency components through a cross-scale conversion collaborative module and strengthening key areas through a self-selected compensation mechanism; S4, Feature Information Reconstruction: Adopting the adaptive information fusion mechanism, the multi-stage output of the decoding part is adaptively fused and spliced with the final output of the decoding part to integrate the feature representations of different levels; S5. Loss function correction: A composite loss function is used, including the generator's adversarial loss, perception-related learning loss, and perceptual loss, as well as the discriminator's adversarial loss, to optimize model convergence. S6. Evaluate network performance: Evaluate translation quality through peak signal-to-noise ratio, structural similarity, and perceptual image similarity metrics, and iteratively optimize the model.
2. The image translation method according to claim 1, wherein: In step S1, the preprocessing operations include data cleaning, normalization, and noise reduction; the data enhancement operations include geometric transformation and illumination change enhancement.
3. The image translation method according to claim 1, wherein: In step S2, the global feature extractor implicitly captures the global context information in the image by calculating the cross-covariance relationship between feature channels, modeling the dependencies and mutual influences between channels; The local feature extractor constructs feature association relationships at different scales by capturing fine-grained texture and edge information in a local area; The self-gated feedforward network guides and optimizes the information flow of features between layers, enhances the interaction between spatial neighborhood pixels, and enables the model to focus more on detail restoration at different stages.
4. The image translation method according to claim 1, wherein: In step S3, the cross-scale collaborative compensation module includes a cross-scale conversion collaboration module and a self-selection compensation mechanism; The cross-scale conversion collaboration module captures the multi-level structural contours and local details of the image by extracting the characteristic components of each frequency. The self-selection compensation mechanism adaptively enhances the key areas to ensure that the characteristics of each area have global contrast.
5. The image translation method according to claim 1, wherein: In step S4, the information adaptive fusion mechanism fully integrates the feature information extracted at different scales and different receptive fields to compensate for the local loss or weakening of the features in a single stage.
6. The image translation method according to claim 1, wherein: In step S5, the loss functions used by the generator include adversarial loss, perceptual-related learning loss, and perceptual loss; the discriminator uses adversarial loss; The perception-related learning loss uses a pre-trained perception network to extract deep semantic features and models the perceptual similarity between the input image and the translated image at this feature level; The query sample, positive sample and Negative samples are mapped to dimensional vectors, denoted as 、 and ; established a For the classification problem of the class, the probability of selecting positive samples by excluding negative samples using cross entropy is as follows: ; in, To query the sample, is a positive sample, is a negative sample, is the temperature coefficient; The Kolmogorov-Arnold Network is designed to replace the multi-layer perceptron. The input is mapped into a combination of basis functions and spline functions, and nonlinear transformation is applied to enhance the feature expression capability. Perception Network Output passes through the comparison network Generate a more discriminative feature stack from the perception network Select The output of the layer is used as input and fed into the contrast network To get the feature set ,in, Representation-aware network No. The output of the layer; each spatial position in each layer is recorded as ,in, is the number of spatial locations per layer; the positive sample of each query sample is , the negative sample is ,in, is the number of channels; the translated image is represented as , the formula for perceptual related learning loss is as follows: ; in, To query sample features, is the positive sample feature, is the negative sample feature, is the number of feature layers, is the number of spatial positions. The perceptual loss optimizes the visual consistency between the translated image and the ground truth through high-level feature constraints. The formula of the perceptual loss is as follows: in, The first The features extracted by the layer, and are the height and width of the feature map, For the generated translation image, the adversarial loss formula is as follows: ,in, is the input SAR image, The generated translation image.
7. An image translation system for executing the image translation method according to any one of claims 1 to 6, characterized in that: include: Data acquisition and preprocessing module, used to collect SAR and optical images and perform data cleaning and enhancement operations; Feature extraction and data construction module, extracting spatial texture features of SAR images and constructing cross-modal mapping relationships; A deep learning model building module that integrates a global-local cascade extractor and a cross-scale collaborative compensation module to achieve SAR to optical image translation; System evaluation and feedback module, which evaluates translation results through subjective and objective indicators and provides optimization suggestions; Model iterative optimization module, adjusting the network structure and training strategy based on feedback; The result visualization module displays the translation results through a graphical interface and verifies the applicability of the task.
8. An electronic device for image translation, characterized in that: The method comprises a memory, a processor, a communication interface and a bus; wherein the memory, the processor and the communication interface are connected to each other via the bus; and the processor is used to execute the image translation method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Digestive tract endoscope image enhancement method, system, equipment and medium
CN116957968A
Liver multi-modal image registration method based on deep learning
CN117314983A