Image region analysis method based on entropy driving feature enhancement
By introducing an entropy-driven feature enhancement method, the problems of regional information entropy differences and pseudo-feature interference in high-resolution image region analysis are solved, achieving high-precision and efficient image region analysis and improving the discriminative power of feature representation and the robustness of the model.
Patent Information
- Application Number
- CN202511623180.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-07
- Publication Date
- 2025-12-05
- Estimated Expiration
- 2045-11-07
AI Technical Summary
Existing image region analysis methods face challenges when processing high-resolution images, including significant differences in regional information entropy, severe interference from spurious features, lack of uncertainty modeling and saliency identification, and high parameter size and training costs, making it difficult to meet the high accuracy and efficiency requirements in complex scenarios.
An entropy-driven feature enhancement method is adopted. By calculating the multi-scale feature entropy distribution, an entropy weight mask is generated to strengthen high-entropy discriminative features and suppress low-entropy redundant features. An entropy-focused cross-space attention module is constructed to allocate attention and enhance the robust representation of real semantic regions.
It significantly improves the accuracy and robustness of image region analysis, enhances the discriminative power of feature representation and the robustness of the model, effectively suppresses pseudo-feature interference in complex scenes, and improves computational efficiency and resource utilization.
Smart Images

Figure CN121074346A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of image analysis and feature enhancement, and particularly relates to an image region analysis method based on entropy-driven feature enhancement. BACKGROUND
[0002] As an important part of computer vision and intelligent image understanding, image region analysis is the core foundation link of target recognition, scene segmentation, anomaly detection, target tracking and semantic understanding. With the popularity of high-resolution and multi-modal images, the spatial structure, texture details and semantic information contained in the image are increasingly rich, providing a more sufficient basis for fine-grained region recognition, structure analysis and feature mining. However, how to efficiently and accurately complete image region analysis in a complex background has become one of the key technical problems that need to be broken through in the field of computer vision.
[0003] The core goal of image region analysis is to accurately extract discriminative region features from large-scale complex images to support a variety of intelligent perception and decision-making tasks downstream. Its research significance lies in: providing accurate region boundaries and semantic information for automated recognition, segmentation and detection; helping to mine potential spatial relationships and structural features in high-dimensional and diversified images, enhancing the comprehensive understanding ability of the model to the global and local information of the scene; by accurately identifying key regions and suppressing redundant information, the computational efficiency, robustness and generalization performance of subsequent tasks can be effectively improved, and the resource consumption of large-scale system deployment can be reduced.
[0004] In the prior art, convolutional neural networks (CNN) and self-attention models (Transformer) have made significant progress in multi-scale feature extraction and global context modeling, but existing image region analysis methods still face significant challenges: first, the region information is highly heterogeneous, and the information entropy of different regions is significantly different. Unified processing strategy weakens the representation ability of high-entropy key regions; second, external light, sensor differences, noise and acquisition conditions introduce a large number of pseudo-features, which interfere with the true region features, reducing the accuracy and robustness of the analysis; third, existing models lack uncertainty modeling and saliency recognition based on information quantity, CNN has limitations in modeling long-range dependencies, and Transformer has global capabilities but attention allocation is uniform, making it difficult to highlight high-entropy regions and consume excessive resources in low-entropy regions; fourth, the model still has deficiencies in parameter size, training cost and multi-scale feature consistency, making it difficult to meet the high-precision and high-efficiency requirements of region analysis in complex scenarios.
[0005] In view of the above problems, there is an urgent need for a method that can explicitly quantify the region information entropy and perform differential modeling based on the saliency level, further improving the accuracy, efficiency and robustness of image region analysis. SUMMARY
[0006] In view of the problem that the region recognition accuracy is insufficient due to uneven distribution of feature information and interference of false features in existing high-resolution image region analysis, the application provides an image region analysis method based on entropy-driven feature enhancement, which introduces an entropy-driven mechanism in the links of feature extraction, weight distribution and multi-scale fusion, strengthens the feature expression of high information amount regions and suppresses the redundant interference of low information amount regions, so as to realize accurate, efficient and robust analysis of image regions.
[0007] The application adopts the following technical scheme: an image region analysis method based on entropy-driven feature enhancement, comprising the following steps: Step 1, data acquisition: collect multi-source image data as input samples and perform a preprocessing operation; Step 2, feature extraction and representation: a deep convolutional neural network or a transformation network structure is used to extract features from the preprocessed image, to obtain a multi-layer feature map and realize hierarchical feature representation of the image; Step 3, construct an entropy-driven feature adaptive enhancement module, calculate the multi-scale feature entropy distribution, generate an entropy weight mask, strengthen high-entropy discriminative features, suppress low-entropy redundant features, and improve the discriminability and semantic representativeness of the feature layer; Step 4, construct an entropy-focused cross-space attention module, combine entropy weight and multi-scale feature interaction, and perform entropy-prior-based attention distribution in the spatial dimension, focus on regions with significant information amount, and enhance the consistency perception and robust representation of real semantic regions; Step 5, output a black and white binary prediction map through an image transformation detection system, perform image region analysis, and black represents a region where no change is recognized, and white represents a region where change is recognized.
[0008] Preferably, the multi-source image data includes remote sensing images, medical images and natural scene image data.
[0009] Uniform preprocessing operations are performed on different types of image data, including normalization and size standardization, channel alignment and color space conversion.
[0010] Preferably, in the multi-layer feature map, the low-layer feature map contains texture, edge, brightness local detail information, and the high-layer feature map contains semantic structure and spatial context information.
[0011] Preferably, the entropy-driven feature adaptive enhancement module calculates the channel information entropy of the image data to quantify the information density distribution of the image features, and the method is as follows: The feature values of each channel are subjected to min-max normalization processing and linearly mapped to a uniform numerical interval; The feature map is multiplied by the entropy weight mask to obtain an enhanced feature map; The spatial dimension is flattened into a one-dimensional vector, and the information uncertainty and complexity of the feature region are evaluated by quantifying the frequency of each eigenvalue; The higher the entropy value, the greater the information uncertainty of the corresponding channel, the richer the detail and variation information; the lower the entropy value, the more uniform the feature distribution of the corresponding channel, the higher the information redundancy, and the more single the feature mode.
[0012] Further, based on the calculated entropy values of each channel, the information complexity contained in different channels is quantified, and the feature map is adaptively weighted; The normalized entropy value is multiplied by the original feature map as a saliency weight to strengthen the discriminative features of low-entropy flat regions and suppress the redundant background interference of high-entropy complex regions, highlighting the key regions.
[0013] Preferably, the entropy focusing cross-space attention module uses local entropy prior as a feature level guide signal, embeds feature entropy values into attention weight generation and gating calculation process, and fuses with deep semantic features to locate discriminative features in key regions, as follows: Step 4.1, input the de-noised and information entropy weighted dual-phase feature map, perform pooling processing, and perform maximum pooling and average pooling operations on each phase feature respectively, to balance between high-entropy local regions and overall spatial distribution; Step 4.2, concatenate the maximum pooled features and the average pooled features of the dual-phase, and output the concatenated four-channel feature map; Step 4.3, apply a convolution operation on the concatenated four-channel feature map to capture the correlation of neighborhood features; Step 4.4, perform nonlinear transformation on the convolution result through Sigmoid function to generate a spatial attention weight map representing the importance of spatial position; Step 4.5, multiply the spatial attention weight map with the original entropy weighted feature element by element to achieve adaptive weighting of image spatial features.
[0014] Preferably, in step 4.1, maximum pooling is used to retain significant responses in local regions and highlight the saliency of high-entropy regions; average pooling is used to extract global statistical features to reflect overall information distribution.
[0015] Preferably, in step 4.1, the pooling processing further includes: Norm pooling is used to enhance feature expression diversity.
[0016] Preferably, in step 4.3, the convolution kernel size is 7x7 or 3x3 or 5x5, or a dilated convolution is used.
[0017] Preferably, in step 4.4, the spatial position is normalized weight distribution by replacing the Sigmoid function with the Softmax function.
[0018] Preferably, the image region analysis is performed by the image transformation detection system, and the method is as follows: Step 5.1, inputting dual-phase picture data, constructing a dual-branch twin network using the first four stages of EfficientNet-V2-S, and extracting dual-phase picture features through a decoder; Step 5.2, inputting the dual-phase picture features into the entropy-driven feature adaptive enhancement module, calculating the entropy values of the multi-scale features, enhancing the discriminative features related to the real changes, and suppressing the feature interference corresponding to the pseudo changes; Step 5.3, distinguishing the object edge and morphological change features through the entropy focusing cross-space attention module; Step 5.4, inputting the feature map into the Uni-FIRE module (unified fusion and refinement module) for interpolation processing: The main branch sequentially passes through a first convolutional layer, a SiLU activation layer and a batch normalization layer, and then passes through a second convolutional layer and a batch normalization layer for processing, and outputs a prediction map; The skip branch adopts 1x1 convolution for channel compression and fusion; The outputs of the skip branch and the main branch are added to generate a multi-scale fusion feature map.
[0019] Compared with the prior art, the above technical scheme has the following technical effects: 1. The method of the present application realizes the self-enhancement weighting and suppression of discriminative features by introducing information entropy as a feature saliency measure, significantly improving the discriminability and interpretability of feature expression.
[0020] 2. The method of the present application strengthens the response to key region features at the feature level with the help of the entropy-driven cross-layer attention mechanism, effectively suppresses the pseudo feature interference introduced by factors such as acquisition conditions, noise and background complexity, and enhances the region analysis accuracy and robustness of the model in complex scenes.
[0021] 3. The method of the present application significantly improves the feature quality, model discriminability and overall interpretability in the image region analysis task by introducing a physically guided adaptive weighting and interaction mechanism in feature representation, providing effective technical support for high-resolution image intelligent analysis and understanding. BRIEF DESCRIPTION OF DRAWINGS
[0022] Figure 1 is a structure diagram of the entropy-driven feature adaptive enhancement module in the present application; Figure 2 is a structure diagram of the entropy focusing cross-space attention module in the present application; Figure 3 The schematic diagram for calculating the entropy value of the features in the application; Figure 4 The feature self-adaptive enhancement result comparison analysis diagram of two different images according to the entropy value of the features in the application; Figure 5 The overall structure diagram of the image transformation detection system network built for verifying the image region analysis effect based on the entropy-driven feature enhancement in the application; Figure 6 The structure diagram of the Uni-FIRE module used in the network; Figure 7 The transformation detection visual comparison result diagram of the application and mainstream methods under different data sets. DETAILED DESCRIPTION
[0023] In order to make the purpose, technical scheme and advantages of the application more clear, the technical scheme of the application will be further described in detail below in combination with the drawings. The described embodiments are only a part of the embodiments involved in the application. All non-innovative embodiments of other researchers in the field on the embodiments belong to the protection scope of the application. Meanwhile, the step numbers in the embodiments are only set for the convenience of description and explanation, and the order between the steps is not limited in any way. The execution order of each step in the embodiments can be adaptively adjusted according to the understanding of those skilled in the art.
[0024] In an embodiment of the application, for the key problems in high-resolution image region analysis, such as feature expression ambiguity, significant pseudo-feature interference and difficulty in effectively extracting discriminative features caused by complex scene structure, highly heterogeneous spatial distribution and acquisition condition difference, an image region analysis method based on feature driving and adaptive weighting is proposed.
[0025] Firstly, a feature saliency quantification mechanism based on information entropy is established to realize feature-level differentiation and adaptive weighted enhancement of high and low information regions, and to suppress the redundancy and interference of low feature information regions. Then, a cross-layer attention fusion framework is designed to deeply fuse multi-scale structural information at the feature level and weaken the pseudo-feature response caused by acquisition condition difference and complex background. Finally, a region analysis model with physical interpretability for key regions and high robustness for complex environments is constructed to improve the clarity of feature expression and the accuracy of region recognition.
[0026] The specific steps of the image region analysis method in this embodiment are as follows:
[0027] Step 1: Data acquisition and preprocessing.
[0028] Collect multi-source image data as input samples, such as remote sensing images, medical images, or natural scene images, etc. For different types of image data, perform unified preprocessing operations, mainly including: Image normalization and size standardization; Channel alignment and color space conversion (such as RGB→Lab or HSV).
[0029] In this embodiment, the collected image data is multi-temporal high-resolution remote sensing image data. By calculating the local information entropy of each pixel channel and performing explicit quantization, the adaptive weighting and strengthening of image feature information are realized. The analysis results obtain the feature map with enhanced saliency and the change response result, in which the features of high information-intensive areas are highlighted, and the low information and noise areas are effectively suppressed, which can provide data basis for subsequent remote sensing image change detection, feature classification and target recognition applications.
[0030] Step 2: Feature extraction and representation.
[0031] Deep convolutional neural network (CNN) or transform network (Transformer) structure is used to extract features from the preprocessed image. After network extraction, multi-layer feature maps are obtained, and through this process, hierarchical feature representation of the image is realized, providing basic feature data for subsequent information entropy calculation.
[0032] In this embodiment, the multi-layer feature map includes low-level feature map and high-level feature map, wherein the low-level feature map mainly contains texture, edge, brightness and other local detail information, and the high-level feature map contains semantic structure and spatial context information.
[0033] Step 3: Construct Entropy-Driven Feature Adaptive Enhancement Module (EFAE).
[0034] The channel information entropy theory is introduced to calculate the saliency prior at the feature level and realize adaptive weighting, and the physical information entropy and deep learning feature selection mechanism are deeply integrated. This module can explicitly enhance the representation ability of high-entropy discriminative features, suppress the interference of low-entropy redundant features on the model, provide interpretable and quantifiable feature importance evaluation basis for the network, and significantly improve the expression ability and robustness of key change features.
[0035] Specifically, different regions in the image (such as texture complex target area and homogeneous background area) show obvious distribution difference in feature response: the more uniform the feature response, the richer the information of the corresponding spatial position, the higher the entropy value, reflecting that the features of this area have higher complexity and uncertainty.
[0036] Therefore, local information entropy can serve as an effective physical quantification index of image spatial information density, providing a theoretical basis for feature-level saliency modeling and adaptive weighting.
[0037] The structure of the entropy-driven feature adaptive enhancement module constructed in this implementation is as follows: Figure 1 As shown, the specific processing is as follows:
[0038] First, the feature values of each channel are normalized using a min-max method, linearly mapping them to a uniform numerical range (usually [0,1]). This transformation is achieved using the following formula: ; in, The input are the original feature values. For feature map The minimum value of each channel. For feature map The maximum value of each channel. For smoothing terms, The output is the normalized eigenvalues.
[0039] Furthermore, to facilitate probability-based entropy calculation, such as Figure 2 As shown, the spatial dimension of the feature map is flattened into a one-dimensional vector. The calculation of entropy depends on the probability distribution of the feature values in the whole. By quantifying the frequency of each feature value, the information uncertainty and complexity of the feature region can be evaluated.
[0040] The probability estimation uses the following formula: ; in, Represents the th element in the characteristic matrix The probability of each position. The flattened feature matrix is The value at that location.
[0041] The entropy value is calculated using the following formula: ; in, Indicates the first Entropy of the channel The smoothing constant term is usually set to... , Scale-invariant entropy weights are used to eliminate the impact of feature map size changes on entropy calculation, ensuring that entropy measurements are consistent and comparable across different scales.
[0042] It is particularly noted that the higher the entropy value, the greater the information uncertainty of the corresponding channel, and the richer the detail and variation information. Conversely, a lower entropy value indicates that the channel feature distribution is relatively uniform, the information redundancy is higher, and the feature mode contained is relatively single.
[0043] By calculating the entropy value of each channel, the information complexity contained in different channels is quantified, and then adaptive weighting processing of the feature map is realized. The weighting method is realized according to the following formula: ; Among them, is the weighted feature value.
[0044] Further, by multiplying the normalized entropy value as the saliency weight with the original feature map, adaptive enhancement of the high-entropy region and corresponding inhibition of the low-entropy region are realized.
[0045] The feature enhancement mechanism has a clear physical meaning. Its essence is the quantitative evaluation of the spatial feature information content, which can guide the subsequent network module to focus on the information-rich area, thereby improving the feature expression ability and scene adaptability in the subsequent task based on image analysis.
[0046] Step 4: Constructing an entropy-focused cross-spatial attention module (EnFocSA).
[0047] Based on the aforementioned feature saliency prior module based on information entropy, the EnFocSA module is further proposed in this embodiment to use local entropy prior as a guiding signal at the feature level, and to deeply fuse it with deep semantic features, embed feature saliency information in attention weight generation, and realize attention allocation under physical constraints through cross-temporal feature interaction and spatial correlation modeling.
[0048] Compared with the traditional channel aggregation method, this module can strengthen the response to real change areas at the feature level, effectively suppress the interference of pseudo-change features caused by imaging differences and other factors, and alleviate the problems of feature boundary blur and local structure information loss caused by single channel weighting. While maintaining physical interpretability, it significantly improves the fine-grained description ability and structural integrity of the features of the target area, improves the perception and discrimination accuracy of the model for real changes, and thus realizes accurate positioning and adaptive enhancement of discriminative features in key areas, providing more accurate and robust feature representation for image region analysis.
[0049] Specifically, as Figure 3As shown, first, the dual-phase feature maps after de-noising and information entropy weighting are input into the module. For the features of each phase, maximum pooling and average pooling operations are respectively performed.
[0050] Among them, the maximum pooling retains the most significant response in the local region, highlighting the significance of the high-entropy region; the average pooling extracts global statistical features, reflecting the overall information distribution. The combination of the two can balance between high-entropy local regions and overall spatial distribution.
[0051] Subsequently, the maximum-pooled features and the average-pooled features of the dual phase are spliced to obtain a four-channel tensor: Among them, respectively represent the maximum-pooled results and the average-pooled results of the dual phase, is the four-channel feature map after splicing.
[0052] To further model the local context information, the embodiment applies a convolution operation on the spliced feature map: Among them, is the convolution kernel, represents the convolution operation, is the convolution result.
[0053] Preferably, the size of the convolution kernel is 7x7. This operation can capture the correlation of neighborhood features, strengthen edge information and local structure expression, and avoid the boundary blur phenomenon after feature fusion.
[0054] Next, the convolution result is nonlinearly transformed by a Sigmoid function to generate a spatial attention weight map: Among them, is the Sigmoid function, is the spatial attention weight map, ranging from 0 to 1, and each position of the weight map represents the importance of the spatial position, which has clear physical interpretability.
[0055] Finally, the embodiment element-wise multiplies the spatial attention weight map and the original entropy-weighted features, thereby realizing adaptive weighting of image spatial features. Through this mechanism, the features of high-entropy complex regions are enhanced, while low-entropy or noise interference regions are suppressed, achieving the highlighting of key regions and the suppression of redundant information.
[0056] Further, visual analysis is performed.
[0057] In this embodiment, a comparative analysis of the adaptive feature enhancement results of two different images based on the entropy value of the features is performed in the MATLAB environment.
[0058] First, a pre-trained ResNet-18 network is used to extract high-level feature maps from two images. Then, the local entropy value of each feature location is calculated using a sliding window to generate an entropy weight matrix representing the complexity of the region. Next, the normalized original features are multiplied by the entropy weights to achieve adaptive adjustment of the feature values—background suppression is performed on high-entropy complex regions, and feature enhancement is performed on low-entropy flat regions. Finally, by calculating the difference between the weighted features, the signal-to-noise ratio of change detection is significantly improved, allowing real change regions to stand out against the complex background.
[0059] like Figure 4 As shown, Figure 4 (a) in the image is a comparison of the original features and the enhanced features of the T1 image; Figure 4 (b) in the figure is a comparison of the original features and the enhanced features of the T2 image; Figure 4 (c) in the figure is a weighted comparison chart of differences.
[0060] As can be seen from the figure, adaptive enhancement of features is achieved through local entropy weighting. The weight matrix is dynamically generated by utilizing the spatial complexity of deep features to enhance features in low-entropy smooth regions and suppress backgrounds in high-entropy complex regions. This significantly improves the signal-to-noise ratio of changing signals, allowing real changes to be accurately highlighted in backgrounds of varying complexity.
[0061] It should be noted that the 7×7 kernel size mentioned above is not the only option. In other embodiments, 3×3 or 5×5 kernels can be used to reduce computational complexity, or dilated convolution can be used to expand the receptive field. Regarding pooling methods, in addition to max pooling and average pooling, statistical histogram pooling can also be used. Methods such as norm pooling can be used to enhance the diversity of feature representation; the sigmoid function can also be replaced by the softmax function to achieve normalized weight allocation for spatial locations.
[0062] Through the above design, this embodiment's entropy-driven cross-spatial attention module can effectively fuse prior salient features and spatial distribution features in an image to generate a physically meaningful dense spatial attention map. This mechanism not only provides information-theoretic interpretability for feature selection and fusion, but also achieves precise localization and fine-grained enhancement of discriminative features in key regions through multi-scale feature compression and local correlation modeling, thereby significantly improving the accuracy, robustness, and clarity of feature representation in image region analysis.
[0063] Step 5: Construct an image transformation detection system and perform image region analysis.
[0064] To verify the effect of image region analysis based on entropy-driven feature enhancement, the embodiment constructs an image transformation detection system based on the region analysis results, and the system structure is as shown in Figure 5 .
[0065] The experimental results and analysis are as follows: Let the dual-time image be , and the size is 3xHxW, where H and W represent the height and width of the two images, respectively. As can be seen from the figure, the detection network uses the first four stages of EfficientNet-V2-S to construct a dual-branch twin network, which extracts features of the dual-time images through the decoder, respectively. Subsequently, the extracted image features enter the entropy-driven feature enhancement module to realize the region analysis of the image, and on the basis of the image analysis, the subsequent change detection is realized, and the entropy focusing cross-space attention module is used to distinguish the object edge and morphological change features.
[0066] Finally, the feature map is input into the Uni-FIRE module, and after interpolation processing, it is sequentially input into the first convolutional layer, the SiLU activation layer and the batch normalization layer, and then processed by the second convolutional layer and the batch normalization layer, and the prediction map is output.
[0067] As shown in Figure 6 , after splicing each input image, it is interpolated to 3xHxW, where H and W represent the height and width of the target interpolation, respectively.
[0068] Subsequently, the interpolated image is input into the main branch and the jump connection branch: In the main branch, it is sequentially input into the first convolutional layer, the SiLU activation layer and the batch normalization layer, and then processed by the second convolutional layer and the batch normalization layer.
[0069] In the jump connection branch, 1x1 convolution is used for channel compression and fusion.
[0070] Finally, the jump branch and the main branch output are added to generate a multi-scale fusion feature map.
[0071] Through the above steps, the embodiment can finally obtain a high saliency feature representation map optimized by entropy-driven feature enhancement and cross-layer attention fusion.
[0072] The results have the following characteristics and uses:
[0073] (1) Result characteristics The output feature map significantly enhances the response of high information density regions in the spatial dimension, and suppresses low information and noise regions; In the semantic dimension, the discriminability of key target regions and backgrounds is enhanced, making the region boundary clearer and the response more concentrated; The feature distribution has higher separability and information integrity, and can provide robust input features for subsequent visual tasks.
[0074] (2) Result use: It can be used as input features or intermediate representations for tasks such as saliency detection, target recognition, change detection, image segmentation, medical lesion extraction, etc.; in remote sensing image analysis, it can be used to highlight high information areas such as building groups and road intersections, to achieve more accurate geospatial change detection; in medical image analysis, it can be used to strengthen the features of abnormal areas (such as lesions and tumors), to improve the accuracy of automatic diagnosis.
[0075] Further, the method of the present application is compared with a variety of representative RSCD methods as baseline models, including FC-Siam-Conc, FC-Siam-Diff, SNUNet, ChangeFormer, DASNet, STANet, BiT and SEIFNet, all baseline methods are re-implemented under the same network configuration to ensure fair and consistent comparison.
[0076] The results are as follows:
[0077] (1) Effectiveness and complementarity of entropy plus cross-attention mechanism: When the saliency prior module of information entropy is used alone, the F1 value on LEVIR-CD and WHU-CD datasets reaches 92.04% / 93.49%, respectively, which is 0.15% / 0.49% higher than the baseline network, and the entropy module and the cross-space attention module work together to further improve to 92.25% / 94.27%, respectively, which is 0.21% / 0.78% higher than the baseline network, verifying the effective cooperation of the two modules in image analysis.
[0078] (2) Visual verification is intuitive and effective: The present application shows significant advantages in complex scenes. The entropy-driven feature adaptive enhancement module can effectively enhance the discriminative features related to real changes, while suppressing the feature interference of pseudo changes; combined with the entropy-focused cross-space attention mechanism, the model performs outstandingly in expressing the continuity of features at the change boundary and depicting fine-grained features, and can clearly distinguish the features related to building edge and shape changes, fully verifying the effectiveness and application potential of the method in complex scenes for high-discriminative feature extraction and interpretable change detection.
[0079] The experimental results on the LEVIR-CD and WHU-CD datasets are summarized in Tables I and II, respectively, and the results with the best performance are highlighted in bold. The results show that the proposed network achieves F1 values of 92.25% and 94.27% on the LEVIR-CD and WHU-CD datasets, respectively, which are superior to a variety of latest comparison methods.
[0080] Table I:
[0081] Table II:
[0082] Specifically, on the WHU-CD dataset, the proposed model outperforms other models in all evaluation indicators; on the LEVIR-CD dataset, although the recall (Rec) is slightly lower than that of the BIT model, the Pre, F1, OA and IoU indicators all reach the highest value, which reflects the excellent overall performance.
[0083] To further verify the effectiveness of each key module in the application, the present embodiment carries out ablation experiments on the LEVIR-CD and WHU-CD two public datasets, as shown in the following Table III and Table IV.
[0084] Table III:
[0085] Table IV:
[0086] The baseline model of the experiment is composed of an encoder based on a high-efficiency convolutional neural network and a decoder and prediction head composed of a reusable unified fusion and refinement module. The innovative modules of the application are gradually introduced on the basis of the baseline model to evaluate their independent contribution and synergistic effect.
[0087] In the first group of experiments, the saliency adaptive enhancement module based on information entropy is applied to the skip connection, so that the features output by the encoder can highlight the high information area and suppress the low feature information area through the entropy-driven weighting mechanism when transmitting across scales. The results show that this module can significantly improve the response capability of the model in the real change area, while effectively reducing the false detection caused by noise and pseudo differences, and the performance indicators are significantly improved compared with the baseline model.
[0088] In the second group of experiments, the entropy focusing cross-space attention module is further introduced on the basis of the baseline model. This module uses entropy prior to guide the generation process of spatial attention weights, promotes the cross-space interaction and fusion between deep features of dual-phase images, and thus enhances the positioning ability of the model to change-related discriminative features. The experimental results show that this module brings significant improvement in feature boundary continuity maintenance, fine-grained change feature description, and feature robustness in complex background, fully proving the role of entropy information in feature optimization and enhancement in the spatial attention mechanism.
[0089] In addition, in order to explore the performance of the cross-space attention module based on entropy driving under different convolution kernel sizes, the embodiment performs additional comparative experiments. The results show that when the convolution kernel size is set to , the model performs best in terms of detection accuracy and boundary preservation. This conclusion shows that a larger receptive field can better capture neighborhood correlation and preserve structural details, so as to be preferred, the convolution kernel size of the module is set to 7.
[0090] Through the above ablation experiments, it can be seen that the feature adaptive saliency enhancement module based on information entropy and the cross-space attention module based on entropy focusing proposed in the present application can effectively improve the feature discrimination ability when used alone, and when working together, they can complement each other at the feature level: on the one hand, the expression of high-entropy discriminative features is strengthened, and on the other hand, the interaction and attention distribution of cross-temporal features are optimized, so as to realize stable detection and accurate positioning of change-related features in complex scenes. The experimental results verify the rationality of the design idea and the effectiveness of the overall architecture from the feature dimension.
[0091] Figure 7 The change detection results of the system on different image samples and the comparison with other models are shown. It can be seen that the method proposed in the present application has significant advantages in the recognition and enhancement of key area features.
[0092] Specifically, the module based on information entropy, as a feature prior mechanism with physical interpretability, can effectively capture discriminative features of potential change regions and achieve high-fidelity boundary description in information-intensive regions (such as (c), (e) in Figure 7 , significantly improving the semantic expression quality of change features. At the same time, the cross-space attention mechanism strengthens the feature response of the real change region and suppresses the low information or noise interference (such as (b), (d) in Figure 7 , effectively reducing the false alarm rate and improving the reliability of feature selection. In addition, the framework has excellent discriminative ability in edge structure preservation and small-scale region feature extraction.
[0093] The visualization results further show that the proposed method can generate change maps with coherent features and complete structures, highlighting the robustness of regional feature perception and the accuracy of change detection in complex image scenes.
[0094] In summary, through the entropy-driven feature enhancement and saliency-guided fusion mechanism of the present application, the system can generate image region analysis results with high information fidelity, high saliency response and strong discriminative ability, thereby providing more physically interpretable and robust feature support for complex visual tasks.
[0095] The above merely describes the preferred embodiments of the present application, and it should be pointed out that those skilled in the art can make several improvements and refinements without departing from the principles of the present application, and these improvements and refinements should also be considered as the protection scope of the present application.
Claims
1. An image region analysis method based on entropy-driven feature enhancement, characterized in that, The method comprises the following steps: Step 1, data acquisition: collect multi-source image data as input samples and perform preprocessing operations; Step 2, feature extraction and representation: use a deep convolutional neural network or a transformation network structure to extract features from the preprocessed image, obtain a multi-layer feature map, and realize hierarchical feature representation of the image; Step 3, construct an entropy-driven feature adaptive enhancement module, calculate the multi-scale feature entropy distribution, generate an entropy weight mask, strengthen high-entropy discriminative features, suppress low-entropy redundant features, and improve the discriminability and semantic representativeness of the feature layer; Step 4, construct an entropy-focused cross-space attention module, combine entropy weights and multi-scale feature interactions to perform entropy-prior-based attention allocation in the spatial dimension, focus on areas with significant information, and enhance consistent perception and robust representation of real semantic regions; Step 5, output a black and white binary prediction map through the image transformation detection system, and perform image region analysis. Black represents areas where changes are not identified, and white represents areas where changes are identified.
2. The entropy-driven feature augmentation based image region analysis method of claim 1, wherein, The multi-source image data includes remote sensing images, medical images, and natural scene image data; Uniform preprocessing operations are performed on different types of image data, including normalization and size standardization, channel alignment, and color space conversion.
3. The entropy-driven feature augmentation based image region analysis method of claim 1, wherein, In the multi-layer feature map of step 2, the low-level feature map contains texture, edge, and local brightness detail information, and the high-level feature map contains semantic structure and spatial context information.
4. The entropy-driven feature augmentation based image region analysis method of claim 1, wherein, In step 3, the entropy-driven feature adaptive enhancement module calculates the channel information entropy of the image data to quantify the information density distribution of the image features. The method is as follows: Perform min-max normalization processing on each channel feature value and linearly map it to a unified numerical interval. The formula is: ; wherein, is the input raw feature map, is the feature map is the minimum value of each channel in the feature map, is the feature map is the maximum value of each channel in the feature map, is the smoothing term, is the normalized feature value output; The spatial dimension of the feature map is flattened into a one-dimensional vector, and the information uncertainty and complexity of the feature region are evaluated by quantifying the frequency of each feature value. The probability estimation formula is: ; wherein, denotes the probability of the th position in the feature matrix, is the value of the flattened feature matrix at , and denote the height and width of the image; Calculate the channel information entropy value. The formula is: ; wherein, denotes the entropy of the channel, is a smoothing constant term, is a scale-invariant entropy weight used to eliminate the influence of feature map size variation on the entropy value calculation.
5. The entropy-driven feature enhancement based image region analysis method according to claim 4, characterized in that, The higher the entropy value, the greater the uncertainty of the corresponding channel information, the richer the detail and variation information. The lower the entropy value, the more uniform the corresponding channel feature distribution, the higher the information redundancy, and the more single the included feature mode; Based on the calculated channel entropy values, the complexity of the information contained in different channels is quantified, and the feature map is adaptively weighted. The method is as follows: ; wherein, is the weighted feature value; Multiply the normalized entropy value by the original feature map as the saliency weight to strengthen the discriminative features in low-entropy flat areas and suppress the redundant background interference in high-entropy complex areas, highlighting the key areas.
6. The entropy-driven feature augmentation based image region analysis method of claim 1, wherein, In step 4, the entropy-focused cross-space attention module uses local entropy prior as a feature layer guide signal, embeds feature entropy values into attention weight generation and gating calculation processes, fuses deep semantic features, and locates discriminative features in key areas. The method is as follows: Step 4.1, input the de-noised and information entropy weighted dual-phase feature map, perform pooling processing, and perform maximum pooling and average pooling operations on each phase feature to balance between high-entropy local areas and overall spatial distribution; Step 4.2, splice the maximum pooled features and average pooled features of the dual-phase to obtain a four-channel tensor: ; wherein, respectively represent the maximum pooling result and the average pooling result of the double time phase, is the four-channel feature map after splicing, is a feature splicing function; Step 4.3, apply a convolution operation on the four-channel feature map after splicing to capture the correlation of neighborhood features: ; wherein, is a convolution kernel, denotes a convolution operation, is a convolution result; Step 4.4, the convolution result is nonlinearly transformed by a Sigmoid function to generate a spatial attention weight map representing the importance of spatial positions: ; wherein, and denote the height and width of the image, is a Sigmoid function, is a spatial attention weight map, ranging between 0 and 1; Step 4.5, multiply the spatial attention weight map with the original entropy-weighted features element by element to achieve adaptive weighting of image spatial features.
7. The entropy-driven feature enhancement based image region analysis method according to claim 6, characterized in that, In step 4.1, the max-pooling is used to retain significant responses in local regions, highlighting the significance of high-entropy regions; the average pooling is used to extract global statistical features, reflecting the overall information distribution; The pooling process further includes: statistical histogram pooling, Norm pooling is used to enhance the diversity of feature representation.
8. The entropy-driven feature enhancement based image region analysis method of claim 6, wherein, In step 4.3, the convolution kernel with a size of 7x7 or 3x3 or 5x5, or with a dilated convolution.
9. The entropy-driven feature enhancement based image region analysis method of claim 6, wherein, In step 4.4, replace the Sigmoid function with a Softmax function to normalize the weight distribution of spatial positions.
10. The entropy-driven feature augmentation based image region analysis method of claim 1, wherein, The image transformation detection system in step 5 processes as follows: Step 5.1, input dual-phase picture data, construct a dual-branch twin network using the first four stages of EfficientNet-V2-S, and extract dual-phase picture features through decoders; Step 5.2, input the dual-phase picture features into the entropy-driven feature adaptive enhancement module, calculate the entropy values of multi-scale features, enhance the discriminative features related to real changes, and suppress the feature interference corresponding to false changes; Step 5.3, distinguish object edge and morphological change features through the entropy-focused cross-space attention module; Step 5.4, input the feature map into the unified fusion and refinement module for interpolation processing: The main branch passes through a first convolution layer, a SiLU activation layer, and a batch normalization layer in turn, and then passes through a second convolution layer and a batch normalization layer for processing, outputting a prediction map; The skip branch uses 1x1 convolution for channel compression and fusion; The outputs of the skip branch and the main branch are added to generate a multi-scale fusion feature map.
Citation Information
Patent Citations
Fine-grained image recognition method based on saliency attention mechanism
CN113642571A
Image sensitive character desensitization method and device, equipment and medium
CN120198925A
Transform-CNN medical image segmentation method and system based on multi-scale fusion semantic enhancement
CN120318256A
Wetland semantic change detection system and method based on frequency domain information and multi-scale feature extraction
CN120544039A
Remote sensing sewage area identification method and system based on graph structure and multi-stage enhancement
CN120726484A
Cited By
Visible light image saliency prediction method and device, equipment and storage medium
CN122090085A
Detection method and device for animal excrement flushing calculation based on image entropy upsampling
CN122223756A
Detection method and device for animal fecal flushing calculation based on image entropy up-sampling
CN122223756B