A remote sensing image semantic segmentation method based on multi-level semantic reasoning
Patent Information
- Application Number
- CN202610854080.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-12
- Publication Date
- 2026-09-25
AI Technical Summary
然而,该类方法主要依赖卷积或局部窗口自注意力机制进行表征,卷积核局部建模难以有效捕获方向一致性与周期性纹理等特征,使得屋顶边缘、道路等结构易出现粘连、模糊或断裂等情形;且Swin Transformer虽然具备长距离建模能力,但其窗口机制并未包含对尺度变化的自适应选择能力,导致模型在细粒度类别之间的识别能力不足
1. 本发明中的跨尺度语义路由模块突破了传统多尺度特征融合依赖固定权重或全局共享策略的局限,通过构建像素级的动态尺度选择机制,使网络能够根据局部纹理复杂度、结构形态及语义需求自动选择最合适的上下文尺度。在处理小目标、细粒度边界或高纹理变化区域时,模块能够优先路由至细尺度特征,从而提升轮廓清晰度与细节辨识能力;在处理大面积同质区域或低纹理背景时,则可自适应选择较大尺度的语义上下文,以增强特征平滑性与稳定性。该机制有效避免了小目标被大尺度特征淹没、边缘被过度平滑、不同尺度之间语义响应不一致等问题,使得整体特征表达更加符合空间结构特性,显著提升了遥感影像的跨尺度建模能力。
Smart Images

Figure CN122821115A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of deep learning and image processing, specifically to a method for semantic segmentation of remote sensing images based on multi-level semantic reasoning. Background Technology
[0002] High-resolution remote sensing images are characterized by complex scene structures, drastic changes in target scale, blurred category boundaries, and high global semantic coupling, posing significant challenges to existing semantic segmentation models in several key aspects. First, remote sensing images simultaneously contain targets of different scales, such as buildings, small vehicles, vegetation, and roads. Deep networks often employ fixed or globally shared fusion strategies when decoding multi-scale information, failing to select appropriate contextual scales based on the semantic needs of different regions. This leads to problems such as weak responses to small targets, blurred and adhered edges, and inconsistent semantic representations across different scales. Second, structures in remote sensing scenes often exhibit clear directional features, such as the linear orientation of roads and the horizontal or vertical structure of building edges. However, existing convolutional or window attention mechanisms struggle to model the collaborative relationships between features of different directions, making them prone to insufficient direction perception, structural breaks, and edge instability when features are occluded, subject to noise interference, or structural discontinuities. Furthermore, remote sensing images contain extensive regional semantic relationships. For example, although the same type of land cover is spatially dispersed, it shares a similar semantic distribution. Different categories may also contain, mix, or conflict with each other. Therefore, relying solely on local convolution or local window attention mechanisms is insufficient to capture such long-distance regional dependencies, making the model prone to phenomena such as category fragmentation, inconsistent regional representation, and category confusion in complex scenes.
[0003] Zeng Junying (“Multi-level Branch Cross-scale Fusion Semantic Segmentation Network for Remote Sensing Images”, Advances in Laser & Optoelectronics, 2024, 1-20) constructed a multi-level branch network structure, designing a shallow Swin Transformer feature extraction module, spatial branches, semantic branches, and boundary branches. Each branch focuses on extracting feature information at a specific level, and a multi-scale decoding module with a large receptive field is introduced to transmit feature information at different scales. However, this type of method mainly relies on convolution or local window self-attention mechanisms for representation. Local modeling by convolutional kernels is difficult to effectively capture features such as directional consistency and periodic texture, making structures such as roof edges and roads prone to adhesion, blurring, or breakage. Moreover, although Swin Transformer has long-distance modeling capabilities, its window mechanism does not include adaptive selection capabilities for scale changes, resulting in insufficient recognition ability between fine-grained categories. In mixed urban and rural areas, and scenes with complex textures and large scale spans, these limitations further lead to problems such as unclear category boundaries, unstable regional representation, and missed detection of small targets.
[0004] It should be noted that the information disclosed in the background section above is only used to enhance the understanding of the background of the present invention, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention
[0005] To address the problems of drastic changes in target scale, easy breakage of directional structure, and fragmentation of categories in remote sensing images, this invention provides a semantic segmentation method for remote sensing images based on multi-level semantic reasoning.
[0006] Other features and advantages of the invention will become apparent from the following detailed description, or may be learned in part by practice of the invention.
[0007] According to a first aspect of the present invention, a method for semantic segmentation of remote sensing images based on multi-level semantic reasoning is provided, the method comprising: Remote sensing images are input into a backbone feature extraction network to obtain multi-scale feature maps; the multi-scale feature maps include first-stage feature maps. Second-stage feature map Third-stage feature map and the fourth stage feature map ; The characteristics of the fourth stage After being processed by the CBRU module, it is recorded as follows: ; The third stage feature map New features are obtained by connecting the CSSR module and the residual. ,Will and Adding them together gives us new features. ; The second-stage feature map The input is the Cross-Scale Semantic Routing (CSSR) module, which is then connected to the residual to obtain new features. ,Will and The feature maps obtained by addition are input into the CBRU module to obtain a new feature map, denoted as . The new feature map is denoted as Compared with the first stage feature map Direct element-wise addition, followed by the CBR module, yields a new feature, denoted as... ; New features As a three-way parallel input direction cooperative reasoning graph module, DCRG extracts directional semantic nodes through multiple directional convolutional kernels, constructs an inter-directional cooperative graph and performs information transmission, and outputs directional consistency enhancement features. The direction consistency enhancement feature is input into the semantic graph self-evolution module SGSE. A region semantic graph is constructed through a bidirectional attention mechanism between learnable semantic nodes and pixel features. After multiple graph evolutions, semantic information is back-transmitted to the pixel space, and global consistency enhancement features are output. Based on the global consistency enhancement features, the semantic segmentation results of the remote sensing image are output through the classification layer.
[0008] In some exemplary embodiments, the execution of the cross-scale semantic routing module (CSSR) includes: The input features are convolutionally reduced to obtain the dimensionality-reduced features; On the dimensionality reduction features, three different window sizes of local self-attention mechanisms are applied to generate feature maps corresponding to different scales; The feature maps of different scales are stitched together; Pixel-level dynamic routing weights are generated for each scale using convolutional layers and the softmax function; The feature maps of different scales are weighted and fused according to the pixel-level dynamic routing weights to obtain fused features; The fused features are residually concatenated with the processed original input features and upsampled to restore the original resolution, outputting the enhanced cross-scale features.
[0009] In some exemplary embodiments, the execution of the Directional Collaborative Reasoning Graph (DCRG) module includes: The input features are processed by four convolutional kernels or attention mechanisms with different directions to extract feature components in different directions and obtain multi-directional feature maps. Global pooling is performed on the directional feature map of each path to obtain the corresponding directional node representation; Multiple directional node representations are stacked to form a directional node matrix; Calculate the similarity between the directional nodes and construct a normalized adjacency matrix between the directional nodes; Information propagation and updating are performed between nodes in the stated direction based on the normalized adjacency matrix. Channel-level weights are generated based on the updated node information, and the multi-directional feature maps are weighted and fused to output directional enhanced features.
[0010] In some exemplary embodiments, the four-way convolutional kernels or attention mechanisms with different directions include a first path for extracting vertical features, a second path for extracting horizontal features, a third path for extracting features along the main diagonal, and a fourth path for extracting features along the anti-diagonal.
[0011] In some exemplary embodiments, the execution of the semantic graph self-evolution module SGSE includes: The input features are transformed and flattened to obtain pixel feature representations; Initialize a set of learnable semantic nodes; Calculate the attention weights between the pixel feature representation and the learnable semantic node, and obtain the node-level representation by aggregation; Based on the node-level representation, self-attention calculation and information update are performed between nodes to complete a self-evolution on the graph structure. The evolved and updated node representations are projected back into the pixel space and fused with the original pixel features to enhance the output features with improved semantic consistency.
[0012] In some exemplary embodiments, the step of projecting the evolved and updated node representation back into the pixel space specifically involves: mapping the updated node representation into pixel-level semantic enhancement weights through a linear transformation, multiplying them element-wise with the flattened original pixel features, and then outputting them after shape restoration.
[0013] In some exemplary embodiments, the method further includes a training process in which a weighted sum of a cross-entropy loss function and a Dice loss function is calculated based on the final output feature map, and the model is trained and optimized under the constraint of the total loss function until convergence.
[0014] In some exemplary embodiments, the CBRU module refers to a concatenated operation of convolution transformation, batch normalization, ReLU nonlinear activation, and upsampling.
[0015] In some exemplary embodiments, the CBR module is based on the CBRU module with the final upsampling operation removed.
[0016] According to a second aspect of the present invention, a storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the remote sensing image semantic segmentation method based on multi-level semantic reasoning as described in the first aspect.
[0017] According to a third aspect of the present invention, a computer program product is provided, on which a computer program is stored, wherein when the computer program is executed by a processor, it implements the remote sensing image semantic segmentation method based on multi-level semantic reasoning described in the first aspect above.
[0018] According to a fourth aspect of the present invention, an electronic device is provided, comprising: Processor; and Memory for storing the executable instructions of the processor; The processor is configured to implement the remote sensing image semantic segmentation method based on multi-level semantic reasoning as described in the first aspect by executing the executable instructions.
[0019] The remote sensing image semantic segmentation method based on multi-level semantic reasoning provided by the embodiments of the present invention has the following advantages compared with the prior art: 1. The cross-scale semantic routing module in this invention overcomes the limitations of traditional multi-scale feature fusion strategies that rely on fixed weights or global sharing. By constructing a pixel-level dynamic scale selection mechanism, the network can automatically select the most suitable context scale based on local texture complexity, structural morphology, and semantic requirements. When processing small targets, fine-grained boundaries, or regions with high texture variation, the module can prioritize routing to fine-scale features, thereby improving contour clarity and detail recognition capabilities. When processing large homogeneous regions or low-texture backgrounds, it can adaptively select a larger-scale semantic context to enhance feature smoothness and stability. This mechanism effectively avoids problems such as small targets being overwhelmed by large-scale features, edges being over-smoothed, and inconsistent semantic responses between different scales, making the overall feature representation more consistent with spatial structural characteristics and significantly improving the cross-scale modeling capability of remote sensing images.
[0020] 2. The inter-directional collaborative reasoning graph module of this invention constructs multi-directional semantic nodes containing horizontal, vertical, and diagonal directions, and establishes collaborative reasoning relationships between nodes, enabling the network to explicitly capture and model directional consistency information in the scene. This module can repair directional breaks caused by occlusion, noise, or texture discontinuities through inter-node information propagation, allowing structures with obvious directions, such as roads and building edges, to maintain continuous and stable responses. Simultaneously, the inter-directional reasoning mechanism can dynamically strengthen and suppress directional features, maintaining directional consistency in complex regions with irregular structures and numerous transitions. Compared to traditional convolutional or local attention methods, this module significantly improves the network's ability to represent linear targets, regular edges, and directional textures, effectively reducing problems such as boundary blurring, orientation deviation, and structural misclassification.
[0021] 3. The semantic graph self-evolution module proposed in this invention models semantic relationships from a regional global perspective. By establishing a learnable semantic node-pixel region bidirectional semantic transfer mechanism, the network can unify category semantic representation globally. The module captures cross-regional semantic commonalities through node aggregation, updates the semantic relationships between categories through graph reasoning between nodes, and finally injects the updated semantic information into the pixel space to enhance semantic consistency. This mechanism can effectively reduce category fragmentation in complex backgrounds, enabling similar regions to maintain consistent semantic feature expression in wide-area scenes. Furthermore, for easily confused categories (such as bare soil and farmland, roads and building edges), the module can significantly improve category discrimination ability through cross-regional semantic constraints. Compared to traditional methods relying on local convolution or local window attention, this module exhibits more stable, robust, and consistent discrimination ability in large-scale, mixed-category, and textured regions.
[0022] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit the invention. Attached Figure Description
[0023] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention. It is obvious that the drawings described below are merely some embodiments of the invention, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort.
[0024] Figure 1 This is a schematic diagram of the method flow of the present invention; Figure 2 This is a schematic diagram of the structure of the HSRNet network involved in the present invention; Figure 3 This is a schematic diagram of the CSSR module involved in the present invention; Figure 4 This is a schematic diagram of the DCRG module involved in the present invention; Figure 5 This is a schematic diagram of the SGSE module involved in the present invention; Figure 6 This is a schematic diagram illustrating the visualization results of the present invention. Detailed Implementation
[0025] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, they are provided so that the invention will be more comprehensive and complete, and will fully convey the concept of the exemplary embodiments to those skilled in the art. The described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.
[0026] Furthermore, the accompanying drawings are merely illustrative of the invention and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and therefore repeated descriptions of them will be omitted. Some block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.
[0027] To address the shortcomings and deficiencies of existing technologies, this invention proposes a multi-level semantic reasoning system composed of three innovative modules: a cross-scale semantic routing module, an inter-directional collaborative reasoning graph module, and a semantic graph self-evolution module. 1. The Cross-Scale Semantic Routing module (CSSR) is used to solve the problem of inconsistent semantic representation across multiple scales. Through a pixel-level dynamic scale selection mechanism, it enables small targets, edge regions, and large scene areas to obtain the most suitable context scale. 2. The Directional Cooperative Reasoning Graph module (DCRG) aims to enhance the stability of directional structures by constructing cooperative reasoning relationships between semantic nodes in multiple directions, thereby effectively restoring the consistency of structural orientation and the continuity of linear goals. 3. The Semantic Graph Self-Evolution module (SGSE) further enhances semantic consistency from a regional global perspective by constructing bidirectional semantic transfer between semantic nodes and regional features, thereby suppressing category fragmentation and improving the semantic consistency of the entire graph.
[0028] Based on the aforementioned multi-level semantic reasoning system, this example implementation provides a remote sensing image semantic segmentation method based on multi-level semantic reasoning, referencing... Figure 1 As shown, the specific steps may include: Step S1: Input the input remote sensing image into the backbone feature extraction network to obtain multi-scale feature maps; Step S2: In the second and third stages, the cross-scale semantic routing module CSSR performs pixel-level dynamic scale selection and fusion of multi-scale contexts, and outputs scale-adaptive enhanced features. Step S3: Input the scale-adaptive enhanced features into the inter-directional collaborative inference graph module DCRG, extract directional semantic nodes through multiple directional convolutional kernels, construct an inter-directional collaborative graph and perform information transmission, and output directional consistency enhanced features; Step S4: Input the direction consistency enhancement feature into the semantic graph self-evolution module SGSE, construct the region semantic graph through the bidirectional attention mechanism between learnable semantic nodes and pixel features, and after multiple graph evolutions, transmit semantic information back to the pixel space to output the global consistency enhancement feature. Step S5: Based on the global consistency enhancement features, output the semantic segmentation results of the remote sensing image through the classification layer.
[0029] The steps in this exemplary embodiment will now be described in more detail with reference to the accompanying drawings and embodiments.
[0030] In step S1, the input remote sensing image is input into the backbone feature extraction network to obtain multi-scale feature maps, specifically: Given an input remote sensing image: (1) Among them, parameters This represents the number of images input at one time during training and testing, i.e., the batch size. Parameter and These represent the height and width of the input image, respectively, with a channel count of 3, i.e., a standard RGB three-channel image. The input image is then fed into the backbone feature extraction network to obtain feature maps at different scales. (2) Among them, feature map It has high resolution and contains rich edge, texture, and detail information; feature map It has strong semantic abstraction capabilities and includes global contextual information; and It is at the middle level, possessing both spatial detail and semantic expressive capabilities.
[0031] In step S2, during the second and third stages, the cross-scale semantic routing module (CSSR) performs pixel-level dynamic scale selection and fusion of multi-scale contexts, outputting scale-adaptive enhanced features; specifically: To address the issue of inconsistent semantic representation across multiple scales, this invention designs a Cross-Scale Semantic Routing module (CSSR).
[0032] As attached Figure 2 As shown, a cross-scale semantic routing module (CSSR) is introduced in the second and third stages to enhance the scale-adaptive perception capability for different spatial regions. For any input feature map: (3) Among them, parameters Given the number of channels in the feature map for this stage, we first obtain the dimensionality-reduced features through a lightweight convolutional projection: (4) in, This represents the convolution projection operation, with parameters... The number of channels after dimensionality reduction, parameter and This refers to spatial resolution.
[0033] As attached Figure 3 As shown, CSSR transforms multi-scale context fusion into a pixel-level scale routing problem, enabling each spatial location to select the most suitable context scale according to its own semantic needs.
[0034] First, regarding dimensionality reduction features Self-attention operations were performed using three different window sizes: (5) Among them, the function Indicates the window size is Local window self-attention operation, These correspond to small-scale, medium-scale, and large-scale contextual features, respectively. The small-scale window enhances small objects, edge regions, and local texture details; the medium-scale window captures local structural relationships; and the large-scale window models semantic consistency and contextual dependencies within a larger region. Subsequently, the features at the three scales are concatenated along the channel dimension and processed using a method of size [size missing]. The convolutional mapping yields routing features: (6) in, Indicates by A mapping function consisting of convolution, batch normalization, and nonlinear activation.
[0035] For spatial location Calculate the routing scores for the three scales respectively: (7) Among them, parameters and They represent the first Learnable weights and biases for each scale branch. Softmax normalization is applied to the scores at the three scales to obtain pixel-level scale routing probabilities: (8) And satisfy: (9) in, Indicates spatial location For scale The selection probability. Based on the scale routing probability, adaptive fusion of the three scale features is performed: (10) Finally, through channel recovery and spatial upsampling operations, the routed features are restored to the original input feature dimensions, and residual connections are used to obtain the CSSR output: (11) Among them, the function This indicates operations such as channel restoration and spatial resolution restoration.
[0036] Through the above process, CSSR can allocate more small-scale context to small targets and boundary regions, and more large-scale context to large homogeneous regions, thereby effectively alleviating the problem of inconsistent scale response in traditional multi-scale fusion.
[0037] In step S3, the scale-adaptive enhanced features are input into the inter-directional collaborative inference graph module DCRG. Multiple directional convolutional kernels are used to extract directional semantic nodes, constructing an inter-directional collaborative graph and performing information transmission, outputting directional consistency enhanced features; specifically: Buildings and roads in remote sensing images exhibit significant directionality; however, convolutional and window self-attention mechanisms lack cross-directional collaborative modeling capabilities, performing poorly when encountering occlusion, structural misalignment, and other issues. To address this problem, this invention proposes a Directional Cooperative Reasoning Graph Module (DCRG). (See attached diagram) Figure 4 As shown, this invention first designs four parallel directional convolution kernels to explicitly decouple the directional components that depend on local space, and treats the features of the four directions as nodes on the graph, thereby improving the consistency of the structure through information propagation between directions.
[0038] First path: Convolutional kernel dimension is... The direction is from top to bottom. This approach focuses on extracting contextual information in the vertical direction. After the feature map undergoes a cross-attention mechanism, each column of pixels is scanned using an asymmetric convolution kernel to calculate the vertical dependencies of pixels within that column. It responds strongly to vertical edges and vertically arranged objects, capturing intensity variations and consistency within a "column".
[0039] Second path: Convolutional kernel dimension is The direction is from top to bottom. This approach focuses on extracting horizontal contextual information. After the feature map undergoes a cross-attention mechanism, each row of pixels is scanned using an asymmetric convolution kernel to calculate the horizontal dependencies of pixels within that row. It responds strongly to horizontal edges and horizontally arranged objects, capturing intensity variations and consistency within a "row".
[0040] Third path: Convolutional kernel dimension is The direction is from the top left to the bottom right. It focuses on extracting contextual information along the main diagonal direction. This convolutional path scans a specific diagonal region along the top-left to bottom-right diagonal direction, calculating dependencies along that direction. It responds strongly to the edges of the main diagonal and structures arranged along this direction, effectively capturing continuity and corresponding patterns in that direction.
[0041] Fourth path: Convolution kernel dimension is The direction is from the top left to the bottom right. The direction is also from top left to bottom right, but because the dimension of the convolution kernel is... Therefore, it captures structural information in another diagonal direction. It responds strongly to structures in the anti-diagonal direction, working in conjunction with the third path to cover the main diagonal local context and corresponding dependencies.
[0042] This yields four feature maps, represented as follows: Among them, subscript This represents four convolutional kernels operating in different directions. Each directional branch is then subjected to global average pooling to obtain a node-level representation. (12) The 4-way node-level representation is stacked to form a directional node matrix: (13) After another linear layer mapping, we obtain the linear projection of the directional node: (14) To construct a directional collaboration graph and achieve information transfer, this invention further calculates the similarity of directional nodes and constructs a normalized adjacency matrix between directional nodes: (15) (16) The following steps are performed to transmit information and update the residuals: (17) Channel-level weights are generated through Sigmoid function mapping: (18) Finally, weights are applied to the directional features and the output is concatenated: (19) (20) in, This indicates a splicing operation. express Convolution, Batch Normalization, and ReLU nonlinear activation operations.
[0043] Through the DCRG module, the network can explicitly model the synergistic relationships between horizontal, vertical and diagonal directions, improving the continuity and stability of directional structures such as roads and building edges.
[0044] In step S4, the direction consistency enhancement feature is input into the semantic graph self-evolution module SGSE. A region semantic graph is constructed through a bidirectional attention mechanism between learnable semantic nodes and pixel features. After multiple graph evolutions, semantic information is backpropagated to the pixel space, and a global consistency enhancement feature is output. Specifically: Remote sensing images exhibit significant intra-class heterogeneity and inter-class similarity, with gradual transitions in feature boundaries. Under these conditions, relying solely on local features is insufficient to guarantee global consistency in segmentation results. To address this issue, this invention proposes a Semantic Graph Self-Evolution Module (SGSE), which utilizes learnable semantic nodes to construct a regional-level structural graph. Through self-evolution among nodes, the semantic consistency of the entire graph is improved.
[0045] Let the module input be... After convolution and pixel flattening, the result is: (twenty one) Among them, parameters ,get: (twenty two) in, It is a linear fully connected layer with parameters. This represents the number of channels after passing through a linear fully connected layer.
[0046] Furthermore, introduce One learnable query node: (twenty three) Each semantic node query can be viewed as a region-level semantic anchor point, used to aggregate corresponding region semantic information from pixel features. The calculation of the... The semantic node and the first Attention weights between pixel tokens: (twenty four) Based on this attention weight, the pixel value representation is weighted and summed to obtain the initial semantic node: (25) All semantic nodes form the initial node matrix: (26) This step is equivalent to aggregating K region nodes in the entire graph using K learnable semantic anchors, implicitly completing the region partitioning and node graph construction.
[0047] Subsequently, self-evolutionary graph reasoning is performed in the semantic node space. For the th Sub-evolution, computing node-level Query, Key, and Value: (27) Based on this, the semantic relationships between nodes can be calculated: (28) Message aggregation is performed based on the semantic relationships between nodes, and node representations are updated using residuals: (29) go through After the semantic graph evolves, the final semantic nodes are obtained: .
[0048] Then, the evolved semantic nodes are mapped back to the pixel feature channel dimension: (30) Using the aforementioned pixel-node attention weights, global semantic node information is re-injected into the pixel space: (31) Finally, the enhanced pixel sequence is restored to a spatial feature map: (32) Through the SGSE module, the network is able to model region-level semantic relationships globally, enhance consistency between spatially dispersed but semantically related regions, and suppress category fragmentation.
[0049] In step S5, based on the global consistency enhancement features, the semantic segmentation result of the remote sensing image is output through the classification layer.
[0050] The following uses the Vaihingen dataset as an example, combined with the appendix. Figure 2 Appendix Figure 3 Appendix Figure 4 Appendix Figure 5 The specific implementation of the present invention, HSRNet (Hierarchical Semantic Reasoning Network), is explained below: This invention, HSRNet, uses ConvNeXt-tiny as its backbone network to perform preliminary feature extraction on images, as shown in the attached diagram. Figure 2 As shown in the image above, input a remote sensing image, denoted as... The input image size is The width and height are both 512, and the number of channels is 3. After preliminary feature extraction by the backbone network ConvNeXt-tiny, four feature maps of different dimensions and sizes in four stages are obtained, denoted as:
[0051] To balance the effectiveness of the module with the increase in the number of parameters and computational cost it brings, feature maps are only used in the second and third stages. Then, a CSSR module is inserted to create a residual connection. Fourth-stage features. After being processed by the CBRU module, it is recorded as follows: The CBRU module refers to the concatenated operation of convolution transformation, batch normalization, ReLU nonlinear activation, and upsampling. (Third-stage feature map) New features are obtained by connecting the CSSR module and the residual. , and Adding them together gives us new features. Second-stage feature map Similarly, new features are obtained by connecting the CSSR module and the residual. , and The feature maps obtained by addition are input into the CBRU module to obtain a new feature map, denoted as . , and the first stage feature map Direct element-wise addition, followed by the CBR module, yields a new feature, denoted as... The CBR module is based on the CBRU module but removes the final upsampling operation.
[0052] New features The three inputs are fed in parallel into the DCGR Block module, and the output of this module is then passed through the SGSE module to obtain the final result.
[0053] The training process calculates a weighted sum of the cross-entropy loss function and the Dice loss function based on the final output feature map. Under the constraint of the above total loss function, the model is trained and optimized until convergence.
[0054] Based on the above experimental steps, experiments were conducted on remote sensing image datasets such as Vaihingen, Potsdam, and LoveDA. The mean Intersection over Union (mIoU) value for each class was used as an indicator of the network model's performance in semantic segmentation tasks on a given dataset. Numerical and visualization results are shown below. UNetFormer is a method proposed in a 2022 paper published in the top remote sensing journal ISPRS Journal of Photogrammetry and Remote Sensing, while PPMamba is a method proposed in a 2025 paper published in the renowned remote sensing journal IEEE Geoscience and Remote Sensing Letters.
[0055] Visualization results as follows Figure 6 As shown in the figure, quantitative analysis reveals that the mIoU value of this invention on the Vaihingen dataset is significantly higher than that of the UNetFormer and PPMamba methods. Visualization shows that this invention achieves higher completeness and semantic consistency in recognizing blue buildings (within red boxes), cyan Low Vegetation category targets, and yellow car targets compared to the UNetFormer and PPMamba methods. Therefore, the quantitative and visualization results demonstrate that the proposed method possesses certain advantages and superiority.
[0056] Furthermore, the above figures are merely illustrative of the processes included in the method according to exemplary embodiments of the present invention, and are not intended to be limiting. It is readily understood that the processes shown in the above figures do not indicate or limit the temporal order of these processes. Additionally, it is readily understood that these processes may be executed synchronously or asynchronously, for example, in multiple modules.
[0057] Other embodiments of the invention will readily occur to those skilled in the art upon consideration of the specification and practice of the invention herein. This application is intended to cover any variations, uses, or adaptations of the invention that follow the general principles of the invention and include common knowledge or customary techniques in the art not disclosed herein. The specification and embodiments are to be considered exemplary only, and the true scope and spirit of the invention are indicated by the claims.
[0058] It should be understood that the present invention is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of the invention is defined only by the appended claims.
Claims
1. A semantic segmentation method for remote sensing images based on multi-level semantic reasoning, characterized in that, Includes the following steps: Remote sensing images are input into a backbone feature extraction network to obtain multi-scale feature maps; the multi-scale feature maps include first-stage feature maps. Second-stage feature map Third-stage feature map and the fourth stage feature map ; The characteristics of the fourth stage After being recorded by the CBRU module ; The third stage feature map New features are obtained by connecting the CSSR module and the residual. ,Will and Adding them together gives us new features. ; The second-stage feature map The input is the Cross-Scale Semantic Routing (CSSR) module, which is then connected to the residual to obtain new features. ,Will and The feature maps obtained by addition are input into the CBRU module to obtain a new feature map, denoted as . The new feature map is denoted as Compared with the first stage feature map Direct element-wise addition, followed by the CBR module, yields a new feature, denoted as... ; New features As a three-way parallel input direction cooperative reasoning graph module, DCRG extracts directional semantic nodes through multiple directional convolutional kernels, constructs an inter-directional cooperative graph and performs information transmission, and outputs directional consistency enhancement features. The direction consistency enhancement feature is input into the semantic graph self-evolution module SGSE. A region semantic graph is constructed through a bidirectional attention mechanism between learnable semantic nodes and pixel features. After multiple graph evolutions, semantic information is back-transmitted to the pixel space, and global consistency enhancement features are output. Based on the global consistency enhancement features, the semantic segmentation results of the remote sensing image are output through the classification layer.
2. The method according to claim 1, characterized in that, The execution of the cross-scale semantic routing module (CSSR) includes: The input features are convolutionally reduced to obtain the dimensionality-reduced features; On the dimensionality reduction features, three different window sizes of local self-attention mechanisms are applied to generate feature maps corresponding to different scales; The feature maps of different scales are stitched together; Pixel-level dynamic routing weights are generated for each scale using convolutional layers and the softmax function; The feature maps of different scales are weighted and fused according to the pixel-level dynamic routing weights to obtain fused features; The fused features are residually concatenated with the processed original input features and upsampled to restore the original resolution, outputting the enhanced cross-scale features.
3. The method according to claim 1, characterized in that, The execution of the Directional Collaborative Reasoning Graph (DCRG) module includes: The input features are processed by four convolutional kernels or attention mechanisms with different directions to extract feature components in different directions and obtain multi-directional feature maps. Global pooling is performed on the directional feature map of each path to obtain the corresponding directional node representation; Multiple directional node representations are stacked to form a directional node matrix; Calculate the similarity between the directional nodes and construct a normalized adjacency matrix between the directional nodes; Information propagation and updating are performed between nodes in the stated direction based on the normalized adjacency matrix. Channel-level weights are generated based on the updated node information, and the multi-directional feature maps are weighted and fused to output directional enhanced features.
4. The method according to claim 3, characterized in that, The four paths of convolutional kernels or attention mechanisms with different directions include a first path for extracting vertical features, a second path for extracting horizontal features, a third path for extracting features along the main diagonal, and a fourth path for extracting features along the anti-diagonal.
5. The method according to claim 1, characterized in that, The execution of the semantic graph self-evolution module SGSE includes: The input features are transformed and flattened to obtain pixel feature representations; Initialize a set of learnable semantic nodes; Calculate the attention weights between the pixel feature representation and the learnable semantic node, and obtain the node-level representation by aggregation; Based on the node-level representation, self-attention calculation and information update are performed between nodes to complete a self-evolution on the graph structure. The evolved and updated node representations are projected back into the pixel space and fused with the original pixel features to enhance the output features with improved semantic consistency.
6. The method according to claim 5, characterized in that, The process of projecting the evolved and updated node representation back into the pixel space specifically involves: mapping the updated node representation into pixel-level semantic enhancement weights through a linear transformation, multiplying them element-wise with the flattened original pixel features, and then outputting them after shape restoration.
7. The method according to claim 1, characterized in that, The method also includes a training process, in which a weighted sum of the cross-entropy loss function and the Dice loss function is calculated based on the final output feature map. Under the constraint of the total loss function, the model is trained and optimized until convergence.
8. The method according to claim 1, characterized in that, The CBRU module refers to the concatenated operation of convolution transformation, batch normalization, ReLU nonlinear activation, and upsampling.
9. The method according to claim 8, characterized in that, The CBR module is based on the CBRU module but with the final upsampling operation removed.
10. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the remote sensing image semantic segmentation method based on multi-level semantic reasoning as described in any one of claims 1 to 9.